Files
zerto-ai-rewind/CONTEXT.md
T
justinandClaude Opus 5 bad5a5706f feat(guard): read the Zerto task, and ask before guarding unknown tools
Two changes that both come from the same mistake: assuming an answer
instead of reading one.

1. Read the task after inserting a checkpoint.

   POST /v1/vpgs/{id}/checkpoints returns a TASK ID, not a result. A 200
   only means queued. The outcome lives in GET /v1/tasks/{id} under
   Status.State: 1 InProgress, 4 Failed, 5 Stopped, 6 Completed, with
   4/5/6 terminal.

   tag_vpgs now waits for that task and refuses unless it Completed, and
   reports task_id and task_state. Measured: two inserts fired back to
   back at one VPG give Completed for the first and Failed for the
   second. That is exactly the case an earlier comment in this file
   called "silently dropped" -- it was never silent, we just never read
   the task. Comment corrected.

   Previously a failed insert surfaced only as wait_for_tag timing out
   45s later with a misleading hint about Azure/AWS. Now it says the
   task failed and the operation did not happen.

2. Unknown tools ask the human instead of passing through.

   The catalog is opt-in, so an unlisted tool ran unguarded. But the set
   of mutating tools is unbounded and grows with every MCP installed,
   while the set of read-only ones is small, so a mutating-only list is
   permanently behind and being behind fails open.

   Adds read_only_tools and zerto_check_tool(server, tool, vm) returning
   read_only / mutating / unknown. unknown does not mean safe: it means
   nobody classified it, so the tool hands the agent a question to put
   to the human, and the human decides whether to checkpoint. On yes the
   agent guards; on no it runs and says plainly that Zerto cannot rewind
   it; if they want it remembered, zerto_add_mutating_tool.

   This stays advisory. An MCP server cannot see or block another
   server's tool calls, so real enforcement belongs in a host PreToolUse
   hook. The skill carries the flow.

Known issue, not addressed here: two concurrent tool calls race on
Keycloak token acquisition in the shared client and one gets HTTP 401.

pytest 41 passed (12 new).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-21 14:47:19 -04:00

47 lines
3.2 KiB
Markdown

# Zerto AI Rewind
PoC MCP that teaches an agent to discover Zerto protection, pin a tagged checkpoint before changing a VM, and recover a file from that tag. If the loop works, these tools are the delta to put in official ZVM MCP.
## Language
**Tagged checkpoint**:
A named bookmark in a VPG journal, inserted by `POST /v1/vpgs/{id}/checkpoints` (`startVpgTaggedCheckpointInsert`). Crash-consistent write-order only; not application-quiesced unless someone scripted that separately. `CheckpointName` is the only field the API accepts, so agent and intent go in the name: `ai:<agent> | <action> | vm=<vm> | change=<id> | <utc>`. Inserts are async tasks and are silently dropped if fired back to back at one VPG; insert, then wait until listed.
_Avoid_: user checkpoint, snapshot, backup, restore point (unqualified)
**VPG**:
A Virtual Protection Group. One to many VMs sharing a journal. A VM can belong to at most three VPGs, recovered to different sites.
_Avoid_: job, policy, replication group
**Rewind**:
The agent loop: find protection, tag every protecting VPG, mutate, then bounded recover. Not a Zerto product name.
_Avoid_: failover (that's DR), undo (that's git or Moholo)
**Bounded recover**:
FLR, offsite clone, or failover test. Failover Live is not a rewind tool.
_Avoid_: recover (unqualified), fail back, restore the VPG
**File-level recovery (FLR)**:
Mount a VM from a journal checkpoint and pull files. The VM stays up. 10.9 FLR Operator RBAC is broken; Administrator is the documented workaround.
Runs at the VPG's **recovery** site, so this server supports it only for locally replicated VPGs (protected site == recovery site). Paths are partition-rooted; on Windows the drive letter is the partition.
_Avoid_: file restore (unqualified), instant restore (local-journal VMs only, not v1)
**find_protection**:
Resolve a VM name, hostname, or Zerto vmIdentifier to exactly one VM and every VPG it is in. Zero or two-plus VMs is a hard stop.
_Avoid_: GetVms (that's the raw inventory call)
**Protecting VPG**:
A VPG whose status is MeetingSLA or a NotMeetingSLA variant, and whose substatus is not a sync. Only these get tagged. 10.9 status 0 is Initializing, not Protecting. A resync deletes existing checkpoints.
_Avoid_: healthy, in sync, Protecting (as status 0)
**Mutating catalog**:
The opt-in list of MCP tools that must call `zerto_guard_before_mutate` first. Paired with `read_only_tools`, the list known not to change a guest. A tool in neither is **unknown**, which is the normal case: the agent asks the human whether to checkpoint rather than assuming either way.
_Avoid_: denylist, hold-everything, treating unknown as safe
**Zerto task**:
Write operations return a task id, not a result. A 200 means queued. `GET /v1/tasks/{id}` carries the real outcome in `Status.State`: 1 InProgress, 4 Failed, 5 Stopped, 6 Completed (terminal is 4/5/6). A second tagged-checkpoint insert fired at a VPG while the first runs comes back Failed, which is only visible if the task is read.
_Avoid_: treating HTTP 200 as success
**Official ZVM MCP**:
HPE Zerto 10.9 MCP (`ZVM.MCP`): inventory, VPG settings, failover test. Not in the demo path. This PoC is one server.
_Avoid_: Zerto MCP (unqualified when you mean this repo)