Files
zerto-ai-rewind/README.md
T
justinandClaude Opus 5 f930b84615 feat(flr): make FLR session lifecycle visible and reapable
zerto_recover_file already tore its session down in a finally block, but
three gaps meant a mount could stay up on the recovery site with nothing
tracking it. FLR cannot run during clone, test, live failover or EJC, so
a stuck session blocks the next recovery.

1. An unmount failure was swallowed (`except ZertoError: pass`). The
   caller got ok=true and never learned the mount was still up. The
   teardown result is now reported in the response as `unmount`, with a
   `warning` when it fails. ok stays true when the bytes did land -- the
   recovery genuinely succeeded -- but the caller is told.

2. If start_flr succeeded on the ZVM while its response failed to parse,
   session_id stayed None and the finally block did nothing, leaking a
   session the process never knew the id of. Teardown now snapshots live
   session ids before starting and reaps anything new that appeared,
   leaving other operators' sessions alone.

3. Nothing could see or clear an orphan left by a crashed process, since
   the finally block only runs if the process survives. Two new tools:

   - zerto_list_flr_sessions: every session the ZVM knows about.
     live_only (default true) keeps the ones still holding a mount;
     ended and failed sessions linger as history and hold nothing.
   - zerto_end_flr_session: unmount one. Gated on confirmed=true,
     matching the other destructive tools, because ending a session
     someone else is pulling files from will interrupt them.

Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean
run, the confirm gate refuses without a human yes, a real recovery from
checkpoint 1368 returned 158 bytes and reported
unmount.ok=true with the session id it ended, and 0 live sessions
remained afterwards.

pytest 27 passed (5 new, including fakes covering the swallowed-failure
and orphan-reap paths).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-21 13:40:53 -04:00

2.7 KiB

zerto-ai-rewind

PoC MCP that makes an agent pin a Zerto tagged checkpoint before it changes a VM, then pull a file back from that tag.

If the loop works, these tools are the delta to add to official ZVM MCP (ZVM.MCP, 10.9). This repo is one MCP process for the demo. It is not a second full ZVM catalog.

What it does

  1. zerto_find_protection — VM name, hostname, or vmIdentifier to exactly one VM and every VPG. Zero or two-plus VMs: stop.
  2. zerto_create_tagged_checkpoint / zerto_guard_before_mutate — same tag on every protecting VPG, wait until listed. The name records which agent and what it is doing: ai:<agent> | <action> | vm=<vm> | change=<id> | <utc>.
  3. zerto_recover_file — FLR after a human sets confirmed=true. Reports its own unmount; zerto_list_flr_sessions / zerto_end_flr_session find and reap a mount orphaned by a crashed recovery.
  4. Mutating catalog — opt-in list of MCP tools that must be guarded. Unlisted tools pass through. Users add entries.

Official ZVM MCP already has inventory and failover test. It does not insert tagged checkpoints or run FLR.

Setup

Python 3.12+.

python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cp config.example.json config.json
# edit zerto_url, username, password

Keycloak password-grant, client_id zerto-client on 10.x. Appliance certs are self-signed; verify_tls defaults to false.

stdio MCP (Claude Desktop, VS Code, Cursor, OpenCode):

{
  "mcpServers": {
    "zerto-rewind": {
      "command": "zerto-rewind-mcp",
      "env": {
        "ZERTO_REWIND_CONFIG": "/absolute/path/to/config.json"
      }
    }
  }
}

Copy skills/zerto-rewind/SKILL.md into the client's skill path.

pytest

Demo

Protected app VM. Agent is about to edit a guest config file.

  1. Guard: discover VPG set, insert tagged checkpoint, wait.
  2. Agent writes the bad config.
  3. Human confirms.
  4. zerto_recover_file from that tag.

Git never had the file. RPO is the journal, not last night's backup.

Certified for this PoC

vSphere ZVM 10.x and ZCA on AWS/Azure, same REST paths. HVM is out (separate swagger). Failover Live is not a tool.

Tagged checkpoints cannot be inserted when the protected site is Azure or AWS (Zerto API). Point this server at the vSphere protected ZVM.

Not this product

Moholo Agent Rewind snapshots the agent's laptop tools. This server uses the Zerto journal as the snapshot store. Do not copy their file blobs.

Upstream

Ask Zerto engineering to add to ZVM.MCP: find-by-unique-VM-with-all-VPGs, tagged checkpoint insert that waits, FLR. Keep VPG settings CRUD where it already is.