wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by the VPG's protected site and the spread is enormous: 5s on vSphere, 60s on Azure, 630s on AWS, measured in one estate. A single constant cannot serve all three. Generous for vSphere, where a tag surfaces in about 4s. Impossible for AWS, where it takes ~128s. So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one, and the checkpoint landed a minute after the agent had been told there was no rewind point. That is worse than the failure it guards against. It is silent, it reads as correct in the log, and it blocks legitimate work on every cloud-protected VM in the estate. The wait is now measured: cadence_seconds() takes the median gap of recent checkpoints, tag_wait_budget() turns that into 2x cadence plus headroom, clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is not polled every 1.5s. Visibility does not scale linearly with cadence, because the insert creates its own off-cadence checkpoint, which is why this is a bounded multiple rather than a proportion. wait_for_tag also checks once before sleeping, so an already-present tag returns without a poll cycle. The failure message now names the measured cadence and says a completed Zerto task with no visible checkpoint means the wait was short, not that the insert was rejected. That was the exact wrong conclusion the old message invited. Hook budgets follow: too small a budget there just relocates the false denial from the guard into the hook, since a cancelled hook has its output discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 -> 330, settings timeout 180 -> 360, both above the 300s cap. Verified live against ZVM 10.9.10. Guard on win2019-1 (VPG CMH-AWS-1, AWS-protected) previously denied at 63s; now succeeds in 110.6s with checkpoint 186. jp-ubuntu unchanged at 6.6s, so vSphere pays nothing for this. Through the hook itself: win2019-1 allowed in 111s with cp 187, jp-ubuntu allowed in 9s with cp 23198. pytest 69 passed (6 new, including one asserting the budget exceeds the visibility actually measured on each of the three platforms). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
zerto-ai-rewind
PoC MCP that makes an agent pin a Zerto tagged checkpoint before it changes a VM, then pull a file back from that tag.
If the loop works, these tools are the delta to add to official ZVM MCP (ZVM.MCP, 10.9). This repo is one MCP process for the demo. It is not a second full ZVM catalog.
What it does
zerto_check_tool— before running anything against a VM:read_only(go),mutating(guard first), orunknown(ask the human whether to checkpoint).zerto_find_protection— VM name, hostname, or vmIdentifier to exactly one VM and every VPG. Zero or two-plus VMs: stop.zerto_create_tagged_checkpoint/zerto_guard_before_mutate— same tag on every protecting VPG, wait until listed. The name records which agent and what it is doing:ai:<agent> | <action> | vm=<vm> | change=<id> | <utc>.zerto_recover_file— FLR after a human setsconfirmed=true. Linux and Windows guest paths. Locally replicated VPGs only: FLR runs at the VPG's recovery site. Reports its own unmount;zerto_list_flr_sessions/zerto_end_flr_sessionfind and reap a mount orphaned by a crashed recovery.- Two catalogs —
mutating_tools(guard first) andread_only_tools(safe). Both cover Linux (ssh,ansible) and Windows (winrm,powershell,smb), and both are illustrative, not exhaustive. A tool in neither is unknown, not safe: ask the human.
Official ZVM MCP already has inventory and failover test. It does not insert tagged checkpoints or run FLR.
Setup
Python 3.12+.
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
cp config.example.json config.json
# edit zerto_url, username, password
Keycloak password-grant, client_id zerto-client on 10.x. Appliance certs are self-signed; verify_tls defaults to false.
stdio MCP (Claude Desktop, VS Code, Cursor, OpenCode):
{
"mcpServers": {
"zerto-rewind": {
"command": "zerto-rewind-mcp",
"env": {
"ZERTO_REWIND_CONFIG": "/absolute/path/to/config.json"
}
}
}
}
Copy skills/zerto-rewind/SKILL.md into the client's skill path.
pytest
Demo
Protected app VM. Agent is about to edit a guest config file.
- Guard: discover VPG set, insert tagged checkpoint, wait.
- Agent writes the bad config.
- Human confirms.
zerto_recover_filefrom that tag.
Git never had the file. RPO is the journal, not last night's backup.
Certified for this PoC
vSphere ZVM 10.x and ZCA on AWS/Azure, same REST paths. HVM is out (separate swagger). Failover Live is not a tool.
Tagged checkpoints do work when the protected site is Azure or AWS. The 9.0 API
reference says they cannot be inserted; that is wrong on 10.9.10, where both were
accepted and the task reached Completed.
What differs is latency and granularity, both set by the protected site:
| protected at | journal gap | tag visible after |
|---|---|---|
| vSphere | 5s | ~4s |
| Azure | 60s | ~34s |
| AWS | 630s | ~128s |
The tagged checkpoint is also stamped about 30s after the insert request, so on a cloud-protected VPG a prompt mutation can land inside the checkpoint meant to precede it. Recover from the newest checkpoint that already existed when the guard ran, not from the tag.
Not this product
Moholo Agent Rewind snapshots the agent's laptop tools. This server uses the Zerto journal as the snapshot store. Do not copy their file blobs.
Upstream
Ask Zerto engineering to add to ZVM.MCP: find-by-unique-VM-with-all-VPGs, tagged checkpoint insert that waits, FLR. Keep VPG settings CRUD where it already is.