The repo said, in four places, that a tagged checkpoint cannot be inserted when the protected site is Azure or AWS. That came from the 9.0 API reference. It is wrong on 10.9.10. Measured against the lab: both an AWS-protected and an Azure-protected VPG accepted the insert, the Zerto task reached Completed, and the checkpoint appeared in the journal. This is the second doc claim this repo carried that live 10.9 contradicts, after the VPG status enum. What actually differs is latency and granularity, and both are set by the PROTECTED site, not the recovery site: protected at journal gap tag visible after vSphere 5s ~4s Azure 60s ~34s AWS 630s ~128s There is a sharper consequence than slowness. The tagged checkpoint is stamped about 30s AFTER the insert request, so on a cloud-protected VPG an agent that mutates promptly puts its change inside the checkpoint that was supposed to precede it. Recovering from that tag would restore the broken state. The docs now say to use the newest checkpoint that already existed when the guard ran. Two of the four were message text, not prose. wait_for_tag's timeout message asserted the insert was unsupported when the real cause was its own 45s budget being far too short for a VPG that checkpoints every 630s, so it told operators the wrong thing at exactly the wrong moment. It now says to check the Zerto task before concluding the insert failed. The 45s timeout itself is still wrong for cloud sources and needs to become cadence-aware. That is a behaviour change, so it is not in this commit. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
91 lines
3.6 KiB
Markdown
91 lines
3.6 KiB
Markdown
# zerto-ai-rewind
|
|
|
|
PoC MCP that makes an agent pin a Zerto tagged checkpoint before it changes a VM, then pull a file back from that tag.
|
|
|
|
If the loop works, these tools are the delta to add to official ZVM MCP (`ZVM.MCP`, 10.9). This repo is one MCP process for the demo. It is not a second full ZVM catalog.
|
|
|
|
## What it does
|
|
|
|
0. `zerto_check_tool` — before running anything against a VM: `read_only` (go), `mutating` (guard first), or `unknown` (**ask the human** whether to checkpoint).
|
|
1. `zerto_find_protection` — VM name, hostname, or vmIdentifier to exactly one VM and every VPG. Zero or two-plus VMs: stop.
|
|
2. `zerto_create_tagged_checkpoint` / `zerto_guard_before_mutate` — same tag on every protecting VPG, wait until listed. The name records which agent and what it is doing: `ai:<agent> | <action> | vm=<vm> | change=<id> | <utc>`.
|
|
3. `zerto_recover_file` — FLR after a human sets `confirmed=true`. Linux and Windows guest paths. Locally replicated VPGs only: FLR runs at the VPG's recovery site. Reports its own unmount; `zerto_list_flr_sessions` / `zerto_end_flr_session` find and reap a mount orphaned by a crashed recovery.
|
|
4. Two catalogs — `mutating_tools` (guard first) and `read_only_tools` (safe). Both cover Linux (`ssh`, `ansible`) and Windows (`winrm`, `powershell`, `smb`), and both are illustrative, not exhaustive. A tool in neither is **unknown, not safe**: ask the human.
|
|
|
|
Official ZVM MCP already has inventory and failover test. It does not insert tagged checkpoints or run FLR.
|
|
|
|
## Setup
|
|
|
|
Python 3.12+.
|
|
|
|
```bash
|
|
python3 -m venv .venv
|
|
source .venv/bin/activate
|
|
pip install -e ".[dev]"
|
|
cp config.example.json config.json
|
|
# edit zerto_url, username, password
|
|
```
|
|
|
|
Keycloak password-grant, client_id `zerto-client` on 10.x. Appliance certs are self-signed; `verify_tls` defaults to false.
|
|
|
|
stdio MCP (Claude Desktop, VS Code, Cursor, OpenCode):
|
|
|
|
```json
|
|
{
|
|
"mcpServers": {
|
|
"zerto-rewind": {
|
|
"command": "zerto-rewind-mcp",
|
|
"env": {
|
|
"ZERTO_REWIND_CONFIG": "/absolute/path/to/config.json"
|
|
}
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
Copy `skills/zerto-rewind/SKILL.md` into the client's skill path.
|
|
|
|
```bash
|
|
pytest
|
|
```
|
|
|
|
## Demo
|
|
|
|
Protected app VM. Agent is about to edit a guest config file.
|
|
|
|
1. Guard: discover VPG set, insert tagged checkpoint, wait.
|
|
2. Agent writes the bad config.
|
|
3. Human confirms.
|
|
4. `zerto_recover_file` from that tag.
|
|
|
|
Git never had the file. RPO is the journal, not last night's backup.
|
|
|
|
## Certified for this PoC
|
|
|
|
vSphere ZVM 10.x and ZCA on AWS/Azure, same REST paths. HVM is out (separate swagger). Failover Live is not a tool.
|
|
|
|
Tagged checkpoints **do** work when the protected site is Azure or AWS. The 9.0 API
|
|
reference says they cannot be inserted; that is wrong on 10.9.10, where both were
|
|
accepted and the task reached `Completed`.
|
|
|
|
What differs is latency and granularity, both set by the **protected** site:
|
|
|
|
| protected at | journal gap | tag visible after |
|
|
|---|---|---|
|
|
| vSphere | 5s | ~4s |
|
|
| Azure | 60s | ~34s |
|
|
| AWS | 630s | ~128s |
|
|
|
|
The tagged checkpoint is also stamped about 30s *after* the insert request, so on a
|
|
cloud-protected VPG a prompt mutation can land *inside* the checkpoint meant to
|
|
precede it. Recover from the newest checkpoint that already existed when the guard
|
|
ran, not from the tag.
|
|
|
|
## Not this product
|
|
|
|
[Moholo Agent Rewind](https://github.com/moholo-founder/agent-rewind) snapshots the agent's laptop tools. This server uses the Zerto journal as the snapshot store. Do not copy their file blobs.
|
|
|
|
## Upstream
|
|
|
|
Ask Zerto engineering to add to `ZVM.MCP`: find-by-unique-VM-with-all-VPGs, tagged checkpoint insert that waits, FLR. Keep VPG settings CRUD where it already is.
|