Two changes that both come from the same mistake: assuming an answer
instead of reading one.
1. Read the task after inserting a checkpoint.
POST /v1/vpgs/{id}/checkpoints returns a TASK ID, not a result. A 200
only means queued. The outcome lives in GET /v1/tasks/{id} under
Status.State: 1 InProgress, 4 Failed, 5 Stopped, 6 Completed, with
4/5/6 terminal.
tag_vpgs now waits for that task and refuses unless it Completed, and
reports task_id and task_state. Measured: two inserts fired back to
back at one VPG give Completed for the first and Failed for the
second. That is exactly the case an earlier comment in this file
called "silently dropped" -- it was never silent, we just never read
the task. Comment corrected.
Previously a failed insert surfaced only as wait_for_tag timing out
45s later with a misleading hint about Azure/AWS. Now it says the
task failed and the operation did not happen.
2. Unknown tools ask the human instead of passing through.
The catalog is opt-in, so an unlisted tool ran unguarded. But the set
of mutating tools is unbounded and grows with every MCP installed,
while the set of read-only ones is small, so a mutating-only list is
permanently behind and being behind fails open.
Adds read_only_tools and zerto_check_tool(server, tool, vm) returning
read_only / mutating / unknown. unknown does not mean safe: it means
nobody classified it, so the tool hands the agent a question to put
to the human, and the human decides whether to checkpoint. On yes the
agent guards; on no it runs and says plainly that Zerto cannot rewind
it; if they want it remembered, zerto_add_mutating_tool.
This stays advisory. An MCP server cannot see or block another
server's tool calls, so real enforcement belongs in a host PreToolUse
hook. The skill carries the flow.
Known issue, not addressed here: two concurrent tool calls race on
Keycloak token acquisition in the shared client and one gets HTTP 401.
pytest 41 passed (12 new).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
76 lines
3.1 KiB
Markdown
76 lines
3.1 KiB
Markdown
# zerto-ai-rewind
|
|
|
|
PoC MCP that makes an agent pin a Zerto tagged checkpoint before it changes a VM, then pull a file back from that tag.
|
|
|
|
If the loop works, these tools are the delta to add to official ZVM MCP (`ZVM.MCP`, 10.9). This repo is one MCP process for the demo. It is not a second full ZVM catalog.
|
|
|
|
## What it does
|
|
|
|
0. `zerto_check_tool` — before running anything against a VM: `read_only` (go), `mutating` (guard first), or `unknown` (**ask the human** whether to checkpoint).
|
|
1. `zerto_find_protection` — VM name, hostname, or vmIdentifier to exactly one VM and every VPG. Zero or two-plus VMs: stop.
|
|
2. `zerto_create_tagged_checkpoint` / `zerto_guard_before_mutate` — same tag on every protecting VPG, wait until listed. The name records which agent and what it is doing: `ai:<agent> | <action> | vm=<vm> | change=<id> | <utc>`.
|
|
3. `zerto_recover_file` — FLR after a human sets `confirmed=true`. Linux and Windows guest paths. Locally replicated VPGs only: FLR runs at the VPG's recovery site. Reports its own unmount; `zerto_list_flr_sessions` / `zerto_end_flr_session` find and reap a mount orphaned by a crashed recovery.
|
|
4. Two catalogs — `mutating_tools` (guard first) and `read_only_tools` (safe). Both cover Linux (`ssh`, `ansible`) and Windows (`winrm`, `powershell`, `smb`), and both are illustrative, not exhaustive. A tool in neither is **unknown, not safe**: ask the human.
|
|
|
|
Official ZVM MCP already has inventory and failover test. It does not insert tagged checkpoints or run FLR.
|
|
|
|
## Setup
|
|
|
|
Python 3.12+.
|
|
|
|
```bash
|
|
python3 -m venv .venv
|
|
source .venv/bin/activate
|
|
pip install -e ".[dev]"
|
|
cp config.example.json config.json
|
|
# edit zerto_url, username, password
|
|
```
|
|
|
|
Keycloak password-grant, client_id `zerto-client` on 10.x. Appliance certs are self-signed; `verify_tls` defaults to false.
|
|
|
|
stdio MCP (Claude Desktop, VS Code, Cursor, OpenCode):
|
|
|
|
```json
|
|
{
|
|
"mcpServers": {
|
|
"zerto-rewind": {
|
|
"command": "zerto-rewind-mcp",
|
|
"env": {
|
|
"ZERTO_REWIND_CONFIG": "/absolute/path/to/config.json"
|
|
}
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
Copy `skills/zerto-rewind/SKILL.md` into the client's skill path.
|
|
|
|
```bash
|
|
pytest
|
|
```
|
|
|
|
## Demo
|
|
|
|
Protected app VM. Agent is about to edit a guest config file.
|
|
|
|
1. Guard: discover VPG set, insert tagged checkpoint, wait.
|
|
2. Agent writes the bad config.
|
|
3. Human confirms.
|
|
4. `zerto_recover_file` from that tag.
|
|
|
|
Git never had the file. RPO is the journal, not last night's backup.
|
|
|
|
## Certified for this PoC
|
|
|
|
vSphere ZVM 10.x and ZCA on AWS/Azure, same REST paths. HVM is out (separate swagger). Failover Live is not a tool.
|
|
|
|
Tagged checkpoints cannot be inserted when the **protected** site is Azure or AWS (Zerto API). Point this server at the vSphere protected ZVM.
|
|
|
|
## Not this product
|
|
|
|
[Moholo Agent Rewind](https://github.com/moholo-founder/agent-rewind) snapshots the agent's laptop tools. This server uses the Zerto journal as the snapshot store. Do not copy their file blobs.
|
|
|
|
## Upstream
|
|
|
|
Ask Zerto engineering to add to `ZVM.MCP`: find-by-unique-VM-with-all-VPGs, tagged checkpoint insert that waits, FLR. Keep VPG settings CRUD where it already is.
|