Files
zerto-ai-rewind/skills/zerto-rewind/SKILL.md
T
justinandClaude Opus 5 bad5a5706f feat(guard): read the Zerto task, and ask before guarding unknown tools
Two changes that both come from the same mistake: assuming an answer
instead of reading one.

1. Read the task after inserting a checkpoint.

   POST /v1/vpgs/{id}/checkpoints returns a TASK ID, not a result. A 200
   only means queued. The outcome lives in GET /v1/tasks/{id} under
   Status.State: 1 InProgress, 4 Failed, 5 Stopped, 6 Completed, with
   4/5/6 terminal.

   tag_vpgs now waits for that task and refuses unless it Completed, and
   reports task_id and task_state. Measured: two inserts fired back to
   back at one VPG give Completed for the first and Failed for the
   second. That is exactly the case an earlier comment in this file
   called "silently dropped" -- it was never silent, we just never read
   the task. Comment corrected.

   Previously a failed insert surfaced only as wait_for_tag timing out
   45s later with a misleading hint about Azure/AWS. Now it says the
   task failed and the operation did not happen.

2. Unknown tools ask the human instead of passing through.

   The catalog is opt-in, so an unlisted tool ran unguarded. But the set
   of mutating tools is unbounded and grows with every MCP installed,
   while the set of read-only ones is small, so a mutating-only list is
   permanently behind and being behind fails open.

   Adds read_only_tools and zerto_check_tool(server, tool, vm) returning
   read_only / mutating / unknown. unknown does not mean safe: it means
   nobody classified it, so the tool hands the agent a question to put
   to the human, and the human decides whether to checkpoint. On yes the
   agent guards; on no it runs and says plainly that Zerto cannot rewind
   it; if they want it remembered, zerto_add_mutating_tool.

   This stays advisory. An MCP server cannot see or block another
   server's tool calls, so real enforcement belongs in a host PreToolUse
   hook. The skill carries the flow.

Known issue, not addressed here: two concurrent tool calls race on
Keycloak token acquisition in the shared client and one gets HTTP 401.

pytest 41 passed (12 new).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-21 14:47:19 -04:00

93 lines
4.1 KiB
Markdown

---
name: zerto-rewind
description: Before changing a VM, find its Zerto VPGs and insert a tagged checkpoint. Recover files from that tag with FLR after a human confirms. Use whenever an agent will mutate a guest that might be protected by Zerto.
---
# Zerto rewind
Zerto already journals the VM. This skill makes the agent use that journal. Git does not have the guest file. Official ZVM MCP does not insert tagged checkpoints.
You talk to **one** MCP: `zerto_rewind_mcp`. Do not also require official ZVM MCP.
## Loop (mandatory)
Before **every** tool call that might touch a guest:
1. Take the hostname / VM name / Zerto `vmIdentifier` from the tool args.
2. Call `zerto_check_tool(server, tool, vm)`. Act on the verdict:
| verdict | what you do |
|---|---|
| `read_only` | Run the tool. No checkpoint. |
| `mutating` | Guard first. Do not ask — it is already known to change the guest. |
| `unknown` | **Ask the human.** Do not assume it is safe, and do not silently guard. |
3. For `mutating`, or for `unknown` where the human said yes: call
`zerto_guard_before_mutate` with `change_id` and `action`.
4. If `ok` is not true: **stop**. Do not mutate.
5. Then run the call.
`unknown` is the normal case, not an edge case. The catalogs are short and the
world of tools is not, so most tools are unclassified. Unknown means *nobody has
said this is read-only* — it does not mean safe. Put the decision to the human:
> `winrm/run_ps` is not a known read-only command. It may change `web01`.
> Insert a Zerto tagged checkpoint first so this is reversible?
If they say yes, guard, then run. If they say no, run it and tell them plainly
that it is not reversible through Zerto. If they want it remembered, call
`zerto_add_mutating_tool` (server, tool, `vm_arg`) so it is guarded
automatically next time instead of asking again.
## find_protection outcomes
| outcome | what you do |
|---|---|
| none | Unprotected or unknown. Refuse the change. Say Zerto cannot rewind this. |
| ambiguous | Two or more VMs matched. Ask for a `vmIdentifier`. Do not guess. |
| ok, no taggable VPG | Syncing or not Protecting. Refuse. A resync deletes checkpoints. |
| ok, taggable VPGs | Tag **every** protecting VPG with the same tag. Wait until listed (the tool blocks). |
A VM can be in up to three VPGs (local backup + remote DR is common). Tag all of them.
## Recover
Human must confirm. Pass `confirmed=true` only after they say yes.
- Bad config / dropped file: `zerto_recover_file` from **that tag**. Pass the guest path
(`/home/j/app.yaml` or `C:\Users\j\app.conf`); the server maps it into the FLR
namespace. **Locally replicated VPGs only** -- FLR happens at the recovery site, so a
cloud-replicated VPG must be recovered from that ZCA. The tool refuses and names the site.
- Inspect a whole VM: `zerto_offsite_clone` or `zerto_start_failover_test`.
- Never Failover Live. Never Move. Those are DR, not rewind.
`zerto_recover_file` unmounts its own FLR session and reports the result in
`unmount`. If `unmount.ok` is false, or a previous recovery died mid-flight,
the mount is still up: FLR cannot run during clone, test, live failover or EJC,
so a stuck session blocks the next recovery. Find it with
`zerto_list_flr_sessions` and clear it with `zerto_end_flr_session`.
## Facts that bite
- A tagged checkpoint is crash-consistent, not app-quiesced.
- Tagged checkpoints are not supported when the **protected** site is Azure or AWS. Talk to the vSphere protected ZVM.
- 10.9 FLR Operator RBAC fails; Administrator is the documented workaround.
- FLR cannot run during clone, test, live failover, or EJC.
- Linux FLR: files >1.5GB are a bad idea; some characters in names are refused.
## Tag
The checkpoint name is the only field the Zerto API takes, so it carries the
whole story:
```
ai:<agent> | <action> | vm=<vm> | change=<change-id> | <utc>
ai:claude | edit /etc/nginx/nginx.conf | vm=web01 | change=chg-412 | 20260921T150405Z
```
Always pass `action`: a plain description of the change you are about to make.
An operator scrolling the journal in the Zerto UI should be able to tell which
agent inserted the checkpoint and why, without reading your transcript.
Same string on every VPG for that call.