The repo said, in four places, that a tagged checkpoint cannot be inserted when the protected site is Azure or AWS. That came from the 9.0 API reference. It is wrong on 10.9.10. Measured against the lab: both an AWS-protected and an Azure-protected VPG accepted the insert, the Zerto task reached Completed, and the checkpoint appeared in the journal. This is the second doc claim this repo carried that live 10.9 contradicts, after the VPG status enum. What actually differs is latency and granularity, and both are set by the PROTECTED site, not the recovery site: protected at journal gap tag visible after vSphere 5s ~4s Azure 60s ~34s AWS 630s ~128s There is a sharper consequence than slowness. The tagged checkpoint is stamped about 30s AFTER the insert request, so on a cloud-protected VPG an agent that mutates promptly puts its change inside the checkpoint that was supposed to precede it. Recovering from that tag would restore the broken state. The docs now say to use the newest checkpoint that already existed when the guard ran. Two of the four were message text, not prose. wait_for_tag's timeout message asserted the insert was unsupported when the real cause was its own 45s budget being far too short for a VPG that checkpoints every 630s, so it told operators the wrong thing at exactly the wrong moment. It now says to check the Zerto task before concluding the insert failed. The 45s timeout itself is still wrong for cloud sources and needs to become cadence-aware. That is a behaviour change, so it is not in this commit. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
96 lines
4.3 KiB
Markdown
96 lines
4.3 KiB
Markdown
---
|
|
name: zerto-rewind
|
|
description: Before changing a VM, find its Zerto VPGs and insert a tagged checkpoint. Recover files from that tag with FLR after a human confirms. Use whenever an agent will mutate a guest that might be protected by Zerto.
|
|
---
|
|
|
|
# Zerto rewind
|
|
|
|
Zerto already journals the VM. This skill makes the agent use that journal. Git does not have the guest file. Official ZVM MCP does not insert tagged checkpoints.
|
|
|
|
You talk to **one** MCP: `zerto_rewind_mcp`. Do not also require official ZVM MCP.
|
|
|
|
## Loop (mandatory)
|
|
|
|
Before **every** tool call that might touch a guest:
|
|
|
|
1. Take the hostname / VM name / Zerto `vmIdentifier` from the tool args.
|
|
2. Call `zerto_check_tool(server, tool, vm)`. Act on the verdict:
|
|
|
|
| verdict | what you do |
|
|
|---|---|
|
|
| `read_only` | Run the tool. No checkpoint. |
|
|
| `mutating` | Guard first. Do not ask — it is already known to change the guest. |
|
|
| `unknown` | **Ask the human.** Do not assume it is safe, and do not silently guard. |
|
|
|
|
3. For `mutating`, or for `unknown` where the human said yes: call
|
|
`zerto_guard_before_mutate` with `change_id` and `action`.
|
|
4. If `ok` is not true: **stop**. Do not mutate.
|
|
5. Then run the call.
|
|
|
|
`unknown` is the normal case, not an edge case. The catalogs are short and the
|
|
world of tools is not, so most tools are unclassified. Unknown means *nobody has
|
|
said this is read-only* — it does not mean safe. Put the decision to the human:
|
|
|
|
> `winrm/run_ps` is not a known read-only command. It may change `web01`.
|
|
> Insert a Zerto tagged checkpoint first so this is reversible?
|
|
|
|
If they say yes, guard, then run. If they say no, run it and tell them plainly
|
|
that it is not reversible through Zerto. If they want it remembered, call
|
|
`zerto_add_mutating_tool` (server, tool, `vm_arg`) so it is guarded
|
|
automatically next time instead of asking again.
|
|
|
|
## find_protection outcomes
|
|
|
|
| outcome | what you do |
|
|
|---|---|
|
|
| none | Unprotected or unknown. Refuse the change. Say Zerto cannot rewind this. |
|
|
| ambiguous | Two or more VMs matched. Ask for a `vmIdentifier`. Do not guess. |
|
|
| ok, no taggable VPG | Syncing or not Protecting. Refuse. A resync deletes checkpoints. |
|
|
| ok, taggable VPGs | Tag **every** protecting VPG with the same tag. Wait until listed (the tool blocks). |
|
|
|
|
A VM can be in up to three VPGs (local backup + remote DR is common). Tag all of them.
|
|
|
|
## Recover
|
|
|
|
Human must confirm. Pass `confirmed=true` only after they say yes.
|
|
|
|
- Bad config / dropped file: `zerto_recover_file` from **that tag**. Pass the guest path
|
|
(`/home/j/app.yaml` or `C:\Users\j\app.conf`); the server maps it into the FLR
|
|
namespace. **Locally replicated VPGs only** -- FLR happens at the recovery site, so a
|
|
cloud-replicated VPG must be recovered from that ZCA. The tool refuses and names the site.
|
|
- Inspect a whole VM: `zerto_offsite_clone` or `zerto_start_failover_test`.
|
|
- Never Failover Live. Never Move. Those are DR, not rewind.
|
|
|
|
`zerto_recover_file` unmounts its own FLR session and reports the result in
|
|
`unmount`. If `unmount.ok` is false, or a previous recovery died mid-flight,
|
|
the mount is still up: FLR cannot run during clone, test, live failover or EJC,
|
|
so a stuck session blocks the next recovery. Find it with
|
|
`zerto_list_flr_sessions` and clear it with `zerto_end_flr_session`.
|
|
|
|
## Facts that bite
|
|
|
|
- A tagged checkpoint is crash-consistent, not app-quiesced.
|
|
- Tagged checkpoints work on Azure and AWS protected VPGs, but they appear late:
|
|
~34s (Azure) and ~128s (AWS) versus ~4s on vSphere, and the checkpoint is stamped
|
|
about 30s after you ask for it. On those VPGs the tag can end up *after* your
|
|
change, so treat the newest checkpoint that already existed as the rewind point.
|
|
- 10.9 FLR Operator RBAC fails; Administrator is the documented workaround.
|
|
- FLR cannot run during clone, test, live failover, or EJC.
|
|
- Linux FLR: files >1.5GB are a bad idea; some characters in names are refused.
|
|
|
|
## Tag
|
|
|
|
The checkpoint name is the only field the Zerto API takes, so it carries the
|
|
whole story:
|
|
|
|
```
|
|
ai:<agent> | <action> | vm=<vm> | change=<change-id> | <utc>
|
|
ai:claude | edit /etc/nginx/nginx.conf | vm=web01 | change=chg-412 | 20260921T150405Z
|
|
```
|
|
|
|
Always pass `action`: a plain description of the change you are about to make.
|
|
An operator scrolling the journal in the Zerto UI should be able to tell which
|
|
agent inserted the checkpoint and why, without reading your transcript.
|
|
|
|
Same string on every VPG for that call.
|