Three fixes from measuring FLR against a Windows guest (ad1, VPG 'local')
instead of only the Linux one.
1. Gate FLR to locally replicated VPGs.
FLR is performed at the VPG's RECOVERY site, because that is where the
mount is created. A VPG replicating to a cloud ZCA has to be recovered
through that ZCA's API, not the protected ZVM's. Supporting that
properly means holding credentials for every ZVM/ZCA in an estate and
routing the call, which is a real design decision, not something to
smuggle in. Until then zerto_recover_file refuses a VPG whose
protected site != recovery site and names the site that owns the
operation, instead of failing later as a confusing path or mount error.
2. Windows paths are not symmetrical with Linux.
Linux /home/j/app.yaml -> Volume2-Ext4%2fhome%2fj%2fapp.yaml
Windows C:\Users\j\f.conf -> C%3a%2fUsers%2fj%2ff.conf
On Windows the drive letter IS the partition name, so the guest path
already carries it. The old code prepended unconditionally and built
'C:/C:/Users/j', which could never match. resolve_flr_path now detects
that the path already starts with the partition.
It also returns the raw path exactly as browse reported it. Download
accepts the raw and the decoded form, and returning raw avoids
re-encoding by hand. Decoding is now unquote_plus, not unquote: browse
form-encodes a space as '+' ("Program+Files"), so a basename compare
against "Program Files" never matched. Matching tries exact first and
only then case-insensitively, since Windows is case-insensitive and
Linux is not.
3. Wait for partition enumeration to settle.
A session reports mounted before the ZVM has finished identifying
volumes, and browsing in that window returns a partial, MIS-LABELLED
list. The same Windows VM enumerated as 'Volume4-Unknown' with no C:
drive, then moments later as a browsable 'C%3a' holding the whole
filesystem. Acting on the early list makes a restorable NTFS disk look
permanently unrestorable. wait_partitions_stable polls until the list
stops changing.
Verified against ZVM 10.x:
- CMH-AWS-4 (recovery aws-zca) refused, naming aws-zca
- C:\ad1.keytab recovered, 58 bytes
- C:\Program Files\internet explorer\sqmapi.dll recovered, 47512 bytes --
drive letter, two spaces, nested dirs, the exact case that was broken
- Linux /home/justin/app-config.yaml still recovers, 158 bytes
- every session unmounted, unmount.ok true
Also drops a duplicate get_vpg that shadowed the existing one (F811).
pytest 32 passed (7 new).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
76 lines
3.4 KiB
Markdown
76 lines
3.4 KiB
Markdown
---
|
|
name: zerto-rewind
|
|
description: Before changing a VM, find its Zerto VPGs and insert a tagged checkpoint. Recover files from that tag with FLR after a human confirms. Use whenever an agent will mutate a guest that might be protected by Zerto.
|
|
---
|
|
|
|
# Zerto rewind
|
|
|
|
Zerto already journals the VM. This skill makes the agent use that journal. Git does not have the guest file. Official ZVM MCP does not insert tagged checkpoints.
|
|
|
|
You talk to **one** MCP: `zerto_rewind_mcp`. Do not also require official ZVM MCP.
|
|
|
|
## Loop (mandatory)
|
|
|
|
Before **every** guest-mutating tool call:
|
|
|
|
1. Take the hostname / VM name / Zerto `vmIdentifier` from the tool args.
|
|
2. Call `zerto_guard_before_mutate` with `change_id` and `action` (or `zerto_find_protection` then `zerto_create_tagged_checkpoint`).
|
|
3. If `ok` is not true: **stop**. Do not mutate.
|
|
4. Then run the mutating call.
|
|
|
|
Reads skip the guard.
|
|
|
|
Unlisted MCP tools pass through. If you are about to change a protected VM with a tool that is not in the catalog, call `zerto_add_mutating_tool` (server, tool, `vm_arg`) and then guard.
|
|
|
|
## find_protection outcomes
|
|
|
|
| outcome | what you do |
|
|
|---|---|
|
|
| none | Unprotected or unknown. Refuse the change. Say Zerto cannot rewind this. |
|
|
| ambiguous | Two or more VMs matched. Ask for a `vmIdentifier`. Do not guess. |
|
|
| ok, no taggable VPG | Syncing or not Protecting. Refuse. A resync deletes checkpoints. |
|
|
| ok, taggable VPGs | Tag **every** protecting VPG with the same tag. Wait until listed (the tool blocks). |
|
|
|
|
A VM can be in up to three VPGs (local backup + remote DR is common). Tag all of them.
|
|
|
|
## Recover
|
|
|
|
Human must confirm. Pass `confirmed=true` only after they say yes.
|
|
|
|
- Bad config / dropped file: `zerto_recover_file` from **that tag**. Pass the guest path
|
|
(`/home/j/app.yaml` or `C:\Users\j\app.conf`); the server maps it into the FLR
|
|
namespace. **Locally replicated VPGs only** -- FLR happens at the recovery site, so a
|
|
cloud-replicated VPG must be recovered from that ZCA. The tool refuses and names the site.
|
|
- Inspect a whole VM: `zerto_offsite_clone` or `zerto_start_failover_test`.
|
|
- Never Failover Live. Never Move. Those are DR, not rewind.
|
|
|
|
`zerto_recover_file` unmounts its own FLR session and reports the result in
|
|
`unmount`. If `unmount.ok` is false, or a previous recovery died mid-flight,
|
|
the mount is still up: FLR cannot run during clone, test, live failover or EJC,
|
|
so a stuck session blocks the next recovery. Find it with
|
|
`zerto_list_flr_sessions` and clear it with `zerto_end_flr_session`.
|
|
|
|
## Facts that bite
|
|
|
|
- A tagged checkpoint is crash-consistent, not app-quiesced.
|
|
- Tagged checkpoints are not supported when the **protected** site is Azure or AWS. Talk to the vSphere protected ZVM.
|
|
- 10.9 FLR Operator RBAC fails; Administrator is the documented workaround.
|
|
- FLR cannot run during clone, test, live failover, or EJC.
|
|
- Linux FLR: files >1.5GB are a bad idea; some characters in names are refused.
|
|
|
|
## Tag
|
|
|
|
The checkpoint name is the only field the Zerto API takes, so it carries the
|
|
whole story:
|
|
|
|
```
|
|
ai:<agent> | <action> | vm=<vm> | change=<change-id> | <utc>
|
|
ai:claude | edit /etc/nginx/nginx.conf | vm=web01 | change=chg-412 | 20260921T150405Z
|
|
```
|
|
|
|
Always pass `action`: a plain description of the change you are about to make.
|
|
An operator scrolling the journal in the Zerto UI should be able to tell which
|
|
agent inserted the checkpoint and why, without reading your transcript.
|
|
|
|
Same string on every VPG for that call.
|