Three fixes from measuring FLR against a Windows guest (ad1, VPG 'local')
instead of only the Linux one.
1. Gate FLR to locally replicated VPGs.
FLR is performed at the VPG's RECOVERY site, because that is where the
mount is created. A VPG replicating to a cloud ZCA has to be recovered
through that ZCA's API, not the protected ZVM's. Supporting that
properly means holding credentials for every ZVM/ZCA in an estate and
routing the call, which is a real design decision, not something to
smuggle in. Until then zerto_recover_file refuses a VPG whose
protected site != recovery site and names the site that owns the
operation, instead of failing later as a confusing path or mount error.
2. Windows paths are not symmetrical with Linux.
Linux /home/j/app.yaml -> Volume2-Ext4%2fhome%2fj%2fapp.yaml
Windows C:\Users\j\f.conf -> C%3a%2fUsers%2fj%2ff.conf
On Windows the drive letter IS the partition name, so the guest path
already carries it. The old code prepended unconditionally and built
'C:/C:/Users/j', which could never match. resolve_flr_path now detects
that the path already starts with the partition.
It also returns the raw path exactly as browse reported it. Download
accepts the raw and the decoded form, and returning raw avoids
re-encoding by hand. Decoding is now unquote_plus, not unquote: browse
form-encodes a space as '+' ("Program+Files"), so a basename compare
against "Program Files" never matched. Matching tries exact first and
only then case-insensitively, since Windows is case-insensitive and
Linux is not.
3. Wait for partition enumeration to settle.
A session reports mounted before the ZVM has finished identifying
volumes, and browsing in that window returns a partial, MIS-LABELLED
list. The same Windows VM enumerated as 'Volume4-Unknown' with no C:
drive, then moments later as a browsable 'C%3a' holding the whole
filesystem. Acting on the early list makes a restorable NTFS disk look
permanently unrestorable. wait_partitions_stable polls until the list
stops changing.
Verified against ZVM 10.x:
- CMH-AWS-4 (recovery aws-zca) refused, naming aws-zca
- C:\ad1.keytab recovered, 58 bytes
- C:\Program Files\internet explorer\sqmapi.dll recovered, 47512 bytes --
drive letter, two spaces, nested dirs, the exact case that was broken
- Linux /home/justin/app-config.yaml still recovers, 158 bytes
- every session unmounted, unmount.ok true
Also drops a duplicate get_vpg that shadowed the existing one (F811).
pytest 32 passed (7 new).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2.7 KiB
Zerto AI Rewind
PoC MCP that teaches an agent to discover Zerto protection, pin a tagged checkpoint before changing a VM, and recover a file from that tag. If the loop works, these tools are the delta to put in official ZVM MCP.
Language
Tagged checkpoint:
A named bookmark in a VPG journal, inserted by POST /v1/vpgs/{id}/checkpoints (startVpgTaggedCheckpointInsert). Crash-consistent write-order only; not application-quiesced unless someone scripted that separately. CheckpointName is the only field the API accepts, so agent and intent go in the name: ai:<agent> | <action> | vm=<vm> | change=<id> | <utc>. Inserts are async tasks and are silently dropped if fired back to back at one VPG; insert, then wait until listed.
Avoid: user checkpoint, snapshot, backup, restore point (unqualified)
VPG: A Virtual Protection Group. One to many VMs sharing a journal. A VM can belong to at most three VPGs, recovered to different sites. Avoid: job, policy, replication group
Rewind: The agent loop: find protection, tag every protecting VPG, mutate, then bounded recover. Not a Zerto product name. Avoid: failover (that's DR), undo (that's git or Moholo)
Bounded recover: FLR, offsite clone, or failover test. Failover Live is not a rewind tool. Avoid: recover (unqualified), fail back, restore the VPG
File-level recovery (FLR): Mount a VM from a journal checkpoint and pull files. The VM stays up. 10.9 FLR Operator RBAC is broken; Administrator is the documented workaround. Runs at the VPG's recovery site, so this server supports it only for locally replicated VPGs (protected site == recovery site). Paths are partition-rooted; on Windows the drive letter is the partition. Avoid: file restore (unqualified), instant restore (local-journal VMs only, not v1)
find_protection: Resolve a VM name, hostname, or Zerto vmIdentifier to exactly one VM and every VPG it is in. Zero or two-plus VMs is a hard stop. Avoid: GetVms (that's the raw inventory call)
Protecting VPG: A VPG whose status is MeetingSLA or a NotMeetingSLA variant, and whose substatus is not a sync. Only these get tagged. 10.9 status 0 is Initializing, not Protecting. A resync deletes existing checkpoints. Avoid: healthy, in sync, Protecting (as status 0)
Mutating catalog:
The opt-in list of MCP tools that must call zerto_guard_before_mutate first. Unlisted tools pass through. Users add entries; the starter list is not the whole world.
Avoid: denylist, hold-everything
Official ZVM MCP:
HPE Zerto 10.9 MCP (ZVM.MCP): inventory, VPG settings, failover test. Not in the demo path. This PoC is one server.
Avoid: Zerto MCP (unqualified when you mean this repo)