feat(flr): windows paths, recovery-site gate, stable partition reads

Three fixes from measuring FLR against a Windows guest (ad1, VPG 'local')
instead of only the Linux one.

1. Gate FLR to locally replicated VPGs.

   FLR is performed at the VPG's RECOVERY site, because that is where the
   mount is created. A VPG replicating to a cloud ZCA has to be recovered
   through that ZCA's API, not the protected ZVM's. Supporting that
   properly means holding credentials for every ZVM/ZCA in an estate and
   routing the call, which is a real design decision, not something to
   smuggle in. Until then zerto_recover_file refuses a VPG whose
   protected site != recovery site and names the site that owns the
   operation, instead of failing later as a confusing path or mount error.

2. Windows paths are not symmetrical with Linux.

   Linux   /home/j/app.yaml  -> Volume2-Ext4%2fhome%2fj%2fapp.yaml
   Windows C:\Users\j\f.conf -> C%3a%2fUsers%2fj%2ff.conf

   On Windows the drive letter IS the partition name, so the guest path
   already carries it. The old code prepended unconditionally and built
   'C:/C:/Users/j', which could never match. resolve_flr_path now detects
   that the path already starts with the partition.

   It also returns the raw path exactly as browse reported it. Download
   accepts the raw and the decoded form, and returning raw avoids
   re-encoding by hand. Decoding is now unquote_plus, not unquote: browse
   form-encodes a space as '+' ("Program+Files"), so a basename compare
   against "Program Files" never matched. Matching tries exact first and
   only then case-insensitively, since Windows is case-insensitive and
   Linux is not.

3. Wait for partition enumeration to settle.

   A session reports mounted before the ZVM has finished identifying
   volumes, and browsing in that window returns a partial, MIS-LABELLED
   list. The same Windows VM enumerated as 'Volume4-Unknown' with no C:
   drive, then moments later as a browsable 'C%3a' holding the whole
   filesystem. Acting on the early list makes a restorable NTFS disk look
   permanently unrestorable. wait_partitions_stable polls until the list
   stops changing.

Verified against ZVM 10.x:

- CMH-AWS-4 (recovery aws-zca) refused, naming aws-zca
- C:\ad1.keytab recovered, 58 bytes
- C:\Program Files\internet explorer\sqmapi.dll recovered, 47512 bytes --
  drive letter, two spaces, nested dirs, the exact case that was broken
- Linux /home/justin/app-config.yaml still recovers, 158 bytes
- every session unmounted, unmount.ok true

Also drops a duplicate get_vpg that shadowed the existing one (F811).

pytest 32 passed (7 new).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
This commit is contained in:
2026-09-21 14:05:02 -04:00
co-authored by Claude Opus 5
parent 59d1f71617
commit 6d87fae60d
8 changed files with 295 additions and 50 deletions
+23 -1
View File
@@ -11,7 +11,29 @@ known path (config, dropped file, one directory).
FLR mounts a checkpoint and copies files out. The protected VM stays up.
Official API: `POST /v1/flrs` then browse/download. This MCP writes the file
to `recovery_dir` on the MCP host. Putting it back on the guest is a second
to `recovery_dir` on the MCP host.
**FLR runs at the VPG's recovery site**, because that is where the mount is
created. A VPG replicating to a cloud ZCA must be recovered through that ZCA's
API, not the protected ZVM's. A production server would hold credentials for
every ZVM/ZCA in the estate and route the call; this one does not, so
`zerto_recover_file` is gated to **locally replicated VPGs** (protected site ==
recovery site) and refuses anything else while naming the site that owns the
operation.
Paths are rooted at partitions, and Linux and Windows are not symmetrical:
| | guest path | FLR path |
|---|---|---|
| Linux | `/home/j/app.yaml` | `Volume2-Ext4%2fhome%2fj%2fapp.yaml` |
| Windows | `C:\Users\j\app.conf` | `C%3a%2fUsers%2fj%2fapp.conf` |
On Windows the drive letter **is** the partition name, so nothing is
prepended. Browse form-encodes: `%2f` separator, `%3a` drive colon, and a
space as `+` (`Program+Files`). A session reports mounted before volume
enumeration settles, so the partition list must be polled until it stops
changing -- an early read can show a restorable NTFS disk as
`Volume4-Unknown`. Putting it back on the guest is a second
step (scp/ssh). That copy-back is not Zerto; it is ordinary file transfer.
Do not use FLR when: