Three fixes from measuring FLR against a Windows guest (ad1, VPG 'local')
instead of only the Linux one.
1. Gate FLR to locally replicated VPGs.
FLR is performed at the VPG's RECOVERY site, because that is where the
mount is created. A VPG replicating to a cloud ZCA has to be recovered
through that ZCA's API, not the protected ZVM's. Supporting that
properly means holding credentials for every ZVM/ZCA in an estate and
routing the call, which is a real design decision, not something to
smuggle in. Until then zerto_recover_file refuses a VPG whose
protected site != recovery site and names the site that owns the
operation, instead of failing later as a confusing path or mount error.
2. Windows paths are not symmetrical with Linux.
Linux /home/j/app.yaml -> Volume2-Ext4%2fhome%2fj%2fapp.yaml
Windows C:\Users\j\f.conf -> C%3a%2fUsers%2fj%2ff.conf
On Windows the drive letter IS the partition name, so the guest path
already carries it. The old code prepended unconditionally and built
'C:/C:/Users/j', which could never match. resolve_flr_path now detects
that the path already starts with the partition.
It also returns the raw path exactly as browse reported it. Download
accepts the raw and the decoded form, and returning raw avoids
re-encoding by hand. Decoding is now unquote_plus, not unquote: browse
form-encodes a space as '+' ("Program+Files"), so a basename compare
against "Program Files" never matched. Matching tries exact first and
only then case-insensitively, since Windows is case-insensitive and
Linux is not.
3. Wait for partition enumeration to settle.
A session reports mounted before the ZVM has finished identifying
volumes, and browsing in that window returns a partial, MIS-LABELLED
list. The same Windows VM enumerated as 'Volume4-Unknown' with no C:
drive, then moments later as a browsable 'C%3a' holding the whole
filesystem. Acting on the early list makes a restorable NTFS disk look
permanently unrestorable. wait_partitions_stable polls until the list
stops changing.
Verified against ZVM 10.x:
- CMH-AWS-4 (recovery aws-zca) refused, naming aws-zca
- C:\ad1.keytab recovered, 58 bytes
- C:\Program Files\internet explorer\sqmapi.dll recovered, 47512 bytes --
drive letter, two spaces, nested dirs, the exact case that was broken
- Linux /home/justin/app-config.yaml still recovers, 158 bytes
- every session unmounted, unmount.ok true
Also drops a duplicate get_vpg that shadowed the existing one (F811).
pytest 32 passed (7 new).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
85 lines
3.8 KiB
Markdown
85 lines
3.8 KiB
Markdown
# Recover ladder: FLR vs whole-VM rewind
|
|
|
|
Zerto's journal can rewind a file or a whole VM. The agent picks the smallest
|
|
operation that actually undoes the damage. Failover Live is DR, not the default
|
|
undo.
|
|
|
|
## File-level recovery (FLR)
|
|
|
|
Use when the guest still boots, SSH/WinRM still works, and the damage is a
|
|
known path (config, dropped file, one directory).
|
|
|
|
FLR mounts a checkpoint and copies files out. The protected VM stays up.
|
|
Official API: `POST /v1/flrs` then browse/download. This MCP writes the file
|
|
to `recovery_dir` on the MCP host.
|
|
|
|
**FLR runs at the VPG's recovery site**, because that is where the mount is
|
|
created. A VPG replicating to a cloud ZCA must be recovered through that ZCA's
|
|
API, not the protected ZVM's. A production server would hold credentials for
|
|
every ZVM/ZCA in the estate and route the call; this one does not, so
|
|
`zerto_recover_file` is gated to **locally replicated VPGs** (protected site ==
|
|
recovery site) and refuses anything else while naming the site that owns the
|
|
operation.
|
|
|
|
Paths are rooted at partitions, and Linux and Windows are not symmetrical:
|
|
|
|
| | guest path | FLR path |
|
|
|---|---|---|
|
|
| Linux | `/home/j/app.yaml` | `Volume2-Ext4%2fhome%2fj%2fapp.yaml` |
|
|
| Windows | `C:\Users\j\app.conf` | `C%3a%2fUsers%2fj%2fapp.conf` |
|
|
|
|
On Windows the drive letter **is** the partition name, so nothing is
|
|
prepended. Browse form-encodes: `%2f` separator, `%3a` drive colon, and a
|
|
space as `+` (`Program+Files`). A session reports mounted before volume
|
|
enumeration settles, so the partition list must be polled until it stops
|
|
changing -- an early read can show a restorable NTFS disk as
|
|
`Volume4-Unknown`. Putting it back on the guest is a second
|
|
step (scp/ssh). That copy-back is not Zerto; it is ordinary file transfer.
|
|
|
|
Do not use FLR when:
|
|
|
|
- `EnabledActions.IsFlrEnabled` is false (initial sync, clone, test, live,
|
|
move, or EJC running)
|
|
- The guest cannot boot or accept a file
|
|
- You do not know which files changed (package install, kernel, ransomware)
|
|
- Linux file >1.5GB, or the name has `\ / : * ? " < > |`
|
|
- OS-level dedup volumes
|
|
- 10.9 FLR Operator role (broken; Administrator is the documented workaround)
|
|
|
|
FLR sessions must be unmounted when done. `zerto_recover_file` does that
|
|
itself and reports it in `unmount`, but that cleanup only runs if the MCP
|
|
process survives the call. After a crash, list orphans with
|
|
`zerto_list_flr_sessions` and end them with `zerto_end_flr_session`.
|
|
|
|
A tagged checkpoint must already exist. Initial sync has an empty journal
|
|
(`GET .../checkpoints` returns `[]`). Guard refuses until status is MeetingSLA
|
|
(or NotMeetingSLA) and substatus is not a sync.
|
|
|
|
## Whole-VM rewind
|
|
|
|
Use when FLR cannot put the guest back: OS broken, too many files, services or
|
|
packages, unknown blast radius.
|
|
|
|
| Operation | What it does | When |
|
|
|---|---|---|
|
|
| Failover test | Test VMs from a checkpoint. Prod stays up. | Inspect only |
|
|
| Offsite clone | Copy VMs at recovery, powered off, unprotected | Inspect or graft |
|
|
| Instant restore | Local-journal VPGs only; one VM, journal kept | Same-site local VPG |
|
|
| Failover Live | Production DR. Reverse protection. Stops the protected VM | Site or VM beyond clone/FLR, human confirmed |
|
|
|
|
Failover Live is not the coding-agent oops button. Same-site VPGs still run a
|
|
real failover. Confirm in the client before the tool runs.
|
|
|
|
## Guard vs recover
|
|
|
|
1. `zerto_find_protection` must return exactly one VM and at least one
|
|
taggable VPG.
|
|
2. `zerto_guard_before_mutate` inserts the tagged checkpoint and waits until
|
|
it is listed. If that fails, do not change the guest.
|
|
3. After a bad change: FLR if the path is known and the guest is up; clone or
|
|
test to inspect; Failover Live only with a human yes.
|
|
|
|
Status integers on 10.9 (`GET /v1/vpgs/statuses`): 0 Initializing, 1
|
|
MeetingSLA, 2 NotMeetingSLA. Older notes that said 0=Protecting / 1=Moving are
|
|
wrong on this API.
|