3.8 KiB
Recover ladder: FLR vs whole-VM rewind
Zerto's journal can rewind a file or a whole VM. The agent picks the smallest operation that actually undoes the damage. Failover Live is DR, not the default undo.
File-level recovery (FLR)
Use when the guest still boots, SSH/WinRM still works, and the damage is a known path (config, dropped file, one directory).
FLR mounts a checkpoint and copies files out. The protected VM stays up.
Official API: POST /v1/flrs then browse/download. This MCP writes the file
to recovery_dir on the MCP host.
FLR runs at the VPG's recovery site, because that is where the mount is
created. A VPG replicating to a cloud ZCA must be recovered through that ZCA's
API, not the protected ZVM's. A production server would hold credentials for
every ZVM/ZCA in the estate and route the call; this one does not, so
zerto_recover_file is gated to locally replicated VPGs (protected site ==
recovery site) and refuses anything else while naming the site that owns the
operation.
Paths are rooted at partitions, and Linux and Windows are not symmetrical:
| guest path | FLR path | |
|---|---|---|
| Linux | /home/j/app.yaml |
Volume2-Ext4%2fhome%2fj%2fapp.yaml |
| Windows | C:\Users\j\app.conf |
C%3a%2fUsers%2fj%2fapp.conf |
On Windows the drive letter is the partition name, so nothing is
prepended. Browse form-encodes: %2f separator, %3a drive colon, and a
space as + (Program+Files). A session reports mounted before volume
enumeration settles, so the partition list must be polled until it stops
changing -- an early read can show a restorable NTFS disk as
Volume4-Unknown. Putting it back on the guest is a second
step (scp/ssh). That copy-back is not Zerto; it is ordinary file transfer.
Do not use FLR when:
EnabledActions.IsFlrEnabledis false (initial sync, clone, test, live, move, or EJC running)- The guest cannot boot or accept a file
- You do not know which files changed (package install, kernel, ransomware)
- Linux file >1.5GB, or the name has
\ / : * ? " < > | - OS-level dedup volumes
- 10.9 FLR Operator role (broken; Administrator is the documented workaround)
FLR sessions must be unmounted when done. zerto_recover_file does that
itself and reports it in unmount, but that cleanup only runs if the MCP
process survives the call. After a crash, list orphans with
zerto_list_flr_sessions and end them with zerto_end_flr_session.
A tagged checkpoint must already exist. Initial sync has an empty journal
(GET .../checkpoints returns []). Guard refuses until status is MeetingSLA
(or NotMeetingSLA) and substatus is not a sync.
Whole-VM rewind
Use when FLR cannot put the guest back: OS broken, too many files, services or packages, unknown blast radius.
| Operation | What it does | When |
|---|---|---|
| Failover test | Test VMs from a checkpoint. Prod stays up. | Inspect only |
| Offsite clone | Copy VMs at recovery, powered off, unprotected | Inspect or graft |
| Instant restore | Local-journal VPGs only; one VM, journal kept | Same-site local VPG |
| Failover Live | Production DR. Reverse protection. Stops the protected VM | Site or VM beyond clone/FLR, human confirmed |
Failover Live is not the coding-agent oops button. Same-site VPGs still run a real failover. Confirm in the client before the tool runs.
Guard vs recover
zerto_find_protectionmust return exactly one VM and at least one taggable VPG.zerto_guard_before_mutateinserts the tagged checkpoint and waits until it is listed. If that fails, do not change the guest.- After a bad change: FLR if the path is known and the guest is up; clone or test to inspect; Failover Live only with a human yes.
Status integers on 10.9 (GET /v1/vpgs/statuses): 0 Initializing, 1
MeetingSLA, 2 NotMeetingSLA. Older notes that said 0=Protecting / 1=Moving are
wrong on this API.