Files
zerto-ai-rewind/docs/recover-ladder.md
T

2.7 KiB

Recover ladder: FLR vs whole-VM rewind

Zerto's journal can rewind a file or a whole VM. The agent picks the smallest operation that actually undoes the damage. Failover Live is DR, not the default undo.

File-level recovery (FLR)

Use when the guest still boots, SSH/WinRM still works, and the damage is a known path (config, dropped file, one directory).

FLR mounts a checkpoint and copies files out. The protected VM stays up. Official API: POST /v1/flrs then browse/download. This MCP writes the file to recovery_dir on the MCP host. Putting it back on the guest is a second step (scp/ssh). That copy-back is not Zerto; it is ordinary file transfer.

Do not use FLR when:

  • EnabledActions.IsFlrEnabled is false (initial sync, clone, test, live, move, or EJC running)
  • The guest cannot boot or accept a file
  • You do not know which files changed (package install, kernel, ransomware)
  • Linux file >1.5GB, or the name has \ / : * ? " < > |
  • OS-level dedup volumes
  • 10.9 FLR Operator role (broken; Administrator is the documented workaround)

FLR sessions must be unmounted when done. zerto_recover_file does that itself and reports it in unmount, but that cleanup only runs if the MCP process survives the call. After a crash, list orphans with zerto_list_flr_sessions and end them with zerto_end_flr_session.

A tagged checkpoint must already exist. Initial sync has an empty journal (GET .../checkpoints returns []). Guard refuses until status is MeetingSLA (or NotMeetingSLA) and substatus is not a sync.

Whole-VM rewind

Use when FLR cannot put the guest back: OS broken, too many files, services or packages, unknown blast radius.

Operation What it does When
Failover test Test VMs from a checkpoint. Prod stays up. Inspect only
Offsite clone Copy VMs at recovery, powered off, unprotected Inspect or graft
Instant restore Local-journal VPGs only; one VM, journal kept Same-site local VPG
Failover Live Production DR. Reverse protection. Stops the protected VM Site or VM beyond clone/FLR, human confirmed

Failover Live is not the coding-agent oops button. Same-site VPGs still run a real failover. Confirm in the client before the tool runs.

Guard vs recover

  1. zerto_find_protection must return exactly one VM and at least one taggable VPG.
  2. zerto_guard_before_mutate inserts the tagged checkpoint and waits until it is listed. If that fails, do not change the guest.
  3. After a bad change: FLR if the path is known and the guest is up; clone or test to inspect; Failover Live only with a human yes.

Status integers on 10.9 (GET /v1/vpgs/statuses): 0 Initializing, 1 MeetingSLA, 2 NotMeetingSLA. Older notes that said 0=Protecting / 1=Moving are wrong on this API.