Files
zerto-ai-rewind/docs/recover-ladder.md
T
justinandClaude Opus 5 8b973653d8 feat(flr): make FLR session lifecycle visible and reapable
zerto_recover_file already tore its session down in a finally block, but
three gaps meant a mount could stay up on the recovery site with nothing
tracking it. FLR cannot run during clone, test, live failover or EJC, so
a stuck session blocks the next recovery.

1. An unmount failure was swallowed (`except ZertoError: pass`). The
   caller got ok=true and never learned the mount was still up. The
   teardown result is now reported in the response as `unmount`, with a
   `warning` when it fails. ok stays true when the bytes did land -- the
   recovery genuinely succeeded -- but the caller is told.

2. If start_flr succeeded on the ZVM while its response failed to parse,
   session_id stayed None and the finally block did nothing, leaking a
   session the process never knew the id of. Teardown now snapshots live
   session ids before starting and reaps anything new that appeared,
   leaving other operators' sessions alone.

3. Nothing could see or clear an orphan left by a crashed process, since
   the finally block only runs if the process survives. Two new tools:

   - zerto_list_flr_sessions: every session the ZVM knows about.
     live_only (default true) keeps the ones still holding a mount;
     ended and failed sessions linger as history and hold nothing.
   - zerto_end_flr_session: unmount one. Gated on confirmed=true,
     matching the other destructive tools, because ending a session
     someone else is pulling files from will interrupt them.

Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean
run, the confirm gate refuses without a human yes, a real recovery from
checkpoint 1368 returned 158 bytes and reported
unmount.ok=true with the session id it ended, and 0 live sessions
remained afterwards.

pytest 27 passed (5 new, including fakes covering the swallowed-failure
and orphan-reap paths).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-21 13:42:42 -04:00

63 lines
2.7 KiB
Markdown

# Recover ladder: FLR vs whole-VM rewind
Zerto's journal can rewind a file or a whole VM. The agent picks the smallest
operation that actually undoes the damage. Failover Live is DR, not the default
undo.
## File-level recovery (FLR)
Use when the guest still boots, SSH/WinRM still works, and the damage is a
known path (config, dropped file, one directory).
FLR mounts a checkpoint and copies files out. The protected VM stays up.
Official API: `POST /v1/flrs` then browse/download. This MCP writes the file
to `recovery_dir` on the MCP host. Putting it back on the guest is a second
step (scp/ssh). That copy-back is not Zerto; it is ordinary file transfer.
Do not use FLR when:
- `EnabledActions.IsFlrEnabled` is false (initial sync, clone, test, live,
move, or EJC running)
- The guest cannot boot or accept a file
- You do not know which files changed (package install, kernel, ransomware)
- Linux file >1.5GB, or the name has `\ / : * ? " < > |`
- OS-level dedup volumes
- 10.9 FLR Operator role (broken; Administrator is the documented workaround)
FLR sessions must be unmounted when done. `zerto_recover_file` does that
itself and reports it in `unmount`, but that cleanup only runs if the MCP
process survives the call. After a crash, list orphans with
`zerto_list_flr_sessions` and end them with `zerto_end_flr_session`.
A tagged checkpoint must already exist. Initial sync has an empty journal
(`GET .../checkpoints` returns `[]`). Guard refuses until status is MeetingSLA
(or NotMeetingSLA) and substatus is not a sync.
## Whole-VM rewind
Use when FLR cannot put the guest back: OS broken, too many files, services or
packages, unknown blast radius.
| Operation | What it does | When |
|---|---|---|
| Failover test | Test VMs from a checkpoint. Prod stays up. | Inspect only |
| Offsite clone | Copy VMs at recovery, powered off, unprotected | Inspect or graft |
| Instant restore | Local-journal VPGs only; one VM, journal kept | Same-site local VPG |
| Failover Live | Production DR. Reverse protection. Stops the protected VM | Site or VM beyond clone/FLR, human confirmed |
Failover Live is not the coding-agent oops button. Same-site VPGs still run a
real failover. Confirm in the client before the tool runs.
## Guard vs recover
1. `zerto_find_protection` must return exactly one VM and at least one
taggable VPG.
2. `zerto_guard_before_mutate` inserts the tagged checkpoint and waits until
it is listed. If that fails, do not change the guest.
3. After a bad change: FLR if the path is known and the guest is up; clone or
test to inspect; Failover Live only with a human yes.
Status integers on 10.9 (`GET /v1/vpgs/statuses`): 0 Initializing, 1
MeetingSLA, 2 NotMeetingSLA. Older notes that said 0=Protecting / 1=Moving are
wrong on this API.