feat(poc): rewind MCP, skill, and recover-ladder docs
Initial PoC: find_protection, tagged checkpoints, FLR, mutating catalog. Lab 10.9 status enums (0=Initializing, 1=MeetingSLA). Credentials stay in gitignored config.json.
This commit is contained in:
@@ -0,0 +1,57 @@
|
||||
# Recover ladder: FLR vs whole-VM rewind
|
||||
|
||||
Zerto's journal can rewind a file or a whole VM. The agent picks the smallest
|
||||
operation that actually undoes the damage. Failover Live is DR, not the default
|
||||
undo.
|
||||
|
||||
## File-level recovery (FLR)
|
||||
|
||||
Use when the guest still boots, SSH/WinRM still works, and the damage is a
|
||||
known path (config, dropped file, one directory).
|
||||
|
||||
FLR mounts a checkpoint and copies files out. The protected VM stays up.
|
||||
Official API: `POST /v1/flrs` then browse/download. This MCP writes the file
|
||||
to `recovery_dir` on the MCP host. Putting it back on the guest is a second
|
||||
step (scp/ssh). That copy-back is not Zerto; it is ordinary file transfer.
|
||||
|
||||
Do not use FLR when:
|
||||
|
||||
- `EnabledActions.IsFlrEnabled` is false (initial sync, clone, test, live,
|
||||
move, or EJC running)
|
||||
- The guest cannot boot or accept a file
|
||||
- You do not know which files changed (package install, kernel, ransomware)
|
||||
- Linux file >1.5GB, or the name has `\ / : * ? " < > |`
|
||||
- OS-level dedup volumes
|
||||
- 10.9 FLR Operator role (broken; Administrator is the documented workaround)
|
||||
|
||||
A tagged checkpoint must already exist. Initial sync has an empty journal
|
||||
(`GET .../checkpoints` returns `[]`). Guard refuses until status is MeetingSLA
|
||||
(or NotMeetingSLA) and substatus is not a sync.
|
||||
|
||||
## Whole-VM rewind
|
||||
|
||||
Use when FLR cannot put the guest back: OS broken, too many files, services or
|
||||
packages, unknown blast radius.
|
||||
|
||||
| Operation | What it does | When |
|
||||
|---|---|---|
|
||||
| Failover test | Test VMs from a checkpoint. Prod stays up. | Inspect only |
|
||||
| Offsite clone | Copy VMs at recovery, powered off, unprotected | Inspect or graft |
|
||||
| Instant restore | Local-journal VPGs only; one VM, journal kept | Same-site local VPG |
|
||||
| Failover Live | Production DR. Reverse protection. Stops the protected VM | Site or VM beyond clone/FLR, human confirmed |
|
||||
|
||||
Failover Live is not the coding-agent oops button. Same-site VPGs still run a
|
||||
real failover. Confirm in the client before the tool runs.
|
||||
|
||||
## Guard vs recover
|
||||
|
||||
1. `zerto_find_protection` must return exactly one VM and at least one
|
||||
taggable VPG.
|
||||
2. `zerto_guard_before_mutate` inserts the tagged checkpoint and waits until
|
||||
it is listed. If that fails, do not change the guest.
|
||||
3. After a bad change: FLR if the path is known and the guest is up; clone or
|
||||
test to inspect; Failover Live only with a human yes.
|
||||
|
||||
Status integers on 10.9 (`GET /v1/vpgs/statuses`): 0 Initializing, 1
|
||||
MeetingSLA, 2 NotMeetingSLA. Older notes that said 0=Protecting / 1=Moving are
|
||||
wrong on this API.
|
||||
Reference in New Issue
Block a user