Files
zerto-ai-rewind/skills/zerto-rewind/SKILL.md
T
justinandClaude Opus 5 f930b84615 feat(flr): make FLR session lifecycle visible and reapable
zerto_recover_file already tore its session down in a finally block, but
three gaps meant a mount could stay up on the recovery site with nothing
tracking it. FLR cannot run during clone, test, live failover or EJC, so
a stuck session blocks the next recovery.

1. An unmount failure was swallowed (`except ZertoError: pass`). The
   caller got ok=true and never learned the mount was still up. The
   teardown result is now reported in the response as `unmount`, with a
   `warning` when it fails. ok stays true when the bytes did land -- the
   recovery genuinely succeeded -- but the caller is told.

2. If start_flr succeeded on the ZVM while its response failed to parse,
   session_id stayed None and the finally block did nothing, leaking a
   session the process never knew the id of. Teardown now snapshots live
   session ids before starting and reaps anything new that appeared,
   leaving other operators' sessions alone.

3. Nothing could see or clear an orphan left by a crashed process, since
   the finally block only runs if the process survives. Two new tools:

   - zerto_list_flr_sessions: every session the ZVM knows about.
     live_only (default true) keeps the ones still holding a mount;
     ended and failed sessions linger as history and hold nothing.
   - zerto_end_flr_session: unmount one. Gated on confirmed=true,
     matching the other destructive tools, because ending a session
     someone else is pulling files from will interrupt them.

Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean
run, the confirm gate refuses without a human yes, a real recovery from
checkpoint 1368 returned 158 bytes and reported
unmount.ok=true with the session id it ended, and 0 live sessions
remained afterwards.

pytest 27 passed (5 new, including fakes covering the swallowed-failure
and orphan-reap paths).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-21 13:40:53 -04:00

73 lines
3.1 KiB
Markdown

---
name: zerto-rewind
description: Before changing a VM, find its Zerto VPGs and insert a tagged checkpoint. Recover files from that tag with FLR after a human confirms. Use whenever an agent will mutate a guest that might be protected by Zerto.
---
# Zerto rewind
Zerto already journals the VM. This skill makes the agent use that journal. Git does not have the guest file. Official ZVM MCP does not insert tagged checkpoints.
You talk to **one** MCP: `zerto_rewind_mcp`. Do not also require official ZVM MCP.
## Loop (mandatory)
Before **every** guest-mutating tool call:
1. Take the hostname / VM name / Zerto `vmIdentifier` from the tool args.
2. Call `zerto_guard_before_mutate` with `change_id` and `action` (or `zerto_find_protection` then `zerto_create_tagged_checkpoint`).
3. If `ok` is not true: **stop**. Do not mutate.
4. Then run the mutating call.
Reads skip the guard.
Unlisted MCP tools pass through. If you are about to change a protected VM with a tool that is not in the catalog, call `zerto_add_mutating_tool` (server, tool, `vm_arg`) and then guard.
## find_protection outcomes
| outcome | what you do |
|---|---|
| none | Unprotected or unknown. Refuse the change. Say Zerto cannot rewind this. |
| ambiguous | Two or more VMs matched. Ask for a `vmIdentifier`. Do not guess. |
| ok, no taggable VPG | Syncing or not Protecting. Refuse. A resync deletes checkpoints. |
| ok, taggable VPGs | Tag **every** protecting VPG with the same tag. Wait until listed (the tool blocks). |
A VM can be in up to three VPGs (local backup + remote DR is common). Tag all of them.
## Recover
Human must confirm. Pass `confirmed=true` only after they say yes.
- Bad config / dropped file: `zerto_recover_file` from **that tag**.
- Inspect a whole VM: `zerto_offsite_clone` or `zerto_start_failover_test`.
- Never Failover Live. Never Move. Those are DR, not rewind.
`zerto_recover_file` unmounts its own FLR session and reports the result in
`unmount`. If `unmount.ok` is false, or a previous recovery died mid-flight,
the mount is still up: FLR cannot run during clone, test, live failover or EJC,
so a stuck session blocks the next recovery. Find it with
`zerto_list_flr_sessions` and clear it with `zerto_end_flr_session`.
## Facts that bite
- A tagged checkpoint is crash-consistent, not app-quiesced.
- Tagged checkpoints are not supported when the **protected** site is Azure or AWS. Talk to the vSphere protected ZVM.
- 10.9 FLR Operator RBAC fails; Administrator is the documented workaround.
- FLR cannot run during clone, test, live failover, or EJC.
- Linux FLR: files >1.5GB are a bad idea; some characters in names are refused.
## Tag
The checkpoint name is the only field the Zerto API takes, so it carries the
whole story:
```
ai:<agent> | <action> | vm=<vm> | change=<change-id> | <utc>
ai:claude | edit /etc/nginx/nginx.conf | vm=web01 | change=chg-412 | 20260921T150405Z
```
Always pass `action`: a plain description of the change you are about to make.
An operator scrolling the journal in the Zerto UI should be able to tell which
agent inserted the checkpoint and why, without reading your transcript.
Same string on every VPG for that call.