zerto_recover_file already tore its session down in a finally block, but
three gaps meant a mount could stay up on the recovery site with nothing
tracking it. FLR cannot run during clone, test, live failover or EJC, so
a stuck session blocks the next recovery.
1. An unmount failure was swallowed (`except ZertoError: pass`). The
caller got ok=true and never learned the mount was still up. The
teardown result is now reported in the response as `unmount`, with a
`warning` when it fails. ok stays true when the bytes did land -- the
recovery genuinely succeeded -- but the caller is told.
2. If start_flr succeeded on the ZVM while its response failed to parse,
session_id stayed None and the finally block did nothing, leaking a
session the process never knew the id of. Teardown now snapshots live
session ids before starting and reaps anything new that appeared,
leaving other operators' sessions alone.
3. Nothing could see or clear an orphan left by a crashed process, since
the finally block only runs if the process survives. Two new tools:
- zerto_list_flr_sessions: every session the ZVM knows about.
live_only (default true) keeps the ones still holding a mount;
ended and failed sessions linger as history and hold nothing.
- zerto_end_flr_session: unmount one. Gated on confirmed=true,
matching the other destructive tools, because ending a session
someone else is pulling files from will interrupt them.
Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean
run, the confirm gate refuses without a human yes, a real recovery from
checkpoint 1368 returned 158 bytes and reported
unmount.ok=true with the session id it ended, and 0 live sessions
remained afterwards.
pytest 27 passed (5 new, including fakes covering the swallowed-failure
and orphan-reap paths).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
Zerto's tagged checkpoint insert takes exactly one field. The 10.x
swagger model VpgInsertTagCheckpointDataApi has a single property,
checkpointName, and the 9.0 API reference lists CheckpointName as the
only request value. There is no description field, so who the agent is
and what it is about to do have to live inside the name.
Old name:
ai:claude:chg-412:20260921T170829Z
New name:
ai:claude | edit /home/justin/app-config.yaml | vm=jp-ubuntu |
change=chg-412 | 20260921T170829Z
zerto_create_tagged_checkpoint and zerto_guard_before_mutate take a new
action argument: free text saying what the agent is about to do. The VM
name is filled in from the find result. An operator reading the journal
in the Zerto UI can now see which agent inserted a checkpoint and why,
without the agent transcript.
Field text is sanitised so the name stays one readable line: control
characters and runs of whitespace collapse to single spaces, ';' becomes
',' because Zerto appends "; Used for File Level Restore" to its own
tags, and '|' becomes '/' because ' | ' is our field separator. Capped
at TAG_MAX_LEN (250).
Measured against ZVM 10.x while picking the format:
- names of at least 400 chars are accepted, and spaces, slashes,
parentheses, '=' and '|' all survive the round trip
- tagged checkpoint inserts fired back to back at one VPG are silently
dropped. The POST returns 200 and queues a task, but only the first
checkpoint appears. tag_vpgs already inserts then waits per VPG, so
it is correct; added a comment so nobody turns that loop into an
asyncio.gather().
Verified end to end: guard inserted cp 1197 on VPG jp-ubuntu, the name
read back byte-identical from the journal, and FLR from that checkpoint
returned the 158 byte pre-mutation file.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn