feat(flr): make FLR session lifecycle visible and reapable #3

Merged
claude merged 1 commits from feat/flr-session-lifecycle into main 2026-09-21 13:43:05 -04:00
Contributor

Follow-up to #2, which this is stacked on — base is fix/flr-partition-path, so the diff here is only the new work. Merge #2 first, then this (or retarget to main once #2 lands).

Closes the three FLR-session gaps identified while reviewing #2.

zerto_recover_file already tore its session down in a finally block, but a mount could still end up stranded on the recovery site with nothing tracking it. That matters because FLR cannot run during clone, test, live failover or EJC — a stuck session blocks the next recovery, and you find out at the worst moment.

1. Unmount failures were swallowed

except ZertoError:
    pass

If the teardown itself failed, the caller got ok: true and never learned the mount was still up. The teardown result is now reported:

"unmount": { "attempted": true, "ok": true, "ended": ["b38ac264-…"], "failed": [] }

On failure the response also carries a warning. ok stays true when the bytes landed — the recovery genuinely succeeded — but the caller is told rather than left guessing.

2. A window where cleanup ran against nothing

If start_flr succeeded on the ZVM but its response failed to parse, session_id_from() raised, session_id stayed None, and the finally block no-opped — leaking a session whose id the process never learned.

Teardown now snapshots live session ids before starting and reaps anything new that appeared. It deliberately leaves pre-existing sessions alone, so a concurrent operator's mount is never torn down (covered by a test).

3. Nothing could see or clear an orphan

The finally only runs if the process survives. A crash, disconnect or timeout mid-recovery left a mount with no way to find it. Two new tools:

  • zerto_list_flr_sessions — every session the ZVM knows about. live_only (default true) keeps the ones still holding a mount; ended and failed sessions linger as history and hold nothing.
  • zerto_end_flr_session — unmount one. Gated on confirmed=true like the other destructive tools, since ending a session someone else is pulling files from will interrupt them.

Verification against ZVM 10.x

  • zerto_list_flr_sessions(live_only=true) → count: 0, total_known: 1 after a clean run; live_only=false shows the UnmountCompletedSuccessfully history row
  • zerto_end_flr_session without confirmed → needs_confirm: true, refused
  • real recovery from checkpoint 1368 → 158 bytes, flr_path: Volume2-Ext4/home/justin/app-config.yaml, unmount.ok: true naming the session it ended
  • 0 live sessions remaining afterwards

pytest: 27 passed (5 new, including fakes that cover the swallowed-failure path, the orphan reap, and the leave-other-sessions-alone guarantee). ruff check clean on all changed files.

Docs updated: SKILL.md, README.md, docs/recover-ladder.md.

🤖 Generated with Claude Code

https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn

Follow-up to #2, which this is **stacked on** — base is `fix/flr-partition-path`, so the diff here is only the new work. Merge #2 first, then this (or retarget to `main` once #2 lands). Closes the three FLR-session gaps identified while reviewing #2. `zerto_recover_file` already tore its session down in a `finally` block, but a mount could still end up stranded on the recovery site with nothing tracking it. That matters because **FLR cannot run during clone, test, live failover or EJC** — a stuck session blocks the *next* recovery, and you find out at the worst moment. ## 1. Unmount failures were swallowed ```python except ZertoError: pass ``` If the teardown itself failed, the caller got `ok: true` and never learned the mount was still up. The teardown result is now reported: ```json "unmount": { "attempted": true, "ok": true, "ended": ["b38ac264-…"], "failed": [] } ``` On failure the response also carries a `warning`. `ok` stays `true` when the bytes landed — the recovery genuinely succeeded — but the caller is told rather than left guessing. ## 2. A window where cleanup ran against nothing If `start_flr` succeeded on the ZVM but its response failed to parse, `session_id_from()` raised, `session_id` stayed `None`, and the `finally` block no-opped — leaking a session whose id the process never learned. Teardown now snapshots live session ids *before* starting and reaps anything new that appeared. It deliberately leaves pre-existing sessions alone, so a concurrent operator's mount is never torn down (covered by a test). ## 3. Nothing could see or clear an orphan The `finally` only runs if the process survives. A crash, disconnect or timeout mid-recovery left a mount with no way to find it. Two new tools: - **`zerto_list_flr_sessions`** — every session the ZVM knows about. `live_only` (default true) keeps the ones still holding a mount; ended and failed sessions linger as history and hold nothing. - **`zerto_end_flr_session`** — unmount one. Gated on `confirmed=true` like the other destructive tools, since ending a session someone else is pulling files from will interrupt them. ## Verification against ZVM 10.x - `zerto_list_flr_sessions(live_only=true)` → `count: 0, total_known: 1` after a clean run; `live_only=false` shows the `UnmountCompletedSuccessfully` history row - `zerto_end_flr_session` without `confirmed` → `needs_confirm: true`, refused - real recovery from checkpoint 1368 → 158 bytes, `flr_path: Volume2-Ext4/home/justin/app-config.yaml`, `unmount.ok: true` naming the session it ended - 0 live sessions remaining afterwards `pytest`: 27 passed (5 new, including fakes that cover the swallowed-failure path, the orphan reap, and the leave-other-sessions-alone guarantee). `ruff check` clean on all changed files. Docs updated: `SKILL.md`, `README.md`, `docs/recover-ladder.md`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude changed target branch from fix/flr-partition-path to main 2026-09-21 13:42:37 -04:00
claude added 1 commit 2026-09-21 13:43:02 -04:00
zerto_recover_file already tore its session down in a finally block, but
three gaps meant a mount could stay up on the recovery site with nothing
tracking it. FLR cannot run during clone, test, live failover or EJC, so
a stuck session blocks the next recovery.

1. An unmount failure was swallowed (`except ZertoError: pass`). The
   caller got ok=true and never learned the mount was still up. The
   teardown result is now reported in the response as `unmount`, with a
   `warning` when it fails. ok stays true when the bytes did land -- the
   recovery genuinely succeeded -- but the caller is told.

2. If start_flr succeeded on the ZVM while its response failed to parse,
   session_id stayed None and the finally block did nothing, leaking a
   session the process never knew the id of. Teardown now snapshots live
   session ids before starting and reaps anything new that appeared,
   leaving other operators' sessions alone.

3. Nothing could see or clear an orphan left by a crashed process, since
   the finally block only runs if the process survives. Two new tools:

   - zerto_list_flr_sessions: every session the ZVM knows about.
     live_only (default true) keeps the ones still holding a mount;
     ended and failed sessions linger as history and hold nothing.
   - zerto_end_flr_session: unmount one. Gated on confirmed=true,
     matching the other destructive tools, because ending a session
     someone else is pulling files from will interrupt them.

Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean
run, the confirm gate refuses without a human yes, a real recovery from
checkpoint 1368 returned 158 bytes and reported
unmount.ok=true with the session id it ended, and 0 live sessions
remained afterwards.

pytest 27 passed (5 new, including fakes covering the swallowed-failure
and orphan-reap paths).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude force-pushed feat/flr-session-lifecycle from f930b84615 to 8b973653d8 2026-09-21 13:43:02 -04:00 Compare
claude merged commit 59d1f71617 into main 2026-09-21 13:43:05 -04:00
claude deleted branch feat/flr-session-lifecycle 2026-09-21 13:43:05 -04:00
Sign in to join this conversation.