Follow-up to #2, which this is stacked on — base is fix/flr-partition-path, so the diff here is only the new work. Merge #2 first, then this (or retarget to main once #2 lands).
Closes the three FLR-session gaps identified while reviewing #2.
zerto_recover_file already tore its session down in a finally block, but a mount could still end up stranded on the recovery site with nothing tracking it. That matters because FLR cannot run during clone, test, live failover or EJC — a stuck session blocks the next recovery, and you find out at the worst moment.
1. Unmount failures were swallowed
exceptZertoError:pass
If the teardown itself failed, the caller got ok: true and never learned the mount was still up. The teardown result is now reported:
On failure the response also carries a warning. ok stays true when the bytes landed — the recovery genuinely succeeded — but the caller is told rather than left guessing.
2. A window where cleanup ran against nothing
If start_flr succeeded on the ZVM but its response failed to parse, session_id_from() raised, session_id stayed None, and the finally block no-opped — leaking a session whose id the process never learned.
Teardown now snapshots live session ids before starting and reaps anything new that appeared. It deliberately leaves pre-existing sessions alone, so a concurrent operator's mount is never torn down (covered by a test).
3. Nothing could see or clear an orphan
The finally only runs if the process survives. A crash, disconnect or timeout mid-recovery left a mount with no way to find it. Two new tools:
zerto_list_flr_sessions — every session the ZVM knows about. live_only (default true) keeps the ones still holding a mount; ended and failed sessions linger as history and hold nothing.
zerto_end_flr_session — unmount one. Gated on confirmed=true like the other destructive tools, since ending a session someone else is pulling files from will interrupt them.
Verification against ZVM 10.x
zerto_list_flr_sessions(live_only=true) → count: 0, total_known: 1 after a clean run; live_only=false shows the UnmountCompletedSuccessfully history row
zerto_end_flr_session without confirmed → needs_confirm: true, refused
real recovery from checkpoint 1368 → 158 bytes, flr_path: Volume2-Ext4/home/justin/app-config.yaml, unmount.ok: true naming the session it ended
0 live sessions remaining afterwards
pytest: 27 passed (5 new, including fakes that cover the swallowed-failure path, the orphan reap, and the leave-other-sessions-alone guarantee). ruff check clean on all changed files.
Follow-up to #2, which this is **stacked on** — base is `fix/flr-partition-path`, so the diff here is only the new work. Merge #2 first, then this (or retarget to `main` once #2 lands).
Closes the three FLR-session gaps identified while reviewing #2.
`zerto_recover_file` already tore its session down in a `finally` block, but a mount could still end up stranded on the recovery site with nothing tracking it. That matters because **FLR cannot run during clone, test, live failover or EJC** — a stuck session blocks the *next* recovery, and you find out at the worst moment.
## 1. Unmount failures were swallowed
```python
except ZertoError:
pass
```
If the teardown itself failed, the caller got `ok: true` and never learned the mount was still up. The teardown result is now reported:
```json
"unmount": { "attempted": true, "ok": true, "ended": ["b38ac264-…"], "failed": [] }
```
On failure the response also carries a `warning`. `ok` stays `true` when the bytes landed — the recovery genuinely succeeded — but the caller is told rather than left guessing.
## 2. A window where cleanup ran against nothing
If `start_flr` succeeded on the ZVM but its response failed to parse, `session_id_from()` raised, `session_id` stayed `None`, and the `finally` block no-opped — leaking a session whose id the process never learned.
Teardown now snapshots live session ids *before* starting and reaps anything new that appeared. It deliberately leaves pre-existing sessions alone, so a concurrent operator's mount is never torn down (covered by a test).
## 3. Nothing could see or clear an orphan
The `finally` only runs if the process survives. A crash, disconnect or timeout mid-recovery left a mount with no way to find it. Two new tools:
- **`zerto_list_flr_sessions`** — every session the ZVM knows about. `live_only` (default true) keeps the ones still holding a mount; ended and failed sessions linger as history and hold nothing.
- **`zerto_end_flr_session`** — unmount one. Gated on `confirmed=true` like the other destructive tools, since ending a session someone else is pulling files from will interrupt them.
## Verification against ZVM 10.x
- `zerto_list_flr_sessions(live_only=true)` → `count: 0, total_known: 1` after a clean run; `live_only=false` shows the `UnmountCompletedSuccessfully` history row
- `zerto_end_flr_session` without `confirmed` → `needs_confirm: true`, refused
- real recovery from checkpoint 1368 → 158 bytes, `flr_path: Volume2-Ext4/home/justin/app-config.yaml`, `unmount.ok: true` naming the session it ended
- 0 live sessions remaining afterwards
`pytest`: 27 passed (5 new, including fakes that cover the swallowed-failure path, the orphan reap, and the leave-other-sessions-alone guarantee). `ruff check` clean on all changed files.
Docs updated: `SKILL.md`, `README.md`, `docs/recover-ladder.md`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude
changed target branch from fix/flr-partition-path to main2026-09-21 13:42:37 -04:00
zerto_recover_file already tore its session down in a finally block, but
three gaps meant a mount could stay up on the recovery site with nothing
tracking it. FLR cannot run during clone, test, live failover or EJC, so
a stuck session blocks the next recovery.
1. An unmount failure was swallowed (`except ZertoError: pass`). The
caller got ok=true and never learned the mount was still up. The
teardown result is now reported in the response as `unmount`, with a
`warning` when it fails. ok stays true when the bytes did land -- the
recovery genuinely succeeded -- but the caller is told.
2. If start_flr succeeded on the ZVM while its response failed to parse,
session_id stayed None and the finally block did nothing, leaking a
session the process never knew the id of. Teardown now snapshots live
session ids before starting and reaps anything new that appeared,
leaving other operators' sessions alone.
3. Nothing could see or clear an orphan left by a crashed process, since
the finally block only runs if the process survives. Two new tools:
- zerto_list_flr_sessions: every session the ZVM knows about.
live_only (default true) keeps the ones still holding a mount;
ended and failed sessions linger as history and hold nothing.
- zerto_end_flr_session: unmount one. Gated on confirmed=true,
matching the other destructive tools, because ending a session
someone else is pulling files from will interrupt them.
Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean
run, the confirm gate refuses without a human yes, a real recovery from
checkpoint 1368 returned 158 bytes and reported
unmount.ok=true with the session id it ended, and 0 live sessions
remained afterwards.
pytest 27 passed (5 new, including fakes covering the swallowed-failure
and orphan-reap paths).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Follow-up to #2, which this is stacked on — base is
fix/flr-partition-path, so the diff here is only the new work. Merge #2 first, then this (or retarget tomainonce #2 lands).Closes the three FLR-session gaps identified while reviewing #2.
zerto_recover_filealready tore its session down in afinallyblock, but a mount could still end up stranded on the recovery site with nothing tracking it. That matters because FLR cannot run during clone, test, live failover or EJC — a stuck session blocks the next recovery, and you find out at the worst moment.1. Unmount failures were swallowed
If the teardown itself failed, the caller got
ok: trueand never learned the mount was still up. The teardown result is now reported:On failure the response also carries a
warning.okstaystruewhen the bytes landed — the recovery genuinely succeeded — but the caller is told rather than left guessing.2. A window where cleanup ran against nothing
If
start_flrsucceeded on the ZVM but its response failed to parse,session_id_from()raised,session_idstayedNone, and thefinallyblock no-opped — leaking a session whose id the process never learned.Teardown now snapshots live session ids before starting and reaps anything new that appeared. It deliberately leaves pre-existing sessions alone, so a concurrent operator's mount is never torn down (covered by a test).
3. Nothing could see or clear an orphan
The
finallyonly runs if the process survives. A crash, disconnect or timeout mid-recovery left a mount with no way to find it. Two new tools:zerto_list_flr_sessions— every session the ZVM knows about.live_only(default true) keeps the ones still holding a mount; ended and failed sessions linger as history and hold nothing.zerto_end_flr_session— unmount one. Gated onconfirmed=truelike the other destructive tools, since ending a session someone else is pulling files from will interrupt them.Verification against ZVM 10.x
zerto_list_flr_sessions(live_only=true)→count: 0, total_known: 1after a clean run;live_only=falseshows theUnmountCompletedSuccessfullyhistory rowzerto_end_flr_sessionwithoutconfirmed→needs_confirm: true, refusedflr_path: Volume2-Ext4/home/justin/app-config.yaml,unmount.ok: truenaming the session it endedpytest: 27 passed (5 new, including fakes that cover the swallowed-failure path, the orphan reap, and the leave-other-sessions-alone guarantee).ruff checkclean on all changed files.Docs updated:
SKILL.md,README.md,docs/recover-ladder.md.🤖 Generated with Claude Code
https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
zerto_recover_file already tore its session down in a finally block, but three gaps meant a mount could stay up on the recovery site with nothing tracking it. FLR cannot run during clone, test, live failover or EJC, so a stuck session blocks the next recovery. 1. An unmount failure was swallowed (`except ZertoError: pass`). The caller got ok=true and never learned the mount was still up. The teardown result is now reported in the response as `unmount`, with a `warning` when it fails. ok stays true when the bytes did land -- the recovery genuinely succeeded -- but the caller is told. 2. If start_flr succeeded on the ZVM while its response failed to parse, session_id stayed None and the finally block did nothing, leaking a session the process never knew the id of. Teardown now snapshots live session ids before starting and reaps anything new that appeared, leaving other operators' sessions alone. 3. Nothing could see or clear an orphan left by a crashed process, since the finally block only runs if the process survives. Two new tools: - zerto_list_flr_sessions: every session the ZVM knows about. live_only (default true) keeps the ones still holding a mount; ended and failed sessions linger as history and hold nothing. - zerto_end_flr_session: unmount one. Gated on confirmed=true, matching the other destructive tools, because ending a session someone else is pulling files from will interrupt them. Verified against ZVM 10.x: listing reports 0 live / 1 known after a clean run, the confirm gate refuses without a human yes, a real recovery from checkpoint 1368 returned 158 bytes and reported unmount.ok=true with the session id it ended, and 0 live sessions remained afterwards. pytest 27 passed (5 new, including fakes covering the swallowed-failure and orphan-reap paths). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGnf930b84615to8b973653d8