fix(checkpoints): derive the tag wait from the VPG's cadence #12

Merged
claude merged 1 commits from fix/cadence-aware-tag-wait into main 2026-09-22 21:37:20 -04:00
Contributor

Fixes the known issue recorded in #9: this hook denied every change to a cloud-protected VM.

The bug

wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by the VPG's protected site, and the spread is enormous:

protected at cadence tag visible after old 45s budget
vSphere 5s ~4s fine
Azure 60s ~34s barely
AWS 630s ~128s impossible

So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute after the agent had already been told there was no rewind point — it is in the journal as cp 56.

That is a worse failure than the one the guard exists to prevent: silent, correct-looking in the log, and it blocks legitimate work on every cloud-protected VM in the estate.

The fix

The wait is measured rather than guessed:

  • cadence_seconds() — median gap across recent checkpoints, None when unmeasurable
  • tag_wait_budget() — 2 x cadence + 30, clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is not polled every 1.5s

Visibility does not scale linearly with cadence — the insert creates its own off-cadence checkpoint — which is why this is a bounded multiple rather than a proportion.

vSphere  cadence=5.0    -> wait  45s, poll every  1.5s
Azure    cadence=60.0   -> wait 150s, poll every  6.0s
AWS      cadence=630.0  -> wait 300s, poll every 15.0s
unmeasurable            -> wait  45s, poll every  1.5s

wait_for_tag also checks once before sleeping, so an already-present tag returns without a poll cycle.

The failure message was itself misleading. It asserted the insert was unsupported, inviting exactly the wrong conclusion. It now names the measured cadence and says a completed Zerto task with no visible checkpoint means the wait was short, not that the insert was rejected.

Hook budgets follow

Leaving them low would just relocate the false denial into the hook: a cancelled hook has its output discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 → 330, settings timeout 180 → 360, both clear of the 300s cap.

Verified live against ZVM 10.9.10

Guard, on the VPG that was broken:

AWS-protected  win2019-1: ok in 110.6s  cp=186    (previously DENIED at 63s)
vSphere        jp-ubuntu: ok in   6.6s  cp=23166  (no regression)

Through the hook itself:

win2019-1: ALLOWED, tagged cp 187 on CMH-AWS-1, 111s
jp-ubuntu: ALLOWED, tagged cp 23198,              9s

vSphere pays nothing for this.

pytest: 69 passed (6 new). One asserts the derived budget exceeds the visibility actually measured on each of the three platforms, so a future tweak to the formula cannot silently reintroduce the AWS case.

🤖 Generated with Claude Code

https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn

Fixes the known issue recorded in #9: this hook denied every change to a cloud-protected VM. ## The bug `wait_for_tag` gave up after a hardcoded **45s**. Checkpoint cadence is set by the VPG's **protected** site, and the spread is enormous: | protected at | cadence | tag visible after | old 45s budget | |---|---|---|---| | vSphere | 5s | ~4s | fine | | Azure | 60s | ~34s | barely | | AWS | 630s | ~128s | **impossible** | So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute after the agent had already been told there was no rewind point — it is in the journal as `cp 56`. That is a worse failure than the one the guard exists to prevent: silent, correct-looking in the log, and it blocks legitimate work on every cloud-protected VM in the estate. ## The fix The wait is measured rather than guessed: - `cadence_seconds()` — median gap across recent checkpoints, `None` when unmeasurable - `tag_wait_budget()` — `2 x cadence + 30`, clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is not polled every 1.5s Visibility does **not** scale linearly with cadence — the insert creates its own off-cadence checkpoint — which is why this is a bounded multiple rather than a proportion. ``` vSphere cadence=5.0 -> wait 45s, poll every 1.5s Azure cadence=60.0 -> wait 150s, poll every 6.0s AWS cadence=630.0 -> wait 300s, poll every 15.0s unmeasurable -> wait 45s, poll every 1.5s ``` `wait_for_tag` also checks once before sleeping, so an already-present tag returns without a poll cycle. **The failure message was itself misleading.** It asserted the insert was unsupported, inviting exactly the wrong conclusion. It now names the measured cadence and says a completed Zerto task with no visible checkpoint means the wait was short, not that the insert was rejected. ## Hook budgets follow Leaving them low would just relocate the false denial into the hook: a cancelled hook has its output **discarded** and the call proceeds unguarded. `ZERTO_HOOK_GUARD_TIMEOUT` 150 → 330, settings timeout 180 → 360, both clear of the 300s cap. ## Verified live against ZVM 10.9.10 Guard, on the VPG that was broken: ``` AWS-protected win2019-1: ok in 110.6s cp=186 (previously DENIED at 63s) vSphere jp-ubuntu: ok in 6.6s cp=23166 (no regression) ``` Through the hook itself: ``` win2019-1: ALLOWED, tagged cp 187 on CMH-AWS-1, 111s jp-ubuntu: ALLOWED, tagged cp 23198, 9s ``` vSphere pays nothing for this. `pytest`: 69 passed (6 new). One asserts the derived budget exceeds the visibility **actually measured** on each of the three platforms, so a future tweak to the formula cannot silently reintroduce the AWS case. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude added 1 commit 2026-09-22 20:12:59 -04:00
wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by
the VPG's protected site and the spread is enormous: 5s on vSphere, 60s on
Azure, 630s on AWS, measured in one estate. A single constant cannot serve
all three.

Generous for vSphere, where a tag surfaces in about 4s. Impossible for AWS,
where it takes ~128s. So the guard reported "no checkpoint, refusing the
change" while Zerto was in the middle of creating one, and the checkpoint
landed a minute after the agent had been told there was no rewind point.

That is worse than the failure it guards against. It is silent, it reads as
correct in the log, and it blocks legitimate work on every cloud-protected
VM in the estate.

The wait is now measured: cadence_seconds() takes the median gap of recent
checkpoints, tag_wait_budget() turns that into 2x cadence plus headroom,
clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is
not polled every 1.5s. Visibility does not scale linearly with cadence,
because the insert creates its own off-cadence checkpoint, which is why
this is a bounded multiple rather than a proportion.

wait_for_tag also checks once before sleeping, so an already-present tag
returns without a poll cycle.

The failure message now names the measured cadence and says a completed
Zerto task with no visible checkpoint means the wait was short, not that
the insert was rejected. That was the exact wrong conclusion the old
message invited.

Hook budgets follow: too small a budget there just relocates the false
denial from the guard into the hook, since a cancelled hook has its output
discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 ->
330, settings timeout 180 -> 360, both above the 300s cap.

Verified live against ZVM 10.9.10. Guard on win2019-1 (VPG CMH-AWS-1,
AWS-protected) previously denied at 63s; now succeeds in 110.6s with
checkpoint 186. jp-ubuntu unchanged at 6.6s, so vSphere pays nothing for
this. Through the hook itself: win2019-1 allowed in 111s with cp 187,
jp-ubuntu allowed in 9s with cp 23198.

pytest 69 passed (6 new, including one asserting the budget exceeds the
visibility actually measured on each of the three platforms).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude merged commit 0595c472e1 into main 2026-09-22 21:37:20 -04:00
claude deleted branch fix/cadence-aware-tag-wait 2026-09-22 21:37:20 -04:00
Sign in to join this conversation.