Fixes the known issue recorded in #9: this hook denied every change to a cloud-protected VM.
The bug
wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by the VPG's protected site, and the spread is enormous:
protected at
cadence
tag visible after
old 45s budget
vSphere
5s
~4s
fine
Azure
60s
~34s
barely
AWS
630s
~128s
impossible
So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute after the agent had already been told there was no rewind point — it is in the journal as cp 56.
That is a worse failure than the one the guard exists to prevent: silent, correct-looking in the log, and it blocks legitimate work on every cloud-protected VM in the estate.
The fix
The wait is measured rather than guessed:
cadence_seconds() — median gap across recent checkpoints, None when unmeasurable
tag_wait_budget() — 2 x cadence + 30, clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is not polled every 1.5s
Visibility does not scale linearly with cadence — the insert creates its own off-cadence checkpoint — which is why this is a bounded multiple rather than a proportion.
vSphere cadence=5.0 -> wait 45s, poll every 1.5s
Azure cadence=60.0 -> wait 150s, poll every 6.0s
AWS cadence=630.0 -> wait 300s, poll every 15.0s
unmeasurable -> wait 45s, poll every 1.5s
wait_for_tag also checks once before sleeping, so an already-present tag returns without a poll cycle.
The failure message was itself misleading. It asserted the insert was unsupported, inviting exactly the wrong conclusion. It now names the measured cadence and says a completed Zerto task with no visible checkpoint means the wait was short, not that the insert was rejected.
Hook budgets follow
Leaving them low would just relocate the false denial into the hook: a cancelled hook has its output discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 → 330, settings timeout 180 → 360, both clear of the 300s cap.
Verified live against ZVM 10.9.10
Guard, on the VPG that was broken:
AWS-protected win2019-1: ok in 110.6s cp=186 (previously DENIED at 63s)
vSphere jp-ubuntu: ok in 6.6s cp=23166 (no regression)
pytest: 69 passed (6 new). One asserts the derived budget exceeds the visibility actually measured on each of the three platforms, so a future tweak to the formula cannot silently reintroduce the AWS case.
Fixes the known issue recorded in #9: this hook denied every change to a cloud-protected VM.
## The bug
`wait_for_tag` gave up after a hardcoded **45s**. Checkpoint cadence is set by the VPG's **protected** site, and the spread is enormous:
| protected at | cadence | tag visible after | old 45s budget |
|---|---|---|---|
| vSphere | 5s | ~4s | fine |
| Azure | 60s | ~34s | barely |
| AWS | 630s | ~128s | **impossible** |
So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute after the agent had already been told there was no rewind point — it is in the journal as `cp 56`.
That is a worse failure than the one the guard exists to prevent: silent, correct-looking in the log, and it blocks legitimate work on every cloud-protected VM in the estate.
## The fix
The wait is measured rather than guessed:
- `cadence_seconds()` — median gap across recent checkpoints, `None` when unmeasurable
- `tag_wait_budget()` — `2 x cadence + 30`, clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is not polled every 1.5s
Visibility does **not** scale linearly with cadence — the insert creates its own off-cadence checkpoint — which is why this is a bounded multiple rather than a proportion.
```
vSphere cadence=5.0 -> wait 45s, poll every 1.5s
Azure cadence=60.0 -> wait 150s, poll every 6.0s
AWS cadence=630.0 -> wait 300s, poll every 15.0s
unmeasurable -> wait 45s, poll every 1.5s
```
`wait_for_tag` also checks once before sleeping, so an already-present tag returns without a poll cycle.
**The failure message was itself misleading.** It asserted the insert was unsupported, inviting exactly the wrong conclusion. It now names the measured cadence and says a completed Zerto task with no visible checkpoint means the wait was short, not that the insert was rejected.
## Hook budgets follow
Leaving them low would just relocate the false denial into the hook: a cancelled hook has its output **discarded** and the call proceeds unguarded. `ZERTO_HOOK_GUARD_TIMEOUT` 150 → 330, settings timeout 180 → 360, both clear of the 300s cap.
## Verified live against ZVM 10.9.10
Guard, on the VPG that was broken:
```
AWS-protected win2019-1: ok in 110.6s cp=186 (previously DENIED at 63s)
vSphere jp-ubuntu: ok in 6.6s cp=23166 (no regression)
```
Through the hook itself:
```
win2019-1: ALLOWED, tagged cp 187 on CMH-AWS-1, 111s
jp-ubuntu: ALLOWED, tagged cp 23198, 9s
```
vSphere pays nothing for this.
`pytest`: 69 passed (6 new). One asserts the derived budget exceeds the visibility **actually measured** on each of the three platforms, so a future tweak to the formula cannot silently reintroduce the AWS case.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by
the VPG's protected site and the spread is enormous: 5s on vSphere, 60s on
Azure, 630s on AWS, measured in one estate. A single constant cannot serve
all three.
Generous for vSphere, where a tag surfaces in about 4s. Impossible for AWS,
where it takes ~128s. So the guard reported "no checkpoint, refusing the
change" while Zerto was in the middle of creating one, and the checkpoint
landed a minute after the agent had been told there was no rewind point.
That is worse than the failure it guards against. It is silent, it reads as
correct in the log, and it blocks legitimate work on every cloud-protected
VM in the estate.
The wait is now measured: cadence_seconds() takes the median gap of recent
checkpoints, tag_wait_budget() turns that into 2x cadence plus headroom,
clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is
not polled every 1.5s. Visibility does not scale linearly with cadence,
because the insert creates its own off-cadence checkpoint, which is why
this is a bounded multiple rather than a proportion.
wait_for_tag also checks once before sleeping, so an already-present tag
returns without a poll cycle.
The failure message now names the measured cadence and says a completed
Zerto task with no visible checkpoint means the wait was short, not that
the insert was rejected. That was the exact wrong conclusion the old
message invited.
Hook budgets follow: too small a budget there just relocates the false
denial from the guard into the hook, since a cancelled hook has its output
discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 ->
330, settings timeout 180 -> 360, both above the 300s cap.
Verified live against ZVM 10.9.10. Guard on win2019-1 (VPG CMH-AWS-1,
AWS-protected) previously denied at 63s; now succeeds in 110.6s with
checkpoint 186. jp-ubuntu unchanged at 6.6s, so vSphere pays nothing for
this. Through the hook itself: win2019-1 allowed in 111s with cp 187,
jp-ubuntu allowed in 9s with cp 23198.
pytest 69 passed (6 new, including one asserting the budget exceeds the
visibility actually measured on each of the three platforms).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude
merged commit 0595c472e1 into main2026-09-22 21:37:20 -04:00
claude
deleted branch fix/cadence-aware-tag-wait2026-09-22 21:37:20 -04:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Fixes the known issue recorded in #9: this hook denied every change to a cloud-protected VM.
The bug
wait_for_taggave up after a hardcoded 45s. Checkpoint cadence is set by the VPG's protected site, and the spread is enormous:So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute after the agent had already been told there was no rewind point — it is in the journal as
cp 56.That is a worse failure than the one the guard exists to prevent: silent, correct-looking in the log, and it blocks legitimate work on every cloud-protected VM in the estate.
The fix
The wait is measured rather than guessed:
cadence_seconds()— median gap across recent checkpoints,Nonewhen unmeasurabletag_wait_budget()—2 x cadence + 30, clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is not polled every 1.5sVisibility does not scale linearly with cadence — the insert creates its own off-cadence checkpoint — which is why this is a bounded multiple rather than a proportion.
wait_for_tagalso checks once before sleeping, so an already-present tag returns without a poll cycle.The failure message was itself misleading. It asserted the insert was unsupported, inviting exactly the wrong conclusion. It now names the measured cadence and says a completed Zerto task with no visible checkpoint means the wait was short, not that the insert was rejected.
Hook budgets follow
Leaving them low would just relocate the false denial into the hook: a cancelled hook has its output discarded and the call proceeds unguarded.
ZERTO_HOOK_GUARD_TIMEOUT150 → 330, settings timeout 180 → 360, both clear of the 300s cap.Verified live against ZVM 10.9.10
Guard, on the VPG that was broken:
Through the hook itself:
vSphere pays nothing for this.
pytest: 69 passed (6 new). One asserts the derived budget exceeds the visibility actually measured on each of the three platforms, so a future tweak to the formula cannot silently reintroduce the AWS case.🤖 Generated with Claude Code
https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn