fix(checkpoints): derive the tag wait from the VPG's cadence (#12)

This commit was merged in pull request #12.
This commit is contained in:
2026-09-22 21:37:20 -04:00
parent 5916cdf8a2
commit 0595c472e1
5 changed files with 197 additions and 37 deletions
+25 -23
View File
@@ -37,14 +37,19 @@ venv that has this package installed.
The hook is synchronous: the host waits. That is the point, because the
checkpoint has to exist before the change does.
Budget for the slow path, not the fast one. A successful tag took about 4s
against a healthy vSphere-protected VPG, but a **refusal took 63s**, because
`wait_for_tag` spends 45s before giving up.
Budget for the slow path, not the fast one. How long the guard takes is set by
the VPG's checkpoint cadence, which is set in turn by its protected site:
So:
| protected at | cadence | guard takes |
|---|---|---|
| vSphere | 5s | ~7s |
| Azure | 60s | ~40s |
| AWS | 630s | ~111s |
- `timeout` in settings.json: 180 (seconds)
- `ZERTO_HOOK_GUARD_TIMEOUT`: 150 (seconds), kept under it
The tag wait is derived from that cadence and capped at 300s, so:
- `timeout` in settings.json: 360 (seconds)
- `ZERTO_HOOK_GUARD_TIMEOUT`: 330 (seconds), kept under it
If the host's timeout fires first it cancels the hook and **discards its
output**, and the tool call carries on through the normal permission flow. A
@@ -68,26 +73,23 @@ likely to be running unattended.
`~/.zerto-guard-hook.log`, or `ZERTO_HOOK_LOG`. One line per decision.
## Known issue: false denials on cloud-protected VPGs
## Why the timeouts are derived, not fixed
The `win2019-1` denial above was a **false negative**, and it is worth
understanding before relying on this hook in an estate with cloud-protected
workloads.
This hook used to deny every change to a cloud-protected VM.
`wait_for_tag` gives up after a hardcoded 45s. That is generous for a
`wait_for_tag` gave up after a hardcoded 45s. That is generous on a
vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about
4s. It is far too short elsewhere: a tag takes ~34s to appear on an
Azure-protected VPG and ~128s on an AWS-protected one, because journal cadence
is set by the protected site (5s vSphere, 60s Azure, 630s AWS).
4s, and impossible on an AWS-protected one, where a tag takes ~128s because the
journal only checkpoints every 630s.
So the hook denied the change, and the checkpoint landed anyway. It is in the
journal as `cp 56`, timestamped a minute after the hook reported failure. The
guard told the agent there was no rewind point while Zerto was in the middle of
creating one.
So the guard reported "no checkpoint, refusing the change" while Zerto was in the
middle of creating one. The checkpoint landed a minute later, in the journal,
after the agent had already been told there was no rewind point.
That failure mode is worse than the one the hook guards against, because it is
silent and looks correct: legitimate work is refused on every cloud-protected VM
while the log reads like the guard is doing its job.
That is a worse failure than the one this hook exists to prevent. It is silent,
it looks correct in the log, and it blocks legitimate work on every
cloud-protected VM in the estate.
Until `wait_for_tag` becomes cadence-aware, scope this hook's matcher to
vSphere-protected workloads.
Both budgets are now derived from the VPG's measured cadence rather than
guessed, which is why the numbers above differ by a factor of fifteen between
platforms.
+1 -1
View File
@@ -8,7 +8,7 @@
{
"type": "command",
"command": "/home/you/zerto-ai-rewind/.venv/bin/python /home/you/zerto-ai-rewind/hooks/zerto_guard_hook.py",
"timeout": 180
"timeout": 360
}
]
}
+6 -2
View File
@@ -24,7 +24,7 @@ tagged checkpoint and waits for the Zerto task to reach Completed.
"matcher": "mcp__.*",
"hooks": [{"type": "command",
"command": "/path/to/.venv/bin/python /path/to/hooks/zerto_guard_hook.py",
"timeout": 180}]}]}}
"timeout": 360}]}]}}
"""
from __future__ import annotations
@@ -40,7 +40,11 @@ LOG = os.environ.get("ZERTO_HOOK_LOG", os.path.expanduser("~/.zerto-guard-hook.l
# Seconds the hook will wait for the checkpoint. Must stay under the hook
# timeout configured in settings.json, or the host cancels us and the tool
# call proceeds unguarded through the normal permission flow.
GUARD_TIMEOUT_S = float(os.environ.get("ZERTO_HOOK_GUARD_TIMEOUT", "150"))
# Must exceed the largest tag wait the guard can take. That is now derived from
# the VPG's checkpoint cadence and capped at 300s (MAX_TAG_TIMEOUT_S), because a
# tag takes ~128s to surface on an AWS-protected VPG. Too small a budget here
# just moves the false denial from the guard into the hook.
GUARD_TIMEOUT_S = float(os.environ.get("ZERTO_HOOK_GUARD_TIMEOUT", "330"))
UNKNOWN_DECISION = os.environ.get("ZERTO_HOOK_UNKNOWN", "prompt") # prompt | allow | deny