Files
zerto-ai-rewind/hooks
justinandClaude Opus 5 d86b4814a5 fix(checkpoints): derive the tag wait from the VPG's cadence
wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by
the VPG's protected site and the spread is enormous: 5s on vSphere, 60s on
Azure, 630s on AWS, measured in one estate. A single constant cannot serve
all three.

Generous for vSphere, where a tag surfaces in about 4s. Impossible for AWS,
where it takes ~128s. So the guard reported "no checkpoint, refusing the
change" while Zerto was in the middle of creating one, and the checkpoint
landed a minute after the agent had been told there was no rewind point.

That is worse than the failure it guards against. It is silent, it reads as
correct in the log, and it blocks legitimate work on every cloud-protected
VM in the estate.

The wait is now measured: cadence_seconds() takes the median gap of recent
checkpoints, tag_wait_budget() turns that into 2x cadence plus headroom,
clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is
not polled every 1.5s. Visibility does not scale linearly with cadence,
because the insert creates its own off-cadence checkpoint, which is why
this is a bounded multiple rather than a proportion.

wait_for_tag also checks once before sleeping, so an already-present tag
returns without a poll cycle.

The failure message now names the measured cadence and says a completed
Zerto task with no visible checkpoint means the wait was short, not that
the insert was rejected. That was the exact wrong conclusion the old
message invited.

Hook budgets follow: too small a budget there just relocates the false
denial from the guard into the hook, since a cancelled hook has its output
discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 ->
330, settings timeout 180 -> 360, both above the 300s cap.

Verified live against ZVM 10.9.10. Guard on win2019-1 (VPG CMH-AWS-1,
AWS-protected) previously denied at 63s; now succeeds in 110.6s with
checkpoint 186. jp-ubuntu unchanged at 6.6s, so vSphere pays nothing for
this. Through the hook itself: win2019-1 allowed in 111s with cp 187,
jp-ubuntu allowed in 9s with cp 23198.

pytest 69 passed (6 new, including one asserting the budget exceeds the
visibility actually measured on each of the three platforms).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-22 20:12:57 -04:00
..

PreToolUse guard hook

Makes the rewind guard enforceable instead of advisory.

The MCP server cannot see another server's tool calls, so zerto_check_tool and zerto_guard_before_mutate only work if the model chooses to call them. A model that skips the step is not stopped by anything. A PreToolUse hook runs in the host, where the tool call genuinely pauses, so a failed checkpoint stops the change.

Decisions

situation decision effect
tool in read_only_tools none runs, no checkpoint
mutating, checkpoint confirmed none, plus additionalContext runs, and the model is told which checkpoint to recover from
mutating, checkpoint failed deny the call never happens
mutating, VM unknown to Zerto prompt human decides; nothing to rewind to
mutating, no VM in the arguments prompt catalog's vm_arg did not match
unlisted tool prompt nobody said it was read-only

ZERTO_HOOK_UNKNOWN switches the unlisted case to allow or deny.

Install

cp hooks/settings.example.json /tmp/x        # then merge the hooks block into
                                             # .claude/settings.json
export ZERTO_REWIND_CONFIG=/abs/path/config.json

Both paths in the command must be absolute, and the interpreter must be the venv that has this package installed.

Timeouts, which matter here

The hook is synchronous: the host waits. That is the point, because the checkpoint has to exist before the change does.

Budget for the slow path, not the fast one. How long the guard takes is set by the VPG's checkpoint cadence, which is set in turn by its protected site:

protected at cadence guard takes
vSphere 5s ~7s
Azure 60s ~40s
AWS 630s ~111s

The tag wait is derived from that cadence and capped at 300s, so:

  • timeout in settings.json: 360 (seconds)
  • ZERTO_HOOK_GUARD_TIMEOUT: 330 (seconds), kept under it

If the host's timeout fires first it cancels the hook and discards its output, and the tool call carries on through the normal permission flow. A timeout is therefore a silent failure of the guard, which is why the hook's own budget is the smaller of the two: it would rather deny than be cancelled.

Verified behaviour

Against a live ZVM 10.9.10, all six rows of the table above. Two worth naming:

  • jp-ubuntu (healthy, local VPG) tagged checkpoint 7180 and allowed the call, passing the tag back through additionalContext.
  • win2019-1 (VPG CMH-AWS-1, protected site AWS) was denied. The deny path works, but see the known issue below: that particular denial was wrong.

Both decisions held with permission_mode: bypassPermissions. A hook still blocks when the user has turned permissions off, which is when an agent is most likely to be running unattended.

Log

~/.zerto-guard-hook.log, or ZERTO_HOOK_LOG. One line per decision.

Why the timeouts are derived, not fixed

This hook used to deny every change to a cloud-protected VM.

wait_for_tag gave up after a hardcoded 45s. That is generous on a vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about 4s, and impossible on an AWS-protected one, where a tag takes ~128s because the journal only checkpoints every 630s.

So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute later, in the journal, after the agent had already been told there was no rewind point.

That is a worse failure than the one this hook exists to prevent. It is silent, it looks correct in the log, and it blocks legitimate work on every cloud-protected VM in the estate.

Both budgets are now derived from the VPG's measured cadence rather than guessed, which is why the numbers above differ by a factor of fifteen between platforms.