3.8 KiB
PreToolUse guard hook
Makes the rewind guard enforceable instead of advisory.
The MCP server cannot see another server's tool calls, so zerto_check_tool and
zerto_guard_before_mutate only work if the model chooses to call them. A model
that skips the step is not stopped by anything. A PreToolUse hook runs in the
host, where the tool call genuinely pauses, so a failed checkpoint stops the
change.
Decisions
| situation | decision | effect |
|---|---|---|
tool in read_only_tools |
none | runs, no checkpoint |
| mutating, checkpoint confirmed | none, plus additionalContext |
runs, and the model is told which checkpoint to recover from |
| mutating, checkpoint failed | deny |
the call never happens |
| mutating, VM unknown to Zerto | prompt |
human decides; nothing to rewind to |
| mutating, no VM in the arguments | prompt |
catalog's vm_arg did not match |
| unlisted tool | prompt |
nobody said it was read-only |
ZERTO_HOOK_UNKNOWN switches the unlisted case to allow or deny.
Install
cp hooks/settings.example.json /tmp/x # then merge the hooks block into
# .claude/settings.json
export ZERTO_REWIND_CONFIG=/abs/path/config.json
Both paths in the command must be absolute, and the interpreter must be the venv that has this package installed.
Timeouts, which matter here
The hook is synchronous: the host waits. That is the point, because the checkpoint has to exist before the change does.
Budget for the slow path, not the fast one. How long the guard takes is set by the VPG's checkpoint cadence, which is set in turn by its protected site:
| protected at | cadence | guard takes |
|---|---|---|
| vSphere | 5s | ~7s |
| Azure | 60s | ~40s |
| AWS | 630s | ~111s |
The tag wait is derived from that cadence and capped at 300s, so:
timeoutin settings.json: 360 (seconds)ZERTO_HOOK_GUARD_TIMEOUT: 330 (seconds), kept under it
If the host's timeout fires first it cancels the hook and discards its output, and the tool call carries on through the normal permission flow. A timeout is therefore a silent failure of the guard, which is why the hook's own budget is the smaller of the two: it would rather deny than be cancelled.
Verified behaviour
Against a live ZVM 10.9.10, all six rows of the table above. Two worth naming:
jp-ubuntu(healthy, local VPG) tagged checkpoint 7180 and allowed the call, passing the tag back throughadditionalContext.win2019-1(VPGCMH-AWS-1, protected site AWS) was denied. The deny path works, but see the known issue below: that particular denial was wrong.
Both decisions held with permission_mode: bypassPermissions. A hook still
blocks when the user has turned permissions off, which is when an agent is most
likely to be running unattended.
Log
~/.zerto-guard-hook.log, or ZERTO_HOOK_LOG. One line per decision.
Why the timeouts are derived, not fixed
This hook used to deny every change to a cloud-protected VM.
wait_for_tag gave up after a hardcoded 45s. That is generous on a
vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about
4s, and impossible on an AWS-protected one, where a tag takes ~128s because the
journal only checkpoints every 630s.
So the guard reported "no checkpoint, refusing the change" while Zerto was in the middle of creating one. The checkpoint landed a minute later, in the journal, after the agent had already been told there was no rewind point.
That is a worse failure than the one this hook exists to prevent. It is silent, it looks correct in the log, and it blocks legitimate work on every cloud-protected VM in the estate.
Both budgets are now derived from the VPG's measured cadence rather than guessed, which is why the numbers above differ by a factor of fifteen between platforms.