Files
zerto-ai-rewind/hooks/README.md
T
justinandClaude Opus 5 d86b4814a5 fix(checkpoints): derive the tag wait from the VPG's cadence
wait_for_tag gave up after a hardcoded 45s. Checkpoint cadence is set by
the VPG's protected site and the spread is enormous: 5s on vSphere, 60s on
Azure, 630s on AWS, measured in one estate. A single constant cannot serve
all three.

Generous for vSphere, where a tag surfaces in about 4s. Impossible for AWS,
where it takes ~128s. So the guard reported "no checkpoint, refusing the
change" while Zerto was in the middle of creating one, and the checkpoint
landed a minute after the agent had been told there was no rewind point.

That is worse than the failure it guards against. It is silent, it reads as
correct in the log, and it blocks legitimate work on every cloud-protected
VM in the estate.

The wait is now measured: cadence_seconds() takes the median gap of recent
checkpoints, tag_wait_budget() turns that into 2x cadence plus headroom,
clamped to 45s..300s, and scales the poll interval with it so a 630s VPG is
not polled every 1.5s. Visibility does not scale linearly with cadence,
because the insert creates its own off-cadence checkpoint, which is why
this is a bounded multiple rather than a proportion.

wait_for_tag also checks once before sleeping, so an already-present tag
returns without a poll cycle.

The failure message now names the measured cadence and says a completed
Zerto task with no visible checkpoint means the wait was short, not that
the insert was rejected. That was the exact wrong conclusion the old
message invited.

Hook budgets follow: too small a budget there just relocates the false
denial from the guard into the hook, since a cancelled hook has its output
discarded and the call proceeds unguarded. ZERTO_HOOK_GUARD_TIMEOUT 150 ->
330, settings timeout 180 -> 360, both above the 300s cap.

Verified live against ZVM 10.9.10. Guard on win2019-1 (VPG CMH-AWS-1,
AWS-protected) previously denied at 63s; now succeeds in 110.6s with
checkpoint 186. jp-ubuntu unchanged at 6.6s, so vSphere pays nothing for
this. Through the hook itself: win2019-1 allowed in 111s with cp 187,
jp-ubuntu allowed in 9s with cp 23198.

pytest 69 passed (6 new, including one asserting the budget exceeds the
visibility actually measured on each of the three platforms).

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-22 20:12:57 -04:00

96 lines
3.8 KiB
Markdown

# PreToolUse guard hook
Makes the rewind guard enforceable instead of advisory.
The MCP server cannot see another server's tool calls, so `zerto_check_tool` and
`zerto_guard_before_mutate` only work if the model chooses to call them. A model
that skips the step is not stopped by anything. A `PreToolUse` hook runs in the
host, where the tool call genuinely pauses, so a failed checkpoint stops the
change.
## Decisions
| situation | decision | effect |
|---|---|---|
| tool in `read_only_tools` | none | runs, no checkpoint |
| mutating, checkpoint confirmed | none, plus `additionalContext` | runs, and the model is told which checkpoint to recover from |
| mutating, checkpoint failed | `deny` | **the call never happens** |
| mutating, VM unknown to Zerto | `prompt` | human decides; nothing to rewind to |
| mutating, no VM in the arguments | `prompt` | catalog's `vm_arg` did not match |
| unlisted tool | `prompt` | nobody said it was read-only |
`ZERTO_HOOK_UNKNOWN` switches the unlisted case to `allow` or `deny`.
## Install
```bash
cp hooks/settings.example.json /tmp/x # then merge the hooks block into
# .claude/settings.json
export ZERTO_REWIND_CONFIG=/abs/path/config.json
```
Both paths in the command must be absolute, and the interpreter must be the
venv that has this package installed.
## Timeouts, which matter here
The hook is synchronous: the host waits. That is the point, because the
checkpoint has to exist before the change does.
Budget for the slow path, not the fast one. How long the guard takes is set by
the VPG's checkpoint cadence, which is set in turn by its protected site:
| protected at | cadence | guard takes |
|---|---|---|
| vSphere | 5s | ~7s |
| Azure | 60s | ~40s |
| AWS | 630s | ~111s |
The tag wait is derived from that cadence and capped at 300s, so:
- `timeout` in settings.json: 360 (seconds)
- `ZERTO_HOOK_GUARD_TIMEOUT`: 330 (seconds), kept under it
If the host's timeout fires first it cancels the hook and **discards its
output**, and the tool call carries on through the normal permission flow. A
timeout is therefore a silent failure of the guard, which is why the hook's own
budget is the smaller of the two: it would rather deny than be cancelled.
## Verified behaviour
Against a live ZVM 10.9.10, all six rows of the table above. Two worth naming:
- `jp-ubuntu` (healthy, local VPG) tagged checkpoint 7180 and allowed the call,
passing the tag back through `additionalContext`.
- `win2019-1` (VPG `CMH-AWS-1`, protected site AWS) was **denied**. The deny path
works, but see the known issue below: that particular denial was wrong.
Both decisions held with `permission_mode: bypassPermissions`. A hook still
blocks when the user has turned permissions off, which is when an agent is most
likely to be running unattended.
## Log
`~/.zerto-guard-hook.log`, or `ZERTO_HOOK_LOG`. One line per decision.
## Why the timeouts are derived, not fixed
This hook used to deny every change to a cloud-protected VM.
`wait_for_tag` gave up after a hardcoded 45s. That is generous on a
vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about
4s, and impossible on an AWS-protected one, where a tag takes ~128s because the
journal only checkpoints every 630s.
So the guard reported "no checkpoint, refusing the change" while Zerto was in the
middle of creating one. The checkpoint landed a minute later, in the journal,
after the agent had already been told there was no rewind point.
That is a worse failure than the one this hook exists to prevent. It is silent,
it looks correct in the log, and it blocks legitimate work on every
cloud-protected VM in the estate.
Both budgets are now derived from the VPG's measured cadence rather than
guessed, which is why the numbers above differ by a factor of fifteen between
platforms.