Files
zerto-ai-rewind/hooks/README.md
T

94 lines
3.8 KiB
Markdown

# PreToolUse guard hook
Makes the rewind guard enforceable instead of advisory.
The MCP server cannot see another server's tool calls, so `zerto_check_tool` and
`zerto_guard_before_mutate` only work if the model chooses to call them. A model
that skips the step is not stopped by anything. A `PreToolUse` hook runs in the
host, where the tool call genuinely pauses, so a failed checkpoint stops the
change.
## Decisions
| situation | decision | effect |
|---|---|---|
| tool in `read_only_tools` | none | runs, no checkpoint |
| mutating, checkpoint confirmed | none, plus `additionalContext` | runs, and the model is told which checkpoint to recover from |
| mutating, checkpoint failed | `deny` | **the call never happens** |
| mutating, VM unknown to Zerto | `prompt` | human decides; nothing to rewind to |
| mutating, no VM in the arguments | `prompt` | catalog's `vm_arg` did not match |
| unlisted tool | `prompt` | nobody said it was read-only |
`ZERTO_HOOK_UNKNOWN` switches the unlisted case to `allow` or `deny`.
## Install
```bash
cp hooks/settings.example.json /tmp/x # then merge the hooks block into
# .claude/settings.json
export ZERTO_REWIND_CONFIG=/abs/path/config.json
```
Both paths in the command must be absolute, and the interpreter must be the
venv that has this package installed.
## Timeouts, which matter here
The hook is synchronous: the host waits. That is the point, because the
checkpoint has to exist before the change does.
Budget for the slow path, not the fast one. A successful tag took about 4s
against a healthy vSphere-protected VPG, but a **refusal took 63s**, because
`wait_for_tag` spends 45s before giving up.
So:
- `timeout` in settings.json: 180 (seconds)
- `ZERTO_HOOK_GUARD_TIMEOUT`: 150 (seconds), kept under it
If the host's timeout fires first it cancels the hook and **discards its
output**, and the tool call carries on through the normal permission flow. A
timeout is therefore a silent failure of the guard, which is why the hook's own
budget is the smaller of the two: it would rather deny than be cancelled.
## Verified behaviour
Against a live ZVM 10.9.10, all six rows of the table above. Two worth naming:
- `jp-ubuntu` (healthy, local VPG) tagged checkpoint 7180 and allowed the call,
passing the tag back through `additionalContext`.
- `win2019-1` (VPG `CMH-AWS-1`, protected site AWS) was **denied**. The deny path
works, but see the known issue below: that particular denial was wrong.
Both decisions held with `permission_mode: bypassPermissions`. A hook still
blocks when the user has turned permissions off, which is when an agent is most
likely to be running unattended.
## Log
`~/.zerto-guard-hook.log`, or `ZERTO_HOOK_LOG`. One line per decision.
## Known issue: false denials on cloud-protected VPGs
The `win2019-1` denial above was a **false negative**, and it is worth
understanding before relying on this hook in an estate with cloud-protected
workloads.
`wait_for_tag` gives up after a hardcoded 45s. That is generous for a
vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about
4s. It is far too short elsewhere: a tag takes ~34s to appear on an
Azure-protected VPG and ~128s on an AWS-protected one, because journal cadence
is set by the protected site (5s vSphere, 60s Azure, 630s AWS).
So the hook denied the change, and the checkpoint landed anyway. It is in the
journal as `cp 56`, timestamped a minute after the hook reported failure. The
guard told the agent there was no rewind point while Zerto was in the middle of
creating one.
That failure mode is worse than the one the hook guards against, because it is
silent and looks correct: legitimate work is refused on every cloud-protected VM
while the log reads like the guard is doing its job.
Until `wait_for_tag` becomes cadence-aware, scope this hook's matcher to
vSphere-protected workloads.