Files
zerto-ai-rewind/hooks/README.md
T
justinandClaude Opus 5 735c6c4a64 docs(hooks): correct the cloud-protected denial claim
hooks/README.md said the win2019-1 denial happened because Zerto cannot
insert a tagged checkpoint on an AWS-protected VPG. That is not what
happened, and it is not true on 10.9.10.

The insert succeeded. The checkpoint is in the journal as cp 56, stamped
about a minute after the hook had already reported failure. What actually
happened is that wait_for_tag gave up after its hardcoded 45s, which is
generous on a VPG that checkpoints every 5s and far too short on one that
checkpoints every 630s.

So the documented "verified" example was a false denial. The deny path
works; that particular denial was wrong. Recorded as a known issue,
because a guard that silently refuses legitimate work on every
cloud-protected VM while looking correct is a worse failure than the one
it exists to prevent.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
2026-09-22 18:05:23 -04:00

94 lines
3.8 KiB
Markdown

# PreToolUse guard hook
Makes the rewind guard enforceable instead of advisory.
The MCP server cannot see another server's tool calls, so `zerto_check_tool` and
`zerto_guard_before_mutate` only work if the model chooses to call them. A model
that skips the step is not stopped by anything. A `PreToolUse` hook runs in the
host, where the tool call genuinely pauses, so a failed checkpoint stops the
change.
## Decisions
| situation | decision | effect |
|---|---|---|
| tool in `read_only_tools` | none | runs, no checkpoint |
| mutating, checkpoint confirmed | none, plus `additionalContext` | runs, and the model is told which checkpoint to recover from |
| mutating, checkpoint failed | `deny` | **the call never happens** |
| mutating, VM unknown to Zerto | `prompt` | human decides; nothing to rewind to |
| mutating, no VM in the arguments | `prompt` | catalog's `vm_arg` did not match |
| unlisted tool | `prompt` | nobody said it was read-only |
`ZERTO_HOOK_UNKNOWN` switches the unlisted case to `allow` or `deny`.
## Install
```bash
cp hooks/settings.example.json /tmp/x # then merge the hooks block into
# .claude/settings.json
export ZERTO_REWIND_CONFIG=/abs/path/config.json
```
Both paths in the command must be absolute, and the interpreter must be the
venv that has this package installed.
## Timeouts, which matter here
The hook is synchronous: the host waits. That is the point, because the
checkpoint has to exist before the change does.
Budget for the slow path, not the fast one. A successful tag took about 4s
against a healthy vSphere-protected VPG, but a **refusal took 63s**, because
`wait_for_tag` spends 45s before giving up.
So:
- `timeout` in settings.json: 180 (seconds)
- `ZERTO_HOOK_GUARD_TIMEOUT`: 150 (seconds), kept under it
If the host's timeout fires first it cancels the hook and **discards its
output**, and the tool call carries on through the normal permission flow. A
timeout is therefore a silent failure of the guard, which is why the hook's own
budget is the smaller of the two: it would rather deny than be cancelled.
## Verified behaviour
Against a live ZVM 10.9.10, all six rows of the table above. Two worth naming:
- `jp-ubuntu` (healthy, local VPG) tagged checkpoint 7180 and allowed the call,
passing the tag back through `additionalContext`.
- `win2019-1` (VPG `CMH-AWS-1`, protected site AWS) was **denied**. The deny path
works, but see the known issue below: that particular denial was wrong.
Both decisions held with `permission_mode: bypassPermissions`. A hook still
blocks when the user has turned permissions off, which is when an agent is most
likely to be running unattended.
## Log
`~/.zerto-guard-hook.log`, or `ZERTO_HOOK_LOG`. One line per decision.
## Known issue: false denials on cloud-protected VPGs
The `win2019-1` denial above was a **false negative**, and it is worth
understanding before relying on this hook in an estate with cloud-protected
workloads.
`wait_for_tag` gives up after a hardcoded 45s. That is generous for a
vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about
4s. It is far too short elsewhere: a tag takes ~34s to appear on an
Azure-protected VPG and ~128s on an AWS-protected one, because journal cadence
is set by the protected site (5s vSphere, 60s Azure, 630s AWS).
So the hook denied the change, and the checkpoint landed anyway. It is in the
journal as `cp 56`, timestamped a minute after the hook reported failure. The
guard told the agent there was no rewind point while Zerto was in the middle of
creating one.
That failure mode is worse than the one the hook guards against, because it is
silent and looks correct: legitimate work is refused on every cloud-protected VM
while the log reads like the guard is doing its job.
Until `wait_for_tag` becomes cadence-aware, scope this hook's matcher to
vSphere-protected workloads.