fix(checkpoints): derive the tag wait from the VPG's cadence (#12)
This commit was merged in pull request #12.
This commit is contained in:
+25
-23
@@ -37,14 +37,19 @@ venv that has this package installed.
|
||||
The hook is synchronous: the host waits. That is the point, because the
|
||||
checkpoint has to exist before the change does.
|
||||
|
||||
Budget for the slow path, not the fast one. A successful tag took about 4s
|
||||
against a healthy vSphere-protected VPG, but a **refusal took 63s**, because
|
||||
`wait_for_tag` spends 45s before giving up.
|
||||
Budget for the slow path, not the fast one. How long the guard takes is set by
|
||||
the VPG's checkpoint cadence, which is set in turn by its protected site:
|
||||
|
||||
So:
|
||||
| protected at | cadence | guard takes |
|
||||
|---|---|---|
|
||||
| vSphere | 5s | ~7s |
|
||||
| Azure | 60s | ~40s |
|
||||
| AWS | 630s | ~111s |
|
||||
|
||||
- `timeout` in settings.json: 180 (seconds)
|
||||
- `ZERTO_HOOK_GUARD_TIMEOUT`: 150 (seconds), kept under it
|
||||
The tag wait is derived from that cadence and capped at 300s, so:
|
||||
|
||||
- `timeout` in settings.json: 360 (seconds)
|
||||
- `ZERTO_HOOK_GUARD_TIMEOUT`: 330 (seconds), kept under it
|
||||
|
||||
If the host's timeout fires first it cancels the hook and **discards its
|
||||
output**, and the tool call carries on through the normal permission flow. A
|
||||
@@ -68,26 +73,23 @@ likely to be running unattended.
|
||||
|
||||
`~/.zerto-guard-hook.log`, or `ZERTO_HOOK_LOG`. One line per decision.
|
||||
|
||||
## Known issue: false denials on cloud-protected VPGs
|
||||
## Why the timeouts are derived, not fixed
|
||||
|
||||
The `win2019-1` denial above was a **false negative**, and it is worth
|
||||
understanding before relying on this hook in an estate with cloud-protected
|
||||
workloads.
|
||||
This hook used to deny every change to a cloud-protected VM.
|
||||
|
||||
`wait_for_tag` gives up after a hardcoded 45s. That is generous for a
|
||||
`wait_for_tag` gave up after a hardcoded 45s. That is generous on a
|
||||
vSphere-protected VPG, which checkpoints every 5s and surfaces a tag in about
|
||||
4s. It is far too short elsewhere: a tag takes ~34s to appear on an
|
||||
Azure-protected VPG and ~128s on an AWS-protected one, because journal cadence
|
||||
is set by the protected site (5s vSphere, 60s Azure, 630s AWS).
|
||||
4s, and impossible on an AWS-protected one, where a tag takes ~128s because the
|
||||
journal only checkpoints every 630s.
|
||||
|
||||
So the hook denied the change, and the checkpoint landed anyway. It is in the
|
||||
journal as `cp 56`, timestamped a minute after the hook reported failure. The
|
||||
guard told the agent there was no rewind point while Zerto was in the middle of
|
||||
creating one.
|
||||
So the guard reported "no checkpoint, refusing the change" while Zerto was in the
|
||||
middle of creating one. The checkpoint landed a minute later, in the journal,
|
||||
after the agent had already been told there was no rewind point.
|
||||
|
||||
That failure mode is worse than the one the hook guards against, because it is
|
||||
silent and looks correct: legitimate work is refused on every cloud-protected VM
|
||||
while the log reads like the guard is doing its job.
|
||||
That is a worse failure than the one this hook exists to prevent. It is silent,
|
||||
it looks correct in the log, and it blocks legitimate work on every
|
||||
cloud-protected VM in the estate.
|
||||
|
||||
Until `wait_for_tag` becomes cadence-aware, scope this hook's matcher to
|
||||
vSphere-protected workloads.
|
||||
Both budgets are now derived from the VPG's measured cadence rather than
|
||||
guessed, which is why the numbers above differ by a factor of fifteen between
|
||||
platforms.
|
||||
|
||||
@@ -8,7 +8,7 @@
|
||||
{
|
||||
"type": "command",
|
||||
"command": "/home/you/zerto-ai-rewind/.venv/bin/python /home/you/zerto-ai-rewind/hooks/zerto_guard_hook.py",
|
||||
"timeout": 180
|
||||
"timeout": 360
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
@@ -24,7 +24,7 @@ tagged checkpoint and waits for the Zerto task to reach Completed.
|
||||
"matcher": "mcp__.*",
|
||||
"hooks": [{"type": "command",
|
||||
"command": "/path/to/.venv/bin/python /path/to/hooks/zerto_guard_hook.py",
|
||||
"timeout": 180}]}]}}
|
||||
"timeout": 360}]}]}}
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -40,7 +40,11 @@ LOG = os.environ.get("ZERTO_HOOK_LOG", os.path.expanduser("~/.zerto-guard-hook.l
|
||||
# Seconds the hook will wait for the checkpoint. Must stay under the hook
|
||||
# timeout configured in settings.json, or the host cancels us and the tool
|
||||
# call proceeds unguarded through the normal permission flow.
|
||||
GUARD_TIMEOUT_S = float(os.environ.get("ZERTO_HOOK_GUARD_TIMEOUT", "150"))
|
||||
# Must exceed the largest tag wait the guard can take. That is now derived from
|
||||
# the VPG's checkpoint cadence and capped at 300s (MAX_TAG_TIMEOUT_S), because a
|
||||
# tag takes ~128s to surface on an AWS-protected VPG. Too small a budget here
|
||||
# just moves the false denial from the guard into the hook.
|
||||
GUARD_TIMEOUT_S = float(os.environ.get("ZERTO_HOOK_GUARD_TIMEOUT", "330"))
|
||||
UNKNOWN_DECISION = os.environ.get("ZERTO_HOOK_UNKNOWN", "prompt") # prompt | allow | deny
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user