tag_vpgs now waits for that task and refuses unless it Completed, reporting task_id and task_state.
This corrects a comment I shipped in #2. I'd written that back-to-back inserts are "silently dropped." They aren't — measured on the live ZVM:
seqA: Completed
seqB: ZertoError -> Zerto task 05b202ad-… (InsertTaggedCP) finished as Failed. The operation did not happen.
It was never silent. We just never read the task. Before this change, a failed insert surfaced only as wait_for_tag timing out 45s later with a misleading hint about Azure/AWS.
2. Unknown tools ask the human instead of passing through
The catalog is opt-in, so an unlisted tool ran unguarded. The structural problem: the mutating set is unbounded and grows with every MCP you install, while the read-only set is small and stable. A mutating-only list is permanently behind, and being behind fails open.
Adds read_only_tools and zerto_check_tool(server, tool, vm):
verdict
behaviour
read_only
run it, no checkpoint
mutating
guard first, don't ask
unknown
ask the human
unknown does not mean safe — it means nobody classified it. The tool hands the agent a question to put to the human:
mystery/do_thing is not a known read-only command. It may change jp-ubuntu. Insert a Zerto tagged checkpoint first so this is reversible?
On yes → guard then run. On no → run and say plainly Zerto can't rewind it. If they want it remembered → zerto_add_mutating_tool, so it's guarded automatically next time instead of asking again.
This stays advisory. An MCP server cannot see or block another server's tool calls, so it is not an enforcement point. Real enforcement belongs in a host PreToolUse hook that calls the guard and denies on failure. This PR gives that hook something correct to call, and the skill carries the flow meanwhile.
pytest: 41 passed (12 new, covering task id/state parsing in both nested and bare shapes, Completed/Failed/timeout paths, and three-way classification).
Known issue, deliberately not fixed here
Two concurrent tool calls race on Keycloak token acquisition in the shared client and one comes back HTTP 401. I hit this while testing and it's unrelated to the task work — worth its own fix.
Stacked on #5 (base `feat/windows-mutating-catalog`) — merge #5 first, or I'll retarget to `main`.
Two changes, both from the same root mistake: assuming an answer instead of reading one.
## 1. Read the Zerto task after inserting a checkpoint
`POST /v1/vpgs/{id}/checkpoints` returns a **task id**, not a result. A 200 means *queued*. The outcome is in `GET /v1/tasks/{id}` under `Status.State`:
```
1 InProgress 4 Failed 5 Stopped 6 Completed (terminal = 4/5/6)
```
`tag_vpgs` now waits for that task and refuses unless it Completed, reporting `task_id` and `task_state`.
**This corrects a comment I shipped in #2.** I'd written that back-to-back inserts are "silently dropped." They aren't — measured on the live ZVM:
```
seqA: Completed
seqB: ZertoError -> Zerto task 05b202ad-… (InsertTaggedCP) finished as Failed. The operation did not happen.
```
It was never silent. We just never read the task. Before this change, a failed insert surfaced only as `wait_for_tag` timing out 45s later with a misleading hint about Azure/AWS.
## 2. Unknown tools ask the human instead of passing through
The catalog is opt-in, so an unlisted tool ran unguarded. The structural problem: **the mutating set is unbounded** and grows with every MCP you install, while **the read-only set is small and stable**. A mutating-only list is permanently behind, and being behind fails *open*.
Adds `read_only_tools` and `zerto_check_tool(server, tool, vm)`:
| verdict | behaviour |
|---|---|
| `read_only` | run it, no checkpoint |
| `mutating` | guard first, don't ask |
| `unknown` | **ask the human** |
`unknown` does not mean safe — it means nobody classified it. The tool hands the agent a question to put to the human:
> `mystery/do_thing` is not a known read-only command. It may change `jp-ubuntu`. Insert a Zerto tagged checkpoint first so this is reversible?
On yes → guard then run. On no → run and say plainly Zerto can't rewind it. If they want it remembered → `zerto_add_mutating_tool`, so it's guarded automatically next time instead of asking again.
**This stays advisory.** An MCP server cannot see or block another server's tool calls, so it is not an enforcement point. Real enforcement belongs in a host `PreToolUse` hook that calls the guard and denies on failure. This PR gives that hook something correct to call, and the skill carries the flow meanwhile.
## Verified live (13 tools)
```
ssh/read_file -> read_only (checkpoint_required=False)
winrm/run_ps -> mutating (checkpoint_required=True)
mystery/do_thing -> unknown (ask_user=True, question supplied)
guard -> ok=true cp=2354 task_state="Completed"
```
`pytest`: 41 passed (12 new, covering task id/state parsing in both nested and bare shapes, Completed/Failed/timeout paths, and three-way classification).
## Known issue, deliberately not fixed here
Two **concurrent** tool calls race on Keycloak token acquisition in the shared client and one comes back HTTP 401. I hit this while testing and it's unrelated to the task work — worth its own fix.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
claude
changed target branch from feat/windows-mutating-catalog to main2026-09-21 15:06:42 -04:00
Two changes that both come from the same mistake: assuming an answer
instead of reading one.
1. Read the task after inserting a checkpoint.
POST /v1/vpgs/{id}/checkpoints returns a TASK ID, not a result. A 200
only means queued. The outcome lives in GET /v1/tasks/{id} under
Status.State: 1 InProgress, 4 Failed, 5 Stopped, 6 Completed, with
4/5/6 terminal.
tag_vpgs now waits for that task and refuses unless it Completed, and
reports task_id and task_state. Measured: two inserts fired back to
back at one VPG give Completed for the first and Failed for the
second. That is exactly the case an earlier comment in this file
called "silently dropped" -- it was never silent, we just never read
the task. Comment corrected.
Previously a failed insert surfaced only as wait_for_tag timing out
45s later with a misleading hint about Azure/AWS. Now it says the
task failed and the operation did not happen.
2. Unknown tools ask the human instead of passing through.
The catalog is opt-in, so an unlisted tool ran unguarded. But the set
of mutating tools is unbounded and grows with every MCP installed,
while the set of read-only ones is small, so a mutating-only list is
permanently behind and being behind fails open.
Adds read_only_tools and zerto_check_tool(server, tool, vm) returning
read_only / mutating / unknown. unknown does not mean safe: it means
nobody classified it, so the tool hands the agent a question to put
to the human, and the human decides whether to checkpoint. On yes the
agent guards; on no it runs and says plainly that Zerto cannot rewind
it; if they want it remembered, zerto_add_mutating_tool.
This stays advisory. An MCP server cannot see or block another
server's tool calls, so real enforcement belongs in a host PreToolUse
hook. The skill carries the flow.
Known issue, not addressed here: two concurrent tool calls race on
Keycloak token acquisition in the shared client and one gets HTTP 401.
pytest 41 passed (12 new).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Stacked on #5 (base
feat/windows-mutating-catalog) — merge #5 first, or I'll retarget tomain.Two changes, both from the same root mistake: assuming an answer instead of reading one.
1. Read the Zerto task after inserting a checkpoint
POST /v1/vpgs/{id}/checkpointsreturns a task id, not a result. A 200 means queued. The outcome is inGET /v1/tasks/{id}underStatus.State:tag_vpgsnow waits for that task and refuses unless it Completed, reportingtask_idandtask_state.This corrects a comment I shipped in #2. I'd written that back-to-back inserts are "silently dropped." They aren't — measured on the live ZVM:
It was never silent. We just never read the task. Before this change, a failed insert surfaced only as
wait_for_tagtiming out 45s later with a misleading hint about Azure/AWS.2. Unknown tools ask the human instead of passing through
The catalog is opt-in, so an unlisted tool ran unguarded. The structural problem: the mutating set is unbounded and grows with every MCP you install, while the read-only set is small and stable. A mutating-only list is permanently behind, and being behind fails open.
Adds
read_only_toolsandzerto_check_tool(server, tool, vm):read_onlymutatingunknownunknowndoes not mean safe — it means nobody classified it. The tool hands the agent a question to put to the human:On yes → guard then run. On no → run and say plainly Zerto can't rewind it. If they want it remembered →
zerto_add_mutating_tool, so it's guarded automatically next time instead of asking again.This stays advisory. An MCP server cannot see or block another server's tool calls, so it is not an enforcement point. Real enforcement belongs in a host
PreToolUsehook that calls the guard and denies on failure. This PR gives that hook something correct to call, and the skill carries the flow meanwhile.Verified live (13 tools)
pytest: 41 passed (12 new, covering task id/state parsing in both nested and bare shapes, Completed/Failed/timeout paths, and three-way classification).Known issue, deliberately not fixed here
Two concurrent tool calls race on Keycloak token acquisition in the shared client and one comes back HTTP 401. I hit this while testing and it's unrelated to the task work — worth its own fix.
🤖 Generated with Claude Code
https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
Two changes that both come from the same mistake: assuming an answer instead of reading one. 1. Read the task after inserting a checkpoint. POST /v1/vpgs/{id}/checkpoints returns a TASK ID, not a result. A 200 only means queued. The outcome lives in GET /v1/tasks/{id} under Status.State: 1 InProgress, 4 Failed, 5 Stopped, 6 Completed, with 4/5/6 terminal. tag_vpgs now waits for that task and refuses unless it Completed, and reports task_id and task_state. Measured: two inserts fired back to back at one VPG give Completed for the first and Failed for the second. That is exactly the case an earlier comment in this file called "silently dropped" -- it was never silent, we just never read the task. Comment corrected. Previously a failed insert surfaced only as wait_for_tag timing out 45s later with a misleading hint about Azure/AWS. Now it says the task failed and the operation did not happen. 2. Unknown tools ask the human instead of passing through. The catalog is opt-in, so an unlisted tool ran unguarded. But the set of mutating tools is unbounded and grows with every MCP installed, while the set of read-only ones is small, so a mutating-only list is permanently behind and being behind fails open. Adds read_only_tools and zerto_check_tool(server, tool, vm) returning read_only / mutating / unknown. unknown does not mean safe: it means nobody classified it, so the tool hands the agent a question to put to the human, and the human decides whether to checkpoint. On yes the agent guards; on no it runs and says plainly that Zerto cannot rewind it; if they want it remembered, zerto_add_mutating_tool. This stays advisory. An MCP server cannot see or block another server's tool calls, so real enforcement belongs in a host PreToolUse hook. The skill carries the flow. Known issue, not addressed here: two concurrent tool calls race on Keycloak token acquisition in the shared client and one gets HTTP 401. pytest 41 passed (12 new). Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGnbad5a5706ftof1bb2343b0