docs: tagged checkpoints do work on cloud-protected VPGs
The repo said, in four places, that a tagged checkpoint cannot be inserted when the protected site is Azure or AWS. That came from the 9.0 API reference. It is wrong on 10.9.10. Measured against the lab: both an AWS-protected and an Azure-protected VPG accepted the insert, the Zerto task reached Completed, and the checkpoint appeared in the journal. This is the second doc claim this repo carried that live 10.9 contradicts, after the VPG status enum. What actually differs is latency and granularity, and both are set by the PROTECTED site, not the recovery site: protected at journal gap tag visible after vSphere 5s ~4s Azure 60s ~34s AWS 630s ~128s There is a sharper consequence than slowness. The tagged checkpoint is stamped about 30s AFTER the insert request, so on a cloud-protected VPG an agent that mutates promptly puts its change inside the checkpoint that was supposed to precede it. Recovering from that tag would restore the broken state. The docs now say to use the newest checkpoint that already existed when the guard ran. Two of the four were message text, not prose. wait_for_tag's timeout message asserted the insert was unsupported when the real cause was its own 45s budget being far too short for a VPG that checkpoints every 630s, so it told operators the wrong thing at exactly the wrong moment. It now says to check the Zerto task before concluding the insert failed. The 45s timeout itself is still wrong for cloud sources and needs to become cadence-aware. That is a behaviour change, so it is not in this commit. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
This commit is contained in:
@@ -64,7 +64,22 @@ Git never had the file. RPO is the journal, not last night's backup.
|
|||||||
|
|
||||||
vSphere ZVM 10.x and ZCA on AWS/Azure, same REST paths. HVM is out (separate swagger). Failover Live is not a tool.
|
vSphere ZVM 10.x and ZCA on AWS/Azure, same REST paths. HVM is out (separate swagger). Failover Live is not a tool.
|
||||||
|
|
||||||
Tagged checkpoints cannot be inserted when the **protected** site is Azure or AWS (Zerto API). Point this server at the vSphere protected ZVM.
|
Tagged checkpoints **do** work when the protected site is Azure or AWS. The 9.0 API
|
||||||
|
reference says they cannot be inserted; that is wrong on 10.9.10, where both were
|
||||||
|
accepted and the task reached `Completed`.
|
||||||
|
|
||||||
|
What differs is latency and granularity, both set by the **protected** site:
|
||||||
|
|
||||||
|
| protected at | journal gap | tag visible after |
|
||||||
|
|---|---|---|
|
||||||
|
| vSphere | 5s | ~4s |
|
||||||
|
| Azure | 60s | ~34s |
|
||||||
|
| AWS | 630s | ~128s |
|
||||||
|
|
||||||
|
The tagged checkpoint is also stamped about 30s *after* the insert request, so on a
|
||||||
|
cloud-protected VPG a prompt mutation can land *inside* the checkpoint meant to
|
||||||
|
precede it. Recover from the newest checkpoint that already existed when the guard
|
||||||
|
ran, not from the tag.
|
||||||
|
|
||||||
## Not this product
|
## Not this product
|
||||||
|
|
||||||
|
|||||||
@@ -70,7 +70,10 @@ so a stuck session blocks the next recovery. Find it with
|
|||||||
## Facts that bite
|
## Facts that bite
|
||||||
|
|
||||||
- A tagged checkpoint is crash-consistent, not app-quiesced.
|
- A tagged checkpoint is crash-consistent, not app-quiesced.
|
||||||
- Tagged checkpoints are not supported when the **protected** site is Azure or AWS. Talk to the vSphere protected ZVM.
|
- Tagged checkpoints work on Azure and AWS protected VPGs, but they appear late:
|
||||||
|
~34s (Azure) and ~128s (AWS) versus ~4s on vSphere, and the checkpoint is stamped
|
||||||
|
about 30s after you ask for it. On those VPGs the tag can end up *after* your
|
||||||
|
change, so treat the newest checkpoint that already existed as the rewind point.
|
||||||
- 10.9 FLR Operator RBAC fails; Administrator is the documented workaround.
|
- 10.9 FLR Operator RBAC fails; Administrator is the documented workaround.
|
||||||
- FLR cannot run during clone, test, live failover, or EJC.
|
- FLR cannot run during clone, test, live failover, or EJC.
|
||||||
- Linux FLR: files >1.5GB are a bad idea; some characters in names are refused.
|
- Linux FLR: files >1.5GB are a bad idea; some characters in names are refused.
|
||||||
|
|||||||
@@ -93,7 +93,10 @@ async def wait_for_tag(
|
|||||||
raise ZertoError(
|
raise ZertoError(
|
||||||
f"Tagged checkpoint {tag!r} did not appear on VPG {vpg_identifier} "
|
f"Tagged checkpoint {tag!r} did not appear on VPG {vpg_identifier} "
|
||||||
f"within {timeout_s:.0f}s. Do not mutate. "
|
f"within {timeout_s:.0f}s. Do not mutate. "
|
||||||
"If the protected site is Azure or AWS, tagged checkpoints are not supported."
|
"On a cloud-protected VPG the tag routinely takes longer than this to appear "
|
||||||
|
"(measured ~34s on Azure, ~128s on AWS), so this timeout may simply be too "
|
||||||
|
"short rather than the insert having failed. Check the Zerto task before "
|
||||||
|
"assuming it did not land."
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -137,8 +137,7 @@ def find_from_rows(query: str, rows: list[dict[str, Any]]) -> FindResult:
|
|||||||
outcome="none",
|
outcome="none",
|
||||||
query=query,
|
query=query,
|
||||||
message=(
|
message=(
|
||||||
f"VM {vm.vm_name} ({vm.vm_identifier}) has no VPG. "
|
f"VM {vm.vm_name} ({vm.vm_identifier}) has no VPG. Unprotected: refuse the change."
|
||||||
"Unprotected: refuse the change."
|
|
||||||
),
|
),
|
||||||
vm=vm,
|
vm=vm,
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -133,8 +133,9 @@ async def zerto_create_tagged_checkpoint(
|
|||||||
the Zerto API accepts, so this is the only place that context can live.
|
the Zerto API accepts, so this is the only place that context can live.
|
||||||
|
|
||||||
Name format: ai:<agent> | <action> | vm=<vm> | change=<change_id> | <utc>
|
Name format: ai:<agent> | <action> | vm=<vm> | change=<change_id> | <utc>
|
||||||
Docs: tagged checkpoints are not supported when the protected site is Azure or AWS;
|
Works on Azure and AWS protected VPGs despite what the 9.0 API reference says,
|
||||||
run this against the vSphere protected ZVM.
|
but the tag appears late there (~34s Azure, ~128s AWS) and is stamped after the
|
||||||
|
request, so it may sit after a prompt mutation.
|
||||||
"""
|
"""
|
||||||
try:
|
try:
|
||||||
result = await _find(query)
|
result = await _find(query)
|
||||||
|
|||||||
@@ -127,7 +127,6 @@ def can_tag(status: int | None, substatus: int | None) -> tuple[bool, str | None
|
|||||||
return False, f"VPG status is {status_name(status)}; not MeetingSLA"
|
return False, f"VPG status is {status_name(status)}; not MeetingSLA"
|
||||||
if substatus in SYNCING:
|
if substatus in SYNCING:
|
||||||
return False, (
|
return False, (
|
||||||
f"VPG is {substatus_name(substatus)}; "
|
f"VPG is {substatus_name(substatus)}; checkpoints are not durable until sync ends"
|
||||||
"checkpoints are not durable until sync ends"
|
|
||||||
)
|
)
|
||||||
return True, None
|
return True, None
|
||||||
|
|||||||
Reference in New Issue
Block a user