feat(flr): windows paths, recovery-site gate, stable partition reads
Three fixes from measuring FLR against a Windows guest (ad1, VPG 'local')
instead of only the Linux one.
1. Gate FLR to locally replicated VPGs.
FLR is performed at the VPG's RECOVERY site, because that is where the
mount is created. A VPG replicating to a cloud ZCA has to be recovered
through that ZCA's API, not the protected ZVM's. Supporting that
properly means holding credentials for every ZVM/ZCA in an estate and
routing the call, which is a real design decision, not something to
smuggle in. Until then zerto_recover_file refuses a VPG whose
protected site != recovery site and names the site that owns the
operation, instead of failing later as a confusing path or mount error.
2. Windows paths are not symmetrical with Linux.
Linux /home/j/app.yaml -> Volume2-Ext4%2fhome%2fj%2fapp.yaml
Windows C:\Users\j\f.conf -> C%3a%2fUsers%2fj%2ff.conf
On Windows the drive letter IS the partition name, so the guest path
already carries it. The old code prepended unconditionally and built
'C:/C:/Users/j', which could never match. resolve_flr_path now detects
that the path already starts with the partition.
It also returns the raw path exactly as browse reported it. Download
accepts the raw and the decoded form, and returning raw avoids
re-encoding by hand. Decoding is now unquote_plus, not unquote: browse
form-encodes a space as '+' ("Program+Files"), so a basename compare
against "Program Files" never matched. Matching tries exact first and
only then case-insensitively, since Windows is case-insensitive and
Linux is not.
3. Wait for partition enumeration to settle.
A session reports mounted before the ZVM has finished identifying
volumes, and browsing in that window returns a partial, MIS-LABELLED
list. The same Windows VM enumerated as 'Volume4-Unknown' with no C:
drive, then moments later as a browsable 'C%3a' holding the whole
filesystem. Acting on the early list makes a restorable NTFS disk look
permanently unrestorable. wait_partitions_stable polls until the list
stops changing.
Verified against ZVM 10.x:
- CMH-AWS-4 (recovery aws-zca) refused, naming aws-zca
- C:\ad1.keytab recovered, 58 bytes
- C:\Program Files\internet explorer\sqmapi.dll recovered, 47512 bytes --
drive letter, two spaces, nested dirs, the exact case that was broken
- Linux /home/justin/app-config.yaml still recovers, 158 bytes
- every session unmounted, unmount.ok true
Also drops a duplicate get_vpg that shadowed the existing one (F811).
pytest 32 passed (7 new).
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
This commit is contained in:
@@ -260,6 +260,55 @@ async def zerto_add_mutating_tool(
|
||||
return _dump({"ok": True, "entry": entry.as_dict()})
|
||||
|
||||
|
||||
async def _site_name(client: ZertoClient, identifier: str | None) -> str:
|
||||
if not identifier:
|
||||
return "unknown site"
|
||||
try:
|
||||
local = await client.get_localsite()
|
||||
if str(local.get("SiteIdentifier")) == identifier:
|
||||
return str(local.get("SiteName") or identifier)
|
||||
for peer in await client.get_peersites():
|
||||
if str(peer.get("SiteIdentifier")) == identifier:
|
||||
return str(peer.get("PeerSiteName") or identifier)
|
||||
except ZertoError:
|
||||
pass
|
||||
return identifier
|
||||
|
||||
|
||||
async def _flr_site_gate(client: ZertoClient, vpg_identifier: str) -> dict[str, Any]:
|
||||
"""Refuse FLR on a VPG whose recovery site is not this ZVM.
|
||||
|
||||
FLR only exists at the VPG's RECOVERY site: the mount is created there. A
|
||||
VPG replicating to a cloud ZCA has to be recovered from that ZCA's API, not
|
||||
this one. Until this server can hold credentials for every ZVM/ZCA in an
|
||||
estate and route the call, restrict FLR to local replication (protected
|
||||
site == recovery site) and say plainly where the operation actually lives,
|
||||
rather than letting it fail as a confusing path or mount error.
|
||||
"""
|
||||
try:
|
||||
vpg = await client.get_vpg(vpg_identifier)
|
||||
except ZertoError as exc:
|
||||
return {"ok": False, "message": f"Could not read VPG {vpg_identifier}: {exc}"}
|
||||
protected = (vpg.get("ProtectedSite") or {}).get("identifier")
|
||||
recovery = (vpg.get("RecoverySite") or {}).get("identifier")
|
||||
if protected and recovery and str(protected) == str(recovery):
|
||||
return {"ok": True}
|
||||
where = await _site_name(client, str(recovery) if recovery else None)
|
||||
return {
|
||||
"ok": False,
|
||||
"not_local_replication": True,
|
||||
"vpg_name": vpg.get("VpgName"),
|
||||
"recovery_site": where,
|
||||
"message": (
|
||||
f"FLR for VPG {vpg.get('VpgName')!r} lives at its recovery site "
|
||||
f"({where}), not at this ZVM. This server only supports file "
|
||||
"recovery for locally replicated VPGs (protected site == recovery "
|
||||
"site). Point an MCP instance at that ZVM/ZCA, or use a bounded "
|
||||
"whole-VM operation instead."
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
async def _live_session_ids(client: ZertoClient) -> set[str]:
|
||||
try:
|
||||
rows = session_rows(await client.list_flrs())
|
||||
@@ -327,6 +376,10 @@ async def zerto_recover_file(
|
||||
|
||||
Requires confirmed=true (human yes). Cannot run during clone/test/live/EJC.
|
||||
10.9 FLR Operator role fails; use an Administrator account.
|
||||
|
||||
Locally replicated VPGs only. FLR is performed at the VPG's recovery site,
|
||||
so a VPG replicating to a cloud ZCA must be recovered from that ZCA's API.
|
||||
guest_path is the path on the guest: /home/x/f.conf or C:\\Users\\x\\f.txt.
|
||||
"""
|
||||
if not confirmed:
|
||||
return _dump(
|
||||
@@ -339,6 +392,9 @@ async def zerto_recover_file(
|
||||
}
|
||||
)
|
||||
client = get_client()
|
||||
gate = await _flr_site_gate(client, vpg_identifier)
|
||||
if not gate.get("ok"):
|
||||
return _dump(gate)
|
||||
dest = Path(dest_dir or _settings.get("recovery_dir") or "./recovered")
|
||||
dest.mkdir(parents=True, exist_ok=True)
|
||||
session_id: str | None = None
|
||||
|
||||
Reference in New Issue
Block a user