Improve write performance, avoid roundtrip per fragment

This commit is contained in:
Lucian Petrut
2026-09-07 15:18:21 +00:00
parent f18a8d3d71
commit 0039fdbb3d
3 changed files with 99 additions and 118 deletions
+46 -53
View File
@@ -3,7 +3,8 @@
This document records how `NfcDisk.write` in This document records how `NfcDisk.write` in
`openvixdisklib/nfc_open.py` implements `VixDiskLib_Write` over NFC AIO. `openvixdisklib/nfc_open.py` implements `VixDiskLib_Write` over NFC AIO.
The request layout matches the captured `VixDiskLib_Read` IO message in The request layout matches the captured `VixDiskLib_Read` IO message in
`docs/nfc_read.md`. Open flags and the IO direction field were taken `docs/nfc_read.md` (write requests use the same fragment fields as
read replies). Open flags and the IO direction field were taken
from a `strace` of VDDK 8 writing one sector to a temporary 10 GiB from a `strace` of VDDK 8 writing one sector to a temporary 10 GiB
disk (`docs/reverse_engineering_procedure.md`). disk (`docs/reverse_engineering_procedure.md`).
@@ -48,73 +49,65 @@ required to write or read sectors.
## Request (44 bytes + data) ## Request (44 bytes + data)
Little-endian, after the usual 16-byte AIO header Little-endian, after the usual 16-byte AIO header
(`magic 0xA100DA7A`, type 7, size 44, monotonic `opId`): (`magic 0xA100DA7A`, type 7, size 44, one `opId` per
`VixDiskLib_Write`):
| Offset | Type | `Write(start, n)` | | Offset | Type | `Write(start, n)` |
| ------ | -------- | -------------------------------------------------- | | ------ | -------- | --------------------------------------------------------- |
| 0 | `uint64` | File handle from `OPEN_FILE` | | 0 | `uint64` | File handle from `OPEN_FILE` |
| 8 | `uint64` | `0` (`NFC_AIO_IO_WRITE`; read uses `1`) | | 8 | `uint64` | `0` (`NFC_AIO_IO_WRITE`; read uses `1`) |
| 16 | `uint64` | Byte offset | | 16 | `uint64` | Byte offset of the **whole** write |
| 24 | `uint64` | Byte length | | 24 | `uint32` | Total byte length |
| 32 | `uint32` | Byte length (same value) | | 28 | `uint32` | Byte offset of this fragment (`0`, `65536`, …) |
| 36 | `uint32` | Byte length (same value) | | 32 | `uint32` | This fragment’s uncompressed length |
| 36 | `uint32` | Extra size (same as 32, or FastLZ packed size) |
| 40 | `uint32` | `0` | | 40 | `uint32` | `0` |
FASTLZ writes use the same 44-byte header. The opcode `uint64` high This is the same 44-byte layout as a **read reply** fragment
half is `2`, offset 36 is the compressed size, and FastLZ bytes follow (`docs/nfc_read.md`): writes stream request fragments, reads stream
instead of raw sectors. If compression does not shrink the chunk, VDDK reply fragments. A single-fragment write (≤ 64 KiB) still looks like a
sends type `0` and raw extra (same as an uncompressed write). `uint64` length at offset 24 because the fragment offset is 0.
FASTLZ writes use the same header. The opcode `uint64` high half is
`2`, offset 36 is the compressed size, and FastLZ bytes follow instead
of raw sectors. If compression does not shrink the fragment, VDDK
sends type `0` and raw extra. Each fragment is compressed on its own;
a 32 MiB FastLZ write is 512 independent FastLZ extras, not one.
Sector bytes follow the 44-byte payload and are **not** counted in AIO Sector bytes follow the 44-byte payload and are **not** counted in AIO
`size`. VDDK sends header + payload + data in one `write()`. The `size`. The replacement sends header + payload + extra in one
replacement does the same (`sendall` of those bytes together) and sets `sendall` and sets `TCP_NODELAY` on the NFC socket so a small FastLZ
`TCP_NODELAY` on the NFC socket so a small FastLZ extra is not delayed extra is not delayed behind Nagle / delayed ACK. Captured VDDK often
behind Nagle / delayed ACK. uses two `write()`s (`60` then `65536`) for a 64 KiB fragment and
coalesces only a 512-byte tail (`572` = 16 + 44 + 512).
The server replies with a type-7 header and a 44-byte payload for that
`opId`. There is no extra data on the write reply (unlike reads).
A 1-sector VDDK write was 572 bytes on the wire: 16 + 44 + 512. A 1-sector VDDK write was 572 bytes on the wire: 16 + 44 + 512.
## Client-side split ## Fragments and the single reply
`NfcAioInitSession` advertises a 64 KiB buffer and count 4. VDDK splits `NfcAioInitSession` advertises a 64 KiB buffer. Extra per type-7
writes larger than 64 KiB into 64 KiB chunks (VDDK programming guide) message is at most that size. VDDK does **not** issue a new `opId` per
and keeps several IOs in flight. The Python client does the same: IO chunk, and it does **not** coalesce separate `VixDiskLib_Write` calls
requests of at most `NFC_AIO_BUFFER_SIZE` bytes, up to (eight 8 KiB writes stayed eight IOs). One public write becomes N
`NFC_AIO_BUFFER_COUNT` (4) outstanding `opId`s before waiting for a client type-7 messages with the **same** `opId`, then **one** 44-byte
reply. reply (no extra) after the last fragment:
OPEN_SESSION is 16 zero bytes in both directions, so that count is a ```
VDDK client default (`vixDiskLib.nfcAio.Session.BufCount`), not a C: type=7 opId=14 size=44 total=66048 dest=0 chunk=65536 + 65536 data
server limit. Raising the client window on this lab (32 MiB writes, C: type=7 opId=14 size=44 total=66048 dest=65536 chunk=512 + 512 data
median of three samples) did not close the gap to VDDK: S: type=7 opId=14 size=44 total=66048 dest=0 chunk=66048
```
| Window | Plain write MiB/s | FastLZ write MiB/s | A 32 MiB write is 512 client fragments and one ACK. The Python client
| ------ | ----------------- | ------------------ | does the same. An earlier attempt that used a distinct `opId` per
| 4 | 13.1 | 16.5 | 64 KiB chunk and a sliding window of 4–512 outstanding IOs was waiting
| 16 | 11.3 (noisy) | 22.3 | for one reply per chunk; raising the window did not match VDDK
| 32 | 13.5 | 22.8 | throughput because VDDK pays one RTT per `Write`, not per fragment.
| 128 | 17.9 | 22.9 |
| 512 | 16.9 | — |
FastLZ flattens by window 16. Window 512 (send the whole 32 MiB before `NfcAioFlushCoalescedWrites` is server-side (`nfcAioServer.c`), not a
reading replies) was slower than 256. `SET_SOCK_OPTS` of 12 zero bytes client merge of API writes. OPEN_SESSION is 16 zero bytes both ways, so
returns send/recv sizes `1675000` and a `uint32` flag `1`; requesting the logged AIO buffer count of 4 is a VDDK client default
8 MiB buffers is echoed but did not help at window 128. The remaining (`vixDiskLib.nfcAio.Session.BufCount`), not a server cap.
VDDK FastLZ advantage (about 140–260 MiB/s vs ~23 MiB/s here) is not
the outstanding-IO count.
VDDK logs at `VixDiskLib_InitEx` spawn a Vmacore pool (`IO: 2`,
`Min workers: 4`, `Max workers: 13`) and NFC AIO uses a thread context
(`NfcAioInitThreadCtx`, “Schedule main processing from IO callback”).
Those are process-wide / async completion threads, not extra NFC
sockets or extra 64 KiB buffers. Sync `VixDiskLib_Write` can still
compress and SSL-write on different threads. That may help plain TLS
overlap; it does not explain most of the FastLZ gap (512 × FastLZ of
64 KiB is tens of milliseconds). `aiomgr.numThreads` and
`AsyncWriteImpl` workers are local disk AIO / on-disk compressed VMDKs,
not NBD.
## Python replacement ## Python replacement
+7 -7
View File
@@ -248,13 +248,13 @@ the ticket switched from `NfcGetVmFiles` to `NfcRandomAccessOpenDisk`
`NfcRandomAccessOpenDisk`). Integration tests create a temporary empty `NfcRandomAccessOpenDisk`). Integration tests create a temporary empty
10 GiB VM for the run so writes cannot land on other lab disks. 10 GiB VM for the run so writes cannot land on other lab disks.
The Python client splits writes larger than 64 KiB into AIO chunks and A write larger than 64 KiB is one AIO `opId` with several type-7
keeps four in flight (VDDK's `NfcAioInitSession` buffer count). Larger request fragments (same layout as read *replies*: total length, then
windows were tried; they do not match VDDK throughput. Header and extra fragment offset / length) and a single 44-byte ACK. Separate
go in one `sendall`, with `TCP_NODELAY`. Details: `VixDiskLib_Write` calls are not coalesced. Header and extra go in one
`docs/nfc_write.md`. Proof: write then read in `sendall`, with `TCP_NODELAY`. Details: `docs/nfc_write.md`. Proof:
`tests/integration/test_nfc_read_write.py` and the VDDK cross-check in write then read in `tests/integration/test_nfc_read_write.py` and the
`tests/integration/test_crosscheck.py`. VDDK cross-check in `tests/integration/test_crosscheck.py`.
## Step 11 — NBDSSL: second TLS after `PROXY vpxa-nfcssl` ## Step 11 — NBDSSL: second TLS after `PROXY vpxa-nfcssl`
+32 -44
View File
@@ -34,12 +34,9 @@ NFC_AIO_MAGIC = 0xA100DA7A
NFC_AIO_HDR_SIZE = 16 NFC_AIO_HDR_SIZE = 16
NFC_SECTOR_SIZE = 512 NFC_SECTOR_SIZE = 512
NFC_PROTOCOL_VERSION = 11 NFC_PROTOCOL_VERSION = 11
# Max data bytes in one AIO IO reply fragment (NfcAioInitSession buffer). # Max data bytes in one AIO IO request/reply fragment
# (NfcAioInitSession buffer).
NFC_AIO_BUFFER_SIZE = 65536 NFC_AIO_BUFFER_SIZE = 65536
# Outstanding write IOs kept in flight. Matches VDDK's logged
# ``NfcAioInitSession`` buffer count of 4. Larger depths were tried
# (see ``docs/nfc_write.md``) and did not close the VDDK throughput gap.
NFC_AIO_BUFFER_COUNT = 4
# Classic NFC message types observed on the wire (uint32 at offset 0). # Classic NFC message types observed on the wire (uint32 at offset 0).
NFC_MSG_SESSION_COMPLETE = 4 NFC_MSG_SESSION_COMPLETE = 4
@@ -341,11 +338,10 @@ class NfcDisk:
data: bytes) -> None: data: bytes) -> None:
"""Write ``num_sectors`` starting at ``start_sector``. """Write ``num_sectors`` starting at ``start_sector``.
Matches ``VixDiskLib_Write``: one ``NFC_AIO_MSG_IO`` request per Matches ``VixDiskLib_Write``: one ``NFC_AIO_MSG_IO`` ``opId``
chunk in byte units, with sector bytes sent after the 44-byte for the whole call. Chunks larger than the AIO buffer (64 KiB)
payload. Chunks larger than the AIO buffer (64 KiB) are split. are extra fragments with that same ``opId``; the server replies
Up to ``NFC_AIO_BUFFER_COUNT`` writes stay in flight. FASTLZ once. FASTLZ open compresses each fragment when that shrinks it.
open compresses each chunk when that shrinks it.
Args: Args:
start_sector: Sector offset from the start of the disk. start_sector: Sector offset from the start of the disk.
@@ -358,50 +354,42 @@ class NfcDisk:
if len(data) != length: if len(data) != length:
raise ValueError( raise ValueError(
f"write data is {len(data)} bytes, need {length}") f"write data is {len(data)} bytes, need {length}")
max_sectors = NFC_AIO_BUFFER_SIZE // self.sector_size disk_offset = start_sector * self.sector_size
offset_sectors = start_sector op_id = self._next_op_id()
remaining = data frag_offset = 0
pending: set[int] = set() while frag_offset < length:
while remaining or pending: chunk = data[frag_offset:frag_offset + NFC_AIO_BUFFER_SIZE]
while remaining and len(pending) < NFC_AIO_BUFFER_COUNT: extra = chunk
n_sectors = min( extra_len = len(chunk)
len(remaining) // self.sector_size, max_sectors)
chunk = remaining[:n_sectors * self.sector_size]
pending.add(self._send_write_chunk(offset_sectors, chunk))
offset_sectors += n_sectors
remaining = remaining[n_sectors * self.sector_size:]
if not pending:
break
rtype, rop, _body = self._aio_recv_reply()
if rtype != NFC_AIO_MSG_IO or rop not in pending:
raise NfcProtocolError(
f"AIO IO write reply type={rtype} opId={rop}, "
f"expected type={NFC_AIO_MSG_IO} opId in {pending}")
pending.remove(rop)
def _send_write_chunk(self, start_sector: int, data: bytes) -> int:
length = len(data)
offset = start_sector * self.sector_size
extra = data
ctype = NFC_COMPRESSION_NONE ctype = NFC_COMPRESSION_NONE
extra_len = length if (
if self.compression == NFC_COMPRESSION_FASTLZ and length >= 16: self.compression == NFC_COMPRESSION_FASTLZ
compressed = fastlz.compress(data) and extra_len >= 16):
if compressed and len(compressed) < length: compressed = fastlz.compress(chunk)
if compressed and len(compressed) < extra_len:
extra = compressed extra = compressed
ctype = NFC_COMPRESSION_FASTLZ
extra_len = len(compressed) extra_len = len(compressed)
ctype = NFC_COMPRESSION_FASTLZ
opcode = NFC_AIO_IO_WRITE | (ctype << 32) opcode = NFC_AIO_IO_WRITE | (ctype << 32)
payload = struct.pack( payload = struct.pack(
"<QQQQIII", "<QQQIIIII",
self.handle, self.handle,
opcode, opcode,
offset, disk_offset,
length,
length, length,
frag_offset,
len(chunk),
extra_len, extra_len,
0) 0)
return self._aio_send(NFC_AIO_MSG_IO, payload, extra) self._sock.sendall(
_pack_aio_hdr(NFC_AIO_MSG_IO, len(payload), op_id)
+ payload + extra)
frag_offset += len(chunk)
rtype, rop, _body = self._aio_recv_reply()
if rtype != NFC_AIO_MSG_IO or rop != op_id:
raise NfcProtocolError(
f"AIO IO write reply type={rtype} opId={rop}, "
f"expected type={NFC_AIO_MSG_IO} opId={op_id}")
def close(self) -> None: def close(self) -> None:
"""Close the VMDK, the AIO session, and the classic NFC session.""" """Close the VMDK, the AIO session, and the classic NFC session."""