Document perf improvement trials

This commit is contained in:
Lucian Petrut
2026-09-07 14:11:24 +00:00
parent d0a6de03d2
commit b11e109ec5
4 changed files with 43 additions and 7 deletions
+4 -3
View File
@@ -158,9 +158,10 @@ obtain a file handle or to read sector 0.
### OPEN_SESSION / sockopts / resource pool
VDDK sends 16 zero bytes (`OPEN_SESSION`), 12 zero bytes
(`SET_SOCK_OPTS`; server returns send/recv buffer sizes), then
`uint32` 1 (`SET_RES_POOL`, log: “Setting Resource Pool(1)”).
VDDK sends 16 zero bytes (`OPEN_SESSION`; server replies with 16 zeros),
12 zero bytes (`SET_SOCK_OPTS`; server returns send/recv buffer sizes
and a `uint32` flag), then `uint32` 1 (`SET_RES_POOL`, log: “Setting
Resource Pool(1)”).
### OPEN_FILE
+33 -1
View File
@@ -82,7 +82,39 @@ A 1-sector VDDK write was 572 bytes on the wire: 16 + 44 + 512.
writes larger than 64 KiB into 64 KiB chunks (VDDK programming guide)
and keeps several IOs in flight. The Python client does the same: IO
requests of at most `NFC_AIO_BUFFER_SIZE` bytes, up to
`NFC_AIO_BUFFER_COUNT` outstanding `opId`s before waiting for a reply.
`NFC_AIO_BUFFER_COUNT` (4) outstanding `opId`s before waiting for a
reply.
OPEN_SESSION is 16 zero bytes in both directions, so that count is a
VDDK client default (`vixDiskLib.nfcAio.Session.BufCount`), not a
server limit. Raising the client window on this lab (32 MiB writes,
median of three samples) did not close the gap to VDDK:
| Window | Plain write MiB/s | FastLZ write MiB/s |
| ------ | ----------------- | ------------------ |
| 4 | 13.1 | 16.5 |
| 16 | 11.3 (noisy) | 22.3 |
| 32 | 13.5 | 22.8 |
| 128 | 17.9 | 22.9 |
| 512 | 16.9 | — |
FastLZ flattens by window 16. Window 512 (send the whole 32 MiB before
reading replies) was slower than 256. `SET_SOCK_OPTS` of 12 zero bytes
returns send/recv sizes `1675000` and a `uint32` flag `1`; requesting
8 MiB buffers is echoed but did not help at window 128. The remaining
VDDK FastLZ advantage (about 140–260 MiB/s vs ~23 MiB/s here) is not
the outstanding-IO count.
VDDK logs at `VixDiskLib_InitEx` spawn a Vmacore pool (`IO: 2`,
`Min workers: 4`, `Max workers: 13`) and NFC AIO uses a thread context
(`NfcAioInitThreadCtx`, “Schedule main processing from IO callback”).
Those are process-wide / async completion threads, not extra NFC
sockets or extra 64 KiB buffers. Sync `VixDiskLib_Write` can still
compress and SSL-write on different threads. That may help plain TLS
overlap; it does not explain most of the FastLZ gap (512 × FastLZ of
64 KiB is tens of milliseconds). `aiomgr.numThreads` and
`AsyncWriteImpl` workers are local disk AIO / on-disk compressed VMDKs,
not NBD.
## Python replacement
+3 -2
View File
@@ -249,8 +249,9 @@ the ticket switched from `NfcGetVmFiles` to `NfcRandomAccessOpenDisk`
10 GiB VM for the run so writes cannot land on other lab disks.
The Python client splits writes larger than 64 KiB into AIO chunks and
keeps up to four in flight (`NfcAioInitSession` buffer count). Header
and extra go in one `sendall`, with `TCP_NODELAY`. Details:
keeps four in flight (VDDK's `NfcAioInitSession` buffer count). Larger
windows were tried; they do not match VDDK throughput. Header and extra
go in one `sendall`, with `TCP_NODELAY`. Details:
`docs/nfc_write.md`. Proof: write then read in
`tests/integration/test_nfc_read_write.py` and the VDDK cross-check in
`tests/integration/test_crosscheck.py`.
+3 -1
View File
@@ -36,7 +36,9 @@ NFC_SECTOR_SIZE = 512
NFC_PROTOCOL_VERSION = 11
# Max data bytes in one AIO IO reply fragment (NfcAioInitSession buffer).
NFC_AIO_BUFFER_SIZE = 65536
# Outstanding write IOs VDDK keeps in flight (NfcAioInitSession count 4).
# Outstanding write IOs kept in flight. Matches VDDK's logged
# ``NfcAioInitSession`` buffer count of 4. Larger depths were tried
# (see ``docs/nfc_write.md``) and did not close the VDDK throughput gap.
NFC_AIO_BUFFER_COUNT = 4
# Classic NFC message types observed on the wire (uint32 at offset 0).