Improve write performance, avoid roundtrip per fragment
This commit is contained in:
+46
-53
@@ -3,7 +3,8 @@
|
|||||||
This document records how `NfcDisk.write` in
|
This document records how `NfcDisk.write` in
|
||||||
`openvixdisklib/nfc_open.py` implements `VixDiskLib_Write` over NFC AIO.
|
`openvixdisklib/nfc_open.py` implements `VixDiskLib_Write` over NFC AIO.
|
||||||
The request layout matches the captured `VixDiskLib_Read` IO message in
|
The request layout matches the captured `VixDiskLib_Read` IO message in
|
||||||
`docs/nfc_read.md`. Open flags and the IO direction field were taken
|
`docs/nfc_read.md` (write requests use the same fragment fields as
|
||||||
|
read replies). Open flags and the IO direction field were taken
|
||||||
from a `strace` of VDDK 8 writing one sector to a temporary 10 GiB
|
from a `strace` of VDDK 8 writing one sector to a temporary 10 GiB
|
||||||
disk (`docs/reverse_engineering_procedure.md`).
|
disk (`docs/reverse_engineering_procedure.md`).
|
||||||
|
|
||||||
@@ -48,73 +49,65 @@ required to write or read sectors.
|
|||||||
## Request (44 bytes + data)
|
## Request (44 bytes + data)
|
||||||
|
|
||||||
Little-endian, after the usual 16-byte AIO header
|
Little-endian, after the usual 16-byte AIO header
|
||||||
(`magic 0xA100DA7A`, type 7, size 44, monotonic `opId`):
|
(`magic 0xA100DA7A`, type 7, size 44, one `opId` per
|
||||||
|
`VixDiskLib_Write`):
|
||||||
|
|
||||||
| Offset | Type | `Write(start, n)` |
|
| Offset | Type | `Write(start, n)` |
|
||||||
| ------ | -------- | -------------------------------------------------- |
|
| ------ | -------- | --------------------------------------------------------- |
|
||||||
| 0 | `uint64` | File handle from `OPEN_FILE` |
|
| 0 | `uint64` | File handle from `OPEN_FILE` |
|
||||||
| 8 | `uint64` | `0` (`NFC_AIO_IO_WRITE`; read uses `1`) |
|
| 8 | `uint64` | `0` (`NFC_AIO_IO_WRITE`; read uses `1`) |
|
||||||
| 16 | `uint64` | Byte offset |
|
| 16 | `uint64` | Byte offset of the **whole** write |
|
||||||
| 24 | `uint64` | Byte length |
|
| 24 | `uint32` | Total byte length |
|
||||||
| 32 | `uint32` | Byte length (same value) |
|
| 28 | `uint32` | Byte offset of this fragment (`0`, `65536`, …) |
|
||||||
| 36 | `uint32` | Byte length (same value) |
|
| 32 | `uint32` | This fragment’s uncompressed length |
|
||||||
|
| 36 | `uint32` | Extra size (same as 32, or FastLZ packed size) |
|
||||||
| 40 | `uint32` | `0` |
|
| 40 | `uint32` | `0` |
|
||||||
|
|
||||||
FASTLZ writes use the same 44-byte header. The opcode `uint64` high
|
This is the same 44-byte layout as a **read reply** fragment
|
||||||
half is `2`, offset 36 is the compressed size, and FastLZ bytes follow
|
(`docs/nfc_read.md`): writes stream request fragments, reads stream
|
||||||
instead of raw sectors. If compression does not shrink the chunk, VDDK
|
reply fragments. A single-fragment write (≤ 64 KiB) still looks like a
|
||||||
sends type `0` and raw extra (same as an uncompressed write).
|
`uint64` length at offset 24 because the fragment offset is 0.
|
||||||
|
|
||||||
|
FASTLZ writes use the same header. The opcode `uint64` high half is
|
||||||
|
`2`, offset 36 is the compressed size, and FastLZ bytes follow instead
|
||||||
|
of raw sectors. If compression does not shrink the fragment, VDDK
|
||||||
|
sends type `0` and raw extra. Each fragment is compressed on its own;
|
||||||
|
a 32 MiB FastLZ write is 512 independent FastLZ extras, not one.
|
||||||
|
|
||||||
Sector bytes follow the 44-byte payload and are **not** counted in AIO
|
Sector bytes follow the 44-byte payload and are **not** counted in AIO
|
||||||
`size`. VDDK sends header + payload + data in one `write()`. The
|
`size`. The replacement sends header + payload + extra in one
|
||||||
replacement does the same (`sendall` of those bytes together) and sets
|
`sendall` and sets `TCP_NODELAY` on the NFC socket so a small FastLZ
|
||||||
`TCP_NODELAY` on the NFC socket so a small FastLZ extra is not delayed
|
extra is not delayed behind Nagle / delayed ACK. Captured VDDK often
|
||||||
behind Nagle / delayed ACK.
|
uses two `write()`s (`60` then `65536`) for a 64 KiB fragment and
|
||||||
|
coalesces only a 512-byte tail (`572` = 16 + 44 + 512).
|
||||||
The server replies with a type-7 header and a 44-byte payload for that
|
|
||||||
`opId`. There is no extra data on the write reply (unlike reads).
|
|
||||||
|
|
||||||
A 1-sector VDDK write was 572 bytes on the wire: 16 + 44 + 512.
|
A 1-sector VDDK write was 572 bytes on the wire: 16 + 44 + 512.
|
||||||
|
|
||||||
## Client-side split
|
## Fragments and the single reply
|
||||||
|
|
||||||
`NfcAioInitSession` advertises a 64 KiB buffer and count 4. VDDK splits
|
`NfcAioInitSession` advertises a 64 KiB buffer. Extra per type-7
|
||||||
writes larger than 64 KiB into 64 KiB chunks (VDDK programming guide)
|
message is at most that size. VDDK does **not** issue a new `opId` per
|
||||||
and keeps several IOs in flight. The Python client does the same: IO
|
chunk, and it does **not** coalesce separate `VixDiskLib_Write` calls
|
||||||
requests of at most `NFC_AIO_BUFFER_SIZE` bytes, up to
|
(eight 8 KiB writes stayed eight IOs). One public write becomes N
|
||||||
`NFC_AIO_BUFFER_COUNT` (4) outstanding `opId`s before waiting for a
|
client type-7 messages with the **same** `opId`, then **one** 44-byte
|
||||||
reply.
|
reply (no extra) after the last fragment:
|
||||||
|
|
||||||
OPEN_SESSION is 16 zero bytes in both directions, so that count is a
|
```
|
||||||
VDDK client default (`vixDiskLib.nfcAio.Session.BufCount`), not a
|
C: type=7 opId=14 size=44 total=66048 dest=0 chunk=65536 + 65536 data
|
||||||
server limit. Raising the client window on this lab (32 MiB writes,
|
C: type=7 opId=14 size=44 total=66048 dest=65536 chunk=512 + 512 data
|
||||||
median of three samples) did not close the gap to VDDK:
|
S: type=7 opId=14 size=44 total=66048 dest=0 chunk=66048
|
||||||
|
```
|
||||||
|
|
||||||
| Window | Plain write MiB/s | FastLZ write MiB/s |
|
A 32 MiB write is 512 client fragments and one ACK. The Python client
|
||||||
| ------ | ----------------- | ------------------ |
|
does the same. An earlier attempt that used a distinct `opId` per
|
||||||
| 4 | 13.1 | 16.5 |
|
64 KiB chunk and a sliding window of 4–512 outstanding IOs was waiting
|
||||||
| 16 | 11.3 (noisy) | 22.3 |
|
for one reply per chunk; raising the window did not match VDDK
|
||||||
| 32 | 13.5 | 22.8 |
|
throughput because VDDK pays one RTT per `Write`, not per fragment.
|
||||||
| 128 | 17.9 | 22.9 |
|
|
||||||
| 512 | 16.9 | — |
|
|
||||||
|
|
||||||
FastLZ flattens by window 16. Window 512 (send the whole 32 MiB before
|
`NfcAioFlushCoalescedWrites` is server-side (`nfcAioServer.c`), not a
|
||||||
reading replies) was slower than 256. `SET_SOCK_OPTS` of 12 zero bytes
|
client merge of API writes. OPEN_SESSION is 16 zero bytes both ways, so
|
||||||
returns send/recv sizes `1675000` and a `uint32` flag `1`; requesting
|
the logged AIO buffer count of 4 is a VDDK client default
|
||||||
8 MiB buffers is echoed but did not help at window 128. The remaining
|
(`vixDiskLib.nfcAio.Session.BufCount`), not a server cap.
|
||||||
VDDK FastLZ advantage (about 140–260 MiB/s vs ~23 MiB/s here) is not
|
|
||||||
the outstanding-IO count.
|
|
||||||
|
|
||||||
VDDK logs at `VixDiskLib_InitEx` spawn a Vmacore pool (`IO: 2`,
|
|
||||||
`Min workers: 4`, `Max workers: 13`) and NFC AIO uses a thread context
|
|
||||||
(`NfcAioInitThreadCtx`, “Schedule main processing from IO callback”).
|
|
||||||
Those are process-wide / async completion threads, not extra NFC
|
|
||||||
sockets or extra 64 KiB buffers. Sync `VixDiskLib_Write` can still
|
|
||||||
compress and SSL-write on different threads. That may help plain TLS
|
|
||||||
overlap; it does not explain most of the FastLZ gap (512 × FastLZ of
|
|
||||||
64 KiB is tens of milliseconds). `aiomgr.numThreads` and
|
|
||||||
`AsyncWriteImpl` workers are local disk AIO / on-disk compressed VMDKs,
|
|
||||||
not NBD.
|
|
||||||
|
|
||||||
## Python replacement
|
## Python replacement
|
||||||
|
|
||||||
|
|||||||
@@ -248,13 +248,13 @@ the ticket switched from `NfcGetVmFiles` to `NfcRandomAccessOpenDisk`
|
|||||||
`NfcRandomAccessOpenDisk`). Integration tests create a temporary empty
|
`NfcRandomAccessOpenDisk`). Integration tests create a temporary empty
|
||||||
10 GiB VM for the run so writes cannot land on other lab disks.
|
10 GiB VM for the run so writes cannot land on other lab disks.
|
||||||
|
|
||||||
The Python client splits writes larger than 64 KiB into AIO chunks and
|
A write larger than 64 KiB is one AIO `opId` with several type-7
|
||||||
keeps four in flight (VDDK's `NfcAioInitSession` buffer count). Larger
|
request fragments (same layout as read *replies*: total length, then
|
||||||
windows were tried; they do not match VDDK throughput. Header and extra
|
fragment offset / length) and a single 44-byte ACK. Separate
|
||||||
go in one `sendall`, with `TCP_NODELAY`. Details:
|
`VixDiskLib_Write` calls are not coalesced. Header and extra go in one
|
||||||
`docs/nfc_write.md`. Proof: write then read in
|
`sendall`, with `TCP_NODELAY`. Details: `docs/nfc_write.md`. Proof:
|
||||||
`tests/integration/test_nfc_read_write.py` and the VDDK cross-check in
|
write then read in `tests/integration/test_nfc_read_write.py` and the
|
||||||
`tests/integration/test_crosscheck.py`.
|
VDDK cross-check in `tests/integration/test_crosscheck.py`.
|
||||||
|
|
||||||
## Step 11 — NBDSSL: second TLS after `PROXY vpxa-nfcssl`
|
## Step 11 — NBDSSL: second TLS after `PROXY vpxa-nfcssl`
|
||||||
|
|
||||||
|
|||||||
+32
-44
@@ -34,12 +34,9 @@ NFC_AIO_MAGIC = 0xA100DA7A
|
|||||||
NFC_AIO_HDR_SIZE = 16
|
NFC_AIO_HDR_SIZE = 16
|
||||||
NFC_SECTOR_SIZE = 512
|
NFC_SECTOR_SIZE = 512
|
||||||
NFC_PROTOCOL_VERSION = 11
|
NFC_PROTOCOL_VERSION = 11
|
||||||
# Max data bytes in one AIO IO reply fragment (NfcAioInitSession buffer).
|
# Max data bytes in one AIO IO request/reply fragment
|
||||||
|
# (NfcAioInitSession buffer).
|
||||||
NFC_AIO_BUFFER_SIZE = 65536
|
NFC_AIO_BUFFER_SIZE = 65536
|
||||||
# Outstanding write IOs kept in flight. Matches VDDK's logged
|
|
||||||
# ``NfcAioInitSession`` buffer count of 4. Larger depths were tried
|
|
||||||
# (see ``docs/nfc_write.md``) and did not close the VDDK throughput gap.
|
|
||||||
NFC_AIO_BUFFER_COUNT = 4
|
|
||||||
|
|
||||||
# Classic NFC message types observed on the wire (uint32 at offset 0).
|
# Classic NFC message types observed on the wire (uint32 at offset 0).
|
||||||
NFC_MSG_SESSION_COMPLETE = 4
|
NFC_MSG_SESSION_COMPLETE = 4
|
||||||
@@ -341,11 +338,10 @@ class NfcDisk:
|
|||||||
data: bytes) -> None:
|
data: bytes) -> None:
|
||||||
"""Write ``num_sectors`` starting at ``start_sector``.
|
"""Write ``num_sectors`` starting at ``start_sector``.
|
||||||
|
|
||||||
Matches ``VixDiskLib_Write``: one ``NFC_AIO_MSG_IO`` request per
|
Matches ``VixDiskLib_Write``: one ``NFC_AIO_MSG_IO`` ``opId``
|
||||||
chunk in byte units, with sector bytes sent after the 44-byte
|
for the whole call. Chunks larger than the AIO buffer (64 KiB)
|
||||||
payload. Chunks larger than the AIO buffer (64 KiB) are split.
|
are extra fragments with that same ``opId``; the server replies
|
||||||
Up to ``NFC_AIO_BUFFER_COUNT`` writes stay in flight. FASTLZ
|
once. FASTLZ open compresses each fragment when that shrinks it.
|
||||||
open compresses each chunk when that shrinks it.
|
|
||||||
|
|
||||||
Args:
|
Args:
|
||||||
start_sector: Sector offset from the start of the disk.
|
start_sector: Sector offset from the start of the disk.
|
||||||
@@ -358,50 +354,42 @@ class NfcDisk:
|
|||||||
if len(data) != length:
|
if len(data) != length:
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
f"write data is {len(data)} bytes, need {length}")
|
f"write data is {len(data)} bytes, need {length}")
|
||||||
max_sectors = NFC_AIO_BUFFER_SIZE // self.sector_size
|
disk_offset = start_sector * self.sector_size
|
||||||
offset_sectors = start_sector
|
op_id = self._next_op_id()
|
||||||
remaining = data
|
frag_offset = 0
|
||||||
pending: set[int] = set()
|
while frag_offset < length:
|
||||||
while remaining or pending:
|
chunk = data[frag_offset:frag_offset + NFC_AIO_BUFFER_SIZE]
|
||||||
while remaining and len(pending) < NFC_AIO_BUFFER_COUNT:
|
extra = chunk
|
||||||
n_sectors = min(
|
extra_len = len(chunk)
|
||||||
len(remaining) // self.sector_size, max_sectors)
|
|
||||||
chunk = remaining[:n_sectors * self.sector_size]
|
|
||||||
pending.add(self._send_write_chunk(offset_sectors, chunk))
|
|
||||||
offset_sectors += n_sectors
|
|
||||||
remaining = remaining[n_sectors * self.sector_size:]
|
|
||||||
if not pending:
|
|
||||||
break
|
|
||||||
rtype, rop, _body = self._aio_recv_reply()
|
|
||||||
if rtype != NFC_AIO_MSG_IO or rop not in pending:
|
|
||||||
raise NfcProtocolError(
|
|
||||||
f"AIO IO write reply type={rtype} opId={rop}, "
|
|
||||||
f"expected type={NFC_AIO_MSG_IO} opId in {pending}")
|
|
||||||
pending.remove(rop)
|
|
||||||
|
|
||||||
def _send_write_chunk(self, start_sector: int, data: bytes) -> int:
|
|
||||||
length = len(data)
|
|
||||||
offset = start_sector * self.sector_size
|
|
||||||
extra = data
|
|
||||||
ctype = NFC_COMPRESSION_NONE
|
ctype = NFC_COMPRESSION_NONE
|
||||||
extra_len = length
|
if (
|
||||||
if self.compression == NFC_COMPRESSION_FASTLZ and length >= 16:
|
self.compression == NFC_COMPRESSION_FASTLZ
|
||||||
compressed = fastlz.compress(data)
|
and extra_len >= 16):
|
||||||
if compressed and len(compressed) < length:
|
compressed = fastlz.compress(chunk)
|
||||||
|
if compressed and len(compressed) < extra_len:
|
||||||
extra = compressed
|
extra = compressed
|
||||||
ctype = NFC_COMPRESSION_FASTLZ
|
|
||||||
extra_len = len(compressed)
|
extra_len = len(compressed)
|
||||||
|
ctype = NFC_COMPRESSION_FASTLZ
|
||||||
opcode = NFC_AIO_IO_WRITE | (ctype << 32)
|
opcode = NFC_AIO_IO_WRITE | (ctype << 32)
|
||||||
payload = struct.pack(
|
payload = struct.pack(
|
||||||
"<QQQQIII",
|
"<QQQIIIII",
|
||||||
self.handle,
|
self.handle,
|
||||||
opcode,
|
opcode,
|
||||||
offset,
|
disk_offset,
|
||||||
length,
|
|
||||||
length,
|
length,
|
||||||
|
frag_offset,
|
||||||
|
len(chunk),
|
||||||
extra_len,
|
extra_len,
|
||||||
0)
|
0)
|
||||||
return self._aio_send(NFC_AIO_MSG_IO, payload, extra)
|
self._sock.sendall(
|
||||||
|
_pack_aio_hdr(NFC_AIO_MSG_IO, len(payload), op_id)
|
||||||
|
+ payload + extra)
|
||||||
|
frag_offset += len(chunk)
|
||||||
|
rtype, rop, _body = self._aio_recv_reply()
|
||||||
|
if rtype != NFC_AIO_MSG_IO or rop != op_id:
|
||||||
|
raise NfcProtocolError(
|
||||||
|
f"AIO IO write reply type={rtype} opId={rop}, "
|
||||||
|
f"expected type={NFC_AIO_MSG_IO} opId={op_id}")
|
||||||
|
|
||||||
def close(self) -> None:
|
def close(self) -> None:
|
||||||
"""Close the VMDK, the AIO session, and the classic NFC session."""
|
"""Close the VMDK, the AIO session, and the classic NFC session."""
|
||||||
|
|||||||
Reference in New Issue
Block a user