VDDK always decompresses the chunks it retrieves. However, some callers may be interested in the compressed chunks, which may be forwarded to another service involved in the backup process. We'll add a read flag (skip_decompression). If set, the read operation will return a "ReadResult" container, containing a list of fragments and their compressed / decompressed lengths. If the returned buffer size matches the AIO buffer size (which now becomes configurable), at most one fragment will be returned. While at it, we're adding perf tests that check various AIO buffer sizes, cross checking against VDDK.
9.7 KiB
VDDK NFC disk read
This document records how VMware VDDK reads VMDK sectors over NBD/NFC
after the open in docs/nfc_open.md, and how OpenVixDiskLib
(NfcDisk.read in openvixdisklib/nfc_open.py) reproduces
VixDiskLib_Read. Capture method:
docs/reverse_engineering_procedure.md.
Mapping from VDDK
VixDiskLib_Read(handle, startSector, numSectors, buf) becomes one
NFC_AIO_MSG_IO (type 7) on the NFC socket. Units on the wire are
bytes, not sectors:
offset = startSector * sectorSize
length = numSectors * sectorSize
sectorSize is 512 from the OPEN_FILE reply on this lab disk.
| VDDK call | Wire effect |
|---|---|
VixDiskLib_Read(h, 0, 1, buf) |
IO offset 0, length 512, one 512-byte fragment |
VixDiskLib_Read(h, 1, 1, buf) |
IO offset 512, length 512 |
VixDiskLib_Read(h, 0, 128, buf) |
IO length 65536 (AIO buffer size), one fragment |
VixDiskLib_Read(h, 0, 129, buf) |
One request of 66048; two reply fragments |
VDDK does not split a Read larger than the AIO buffer into
multiple requests. The client sends one AIO message; the server
answers with one or more same-opId replies, each carrying at most
the OPEN_SESSION buffer size (VDDK default 65536).
vixDiskLib.nfcAio.Session.BufSizeIn64KB=32 advertises 2 MiB; a
129-sector read then returns one 66048-byte extra, and a 2 MiB +
512 read returns 2097152 + 512. See docs/nfc_open.md (OPEN_SESSION).
Sparse regions are still transferred as zeros. A read of 8 sectors at LBA 8 on this disk was 4096 zero bytes on the wire, not a skip.
Request (44 bytes)
Little-endian, after the usual 16-byte AIO header
(magic 0xA100DA7A, type 7, size 44, monotonic opId):
| Offset | Type | VDDK Read(start, n) |
|---|---|---|
| 0 | uint64 |
File handle from OPEN_FILE |
| 8 | uint64 |
Direction in low 32 bits; FastLZ type 2 in high 32 bits |
| 16 | uint64 |
Byte offset |
| 24 | uint64 |
Byte length |
| 32 | uint32 |
Byte length (same value) |
| 36 | uint32 |
Byte length, or compressed extra size when type is FastLZ |
| 40 | uint32 |
0 |
An earlier guess that offset 36 was NFC_DISK (2) was wrong: a
1-sector VDDK read puts 512 in both uint32 length fields. A Python
read that sent (512, 2, 0) still worked for one sector; OpenVixDiskLib
now matches VDDK.
VIXDISKLIB_FLAG_OPEN_COMPRESSION_FASTLZ does not change OPEN_FILE
flags. The IO opcode at offset 8 is a uint64: low 32 bits are still
0/1 (write/read), high 32 bits are the NFC compression type
(2 = FastLZ). OPEN still uses handshake PlainText.
| Open flag / wire | Request extra | Reply extra |
|---|---|---|
| No compression flag | Raw length bytes on write |
Raw fragment at offset 32 |
| FASTLZ, data that shrinks | FastLZ bytes; offset 36 = packed size | Opcode type 2; extra is FastLZ of offset 32 |
| FASTLZ, incompressible | Raw bytes; opcode type 0 |
Opcode type 0; extra is raw |
Reads with FASTLZ always request type 2. The server may answer type
2 or fall back to type 0. Decompress into the uncompressed fragment
length at offset 32 and copy to the dest at offset 28.
64 KiB chunks use FastLZ level 2 (first byte has bit 5 set). Smaller
chunks use level 1. VDDK’s URL form is FASTLZ-vpxa-nfc://…; authd
PROXY is unchanged.
Reply
Each fragment is: 16-byte AIO header (same type and opId) + 44-byte
payload + chunkLength data bytes.
Reply payload (handle is zeroed; lengths describe this fragment):
| Offset | Type | Meaning |
|---|---|---|
| 0 | uint64 |
0 |
| 8 | uint64 |
1 (read) |
| 16 | uint64 |
Byte offset of the request on disk |
| 24 | uint32 |
Total request length |
| 28 | uint32 |
Fragment byte offset in this request (0, 65536, …), not disk |
| 32 | uint32 |
This fragment’s uncompressed byte length |
| 36 | uint32 |
Same as offset 32, or compressed extra size when type is FastLZ |
| 40 | uint32 |
0 |
When there is a single fragment, offsets 24–31 look like a uint64
length (the fragment offset is 0). The 129-sector capture shows why
they are two uint32s: fragment 0 has (66048, 0) then chunk 65536;
fragment 1 has (66048, 65536) then chunk 512. 0x00010000 at offset
28 is the byte offset, not a 0-based index. Disk byte address of a
fragment is request offset (payload 16) plus payload 28.
Read loop: receive fragments with that opId until the concatenated
data length equals the request. Use the uint32 at payload offset 32
as the extra-data size for that fragment, and copy it to the byte
offset at payload offset 28 — fragments are not always delivered in
order. Do not treat extra data as part of AIO size (that field stays
44).
129-sector example (one client request, two server fragments):
C: type=7 opId=18 size=44 offset=0 length=66048
S: type=7 opId=18 size=44 dest=0 chunk=65536 + 65536 data
S: type=7 opId=18 size=44 dest=65536 chunk=512 + 512 data
dest in that dump is payload offset 28 (ReadFragment.dest): 0 and
65536 are positions in this 66048-byte read, not sector numbers.
Lab check
Integration tests create an empty 10 GiB thin disk, write a repeating pattern at each captured range (including 129 sectors), and read it back. An unwritten region is zeros.
Writes use the same 44-byte IO payload with opcode 2; see
docs/nfc_write.md.
OpenVixDiskLib
NfcDisk.read(start_sector, num_sectors) in
openvixdisklib/nfc_open.py. Run:
.venv/bin/pytest tests/integration/test_nfc_read_write.py
The integration test writes and then reads the captured VDDK ranges (including a 129-sector transfer that must assemble two read fragments).
Skip decompression (OpenVixDiskLib extension)
VixDiskLib_Read always fills buf with uncompressed sector bytes.
OpenVixDiskLib can skip FastLZ decode so a backup application can
forward the compressed data as-is, avoiding unnecessary re-compression.
NfcDisk.readinto(..., skip_decompression=True) and
VixDiskLibHandle.read(..., skip_decompression=True) still send one
IO request and wait until uncompressed filled == length. They do
not decompress. Extras are packed densely from offset 0 of buf.
ReadResult.fragments describes each extra. Type 2 extras are
FastLZ; type 0 fallbacks are raw. Concatenating extras is not a
valid FastLZ stream; the caller must use the table to split them.
| Field | Meaning |
|---|---|
dest |
Byte offset in this uncompressed read (NFC payload 28). Not a disk LBA or VMDK file offset. |
uncompressed_length |
Uncompressed fragment size (NFC payload 32). |
compression_type |
NFC_COMPRESSION_NONE (0) or NFC_COMPRESSION_FASTLZ (2). |
offset |
Start of this extra in packed buf (receive order, densely from 0). |
length |
Extra size on the wire. |
Disk byte address of a fragment is start_sector * 512 + dest. A
129-sector read from sector 0 or from sector 1000 still reports
dest=0 and dest=65536 when extras are 64 KiB.
buf is sized for the uncompressed request, so it is always large
enough. Default read still decompresses; fragments is empty and
compressed_length is still the extra bytes on the wire.
skip_decompression with a plain (no FASTLZ) open only records raw
extras (compressed_length == uncompressed_length).
This is not VixDiskLib_Read. Do not add an open flag for it;
compression on the wire is already the FASTLZ open flag.
A 32 MiB read at 64 KiB extras is 512 fragments in one result. A
2 MiB OPEN_SESSION extra (aio_buffer_size=2097152) is 16 fragments
for the same read. One dest PUT per extra is not viable.
uncompressed request (offsets in this read, not on disk)
|---------------- 64KiB --|-- 64KiB --|-- ... --|
dest=0 dest=65536
extra (FastLZ or raw) extra (FastLZ or raw)
buf when skip_decompression=True: extras packed densely from offset 0
What is still VDDK-only
- zlib and skipz NBD compression flags
VixDiskLib_ReadAsync(same IO messages, different client threading)VixDiskLib_QueryAllocatedBlocks/ allocation bitmapsVixDiskLib_GetInfocapacity (not required to read a known range)