zerto_recover_file no longer writes the file to this host and hands back a path, because a path here means nothing to a caller elsewhere and letting the caller choose it was an arbitrary write. Both drivers still read rec["path"], so they broke. They now use the returned content. A small recovered_bytes() helper in each driver decodes the text or base64 form, so the Windows copy-back keeps shipping exact bytes rather than letting PowerShell rewrite line endings, which is the bug that put a stray CR in an earlier take. The Linux driver writes the bytes to a local file before scp, since scp needs something on disk to send. Verified against a live recovery: the helper returns 164 bytes whose sha256 matches the one the server reported, and the keys the on-screen show() filter uses are all still present. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
Demo recording harness
Records the rewind loop running against a live ZVM as a narrated MP4. Terminal only: no screen capture, no video editor.
tmux ──▶ asciinema ──▶ agg ──▶ ffmpeg ──▶ mp4
│ ▲
└── left pane: driver, right pane: journal │
│
xAI /v1/tts ──▶ wav per beat ──────────┘ (mux_vo.py)
Pieces
| file | what it does |
|---|---|
windows_driver.py |
Windows demo. Intro over the diagram, then the live loop via WinRM. |
driver.py |
The Linux equivalent, over SSH. |
diagram.py |
Architecture diagram, revealed in four chunks against intro beats i1..i4. |
journal.py |
Right-hand pane. Polls a VPG's checkpoints so the tag appears on camera. |
record_win.sh / record.sh |
Drive tmux + asciinema, then agg and ffmpeg. |
narration.md |
The script. One [beat] per block; the source of truth. |
vo/build_narration.py |
Parses narration.md, synthesises a wav per beat, records durations. |
mux_vo.py |
Aligns the wavs to the recorded beat marks and muxes the audio. |
Running it
cp demo/demo_win.example.json demo/demo_win.json # then fill in the guest creds
export XAI_KEY_FILE=~/xai-api.key XAI_VOICE_ID=<id from /v1/custom-voices>
python3 demo/vo/build_narration.py 1.0 # synthesise, writes durations.json
demo/record_win.sh win1 # record; writes marks.jsonl
python3 demo/mux_vo.py win1 # -> win1_narrated.mp4
demo_win.json holds live guest credentials and is gitignored. So are the
generated .wav, .cast, .gif and .mp4 files.
How the audio stays in sync
The driver writes marks.jsonl as it runs: one line per beat with the real
elapsed time it started. mux_vo.py delays each wav to its recorded mark, so
sync survives a slow API call or an FLR mount that takes longer than usual.
Nothing is predicted.
Two things this depends on:
agg --idle-time-limitmust be larger than the longest pause (the scripts pass 3600). The default is 5 seconds, which compresses idle time, and that silently breaks the mapping between wall clock and video time.- Each beat holds for its narration length.
hold()sleeps out whatever is left after the work finishes, so a line is never cut off mid-sentence.
Check alignment after a mux: the FLR wait should measure near silence.
ffmpeg -v error -i out.mp4 -vn -ac 1 /tmp/a.wav
ffmpeg -hide_banner -ss 150 -t 6 -i /tmp/a.wav -af volumedetect -f null /dev/null 2>&1 | grep mean_volume
Speech sits around -22 dB; a correctly aligned gap reads about -91 dB.
Narration gotchas
- Do not map acronyms to run-together phonetics. The
replacemap takes{"phrase": "pronunciation"}, and{"VM": "vee em"}gets spoken as one word, "vem". Either leave the acronym alone or expand it:{"VM": "virtual machine"}. - Write for speech, not for the page. Short declaratives and fragment stacks read well and sound robotic out loud. Commas and full stops are what the engine uses for pacing, so clauses joined with commas breathe; a wall of four-word sentences marches.
volumedetectreportsn_samples: 0when pointed at a file whose first stream is video. Extract the audio first, then measure, or you will think a working track is silent.