The demo tooling only existed in a session scratch directory, which is
temporary. This puts it in the repo so the video can be rebuilt.
Terminal only: tmux drives a two pane session, asciinema records it, agg
renders it, ffmpeg encodes it. The left pane runs the loop against a live
ZVM, the right pane polls the VPG journal so the tagged checkpoint appears
on camera as it lands.
Narration is synthesised per beat and aligned to marks the driver writes
while it runs, rather than to predicted timings, so a slow API call or an
FLR mount that takes longer than usual does not drift the audio. Two
things that has to respect are written down in the README: agg's
idle-time-limit must exceed the longest pause or it compresses idle time
and breaks the wall-clock mapping, and each beat holds for its narration
length so no line is cut off.
demo_win.json carries live guest credentials, so only an example with
placeholders is committed and the real file is gitignored, along with the
generated wav, cast, gif and mp4.
The xAI key path and voice id come from the environment now instead of
being hardcoded to one machine.
Also records the narration gotchas that cost time: mapping an acronym to
run-together phonetics ({"VM": "vee em"}) is spoken as one word, "vem";
prose written for the page sounds robotic read aloud; and volumedetect
reports no samples when aimed at a file whose first stream is video,
which makes a working audio track look silent.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_016yVfC5nvZowoLFnEGWhLGn
Demo recording harness
Records the rewind loop running against a live ZVM as a narrated MP4. Terminal only: no screen capture, no video editor.
tmux ──▶ asciinema ──▶ agg ──▶ ffmpeg ──▶ mp4
│ ▲
└── left pane: driver, right pane: journal │
│
xAI /v1/tts ──▶ wav per beat ──────────┘ (mux_vo.py)
Pieces
| file | what it does |
|---|---|
windows_driver.py |
Windows demo. Intro over the diagram, then the live loop via WinRM. |
driver.py |
The Linux equivalent, over SSH. |
diagram.py |
Architecture diagram, revealed in four chunks against intro beats i1..i4. |
journal.py |
Right-hand pane. Polls a VPG's checkpoints so the tag appears on camera. |
record_win.sh / record.sh |
Drive tmux + asciinema, then agg and ffmpeg. |
narration.md |
The script. One [beat] per block; the source of truth. |
vo/build_narration.py |
Parses narration.md, synthesises a wav per beat, records durations. |
mux_vo.py |
Aligns the wavs to the recorded beat marks and muxes the audio. |
Running it
cp demo/demo_win.example.json demo/demo_win.json # then fill in the guest creds
export XAI_KEY_FILE=~/xai-api.key XAI_VOICE_ID=<id from /v1/custom-voices>
python3 demo/vo/build_narration.py 1.0 # synthesise, writes durations.json
demo/record_win.sh win1 # record; writes marks.jsonl
python3 demo/mux_vo.py win1 # -> win1_narrated.mp4
demo_win.json holds live guest credentials and is gitignored. So are the
generated .wav, .cast, .gif and .mp4 files.
How the audio stays in sync
The driver writes marks.jsonl as it runs: one line per beat with the real
elapsed time it started. mux_vo.py delays each wav to its recorded mark, so
sync survives a slow API call or an FLR mount that takes longer than usual.
Nothing is predicted.
Two things this depends on:
agg --idle-time-limitmust be larger than the longest pause (the scripts pass 3600). The default is 5 seconds, which compresses idle time, and that silently breaks the mapping between wall clock and video time.- Each beat holds for its narration length.
hold()sleeps out whatever is left after the work finishes, so a line is never cut off mid-sentence.
Check alignment after a mux: the FLR wait should measure near silence.
ffmpeg -v error -i out.mp4 -vn -ac 1 /tmp/a.wav
ffmpeg -hide_banner -ss 150 -t 6 -i /tmp/a.wav -af volumedetect -f null /dev/null 2>&1 | grep mean_volume
Speech sits around -22 dB; a correctly aligned gap reads about -91 dB.
Narration gotchas
- Do not map acronyms to run-together phonetics. The
replacemap takes{"phrase": "pronunciation"}, and{"VM": "vee em"}gets spoken as one word, "vem". Either leave the acronym alone or expand it:{"VM": "virtual machine"}. - Write for speech, not for the page. Short declaratives and fragment stacks read well and sound robotic out loud. Commas and full stops are what the engine uses for pacing, so clauses joined with commas breathe; a wall of four-word sentences marches.
volumedetectreportsn_samples: 0when pointed at a file whose first stream is video. Extract the audio first, then measure, or you will think a working track is silent.