Files
zerto-ai-rewind/demo/README.md
T

76 lines
3.4 KiB
Markdown

# Demo recording harness
Records the rewind loop running against a live ZVM as a narrated MP4. Terminal
only: no screen capture, no video editor.
```
tmux ──▶ asciinema ──▶ agg ──▶ ffmpeg ──▶ mp4
│ ▲
└── left pane: driver, right pane: journal │
│
xAI /v1/tts ──▶ wav per beat ──────────┘ (mux_vo.py)
```
## Pieces
| file | what it does |
|---|---|
| `windows_driver.py` | Windows demo. Intro over the diagram, then the live loop via WinRM. |
| `driver.py` | The Linux equivalent, over SSH. |
| `diagram.py` | Architecture diagram, revealed in four chunks against intro beats i1..i4. |
| `journal.py` | Right-hand pane. Polls a VPG's checkpoints so the tag appears on camera. |
| `record_win.sh` / `record.sh` | Drive tmux + asciinema, then agg and ffmpeg. |
| `narration.md` | The script. One `[beat]` per block; the source of truth. |
| `vo/build_narration.py` | Parses `narration.md`, synthesises a wav per beat, records durations. |
| `mux_vo.py` | Aligns the wavs to the recorded beat marks and muxes the audio. |
## Running it
```bash
cp demo/demo_win.example.json demo/demo_win.json # then fill in the guest creds
export XAI_KEY_FILE=~/xai-api.key XAI_VOICE_ID=<id from /v1/custom-voices>
python3 demo/vo/build_narration.py 1.0 # synthesise, writes durations.json
demo/record_win.sh win1 # record; writes marks.jsonl
python3 demo/mux_vo.py win1 # -> win1_narrated.mp4
```
`demo_win.json` holds live guest credentials and is gitignored. So are the
generated `.wav`, `.cast`, `.gif` and `.mp4` files.
## How the audio stays in sync
The driver writes `marks.jsonl` as it runs: one line per beat with the real
elapsed time it started. `mux_vo.py` delays each wav to its recorded mark, so
sync survives a slow API call or an FLR mount that takes longer than usual.
Nothing is predicted.
Two things this depends on:
- **`agg --idle-time-limit` must be larger than the longest pause** (the scripts
pass 3600). The default is 5 seconds, which compresses idle time, and that
silently breaks the mapping between wall clock and video time.
- **Each beat holds for its narration length.** `hold()` sleeps out whatever is
left after the work finishes, so a line is never cut off mid-sentence.
Check alignment after a mux: the FLR wait should measure near silence.
```bash
ffmpeg -v error -i out.mp4 -vn -ac 1 /tmp/a.wav
ffmpeg -hide_banner -ss 150 -t 6 -i /tmp/a.wav -af volumedetect -f null /dev/null 2>&1 | grep mean_volume
```
Speech sits around -22 dB; a correctly aligned gap reads about -91 dB.
## Narration gotchas
- **Do not map acronyms to run-together phonetics.** The `replace` map takes
`{"phrase": "pronunciation"}`, and `{"VM": "vee em"}` gets spoken as one word,
"vem". Either leave the acronym alone or expand it: `{"VM": "virtual machine"}`.
- **Write for speech, not for the page.** Short declaratives and fragment stacks
read well and sound robotic out loud. Commas and full stops are what the engine
uses for pacing, so clauses joined with commas breathe; a wall of four-word
sentences marches.
- `volumedetect` reports `n_samples: 0` when pointed at a file whose first stream
is video. Extract the audio first, then measure, or you will think a working
track is silent.