# 2. The pipeline, stage by stage

One "reel" is a 25–41 s vertical 1080×1920 cut out of a 60–75 minute 4K talk.
Ten of them were built from two source talks (`kata`, `kata9`) on 2026-10-07.

The pipeline splits into **per-source work** (done once per talk, amortised over
every reel cut from it) and **per-reel work**.

## Per-source stages (once per talk)

| # | Stage | Tool | What it does | Parallel? |
|---|---|---|---|---|
| S1 | Ingest master | `yt-dlp -N 4` (`prep.sh`) | format `401+bestaudio` → **AV1 3840×2160 25 fps + Opus**, muxed to MKV. 1.27 GB (kata) / 889 MB (kata9) | 4 download connections |
| S2 | Ingest proxy | `yt-dlp` fmt `134` | **H.264 640×360 238 kb/s** MP4 — the tracking proxy, 43 MB | — |
| S3 | Audio demux ×2 | ffmpeg | `pcm_s16le` 16 kHz mono (143 MB, for ASR) **and** `pcm_s24le` 48 kHz stereo (1.29 GB, the edit master) | 1 process each, serial |
| S4 | Silence map | ffmpeg `silencedetect=noise=-34dB:d=0.45` | the cut grid used by both the chunker and the editor | serial |
| S5 | Face + shot track | OpenCV + **YuNet ONNX** on CPU (`track_gen.py`) | every 10th frame (0.4 s) of the proxy, 2× upscale → face box; HSV stats classify camera-vs-slide. 101,602 frames → 10,161 samples | measured at **12.4 cores**; ~163 s/talk projected from a 2 000-frame run |
| S6 | Transcribe | **faster-whisper large-v3, int8, CPU** (`analyze.py` → 14× `tr_chunk.py`) | audio split at silences into 14 chunks, 14 worker processes × `cpu_threads=3`, word timestamps, beam 5, VAD | **14 procs × 3 threads = 42 threads on 20 cores** |

S6 for `kata9`: workers launched 14:00:09, last chunk written 14:32:06 →
**~32 min wall**, essentially the whole box saturated. That is the single
largest compute item in the entire system, larger than all ten renders combined.

## Per-reel stages (`build.py` → `lib/`)

Order of execution for one reel:

| # | Stage | Tool | Detail |
|---|---|---|---|
| R1 | Phrase→time resolve | Python | locate the spoken phrases on the word timeline |
| R2 | **Cut + crop + grade + encode**, one ffmpeg per piece (5–8 pieces/reel) | ffmpeg, `libx264 -preset medium -crf 13` | decodes 4K AV1 in software, `crop`→`scale=1080:1920:flags=lanczos`, face-tracking pan, grade chain (`eq`,`curves`,`vignette`,**`noise=c0s=3`**), optional `unsharp`. **Serial — one piece at a time.** |
| R3 | Audio cut extract | ffmpeg | `pcm_s24le` slice per piece with 12.5 ms crossfade pads |
| R4 | Video concat | ffmpeg concat demuxer, `-c copy` | free |
| R5 | Audio chain-crossfade | ffmpeg `acrossfade=d=0.025:c1=qsin` | one filter graph, N inputs |
| R6 | **Speed re-encode** | ffmpeg `setpts=PTS/1.10`, `libx264 -preset medium -crf 14` | **re-encodes the whole assembled picture a second time** just to apply the 10 % speed-up |
| R7 | Voice tempo | ffmpeg `rubberband=tempo=1.10:pitchq=quality` | available on this box (falls back to `atempo`) |
| R8 | Loudness | ffmpeg two-pass `loudnorm` → −14 LUFS | two full passes over the voice track |
| R9 | **Text raster** | **Pillow, pure Python** (`typo.render_layer`) | one 1080×1920 **RGBA PNG per frame**, `compress_level=1`, 25 fps → 630–1020 PNGs/reel written to disk, then read back by ffmpeg. **Single-threaded.** |
| R10 | Music bed | numpy additive synth (`music.py`) | 48 kHz stereo 16-bit WAV, chord pads + sub |
| R11 | **Final composite** | ffmpeg, `libx264 -preset slow -crf 20 -maxrate 14M` | overlays the PNG sequence, side-chain-ducks the music under the voice, limiter, `aac 256k`, `+faststart`. **Third encode of the same picture.** |
| R12 | QC flash scan | ffmpeg `scale=96:160,signalstats` | full decode of the finished file |
| R13 | QC freeze scan | ffmpeg `freezedetect=n=-55dB:d=0.5` | **another** full decode of the finished file |

### Architecture decisions worth naming

1. **Everything is CPU.** There is not one CUDA call in the pipeline. AV1 4K
   software decode, three x264 encodes, YuNet on onnxruntime-CPU, Whisper on
   ctranslate2-CPU. The GB10 GPU idles at 3.77 W from ingest to upload.
2. **Nothing inside a reel runs concurrently.** `edit.sh()` is a blocking
   `subprocess.run`. Five to eight piece-encodes that are perfectly independent
   run one after another; x264's internal threading is the *only* parallelism,
   and it leaves half the box idle (measured: 7.5–9.0 cores of 20).
3. **The picture is encoded three times** (R2 crf 13 → R6 crf 14 → R11 crf 20)
   and fully decoded five times (R2 source, R6, R11, R12, R13).
4. **The text layer is a PNG sequence on disk**, not an in-graph overlay — a
   single-threaded Python loop writing ~1 GB of PNG per reel for ffmpeg to read
   back immediately.
5. **Film grain is added before compression** (`noise=c0s=3:c0f=t+u` in every
   grade), which is exactly the signal x264 is worst at — it inflates both
   encode time and bitrate at every one of the three hops.
6. **Transcription is the hidden giant**, and it is run at 42 threads on 20
   cores — 2.1× oversubscribed — against a GPU that could do it in a fraction
   of the time if ctranslate2 had been built with CUDA.
