# 3. Codecs and formats at every hop

## Video

| Hop | Container | Codec | Resolution | Rate control | Notes |
|---|---|---|---|---|---|
| Source master (ingest) | MKV | **AV1** (yt-dlp fmt 401), decoded by `libdav1d` | 3840×2160 @ 25 fps | — | 1.27 GB / 75 min. AV1 4K software decode is the most expensive decode in the house |
| Tracking proxy | MP4 | H.264 (fmt 134) | 640×360 @ 25 fps, 238 kb/s | — | only used by the face tracker |
| Per-piece cut (R2) | MP4 | **libx264** | 1080×1920 @ 25 fps | `-preset medium -crf 13` | near-lossless intermediate, 5–8 per reel |
| Concatenated base (R4) | MP4 | — | — | `-c copy` | free |
| Speed-adjusted base (R6) | MP4 | **libx264** | 1080×1920 @ 25 fps | `-preset medium -crf 14` | **second generation** |
| Text layer (R9) | PNG sequence | PNG RGBA, `compress_level=1` | 1080×1920 @ 25 fps | lossless | 630–1020 files/reel on disk |
| Delivery (R11) | MP4 `+faststart` | **libx264 High profile** | 1080×1920 @ 25 fps | `-preset slow -crf 20 -maxrate 14M -bufsize 22M` | **third generation**, measured ~5.0 Mb/s |

Nothing in that column is hardware. `h264_nvenc`, `hevc_nvenc`, `av1_nvenc`,
`av1_cuvid` and `h264_cuvid` are all present in this ffmpeg build and all unused.

## Audio

| Hop | Format | Rate | Notes |
|---|---|---|---|
| Source | Opus (in MKV) | 48 kHz | from `bestaudio` |
| ASR feed | `pcm_s16le` mono | 16 kHz | 143 MB per talk, fed to Whisper |
| Edit master | `pcm_s24le` stereo | 48 kHz | 1.29 GB per talk — every cut is sliced from this |
| Piece slices | `pcm_s24le` | 48 kHz | ±12.5 ms crossfade pads |
| Joins | `acrossfade=d=0.025:c1=qsin:c2=qsin` | 48 kHz | 25 ms equal-power |
| Speed | `rubberband=tempo=1.10:pitchq=quality` | 48 kHz | formant-preserving; `atempo` fallback |
| Voice loudness | two-pass `loudnorm` | −14 LUFS | |
| Music bed | numpy synth → 16-bit WAV | 48 kHz stereo | `loudnorm` to −22 LUFS, then side-chain ducked under the voice (`sidechaincompress`), `amix`, `alimiter=0.95` |
| Delivery | **AAC-LC 256 kb/s** | 48 kHz stereo | measured 250 kb/s |

## Generational loss

The delivered picture is a **third-generation x264 encode** of a software-decoded
AV1 4K master, with synthetic film grain injected before the first of the three.
crf 13 → crf 14 → crf 20 is conservative enough that it is not visible, but the
cost is real: two full extra encodes per reel plus the grain penalty, for a
10 % speed change (R6) that could be folded into the cut filter graph
(`setpts` in R2) or into the final composite's graph for free.
