# 5. Bottlenecks, ranked

Ranked by how much they cost, not by how interesting they are.

## B1 — The pipeline never runs two things at once *(biggest)*

`edit.sh()` is a blocking `subprocess.run`. Peak process count measured across
all ten builds: **2** — one Python driver, one `ffmpeg`. Five to eight
piece-encodes that share nothing run strictly one after another, and the only
parallelism in the system is x264's internal threading.

**Evidence:** avg 7.5–9.0 cores of 20 for every reel; 6 cut encodes serial
19.5 s vs 9.3 s when launched together (same total CPU).
**Cost:** ~55 % of the machine, continuously.

## B2 — The picture is encoded three times

crf 13 piece encodes → crf 14 speed re-encode → crf 20 delivery encode. Encoding
is **95 % of a reel's CPU**, and two of the three passes exist only because the
pipeline writes intermediates to disk between steps.

**Evidence:** 360.9 s CPU across the three passes vs **116.8 s** for one fused
graph producing a *better* file (SSIM 0.98305 vs 0.98241, 28 % smaller).
**Cost:** ~3× the encode CPU, i.e. ~2× the whole pipeline.

## B3 — Transcription is the largest single compute item in the system

14 worker processes × `cpu_threads=3` = **42 threads on 20 cores, 2.1×
oversubscribed**, running `large-v3` int8 on CPU. For `kata9` this took
**~32 minutes of a fully loaded box** — more than all ten renders put together
(15.2 min). ctranslate2 in this venv is **compiled without CUDA**, and torch is
not installed, so the GPU cannot be used for ASR at all today.

**Cost:** roughly two thirds of the end-to-end time for a new talk.

## B4 — The text layer is a single-threaded Python loop writing a PNG per frame

`typo.render_layer` draws 1080×1920 RGBA with Pillow and saves one PNG per
frame, 630–1020 per reel, which `ffmpeg` then reads back immediately.

**Evidence:** measured at exactly **1.00 core**; 1000 frames take 10.7 s serial
vs **1.2 s** across 10 processes.
**Cost:** ~20 % of each reel's wall clock spent at 5 % machine utilisation.

## B5 — Film grain is added before compression, three times

Every grade ends with `noise=c0s=3:c0f=t+u`. Grain is the hardest thing for
x264 to code, and it is injected at the *first* of three generations, so it is
re-encoded twice more.

**Evidence:** removing it from the fused graph: **116.8 s → 97.7 s CPU
(−16 %)** and **5.28 MB → 4.43 MB (−16 %)**.
**Side effect:** the grain is random per run, so renders are not reproducible.

## B6 — Write amplification

**5.8 GB written** to produce **182 MB** of delivered video across ten reels —
**32×**. Per 30 s of reel: 78 MB of piece encodes + 78 MB concat + 45 MB base +
~35 MB of PNGs + ~10 MB of 24-bit WAV, nearly all of it read back once and
deleted. Not a bottleneck on an NVMe box today; it is the first thing that
bites on a laptop and the thing that would saturate a network filesystem in a
render-farm setup.

## B7 — Half the work lands on the slow cores

51.5 % of running-thread samples sit on the Cortex-A725 E-cores, which top out
at 2.81 GHz vs 3.90 GHz for the X925 P-cores. Nothing sets affinity. For the
*latency* of a single reel this is pure loss; for *throughput* with many reels
in flight it is the right behaviour, so this one only matters once B1 is fixed
and you care about one reel finishing fast.

## What is NOT a bottleneck

- **Memory.** 1.67–1.72 GB peak per reel on a 120 GB box. Twenty concurrent
  builds would use ~35 GB.
- **Decode.** `libdav1d` does 4K AV1 at 25× realtime; decode is ~3 % of cost.
- **Audio.** Every audio stage together is under 3 s of CPU per reel.
- **Disk throughput.** 5.8 GB over 15 minutes is nothing for the NVMe.
