# 6. What to change, in order of payback

Every number here is measured on this box, not projected, unless the line says
"projected".

## 1. Fuse the three encodes into one ffmpeg graph — **~3× less CPU per reel**

Build one `-filter_complex` that takes the N source cuts as N inputs, applies
`crop/scale/grade` per input, `concat`s them, applies `setpts=PTS/1.10`,
overlays the text layer, mixes the audio, and encodes **once** at delivery
settings.

- Measured: 360.9 s → **116.8 s** CPU for the same 30 s, 32.0 s → **15.8 s** wall.
- Output is *smaller* (7.36 MB → 5.28 MB) and marginally *better*
  (SSIM 0.98241 → 0.98305).
- Removes the 78 MB + 78 MB + 45 MB of intermediates per reel outright.
- Risk: one long filter graph is harder to debug than eight small files. Keep a
  `--debug-stages` flag that falls back to the current path.

## 2. Run reels concurrently — **~2.1× throughput, free**

Nothing in a reel build depends on another reel. Four to six concurrent builds
(`-threads 4` each) fills the box.

- Measured on the cut stage: 19.5 s → **9.3 s** for the same CPU.
- Memory cost: ~1.7 GB per concurrent build.
- Combined with #1, ten reels should go from **914 s to roughly 150–200 s**
  (projected from the two measured factors; the fused graph is filter-bound at
  ~7.4 cores, so ~3 concurrent builds saturate the box).

## 3. Parallelise the text raster — **8.8× on that stage**

Split the frame range across `multiprocessing.Pool(10)`. Measured 10.7 s →
1.2 s for 1000 frames. Better still, once #1 lands, render the text layer
straight into the fused graph as an overlay source and stop writing PNGs
to disk at all.

## 4. Move grain after compression, or drop it — **−16 % CPU, −16 % bitrate**

`noise=c0s=3:c0f=t+u` currently runs before the first of three encodes. Either
drop it, or apply it in the final graph only. Measured on the fused path:
116.8 s → 97.7 s CPU, 5.28 MB → 4.43 MB. Also makes renders reproducible.

## 5. Fix ASR — **the single biggest end-to-end win for a new talk**

Today: `large-v3` int8 on CPU, 14 procs × 3 threads on 20 cores (2.1×
oversubscribed), ~32 min per talk. Three options, cheapest first:

1. **Stop oversubscribing.** 14 × 3 = 42 threads on 20 cores. Either 6 workers ×
   3 threads or 14 workers × 1 thread. Costs nothing to try.
2. **Build ctranslate2 with CUDA** (the pip wheel for aarch64 has none —
   `get_cuda_device_count()` returns 0) and run `large-v3` in `float16` on the
   GB10. This is the real fix; the GPU is at 3.77 W today.
3. **Or switch the ASR to a CUDA stack that already ships aarch64 builds**
   (whisper.cpp with CUDA, or NVIDIA NeMo/Parakeet). Note torch is not
   installed in this venv at all.

Expected gain is large but **not measured here** — no CUDA ASR path exists on
this box yet, so no number is quoted.

## 6. Unblock NVENC — driver, not hardware

`h264_nvenc` fails with *"Required: 13.1 Found: 13.0 … minimum driver 610.00"*
against the installed 580.173.02. Either update the driver or rebuild ffmpeg
against nv-codec-headers 13.0. Once it works, the delivery encode moves off the
CPU entirely and NVDEC (`av1_cuvid`, already working, 21× less CPU than
`libdav1d`) can feed it. This matters most when many reels encode at once.

## 7. Merge the two QC decodes

`flash_scan` and `freeze_scan` each decode the finished file. One `ffmpeg` with
both `signalstats` and `freezedetect` in the graph halves it. Small (~5 s CPU
per reel) but trivial to do.

## 8. Pin latency-critical work to P-cores

`taskset -c 5-9,15-19` for a single-reel interactive render. Only worth it once
#2 lands and you are optimising one reel's latency rather than batch
throughput.

---

## Expected effect, stacked

| | Wall for 10 reels | CPU for 10 reels |
|---|---:|---:|
| Measured today | 914 s | 7 586 s |
| + fused single encode (#1, measured 3.1× CPU) | ~450 s | ~2 450 s |
| + concurrency (#2, measured 2.1× wall) | **~200 s** | ~2 450 s |
| + grain moved (#4, measured −16 %) | ~175 s | ~2 050 s |

The last two rows are **projections composed from measured factors**, not
end-to-end measurements. The honest way to confirm them is to build the fused
path and re-run `measure.py` — the harness is in `raw/`.
