# 0. Summary

**Scope.** Everything the reel pipeline does, from pulling a 4K talk off YouTube
to writing the finished 1080×1920 MP4 — which stages exist, what codecs they
use, what each one costs in CPU, cores, clock, memory and disk, where it is
wasting the machine, and what 100 clips a day would actually take.

**Measured on an NVIDIA DGX Spark (GB10), 2026-10-07.** All ten reels were
re-rendered from scratch under instrumentation; the per-stage and per-alternative
numbers are fresh benchmarks on the same box. Nothing is estimated unless the
line says "projected".

**This is the DGX half.** The MacBook session has to run the same harness
(`raw/measure.py`) on its own machine — that box is the compute-constrained one,
and three numbers here say it will hurt there: 63 GB of ASR worker memory,
2.75 GB of source per talk, 5.8 GB written per ten reels.

---

## The six numbers that matter

| | |
|---|---|
| **2.11 CPU-hours** | to produce **317 seconds** of finished video — 24 CPU-seconds per second of output |
| **8.3 of 20 cores** | average utilisation, every reel, start to finish — the box is **58 % idle throughout** |
| **2 processes** | peak count in the whole build tree: one Python driver and one ffmpeg. Nothing ever runs in parallel |
| **95 %** | of a reel's CPU is x264 — encoding **the same picture three times** |
| **2.43 CPU-hours** | to transcribe **one talk** — more than all ten renders combined |
| **0 W** | of GPU work. The Blackwell GPU idles at 3.77 W from ingest to upload |


## Against the MacBook run (§8)

Same ten reels, same 317 s of video, same wall clock — **914 s on the DGX vs
907 s on the MacBook** — while the DGX spent **3.6× the CPU** (7 586 s vs
2 103 s) and 3.5× the cores. A 20-core workstation ties with a laptop, because
throughput is set by one ffmpeg at a time. Either that run encodes in hardware
or it doesn't encode three times; both point at the same fix.

## What it costs today

| | Wall | CPU |
|---|---:|---:|
| Ten reels, end to end | 914 s | 7 586 s |
| One talk: ASR | ~32 min | 8 750 s |
| One talk: face/shot tracking | ~163 s | 2 030 s |
| One talk: ingest + demux + silence map | ~2.7 min | ~25 s |

## The five things worth changing

| # | Change | Measured effect |
|---|---|---|
| 1 | **Fuse the three encodes into one ffmpeg graph** | 360.9 s → **116.8 s CPU** for the same 30 s, 28 % smaller file, *better* SSIM |
| 2 | **Run reels concurrently** (nothing stops you) | 6 cut encodes 19.5 s → **9.3 s**, same CPU |
| 3 | **6 ASR workers instead of 14** (42 threads on 20 cores today) | ~32 min → **~7 min** per talk, projected from measured per-chunk cost |
| 4 | **Parallelise or inline the text raster** | 1 000 frames 10.7 s → **1.2 s** |
| 5 | **Move the film grain after compression, or drop it** | **−16 % CPU, −16 % bitrate**, and renders become reproducible |

Stacked, ten reels should go from **914 s to roughly 175–200 s** of wall clock.

## The two things that are blocked, not broken

- **NVENC** — `h264_nvenc` refuses: *"Required: 13.1 Found: 13.0 … minimum
  driver 610.00"*, against the installed 580.173.02. Driver update or an ffmpeg
  rebuilt against nv-codec-headers 13.0.
- **GPU ASR** — the aarch64 ctranslate2 wheel is compiled without CUDA
  (`get_cuda_device_count() → 0`) and torch isn't installed. A build job, not a
  hardware limit.

## On "do we need the DGX"

Today, **no** — the pipeline can't use it. 8.3 of 20 cores, 1.7 GB of 120 GB,
no GPU. Any 8-core box with enough RAM would be nearly as fast, because one
ffmpeg at a time is the ceiling.

After the fixes, **one box of this class covers 100 clips/day in ~2.6 h** with
21 hours of headroom — and the DGX becomes genuinely the right machine only once
the two blocked GPU paths are unblocked, because that is where its advantage
actually lives. Adding servers or a laptop render farm is stage three; today a
second machine would also sit 58 % idle. See §7.

---

### Files

| | |
|---|---|
| `index.html` | the same report as a page, with the clip table and charts |
| `01_hardware_and_method.md` | the box, and exactly how each number was taken |
| `02_pipeline_stages.md` | every stage, in order, and the architecture decisions behind them |
| `03_codecs.md` | codecs, containers, bitrates and generational loss at every hop |
| `04_measurements.md` | the finished-clips table, per-stage CPU, and all benchmarks |
| `05_bottlenecks.md` | ranked, with the evidence for each |
| `06_optimizations.md` | what to change, in payback order |
| `08_comparison_dgx_vs_macbook.md` | the DGX-vs-MacBook comparison, and what it does and doesn't prove |
| `07_scaling.md` | the 100-clips-a-day arithmetic, and the server/farm question |
| `data/` | raw JSON: per-clip stats, per-stage profiles, benchmark results |
| `raw/` | the measurement harness itself, and the pipeline sources it measured |
