# 4. Measurements

All numbers below are from a **fresh instrumented re-render of all ten reels**
on 2026-10-07, 15:08–15:23, on an otherwise idle DGX Spark.

## 4.1 Finished clips

| Reel | Length | Speed | Wall | CPU time | Avg cores | Peak cores | On P-cores | Avg GHz | Peak memory | Output |
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| `01_fiction_platforms` | 25.2 s | **7.55 fps** | 83.4 s | 627 s | 7.52 | 18.8 | 50.4 % | 3.37 | 1.68 GB | 16.3 MB |
| `02_retcon` | 35.2 s | **7.47 fps** | 117.9 s | 946 s | 8.03 | 19.7 | 51.6 % | 3.39 | 1.70 GB | 21.0 MB |
| `03_jobber` | 29.4 s | **9.82 fps** | 74.8 s | 671 s | 8.98 | 19.3 | 52.1 % | 3.39 | 1.71 GB | 19.5 MB |
| `04_whoop` | 37.5 s | **8.56 fps** | 109.5 s | 916 s | 8.37 | 19.3 | 50.8 % | 3.38 | 1.70 GB | 22.3 MB |
| `05_out_on_your_back` | 27.6 s | **7.83 fps** | 88.0 s | 693 s | 7.87 | 19.4 | 51.0 % | 3.39 | 1.67 GB | 18.1 MB |
| `06_nobody_cares_about_facts` | 28.5 s | **9.81 fps** | 72.6 s | 643 s | 8.85 | 19.4 | 51.7 % | 3.39 | 1.67 GB | 16.2 MB |
| `09_burn_bridges` | 28.8 s | **8.53 fps** | 84.3 s | 707 s | 8.39 | 20.0 | 52.0 % | 3.39 | 1.72 GB | 12.2 MB |
| `10_infosys` | 25.4 s | **7.02 fps** | 90.4 s | 693 s | 7.66 | 20.0 | 51.6 % | 3.37 | 1.67 GB | 16.6 MB |
| `11_medium_risk` | 39.0 s | **8.98 fps** | 108.5 s | 953 s | 8.78 | 19.5 | 51.2 % | 3.39 | 1.69 GB | 20.1 MB |
| `12_corpus_back` | 40.7 s | **12.03 fps** | 84.5 s | 736 s | 8.71 | 19.7 | 53.0 % | 3.41 | 1.68 GB | 19.8 MB |
| **10 reels** | **317.1 s** | **8.67 fps** | **914 s** | **7586 s** | **8.30** | — | **51.5 %** | **3.39** | **1.72 GB** | **182.1 MB** |

"Speed" = finished frames (25 fps) produced per second of wall clock — the same
sense as the reference dashboard. Energy is omitted because this box exposes no
energy counter (see §1); CPU time is the stand-in.

### What that says

- **2.11 CPU-hours** to produce **317 seconds** of finished video — **24 CPU-seconds per second of output**, 2.88 s of wall per second of output.
- Every reel lands in the same narrow band: **7.5–9.0 of 20 cores**. The box is **55–60 % idle for the whole run** and the idle half never gets used, because the pipeline never runs two things at once.
- Peak cores reaches 20 — x264 does briefly fill the machine, but only inside one ffmpeg invocation, and then the pipeline goes back to a single process.
- **On P-cores 50–53 %**: the scheduler is splitting work evenly across the 10 Cortex-X925 and 10 Cortex-A725 cores. Half of this workload is running on cores that are 28 % slower. Nothing is pinned; there is no `taskset` / `sched_setaffinity` anywhere in the pipeline.
- **Peak memory 1.67–1.72 GB** per reel, on a 120 GB machine. **Memory is not a constraint at any point.** The whole build could run 20-way parallel inside 35 GB.
- Peak process count in the tree: **2** (the Python driver plus exactly one ffmpeg). That single number is the whole problem.
- Disk written during the run: **5.8 GB** for 0.18 GB of delivered video — **32× write amplification** from intermediates and the PNG text layer.

## 4.2 Where the CPU inside one reel actually goes

Two reels re-run with every `ffmpeg` call and every in-process Python stage timed
individually (`stageprof.py`). CPU is `getrusage` delta, so it is exact.

| Stage | `01_fiction_platforms` (25.2 s, 3 b-roll) | | `12_corpus_back` (40.7 s, no b-roll) | |
|---|---:|---:|---:|---:|
| | **CPU** | **share** | **CPU** | **share** |
| R2 cut + crop + grade + encode (x264 medium crf13) | 301.3 s | 47.9 % | 314.7 s | 43.0 % |
| R6 speed-up re-encode (x264 medium crf14) | 115.6 s | 18.4 % | 197.3 s | 26.9 % |
| R11 final composite + audio mix (x264 slow crf20 + aac) | 112.9 s | 17.9 % | 183.6 s | 25.1 % |
| R2b b-roll insert encodes (x264 medium crf14) | 68.0 s | 10.8 % | 0.0 s | 0.0 % |
| R9 text layer raster (Pillow, pure Python) | 17.3 s | 2.7 % | 15.9 s | 2.2 % |
| R12 QC flash scan (full decode) | 4.9 s | 0.8 % | 7.6 s | 1.0 % |
| R13 QC freeze scan (full decode) | 4.2 s | 0.7 % | 6.1 s | 0.8 % |
| R10 music bed synthesis (numpy) | 2.9 s | 0.5 % | 4.1 s | 0.6 % |
| R8 two-pass loudnorm | 1.1 s | 0.2 % | 1.8 s | 0.2 % |
| R7 rubberband tempo | 0.4 s | 0.1 % | 0.6 s | 0.1 % |
| R4 concat (stream copy) | 0.5 s | 0.1 % | 0.4 s | 0.0 % |
| R3 audio slice extract (pcm_s24le) | 0.1 s | 0.0 % | 0.1 s | 0.0 % |
| ffprobe | 0.1 s | 0.0 % | 0.0 s | 0.0 % |
| R5 audio crossfades | 0.0 s | 0.0 % | 0.0 s | 0.0 % |
| misc | 0.0 s | 0.0 % | 0.0 s | 0.0 % |
| **total** | **629 s** | | **732 s** | |

- **Encoding is 95 % of a reel's CPU on clip 01 and 95 % on clip 12.** Everything else is noise by comparison.
- The three encode passes are roughly **50 / 25 / 25** — the piece encodes, then two full re-encodes of the same picture.
- **The text raster is cheap in CPU (16–17 s) but it is 100 % single-threaded**, so it costs ~17 s of *wall* with 19 cores parked. On a 76 s build that is 22 % of the wall clock spent at 1 core.
- QC (flash + freeze scan) decodes the finished file twice for ~10 s of CPU. Cheap, but it is two decodes where one `ffmpeg` with both filters would do.
- Audio is free: every audio stage together is under 3 s of CPU per reel.

## 4.3 Benchmarks — what the alternatives actually cost

Same 30 s of source (six 5 s pieces, the real crop/scale/grade chain), measured
the same way.

### Is decoding the problem? No.

| | Wall | CPU | Avg cores |
|---|---:|---:|---:|
| 4K AV1 software decode (`libdav1d`), 20 s | 0.80 s | 12.6 s | 15.7 |
| 4K AV1 **NVDEC** (`av1_cuvid`), 20 s | 1.80 s | **0.6 s** | 0.35 |

`libdav1d` decodes 4K AV1 at **25× realtime** using 15.7 cores. NVDEC uses 21×
less CPU but is *slower in wall time* here. Both paths produce byte-identical
frame counts (125 frames of 3840×2160, verified). **Decode is ~3 % of the cost
— it is not the bottleneck.** NVDEC only matters once you want many reels
encoding at the same time and need the CPU back.

### Is encoding the problem? Yes — and it is encoded three times.

Today's path, for the same 30 s:

| Pass | Wall | CPU | Output |
|---|---:|---:|---:|
| R2 six cut encodes, `-preset medium -crf 13`, serial | 19.8 s | 171.2 s | 77.9 MB |
| R4 concat, `-c copy` | 0.3 s | 0.3 s | 77.9 MB |
| R6 speed re-encode, `-preset medium -crf 14` | 6.2 s | 104.5 s | 44.9 MB |
| R11 delivery, `-preset slow -crf 20` | 5.8 s | 84.9 s | 7.36 MB |
| **total** | **32.0 s** | **360.9 s** | |

One fused ffmpeg graph instead — six trimmed inputs → crop/scale/grade → concat
→ `setpts` → **one** encode straight to delivery settings:

| | Wall | CPU | Output |
|---|---:|---:|---:|
| **Fused, single encode** (`slow crf20`) | **15.8 s** | **116.8 s** | **5.28 MB** |
| Fused, grain removed from the grade | 15.0 s | 97.7 s | 4.43 MB |
| Fused, `-preset medium` | 15.4 s | 92.3 s | 5.82 MB |

**3.1× less CPU, 2.0× less wall, and a 28 % smaller file.**

### Does collapsing the encodes cost quality? No — it gains a little.

Against a lossless (FFV1) reference of the graded picture, so only generational
loss is being measured:

| Path | SSIM vs reference | PSNR |
|---|---:|---:|
| Three generations (crf13 → crf14 → slow crf20), as shipped | 0.98241 | 47.00 dB |
| **One generation (slow crf20 direct)** | **0.98305** | **47.23 dB** |

### Parallelism is free money

| | Wall | CPU | Avg cores |
|---|---:|---:|---:|
| 6 cut encodes, serial (as shipped) | 19.5 s | 170.9 s | 8.75 |
| 6 cut encodes, all at once, `-threads 4` | **9.3 s** | 166.7 s | **17.98** |
| 6 cut encodes, serial, `-preset veryfast` | 16.1 s | **87.3 s** | 5.43 |
| Text raster, 1000 frames, serial (as shipped) | 10.7 s | 10.7 s | **1.00** |
| Text raster, 1000 frames, 10 processes | **1.2 s** | 12.0 s | 9.84 |

Same work, same CPU, **2.1× the throughput** on the cuts and **8.8×** on the
text layer, purely from not doing one thing at a time.

### NVENC is present but blocked

`h264_nvenc`, `hevc_nvenc` and `av1_nvenc` all exist in this ffmpeg build, and
every attempt to use them fails with:

```
Driver does not support the required nvenc API version. Required: 13.1  Found: 13.0
The minimum required Nvidia driver for nvenc is 610.00 or newer
```

Installed driver is **580.173.02**. So hardware encode on this box is one driver
update (or one ffmpeg rebuilt against nv-codec-headers 13.0) away — not a
hardware limitation. **Until then, NVENC numbers in this report are absent, not
estimated.** NVDEC works today.
## 4.4 Per-source stages (measured once, amortised over every reel from that talk)

| Stage | Wall | CPU | Avg cores | Peak RSS | Note |
|---|---:|---:|---:|---:|---|
| Demux 48 kHz `pcm_s24le` (75 min talk) | 4.5 s | 10.0 s | 2.21 | 0.03 GB | writes 1.29 GB |
| Demux 16 kHz `pcm_s16le` | 5.1 s | 12.3 s | 2.43 | 0.03 GB | writes 143 MB |
| `silencedetect` over the 16 kHz file | 0.4 s | 1.4 s | 3.54 | 0.03 GB | free |
| Face + shot track, measured on 2 000 frames | 3.2 s | 39.7 s | 12.4 | 0.17 GB | → **~163 s wall / ~2 030 s CPU** for a full 101 602-frame talk |
| **ASR, one real 4.6-min chunk, `cpu_threads=3`** (exactly as `tr_chunk.py` runs it) | **180.4 s** | **624.7 s** | 3.46 | **4.53 GB** | 97.3 % on P-cores |
| ASR, same chunk, `cpu_threads=20` | **killed at 1 337 s, unfinished** | 8 500 s | 6.35 | 4.52 GB | see below |

### The ASR numbers, extrapolated honestly

One chunk costs **624.7 CPU-seconds**. A talk is split into 14 chunks of
similar length, so one talk costs **≈ 8 750 CPU-seconds = 2.43 CPU-hours** of
transcription — **more than all ten reel renders put together (2.11 CPU-h)**.

- A single worker with 3 threads actually consumes **3.46 cores**. Fourteen of
  them demand **48 cores on a 20-core box**.
- Perfectly packed, 8 750 CPU-s on 20 cores is **7.3 minutes**. The production
  run took **~32 minutes** — about **23 % efficiency**, the rest lost to
  2.4× oversubscription.
- **Running 6 workers instead of 14 should get most of that back** (6 × 3.46 ≈
  21 cores). That is a one-character change in `analyze.py`.
- **Raising `cpu_threads` is the wrong direction**: at 20 threads the same chunk
  had burned 8 500 CPU-seconds and 22 minutes without finishing, versus 625 s
  and 3 minutes at 3 threads. It was killed, so "≥ 1 337 s" is a floor, not a
  measurement of completion.
- **Each worker holds 4.53 GB.** Fourteen of them is **~63 GB resident** — fine
  on this 120 GB box, impossible on a 16/32 GB laptop. This is the number the
  MacBook session should check first.
