Ten finished reels, re-rendered under instrumentation on an NVIDIA DGX Spark (GB10) · 2026-10-07
This is the DGX half of the picture. The MacBook session needs to run the same harness on its own box — these numbers describe an unconstrained 20-core machine.
Compared against the FINISHED CLIPS table that came with the brief — the other machine's run of the same ten reels. Its wall-clock column is derived (CPU time ÷ average cores), everything else is read straight off it.
| Metric | DGX Spark (GB10) | MacBook | Ratio |
|---|---|---|---|
| Clips | 10 | 10 | — |
| Video delivered | 317.1 s | 317.0 s | 1.00× |
| Frames | 7 928 | 7 925 | 1.00× |
| Wall clock, whole batch | 914 s | 907 s (derived) | 1.01× |
| Throughput | 8.67 fps | 8.74 fps | 0.99× |
| CPU time, whole batch | 7 586 s | 2 103 s | 3.61× more on the DGX |
| CPU per delivered frame | 0.957 s | 0.265 s | 3.61× |
| Average cores | 8.32 | 2.35 | 3.53× |
| Peak cores | 19.5 | 4.5 | 4.33× |
| On P-cores | 51.5 % | 75.7 % | 0.68× |
| Average clock | 3.39 GHz | 2.83 GHz | 1.20× |
| Peak memory | 1.67–1.72 GB | 1.55–1.66 GB | ~1.05× |
| Energy | no counter on this box | 1.15 Wh total | — |
| Speed range | 7.02–12.03 fps | 6.23–11.40 fps | — |
Throughput — a dead heat.
Cores burned to get it — 3.5× apart.
CPU per delivered frame — 3.6× apart.
The DGX did not win. Same ten reels, same 317 seconds of video, same wall clock (914 s vs 907 s) — while spending 3.6× the CPU to get there. A 20-core workstation and a laptop tie, because the pipeline is serialised: throughput is set by one ffmpeg at a time, not by how many cores are available.
Two explanations fit the 3.6×, and the brief's screenshot can't distinguish them:
(a) the MacBook run encodes in hardware (VideoToolbox), which drops CPU while leaving wall clock
alone — consistent with its 1.15 Wh and 2.35 cores; or (b) the MacBook pipeline doesn't do the
triple encode. Note that (b) would be almost exactly the 3.1× measured here for collapsing three
encodes into one. Either way it points at the same fix. The test that settles it: ask that session for its
ffmpeg command lines, or run stageprof.py there.
Caveats worth keeping: the per-clip durations derived from the screenshot (18.5–42.9 s) only roughly overlap the ones measured here (25.2–40.7 s), so the row-to-row mapping is unknown even though the totals agree to 0.03 %. And if that run counted only its parent process rather than every ffmpeg child, the 3.6× is an artifact — though its 2.6–6.3 peak-core column says it was counting children.
| Reel | Length | Speed | Wall | CPU time | Avg cores | Peak cores | On P-cores | Avg GHz | Peak memory | Output |
|---|---|---|---|---|---|---|---|---|---|---|
| 01_fiction_platforms | 25.2 s | 7.55 fps | 83.4 s | 627 s | 7.52 | 18.8 | 50.4 % | 3.37 | 1.68 GB | 16.3 MB |
| 02_retcon | 35.2 s | 7.47 fps | 117.9 s | 946 s | 8.03 | 19.7 | 51.6 % | 3.39 | 1.70 GB | 21.0 MB |
| 03_jobber | 29.4 s | 9.82 fps | 74.8 s | 671 s | 8.98 | 19.3 | 52.1 % | 3.39 | 1.71 GB | 19.5 MB |
| 04_whoop | 37.5 s | 8.56 fps | 109.5 s | 916 s | 8.37 | 19.3 | 50.8 % | 3.38 | 1.70 GB | 22.3 MB |
| 05_out_on_your_back | 27.6 s | 7.83 fps | 88.0 s | 693 s | 7.87 | 19.4 | 51.0 % | 3.39 | 1.67 GB | 18.1 MB |
| 06_nobody_cares_about_facts | 28.5 s | 9.81 fps | 72.6 s | 643 s | 8.85 | 19.4 | 51.7 % | 3.39 | 1.67 GB | 16.2 MB |
| 09_burn_bridges | 28.8 s | 8.53 fps | 84.3 s | 707 s | 8.39 | 20.0 | 52.0 % | 3.39 | 1.72 GB | 12.2 MB |
| 10_infosys | 25.4 s | 7.02 fps | 90.4 s | 693 s | 7.66 | 20.0 | 51.6 % | 3.37 | 1.67 GB | 16.6 MB |
| 11_medium_risk | 39.0 s | 8.98 fps | 108.5 s | 953 s | 8.78 | 19.5 | 51.2 % | 3.39 | 1.69 GB | 20.1 MB |
| 12_corpus_back | 40.7 s | 12.03 fps | 84.5 s | 736 s | 8.71 | 19.7 | 53.0 % | 3.41 | 1.68 GB | 19.8 MB |
| 10 reels | 317.1 s | 8.67 fps | 914 s | 7586 s | 8.30 | — | 51.5 % | 3.39 | 1.72 GB | 182.1 MB |
Speed = finished frames (25 fps) per second of wall clock. There is no Energy column because
this box exposes no energy counter — no RAPL, no powercap, no hwmon energy sensor; nvidia-smi
reports GPU power only. CPU-seconds stands in for it.
Clip 01, every ffmpeg call and Python stage timed separately. Encoding is 95 % of it, spread over three generations of the same picture. The text raster is small in CPU but runs at exactly 1.00 core, so it costs ~20 % of the wall clock with 19 cores parked.
The full track is the whole machine. Peak process count across all ten builds was 2: one Python driver and one ffmpeg. Nothing in the pipeline ever runs two things at once.
Same 30 s of source. The fused graph — cut, grade, concat, speed and delivery encode in one ffmpeg — uses 3.1× less CPU, half the wall clock, writes a 28 % smaller file, and scores better against a lossless reference (SSIM 0.98305 vs 0.98241).
Transcription is the largest single compute item in the system: 2.43 CPU-hours per
talk, more than all ten renders put together. It runs large-v3 int8 on the CPU as 14
workers × 3 threads — 42 threads on 20 cores — holding 4.53 GB each, ~63 GB resident.
Measured: one 4.6-minute chunk costs 180 s wall and 625 s CPU alone. Perfectly packed that is 7.3 minutes
per talk; the production run took ~32. Raising cpu_threads to 20 made it far worse —
killed at 22 minutes, unfinished, after 8 500 CPU-seconds.
| Step | Today | After the measured fixes |
|---|---|---|
| Ingest ×10 talks | 0.4 h | 0.4 h |
| Transcribe ×10 | 5.3 h | 1.2 h |
| Track ×10 | 0.5 h | 0.5 h |
| Render ×100 reels | 2.5 h | ~0.5 h |
| Total wall clock | ≈ 8.8 h | ≈ 2.6 h |
≈ 51 CPU-hours of work against 480 core-hours a day available — about 11 % of the machine, stretched across 8.8 hours because almost nothing runs in parallel. The constraint is the code, not the hardware. The right-hand column is a projection composed from measured factors, not an end-to-end run. What breaks first at 100/day is human review, then ASR memory on anything smaller than this box, then storage at ~28 GB of source per day.
Orange steps are the ones that cost. Nothing in this chain touches the GPU — NVDEC and NVENC are both present in the ffmpeg build and both unused.
One Python driver, one ffmpeg, start to finish. Independent piece encodes run one after another.
Measured: 7.5–9.0 cores of 20 · 6 cuts serial 19.5 s vs 9.3 s together, same CPU.
crf 13 → crf 14 → crf 20, two of them existing only because intermediates go to disk.
Measured: 360.9 s CPU vs 116.8 s for one fused graph producing a better, smaller file.
large-v3 int8 on CPU, 14 workers × 3 threads = 42 threads on 20 cores. ctranslate2 here has no CUDA.
Logged: ~32 min of a saturated box per talk — more than all ten renders combined (15.2 min).
One 1080×1920 RGBA PNG per frame, written to disk and read straight back by ffmpeg.
Measured at 1.00 core · 1000 frames 10.7 s serial vs 1.2 s across 10 processes.
noise=c0s=3 sits in every grade, ahead of the first of three encodes. Also makes renders non-reproducible.
Measured: removing it is −16 % CPU and −16 % bitrate.
5.8 GB written to deliver 182 MB. Fine on NVMe, fatal on a laptop or a network share.
Per 30 s reel: 78 MB cuts + 78 MB concat + 45 MB base + ~35 MB PNG.
Memory (1.7 GB peak per reel of 120 GB), decode (dav1d does 4K AV1 at 25× realtime), audio (<3 s CPU), disk throughput.
h264_nvenc needs driver ≥ 610; this box has 580.173.02 (nvenc API 13.0 vs 13.1 required).
NVDEC works today: 21× less CPU than libdav1d, byte-identical output.
Full write-up, raw measurements and the harness itself are in the
accompanying zip: 00_summary.md … 07_scaling.md, data/, raw/.