Reel pipeline — compute report

Ten finished reels, re-rendered under instrumentation on an NVIDIA DGX Spark (GB10) · 2026-10-07

This is the DGX half of the picture. The MacBook session needs to run the same harness on its own box — these numbers describe an unconstrained 20-core machine.

2.11
CPU-hours
to make 317 s of video
8.67
fps
finished frames per wall second
8.3 / 20
cores used
the box is 58 % idle, always
95 %
of CPU is x264
the same picture, encoded 3×
0 W
GPU work
GB10 idles at 3.77 W throughout
32×
write amplification
5.8 GB written for 182 MB out

DGX Spark vs the MacBook run

Compared against the FINISHED CLIPS table that came with the brief — the other machine's run of the same ten reels. Its wall-clock column is derived (CPU time ÷ average cores), everything else is read straight off it.

MetricDGX Spark (GB10)MacBookRatio
Clips1010—
Video delivered317.1 s317.0 s1.00×
Frames7 9287 9251.00×
Wall clock, whole batch914 s907 s (derived)1.01×
Throughput8.67 fps8.74 fps0.99×
CPU time, whole batch7 586 s2 103 s3.61× more on the DGX
CPU per delivered frame0.957 s0.265 s3.61×
Average cores8.322.353.53×
Peak cores19.54.54.33×
On P-cores51.5 %75.7 %0.68×
Average clock3.39 GHz2.83 GHz1.20×
Peak memory1.67–1.72 GB1.55–1.66 GB~1.05×
Energyno counter on this box1.15 Wh total—
Speed range7.02–12.03 fps6.23–11.40 fps—

The result in three bars

DGX Spark (GB10) MacBook

Throughput — a dead heat.

DGX Spark (GB10)
8.67 fps
MacBook
8.74 fps

Cores burned to get it — 3.5× apart.

DGX Spark (GB10)
8.32 cores
MacBook
2.35 cores

CPU per delivered frame — 3.6× apart.

DGX Spark (GB10)
0.96 s CPU/frame
MacBook
0.27 s CPU/frame

The DGX did not win. Same ten reels, same 317 seconds of video, same wall clock (914 s vs 907 s) — while spending 3.6× the CPU to get there. A 20-core workstation and a laptop tie, because the pipeline is serialised: throughput is set by one ffmpeg at a time, not by how many cores are available.

Two explanations fit the 3.6×, and the brief's screenshot can't distinguish them: (a) the MacBook run encodes in hardware (VideoToolbox), which drops CPU while leaving wall clock alone — consistent with its 1.15 Wh and 2.35 cores; or (b) the MacBook pipeline doesn't do the triple encode. Note that (b) would be almost exactly the 3.1× measured here for collapsing three encodes into one. Either way it points at the same fix. The test that settles it: ask that session for its ffmpeg command lines, or run stageprof.py there.

Caveats worth keeping: the per-clip durations derived from the screenshot (18.5–42.9 s) only roughly overlap the ones measured here (25.2–40.7 s), so the row-to-row mapping is unknown even though the totals agree to 0.03 %. And if that run counted only its parent process rather than every ffmpeg child, the 3.6× is an artifact — though its 2.6–6.3 peak-core column says it was counting children.

Finished clips

ReelLengthSpeedWallCPU timeAvg cores Peak coresOn P-coresAvg GHzPeak memoryOutput
01_fiction_platforms25.2 s7.55 fps83.4 s627 s7.5218.850.4 %3.371.68 GB16.3 MB
02_retcon35.2 s7.47 fps117.9 s946 s8.0319.751.6 %3.391.70 GB21.0 MB
03_jobber29.4 s9.82 fps74.8 s671 s8.9819.352.1 %3.391.71 GB19.5 MB
04_whoop37.5 s8.56 fps109.5 s916 s8.3719.350.8 %3.381.70 GB22.3 MB
05_out_on_your_back27.6 s7.83 fps88.0 s693 s7.8719.451.0 %3.391.67 GB18.1 MB
06_nobody_cares_about_facts28.5 s9.81 fps72.6 s643 s8.8519.451.7 %3.391.67 GB16.2 MB
09_burn_bridges28.8 s8.53 fps84.3 s707 s8.3920.052.0 %3.391.72 GB12.2 MB
10_infosys25.4 s7.02 fps90.4 s693 s7.6620.051.6 %3.371.67 GB16.6 MB
11_medium_risk39.0 s8.98 fps108.5 s953 s8.7819.551.2 %3.391.69 GB20.1 MB
12_corpus_back40.7 s12.03 fps84.5 s736 s8.7119.753.0 %3.411.68 GB19.8 MB
10 reels317.1 s8.67 fps 914 s7586 s8.30— 51.5 %3.391.72 GB 182.1 MB

Speed = finished frames (25 fps) per second of wall clock. There is no Energy column because this box exposes no energy counter — no RAPL, no powercap, no hwmon energy sensor; nvidia-smi reports GPU power only. CPU-seconds stands in for it.

Where a reel's CPU goes

R2 cut + grade + encode
301.3 s CPU
R6 speed re-encode
115.6 s CPU
R11 final composite
112.9 s CPU
R2b b-roll encodes
68.0 s CPU
R9 text raster (1 core)
17.3 s CPU
R12/R13 QC scans
9.1 s CPU
R10 music synth
2.9 s CPU
audio (all stages)
2.1 s CPU

Clip 01, every ffmpeg call and Python stage timed separately. Encoding is 95 % of it, spread over three generations of the same picture. The text raster is small in CPU but runs at exactly 1.00 core, so it costs ~20 % of the wall clock with 19 cores parked.

Cores used per reel, out of 20

01_fiction_platforms
7.5 cores
02_retcon
8.0 cores
03_jobber
9.0 cores
04_whoop
8.4 cores
05_out_on_your_back
7.9 cores
06_nobody_cares_about_facts
8.8 cores
09_burn_bridges
8.4 cores
10_infosys
7.7 cores
11_medium_risk
8.8 cores
12_corpus_back
8.7 cores

The full track is the whole machine. Peak process count across all ten builds was 2: one Python driver and one ffmpeg. Nothing in the pipeline ever runs two things at once.

Three encodes, or one

as shipped measured alternative
Today: 3 encodes
360.9 s CPU
One fused graph
116.8 s CPU
Fused, no grain
97.7 s CPU
Fused, preset medium
92.3 s CPU

Same 30 s of source. The fused graph — cut, grade, concat, speed and delivery encode in one ffmpeg — uses 3.1× less CPU, half the wall clock, writes a 28 % smaller file, and scores better against a lossless reference (SSIM 0.98305 vs 0.98241).

Cost of a talk vs cost of the reels cut from it

Transcribe one talk (14 chunks)
8750.0 s CPU
Render ten reels
7586.0 s CPU
Face + shot track one talk
2030.0 s CPU
Demux + silence map
25.0 s CPU

Transcription is the largest single compute item in the system: 2.43 CPU-hours per talk, more than all ten renders put together. It runs large-v3 int8 on the CPU as 14 workers × 3 threads — 42 threads on 20 cores — holding 4.53 GB each, ~63 GB resident. Measured: one 4.6-minute chunk costs 180 s wall and 625 s CPU alone. Perfectly packed that is 7.3 minutes per talk; the production run took ~32. Raising cpu_threads to 20 made it far worse — killed at 22 minutes, unfinished, after 8 500 CPU-seconds.

A 100-clip day

StepTodayAfter the measured fixes
Ingest ×10 talks0.4 h0.4 h
Transcribe ×105.3 h1.2 h
Track ×100.5 h0.5 h
Render ×100 reels2.5 h~0.5 h
Total wall clock≈ 8.8 h≈ 2.6 h

≈ 51 CPU-hours of work against 480 core-hours a day available — about 11 % of the machine, stretched across 8.8 hours because almost nothing runs in parallel. The constraint is the code, not the hardware. The right-hand column is a projection composed from measured factors, not an end-to-end run. What breaks first at 100/day is human review, then ASR memory on anything smaller than this box, then storage at ~28 GB of source per day.

The pipeline

yt-dlp
AV1 4K + Opus
→ demux
pcm_s16le 16k · pcm_s24le 48k
→ Whisper large-v3
int8 CPU, 14×3 threads, ~32 min
→ YuNet face track
ONNX CPU, 12.4 cores, ~163 s
→ cut + grade + x264 crf13
one piece at a time
→ concat copy→ speed x264 crf14
2nd generation
→ text raster
Pillow, 1 core, PNG per frame
→ composite x264 slow crf20
3rd generation + AAC 256k
→ QC scans ×2

Orange steps are the ones that cost. Nothing in this chain touches the GPU — NVDEC and NVENC are both present in the ffmpeg build and both unused.

Bottlenecks

B1 · Nothing runs in parallel

One Python driver, one ffmpeg, start to finish. Independent piece encodes run one after another.

Measured: 7.5–9.0 cores of 20 · 6 cuts serial 19.5 s vs 9.3 s together, same CPU.

B2 · The picture is encoded three times

crf 13 → crf 14 → crf 20, two of them existing only because intermediates go to disk.

Measured: 360.9 s CPU vs 116.8 s for one fused graph producing a better, smaller file.

B3 · Transcription dwarfs everything

large-v3 int8 on CPU, 14 workers × 3 threads = 42 threads on 20 cores. ctranslate2 here has no CUDA.

Logged: ~32 min of a saturated box per talk — more than all ten renders combined (15.2 min).

B4 · Text layer is one Python core

One 1080×1920 RGBA PNG per frame, written to disk and read straight back by ffmpeg.

Measured at 1.00 core · 1000 frames 10.7 s serial vs 1.2 s across 10 processes.

B5 · Grain is added before compression

noise=c0s=3 sits in every grade, ahead of the first of three encodes. Also makes renders non-reproducible.

Measured: removing it is −16 % CPU and −16 % bitrate.

B6 · 32× write amplification

5.8 GB written to deliver 182 MB. Fine on NVMe, fatal on a laptop or a network share.

Per 30 s reel: 78 MB cuts + 78 MB concat + 45 MB base + ~35 MB PNG.

Not bottlenecks

Memory (1.7 GB peak per reel of 120 GB), decode (dav1d does 4K AV1 at 25× realtime), audio (<3 s CPU), disk throughput.

NVENC is blocked, not absent

h264_nvenc needs driver ≥ 610; this box has 580.173.02 (nvenc API 13.0 vs 13.1 required).

NVDEC works today: 21× less CPU than libdav1d, byte-identical output.

Downloads

Everything behind this page, including the measurement harness — run measure.py on the MacBook and the two sides become the same instrument.

Full write-up, raw measurements and the harness itself are in the accompanying zip: 00_summary.md … 07_scaling.md, data/, raw/.