# 1. The machine, and how these numbers were taken

## Machine under test — NVIDIA DGX Spark (GB10)

| | |
|---|---|
| Arch | aarch64, Linux 6.17.0-1029-nvidia |
| CPU | 20 cores, no SMT — **10× Cortex-X925 @ 3.90 GHz (P)** + **10× Cortex-A725 @ 2.81 GHz (E)** |
| P-core ids | 5–9, 15–19 · E-core ids 0–4, 10–14 |
| GPU | GB10 Blackwell, driver 580.173.02, CUDA 13.0 — **idle at 3.77 W throughout every render** |
| Video engines | NVDEC (`av1_cuvid`, `h264_cuvid`), NVENC (`h264_nvenc`, `hevc_nvenc`, `av1_nvenc`) — **all unused by the pipeline** |
| ffmpeg | N-126207-g21bbd98e7b (2026-08-18), gcc 15.2 |
| Python | 3.x venv at `work/.venv` — numpy 2.5.3, Pillow 12.3.0, opencv-headless 5.0.0, onnxruntime 1.30.0 (CPU), faster-whisper 1.2.1, ctranslate2 4.8.2 (**no CUDA support compiled in**), **no torch** |

This is the **DGX report**. The MacBook session has to run the same measurement
on its own box: the numbers below are what an unconstrained 20-core machine
does, so they describe the *shape* of the workload, not the laptop's ceiling.

## What was measured, and how

Every reel was **re-rendered from scratch** under instrumentation
(`measure.py`). Nothing here is estimated from the earlier production run except
where a row explicitly says "logged".

| Metric | Source | Honest caveat |
|---|---|---|
| Wall | `time()` around the build process | +5–10 % vs an unobserved run (sampler overhead, see below) |
| CPU time | `getrusage(RUSAGE_CHILDREN)` delta — exact, includes every reaped `ffmpeg` | — |
| Avg cores | CPU time ÷ wall | — |
| Peak cores | max over 0.5 s windows of the process tree's CPU-time derivative, sampled at 20 Hz | a <0.5 s spike reads lower than it was |
| On P-cores | share of *running-thread samples* found on cpu 5–9/15–19 (`/proc/<pid>/task/<tid>/stat` field 39) at 20 Hz | sample-count weighted, not cycle weighted — ±2 pts |
| Avg GHz | per-cpu `scaling_cur_freq` weighted by that cpu's busy-tick delta from `/proc/stat` | system-wide; the box was otherwise idle (load avg 0.12 before start) |
| Peak memory | max over samples of **summed RSS across the whole process tree** | shared pages counted once per process |
| Energy | **not available on this box** | there is no RAPL / powercap / hwmon energy counter on GB10 Linux; `nvidia-smi` exposes GPU power only (3.77 W idle). The reference screenshot's "Energy (Wh)" column comes from Apple's `powermetrics`, which has no equivalent here. Rather than invent it, the column is dropped and **CPU-seconds is used as the energy proxy** — on a fixed-clock box it is proportional. |

**Sampler overhead:** the 20 Hz sampler walks every thread of the tree each tick
and costs ~0.4 of a core (~2 % of the box). Clip 01 took 83.4 s observed vs
76.7 s logged in production — the gap is sampler overhead plus run-to-run
variance. Treat wall times as ~5–10 % pessimistic and CPU times as exact.

**One production number is not comparable:** `02_retcon` was logged at 323.9 s
in the original batch but measures 117.9 s here. The original run had other work
on the box. The measured number is the real cost.
