# 8. DGX Spark vs the MacBook run

The brief arrived with a **FINISHED CLIPS** screenshot — the other machine's run
of the same ten reels. That table is the comparison baseline here. Its columns
are read straight off the image; **wall clock is derived** as CPU time ÷ average
cores, because the screenshot doesn't carry a wall column.

## Side by side

| Metric | DGX Spark (GB10) | MacBook | Ratio |
|---|---:|---:|---|
| Clips | 10 | 10 | — |
| Video delivered | 317.1 s | 317.0 s | 1.00× |
| Frames | 7 928 | 7 925 | 1.00× |
| Wall clock, whole batch | 914 s | 907 s *(derived)* | 1.01× |
| Throughput | 8.67 fps | 8.74 fps | 0.99× |
| **CPU time, whole batch** | **7 586 s** | **2 103 s** | **3.61× more on the DGX** |
| CPU per delivered frame | 0.957 s | 0.265 s | 3.61× |
| Average cores | 8.32 | 2.35 | 3.53× |
| Peak cores | 19.5 | 4.5 | 4.33× |
| On P-cores | 51.5 % | 75.7 % | 0.68× |
| Average clock | 3.39 GHz | 2.83 GHz | 1.20× |
| Peak memory | 1.67–1.72 GB | 1.55–1.66 GB | ~1.05× |
| Energy | no counter on this box | 1.15 Wh total | — |
| Speed range | 7.02–12.03 fps | 6.23–11.40 fps | — |

## The result

**The DGX did not win.** Same ten reels, same 317 seconds of video, the same
wall clock to within 1 % — while burning **3.6× the CPU** to get there.

A 20-core workstation and a laptop finish in a dead heat because **the pipeline
is serialised**. Throughput is set by how fast one ffmpeg can run, not by how
many cores the box has. The DGX's extra cores show up in the *peak cores* column
(19.5 vs 4.5) and nowhere in the result.

The peak-memory columns agreeing to ~5 % (1.67–1.72 GB vs 1.55–1.66 GB) is the
strongest evidence the two runs really are the same workload with the same
working set.

## Why 3.6×, and how to settle it

Two explanations fit, and the screenshot cannot distinguish them:

**(a) The MacBook run encodes in hardware.** VideoToolbox moves the encode to
the media engine: CPU collapses, wall clock barely moves. Its 2.35 average cores
and 1.15 Wh for the whole batch fit this well. If so, that session already did
what §6 recommends — and this box can't, because NVENC is blocked by the driver
version check.

**(b) The MacBook pipeline doesn't encode three times.** Collapsing three
generations into one was measured here at **3.1×** — close enough to 3.6× that
this alone could account for it.

Either way it points at the same fix, which is reassuring rather than
ambiguous. To settle it: ask that session for its ffmpeg command lines, or run
`raw/stageprof.py` there and compare the per-stage split.

## Caveats, stated plainly

- **Wall clock for the MacBook is derived**, not measured. Everything built on
  it (throughput, CPU-per-frame ratios) inherits that.
- **The row-to-row mapping is unknown.** Per-clip durations derived from the
  screenshot span 18.5–42.9 s against 25.2–40.7 s measured here, even though the
  totals agree to 0.03 %. Treat per-clip comparisons as unsafe; the batch
  totals are the solid part.
- **Accounting methods may differ.** If that run counted only its parent process
  and not every ffmpeg child, the 3.6× is an artifact. Its peak-cores column
  (2.6–6.3) says it was seeing ffmpeg's threads, so this is unlikely — but
  running `raw/measure.py` on the MacBook would remove the doubt entirely, since
  both sides would then be the same instrument.
- **The screenshot may be cropped.** Ten rows are visible; if the real table is
  longer, the totals change.
