# 7. What happens at 100 clips a day

The brief: *"if you have not even tried to push 100 clips a day then you have not
explored any constraints."* Fair. Here is the arithmetic from the measured
numbers, so the first real attempt starts from a budget instead of a guess.

Assume the shape we actually ran: **10 reels per source talk**, so 100 reels/day
means **10 talks/day**.

## Cost per unit, measured

| Unit | Wall (as it runs today) | CPU |
|---|---:|---:|
| One reel | 91 s | 759 s |
| One talk — ASR | ~32 min (23 % efficient) | 8 750 s |
| One talk — face/shot track | ~163 s | 2 030 s |
| One talk — demux + silence map | ~10 s | 24 s |
| One talk — ingest (4K master + proxy) | ~2.5 min | network-bound, ~0 CPU |

## A 100-clip day on this box, today

| | Wall |
|---|---:|
| Ingest ×10 | 0.4 h |
| ASR ×10 | **5.3 h** |
| Track ×10 | 0.5 h |
| Render ×100 | **2.5 h** |
| **Total, back to back** | **≈ 8.8 h** |

**Total CPU: ≈ 51 CPU-hours per day.** The DGX has 20 cores × 24 h = 480
core-hours available, so the *work* fits in about 11 % of the machine's day —
but because the pipeline serialises almost everything, it occupies **8.8 hours of
wall clock**. That is the whole finding: **the constraint is the code, not the
hardware.**

And 8.8 h assumes every reel is right the first time. It never is — the last two
weeks of reel work went to v4, v11 rounds. At three render rounds per reel the
render column alone goes to 7.5 h and the day stops fitting.

## The same day after the measured fixes

| | Wall | Basis |
|---|---:|---|
| Ingest ×10 | 0.4 h | unchanged, network-bound |
| ASR ×10 at 6 workers instead of 14 | **1.2 h** | measured per-chunk cost ÷ 20 cores |
| Track ×10 | 0.5 h | unchanged |
| Render ×100, fused graph + 3 concurrent | **~0.5 h** | measured 3.1× CPU and 2.1× wall factors |
| **Total** | **≈ 2.6 h** | |

Those render and ASR rows are **projections composed from measured factors**, not
end-to-end runs. They are the number to falsify first.

## What actually breaks first at 100/day

Ranked by what hits soonest:

1. **Human review, not compute.** 100 reels/day is 100 QC passes, 100 approvals,
   100 uploads. Nothing in this report helps with that.
2. **ASR memory, if it ever leaves this box.** 14 workers × 4.53 GB = **~63 GB
   resident**. The DGX shrugs; a 16 or 32 GB MacBook cannot run the pipeline as
   written at all. **This is the first thing the MacBook session should measure.**
3. **Storage growth.** Per talk: 1.27 GB master + 1.29 GB 48 kHz WAV + 143 MB
   16 kHz WAV + 43 MB proxy ≈ **2.75 GB**. Ten talks/day = **~28 GB/day,
   ~10 TB/year** of source alone, before any intermediates. The 120–150 TB in
   the server room covers several years; a laptop does not cover a week.
4. **Transient disk churn.** 5.8 GB written per 10 reels today → **58 GB/day**
   of intermediates that are written, read once and deleted. On a network
   filesystem in a render-farm setup this is the thing that falls over. Fixing
   B2 (fuse the encodes) removes most of it.
5. **Only then, cores.**

## Do we need the DGX, a server, or a render farm?

Three honest answers from the measurements:

**Today the DGX is overkill, because the pipeline can't use it.** 8.3 of 20
cores, 1.7 GB of 120 GB, 0 % GPU. Any 8-core machine with enough RAM would
produce nearly the same wall-clock times, because the limiting factor is that
one ffmpeg runs at a time.

**After the fixes, one box of this class is enough for 100 clips/day** — 2.6 h
of wall clock, with 21 h of headroom. You would reach ~400 reels/day before the
machine, rather than the operator, is the limit.

**Where the DGX is genuinely the right machine is the part that isn't built
yet.** The GPU is idle at 3.77 W from ingest to upload. The two workloads that
*want* a GPU — ASR (2.43 CPU-hours per talk today) and the encodes — are both
blocked by packaging, not by hardware:

- ctranslate2's aarch64 wheel has **no CUDA** (`get_cuda_device_count() → 0`),
  and torch isn't installed. A CUDA ASR path is a build job.
- NVENC fails on a **driver version check** (needs ≥ 610, box has 580.173.02),
  while NVDEC already works and uses 21× less CPU than `libdav1d`.

Fix those two and the DGX stops being a fast CPU box and starts being the thing
it was bought as.

### On the server room and distributed rendering

Both halves of this workload shard cleanly, which is the good news:

- **ASR** is already chunked at silence boundaries — 14 independent chunks per
  talk, each needing the 16 kHz WAV slice (~10 MB) and nothing else. Perfectly
  suited to the three servers (~120 cores, 192 GB). Ten talks' ASR spread over
  120 cores is **~12 minutes**, versus 5.3 hours today.
- **Rendering** is independent per reel. What a render node needs is the 4K
  master (1.27 GB) or, much better, **only the source ranges it will cut**
  (~30 s of 4K ≈ 60 MB) plus the 48 kHz audio slice and the cut JSON. Ship cut
  lists, never masters.
- **Do not** put the intermediates on a network share — that is the 58 GB/day of
  churn. Keep scratch node-local; ship only the finished MP4 back.

A RenderMan-style farm across idle laptops is the right *shape*, but it is
stage three. Stage one is the fused encode and the worker count — both measured,
both worth ~3–5× each, both changes to one file. Stage two is the GPU
packaging. Only after those does adding machines buy anything, because today a
second machine would also sit 58 % idle.
