How to make this faster — in pictures

Four diagrams. Every number was measured on your box today.

Picture 1

A reel is never right first time — you go round six times. This is the waiting.

Today — six full renders 9 min 06 s

With a proxy — six previews, then one real render 2 min 19 s

05 min9 min
full render · 91 s each
proxy preview · ~8 s each
one final render, at the end

A proxy is the same edit at 540×960, no grain, fastest preset. You judge the cut, the timing and the words on it. Full quality runs once, on the version you approved.

The number that matters isn't the total — it's the gap between "change that" and "let me see": 91 seconds becomes 8.

Picture 2

You change one word of text. This is what the machine redoes.

Today 629 s of CPU

What actually needed doing 139 s — 22 %

re-cut & re-grade the video · 301 s
b-roll re-encode · 68 s
speed pass · 116 s
the new text · 17 s
putting it together · 113 s
quality checks · 9 s

Measured on 01_fiction_platforms. The three red bars are the picture being cut from the 4K file, graded and encoded again — identically. 78 % of the work is thrown away.

The fix: name each cut after its own ingredients — which seconds of the talk, which crop, which grade. Same ingredients, reuse the file already on disk. Most of your rounds change words and timing, not the picture, so most rounds stop touching video at all.

Picture 3

Why it redoes all that: three different jobs are welded into one script.

Today — one blob build.py

understand the talk

transcribe · track the face

decide the reel

which seconds · what text · which b-roll

render it

cut · grade · encode ×3

Touch any part — the whole thing runs.

Better — three parts each runs only when its own input changes

1 · Understand

once per talk, then never again
~40 min → a talk-index.json you keep forever

2 · Decide

the only part you iterate on
a short spec: ranges, words, b-roll, grade. Editing it costs nothing

3 · Render

proxy while deciding, full once at the end
same spec in → same video out, every time

Why "same spec → same video" matters: the grade currently sprinkles random film grain before compression, so the same reel rendered twice gives two different files. Nothing can be cached, and you can't tell whether a change changed anything. One-line fix — and it's what makes Picture 2 possible.

Picture 4

Which machine does what.

DGX Spark once per talk + model work

Transcribe, track faces — and the thing it's actually for: a local model that checks the b-roll matches what he's talking about, before you see it.

Runs 10× a day, not 100×.

Whatever you're sitting at the loop you live in

Proxies, 8 seconds each. Any laptop does this — it's 540×960 on the cheapest settings.

Runs hundreds of times a day.

The 3 servers you own ~120 cores, already paid for

Final renders, one reel per worker. 100 reels over 120 cores is minutes, not hours.

Send them the spec, not the footage.

Nothing here needs buying. The 4K masters stay put — a render worker only needs the 30-second stretches it will actually cut (~60 MB), not the 1.27 GB file.

In one line

91 s → 8 s

to see a change, by reviewing on proxies

78 % → 0 %

of a text change wasted re-rendering video that didn't change

Both come from two moves: make renders repeatable, then only redo what changed. Neither needs new hardware, and both matter more than every core-count argument in the report — the thing you're short of is rounds, not cores.

← back to the measurements