Jonathan Earp Projects / 01
● Project 01 / 2026 ~4,100 lines of Python Runs off one command

PPipeline for EEdge TTracing and EExpression RRendering

Desmos was never meant to animate anything. This renders video inside it anyway, as equations. One command goes and finds a clip, traces it, rebuilds it in the calculator and posts the result.

4,494,696
Plays
2,594,952
Reach
294,155
Likes
99,914
Shares
01

What it does

Feed it a Family Guy clip. You get back a vertical video with a Desmos graph on top drawing the scene as line art, expression list and all, and the original clip playing underneath.

It runs off one command. autopilot.py --count 10 finds ten clips, works out which episode each is from, throws out the ones it can't caption, renders them, uploads them to YouTube and leaves the files and captions in a folder for me.

About 4,100 lines of Python across 10 modules. No framework, mostly because I didn't realise I'd want one until it was too annoying to add. It leans on yt-dlp, faster-whisper, playwright, opencv, numpy and ffmpeg to do the heavy parts.

01
Harvest

Pull candidate clips from search and a channel list.

02
Identify

Transcribe it, match the dialogue to an episode.

03
Dedupe

Drop anything I've already seen or posted.

04
Render

Trace to polylines, pack into expressions, screenshot.

05
Measure

Time the render, save where the clip came from.

06
Upload

Push to YouTube, then check it's actually live.

07
Stage

Write the files and captions for the manual IG post.

The next four sections are the render path, in order, with a demo each. Most of this was easier to understand once I could see it move.

02

Tracing a frame

Before Desmos can draw anything, a frame has to stop being pixels and start being lines.

Each frame gets decoded to grayscale, run through a Sobel filter, thresholded, then I walk the edge pixels one at a time to get outlines in order: a list of points you could draw without lifting the pen. Even a simple frame comes back with thousands of them, far more than Desmos will take, so the last step throws most of them away.

Try it Frame → polylines Step through the stages
Points — —

Those are the real algorithms running in your browser, on a cat I drew in code so there's nothing copyrighted in the repo. Same Sobel operator, same border-following, same Ramer–Douglas–Peucker simplification.

Push the tolerance slider around and you'll see the tradeoff I spent a while stuck on. Too tight and the expression lists get enormous and Desmos crawls. Too loose and the cat stops looking like a cat. I ended up around 1.5 px, which drops roughly 90% of the points and still keeps the shape.

03

Packing it into Desmos

Now I've got a pile of polylines per frame, and I need Desmos to show exactly one frame at a time.

The whole thing is list-valued parametric expressions with a single slider n choosing the frame. Except Desmos caps how long a list can be, so I can't dump every frame into one list. Frames get packed into chunks of about ten.

Which creates a new problem. With the slider on frame 14, chunks 1 and 3 still get evaluated, and if they index past the end of their list the whole graph errors and nothing renders.

Gating fixed it. Each chunk gets an expression that is 1 inside its own frame range and undefined everywhere else, and undefined doesn't error in Desmos, it just makes the curve vanish. The index is clamped alongside it so it always lands somewhere legal. No boundary special-casing anywhere.

G₁ = {1 ≤ n ≤ 10 : 1}
The gate. Equals 1 while the slider is inside this chunk. Undefined otherwise, which quietly blanks the curve instead of erroring.
J₁ = min(max(n − 0, 1), 10)
The clamped index. Always lands somewhere valid, even when n is nowhere near this chunk. This is what stops the out-of-range error.
R₁ = [O₁[J₁] … E₁[J₁]]
The slice. Pulls this frame's strokes out of the chunk's offset and end lists.
Try it One slider, three chunks Drag n across a boundary
Slider n n = 14
Chunk C₁ — frames 1–10
G₁ = · J₁ = ·
·
Chunk C₂ — frames 11–20
G₂ = · J₂ = ·
·
Chunk C₃ — frames 21–30
G₃ = · J₃ = ·
·

Exactly one chunk is defined at a time. The other two evaluate to nothing and disappear, still doing perfectly legal arithmetic while they do it.

04

The screenshot problem

Once the graph exists, I screenshot it frame by frame in a headless Chromium and stack it over the source clip with ffmpeg. This is where all the time goes, and where the worst bug was hiding.

I assumed the slow part was tracing. My tracer runs 8 worker processes on a 15-core machine, so the obvious move was more workers. I profiled it to work out how many to add.

One 13.8 s clip, 331 frames. Tracing is 4%. If I'd made it take literally zero time I'd have saved seven seconds out of 171.

Screenshotting was 94% of it, one frame at a time, 486 ms each, in a single browser page. So I parallelised the capture instead of the tracing and the same clip went from 171.6 s to 110.4 s.

Then the frames came out different.

The frames had been wrong the whole time

Obviously the parallel version was broken. Serial had been running for months. So I tested serial against itself: same clip twice on an idle machine, every frame diffed.

Zero of 40 frames differed. Perfectly deterministic, which I took as proof it was correct. All I'd shown was that it was consistent, and something broken the same way every time is consistent too.

The bug was in the settle logic. After moving the slider I wait for the canvas to stop changing before I screenshot it, and "stopped changing" meant 2 unchanged animation frames, about 33 ms. Desmos pauses for longer than that in the middle of an update. So I was catching the background redrawn and the character not.

Try it Where the screenshot lands Start at 2, then drag it up
Unchanged frames required stable = 2
The graph updating
What got screenshotted

To find the real answer I raised the threshold until the output stopped changing. stable=12 and stable=24 come out byte-identical, so 12 is converged and anything at or above 4 is fine.

Frames differing from the converged output, across 40 sampled frames
Settle thresholdFrames wrongVerdict
stable = 2 what I shipped 39 of 40 Months of renders with missing linework
stable = 40 of 40Correct
stable = 60 of 40Correct
stable = 120 of 40Converged — same bytes as stable = 24

Every video I'd made before this had incomplete linework on nearly every frame and nobody noticed, me included. The fix corrected the linework and made the render faster, since the parallel capture that exposed the bug was also the speed-up.

What a render costs

Render time scales with clip length at roughly 11–12×, but the spread is wide: 7.3× to 15.5× in the first batch, depending on how much the tracer reuses between frames. My first five measurements all landed between 10× and 13×, which made me think the constant was far tighter than it is.

Run 1 — before the capture fix Run 2 — after
20 clips against a 12× reference line. The two batches are different clips, so the gap between them isn't a clean measure of the speedup. The 1.55× comes from the same-clip test above. In practice: ten 15-second clips take about half an hour. Ten 45-second clips take an hour and a half. How long I'm sat waiting comes down almost entirely to how long the clip is.
05

Version by version

None of this was designed up front. Each version came out of a measurement that disagreed with what I expected.

Try it What each upgrade actually bought Click through, or use arrow keys

Both of the big wins came from measuring something I was already sure I understood. Tracing was the bottleneck; serial capture was correct. Neither survived a number being put on it.

06

Finding clips

Rendering is only half of it. Something has to go find things worth rendering, and that half had its own set of problems.

Shorts were invisible

yt-dlp --flat-playlist reports duration: None for every single YouTube Short. My length filter read:

if not vid or not dur: continue

which threw away 100% of them. Shorts are under 60 seconds by definition and I was searching for 8 to 45, so my best source was being dropped entirely while the log printed 40 results, 0 in band, which looks exactly like a source with nothing good in it.

Fixed by letting unknown-duration entries through to the full metadata fetch. Candidate pool went from 81 to 120 per run, and the first run after the fix pulled back 46 Shorts.

Nothing joined to anything

Performance data was keyed by Instagram's re-encoded copy (ig_<media_id>), source data by render tag (auto_s03e09), and nothing joined them. So I couldn't answer the question I'd built the whole thing for: does the original's view count predict anything?

Duration looks like the obvious join key. It isn't: 26 matches out of 68, and tightening the tolerance makes it worse (12 matches at ±0.05 s), because Instagram's re-encode shifts the duration by more than a frame.

What worked was dialogue. Transcribe both sides, match on n-gram overlap, and require a margin over the runner-up so near-ties get recorded as ambiguous instead of guessed. 56 of 68 linked, median overlap 0.95, median margin 0.94, zero ambiguous.

The limit that doesn't say "quota"

YouTube has two separate daily limits: an API project quota and a per-channel upload cap around 26 a day. The second comes back as HTTP 400 with the word "quota" nowhere in it. My handler checked for "quota", missed, and burned 30 doomed uploads in a row.

How long a clip can safely be

2 of 140 uploads got Content-ID blocked. Both were over 89 seconds. Nothing under 80 was ever touched. So I capped clips at 45 seconds, which costs nothing: the entire top ten by plays is under 32 seconds anyway.

07

The picker that didn't work

The original plan was a model that picks which clips to render, ranked by predicted views. I built it and it doesn't predict well enough to use. Here's why.

A retention model is real and holds up out of sample, cross-validated R² of 0.49. The problem is it's duration wearing a hat, and duration doesn't separate winners from losers: the top ten posts by plays run 8 to 32 seconds, the bottom ten 6 to 40.

Four correlations across 68 posts. Ranking clips by predicted retention would order them about as well as shuffling them.
Correlation with plays, and what happens if you try to predict them directly
RelationshiprWhat that means
Predicted retention → plays +0.02 Nothing at all, not even a weak signal
Measured retention → plays+0.40Real, but I only get it after posting
Duration → plays−0.15Noise
Duration → retention−0.74Strong, which is why the model is really just duration

Predicting plays directly is worse than useless: cross-validated R² of −0.89, so I'd have done better guessing the average every time. And the part of retention that does predict plays is only measurable once the post is live, which is too late to decide what to render.

So it doesn't forecast anything. It filters on length, on whether I can caption it, and on whether I've already seen it, which is where the real work turned out to be. It records one guess, how many views the original uploader got, labelled as an untested hypothesis with the test written and currently reporting n = 0. Past that it does volume: the hit rate is about 4.9%, so ten clips is roughly a 40% shot at one that lands.
08

Still broken

  • About a week of data, one account, one show. The correlations up there come from a tiny sample of one very specific kind of video, and none of it generalises.
  • The source-popularity signal is untested. The test is written and reports n = 0. I have not shown it works.
  • Instagram posting is still manual. YouTube is fully automatic. The platform that produced all 12,837 hours is the one I still upload by hand.
  • yt-dlp is a single point of failure. Any time YouTube changes something, harvest breaks. At least it breaks loudly and skips the run instead of posting garbage.
  • The more it posts, the more Content-ID hits I take. Two of 140 got blocked so far. The 45-second cap helps, but it doesn't make the problem go away.