PPipeline for EEdge TTracing and EExpression RRendering
Desmos was never meant to animate anything. This renders video inside it anyway, as equations. One command goes and finds a clip, traces it, rebuilds it in the calculator and posts the result.
What it does
Feed it a Family Guy clip. You get back a vertical video with a Desmos graph on top drawing the scene as line art, expression list and all, and the original clip playing underneath.
It runs off one command. autopilot.py --count 10 finds ten clips,
works out which episode each is from, throws out the ones it can't caption,
renders them, uploads them to YouTube and leaves the files and captions in a
folder for me.
About 4,100 lines of Python across 10 modules. No framework,
mostly because I didn't realise I'd want one until it was too annoying to add.
It leans on yt-dlp, faster-whisper,
playwright, opencv, numpy and
ffmpeg to do the heavy parts.
Pull candidate clips from search and a channel list.
Transcribe it, match the dialogue to an episode.
Drop anything I've already seen or posted.
Trace to polylines, pack into expressions, screenshot.
Time the render, save where the clip came from.
Push to YouTube, then check it's actually live.
Write the files and captions for the manual IG post.
The next four sections are the render path, in order, with a demo each. Most of this was easier to understand once I could see it move.
Tracing a frame
Before Desmos can draw anything, a frame has to stop being pixels and start being lines.
Each frame gets decoded to grayscale, run through a Sobel filter, thresholded, then I walk the edge pixels one at a time to get outlines in order: a list of points you could draw without lifting the pen. Even a simple frame comes back with thousands of them, far more than Desmos will take, so the last step throws most of them away.
Those are the real algorithms running in your browser, on a cat I drew in code so there's nothing copyrighted in the repo. Same Sobel operator, same border-following, same Ramer–Douglas–Peucker simplification.
Push the tolerance slider around and you'll see the tradeoff I spent a while stuck on. Too tight and the expression lists get enormous and Desmos crawls. Too loose and the cat stops looking like a cat. I ended up around 1.5 px, which drops roughly 90% of the points and still keeps the shape.
Packing it into Desmos
Now I've got a pile of polylines per frame, and I need Desmos to show exactly one frame at a time.
The whole thing is list-valued parametric expressions with a single slider
n choosing the frame. Except Desmos caps how long a list can be,
so I can't dump every frame into one list. Frames get packed into
chunks of about ten.
Which creates a new problem. With the slider on frame 14, chunks 1 and 3 still get evaluated, and if they index past the end of their list the whole graph errors and nothing renders.
Gating fixed it. Each chunk gets an expression that is 1 inside its
own frame range and undefined everywhere else, and undefined doesn't error in
Desmos, it just makes the curve vanish. The index is clamped alongside it so it
always lands somewhere legal. No boundary special-casing anywhere.
n is nowhere near this chunk. This is what stops the out-of-range error.Exactly one chunk is defined at a time. The other two evaluate to nothing and disappear, still doing perfectly legal arithmetic while they do it.
The screenshot problem
Once the graph exists, I screenshot it frame by frame in a headless Chromium and stack it over the source clip with ffmpeg. This is where all the time goes, and where the worst bug was hiding.
I assumed the slow part was tracing. My tracer runs 8 worker processes on a 15-core machine, so the obvious move was more workers. I profiled it to work out how many to add.
Screenshotting was 94% of it, one frame at a time, 486 ms each, in a single browser page. So I parallelised the capture instead of the tracing and the same clip went from 171.6 s to 110.4 s.
Then the frames came out different.
The frames had been wrong the whole time
Obviously the parallel version was broken. Serial had been running for months. So I tested serial against itself: same clip twice on an idle machine, every frame diffed.
Zero of 40 frames differed. Perfectly deterministic, which I took as proof it was correct. All I'd shown was that it was consistent, and something broken the same way every time is consistent too.
The bug was in the settle logic. After moving the slider I wait for the canvas to stop changing before I screenshot it, and "stopped changing" meant 2 unchanged animation frames, about 33 ms. Desmos pauses for longer than that in the middle of an update. So I was catching the background redrawn and the character not.
To find the real answer I raised the threshold until the output stopped
changing. stable=12 and stable=24 come out
byte-identical, so 12 is converged and anything at or above 4 is fine.
| Settle threshold | Frames wrong | Verdict |
|---|---|---|
| stable = 2 what I shipped | 39 of 40 | Months of renders with missing linework |
| stable = 4 | 0 of 40 | Correct |
| stable = 6 | 0 of 40 | Correct |
| stable = 12 | 0 of 40 | Converged — same bytes as stable = 24 |
Every video I'd made before this had incomplete linework on nearly every frame and nobody noticed, me included. The fix corrected the linework and made the render faster, since the parallel capture that exposed the bug was also the speed-up.
What a render costs
Render time scales with clip length at roughly 11–12×, but the spread is wide: 7.3× to 15.5× in the first batch, depending on how much the tracer reuses between frames. My first five measurements all landed between 10× and 13×, which made me think the constant was far tighter than it is.
Version by version
None of this was designed up front. Each version came out of a measurement that disagreed with what I expected.
Both of the big wins came from measuring something I was already sure I understood. Tracing was the bottleneck; serial capture was correct. Neither survived a number being put on it.
Finding clips
Rendering is only half of it. Something has to go find things worth rendering, and that half had its own set of problems.
Shorts were invisible
yt-dlp --flat-playlist reports duration: None for
every single YouTube Short. My length filter read:
if not vid or not dur: continue
which threw away 100% of them. Shorts are under 60 seconds by definition and I
was searching for 8 to 45, so my best source was being dropped entirely while the
log printed 40 results, 0 in band, which looks exactly like a source
with nothing good in it.
Fixed by letting unknown-duration entries through to the full metadata fetch. Candidate pool went from 81 to 120 per run, and the first run after the fix pulled back 46 Shorts.
Nothing joined to anything
Performance data was keyed by Instagram's re-encoded copy
(ig_<media_id>), source data by render tag
(auto_s03e09), and nothing joined them. So I couldn't answer the
question I'd built the whole thing for: does the original's view count predict
anything?
Duration looks like the obvious join key. It isn't: 26 matches out of 68, and tightening the tolerance makes it worse (12 matches at ±0.05 s), because Instagram's re-encode shifts the duration by more than a frame.
What worked was dialogue. Transcribe both sides, match on n-gram overlap, and require a margin over the runner-up so near-ties get recorded as ambiguous instead of guessed. 56 of 68 linked, median overlap 0.95, median margin 0.94, zero ambiguous.
The limit that doesn't say "quota"
YouTube has two separate daily limits: an API project quota and a per-channel upload cap around 26 a day. The second comes back as HTTP 400 with the word "quota" nowhere in it. My handler checked for "quota", missed, and burned 30 doomed uploads in a row.
How long a clip can safely be
2 of 140 uploads got Content-ID blocked. Both were over 89 seconds. Nothing under 80 was ever touched. So I capped clips at 45 seconds, which costs nothing: the entire top ten by plays is under 32 seconds anyway.
The picker that didn't work
The original plan was a model that picks which clips to render, ranked by predicted views. I built it and it doesn't predict well enough to use. Here's why.
A retention model is real and holds up out of sample, cross-validated R² of 0.49. The problem is it's duration wearing a hat, and duration doesn't separate winners from losers: the top ten posts by plays run 8 to 32 seconds, the bottom ten 6 to 40.
| Relationship | r | What that means |
|---|---|---|
| Predicted retention → plays | +0.02 | Nothing at all, not even a weak signal |
| Measured retention → plays | +0.40 | Real, but I only get it after posting |
| Duration → plays | −0.15 | Noise |
| Duration → retention | −0.74 | Strong, which is why the model is really just duration |
Predicting plays directly is worse than useless: cross-validated R² of −0.89, so I'd have done better guessing the average every time. And the part of retention that does predict plays is only measurable once the post is live, which is too late to decide what to render.
n = 0. Past that it does volume:
the hit rate is about 4.9%, so ten clips is roughly a 40% shot at one that
lands.
Still broken
- About a week of data, one account, one show. The correlations up there come from a tiny sample of one very specific kind of video, and none of it generalises.
- The source-popularity signal is untested. The test is written and reports
n = 0. I have not shown it works. - Instagram posting is still manual. YouTube is fully automatic. The platform that produced all 12,837 hours is the one I still upload by hand.
yt-dlpis a single point of failure. Any time YouTube changes something, harvest breaks. At least it breaks loudly and skips the run instead of posting garbage.- The more it posts, the more Content-ID hits I take. Two of 140 got blocked so far. The 45-second cap helps, but it doesn't make the problem go away.