The Explainer Video Is the New TL;DR
An agent researched, wrote, voiced, animated and checked a 4:25 explainer on how coding agents work. Here is the complete pipeline behind it, what it cost, and what it still cannot check.
The video above is the experiment: 4:25, from pressing Enter to KV caches and reinforcement learning, in this blog's colors. Nobody recorded a word or drew a frame. An agent researched the topic, wrote the narration, voiced it on my laptop, timed every caption to the spoken word, animated it in React and checked the result. This article is the factory that made it.
| β‘ | TL;DR What. A small video studio that turns a topic into a narrated, animated explainer with burned-in captions. The first output is "How AI Coding Agents Actually Work", 4:25, five levels from ELI5 to PhD. Why. Andrej Karpathy calls bespoke explainer videos the output format he is "most bullish on". I think the reading interface is breaking: people ask one AI to summarize what another AI wrote, and the human drops out. How. One Claude Code session on Claude Opus 5.5 at max effort. Sourced script in near-ASD-STE100 English, a free local voice (Kokoro), word timing with Whisper, Remotion for the animation, procedural sound, automated QA. Cost. $0 for the voice and the renderer licence, plus local compute. ElevenLabs is implemented and unused. What I did not prove. That anyone understands more from the video than from a page. No viewer study exists here yet. |
| π‘ | ELI5 You can ask a robot to explain something. Usually it gives you a long page, and you ask another robot to make the page shorter. Now the first robot makes a short cartoon instead, with a voice and subtitles. Each sentence is short. Each picture appears exactly when its word is spoken. Change one sentence, and only that part gets a new voice. |
A reader who only asks for the TL;DR
Here is the pattern I keep seeing, in my own work too. An agent writes a report, because reports are cheap now. The person who should read it asks a second agent for a TL;DR. The TL;DR gets an ELI5. Three generations later, the person who owns the decision has never opened the reasoning. Guides, docs, PDFs, chat messages and even interactive HTML pages share the same weakness: they wait for a reader who increasingly does not come.
That is my view, not a finding. The video is the test of a different interface.
On October 2 (UTC), Andrej Karpathy posted a ladder of output formats for understanding what language models produce. It drew more than six million views in under two days. Rung one: ask the model to write in ASD-STE100, a controlled language from aerospace maintenance, or "80% of the way to ASD-STE100". Rung two: ask for a diagram. Rung three: ask for HTML. Top rung: "The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic." His example prompt: "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration." His verdict: "This is actually starting to work!"
His summary is the part I care about: as models do more of the legwork, "a lot more of our work will rise up the abstractions into oversight and understanding." Oversight needs an interface people will actually use. A four-minute video might be one.
The post names no model and no animation tool. Three days earlier, Latent Space's AINews had already run the headline "Opus 5.5 is good at explainer videos" (September 29, 2026). Anthropic's own Opus 5.5 announcement claims nothing about animation; its only graphics remark is that in one tester's game-building test "Opus 5.5 scored higher than any other model on the strength of its graphics and polish." So treat "Opus 5.5 makes great animations" as a community observation. The video above is my own data point.
Rung one: ASD-STE100, and why the script uses the soft setting
ASD-STE100 Simplified Technical English is a real standard, owned by ASD and maintained by its STE Maintenance Group. The current version is Issue 9, dated January 15, 2025. It has 53 writing rules and a dictionary of 875 approved words; the PDF is free after a short form. The rules read like a style guide written by someone who has seen an engine fail: procedural sentences of at most 20 words, descriptive sentences of at most 25, no more than six sentences per paragraph, active voice, one instruction per sentence, the same word for the same thing. Its copyright notice forbids reproducing the standard, so I paraphrase the rules and copy none of the dictionary.
The first warning sign came from the post itself. Karpathy attached a slick one-sheet about the standard. Every number on it is right. Two dictionary rows are wrong when you check them against the Issue 9 PDF: the sheet marks "approximately" as not approved and tells you to use ABOUT, but Issue 9 approves APPROXIMATELY and reserves ABOUT for "concerned with"; and the sheet marks TEST as an approved verb, while Issue 9 approves TEST only as a noun. A clean diagram can still teach the wrong row. The STE Maintenance Group says it bluntly: "Plausibility must not be confused with verified compliance."
The second warning is about content. Lucian Ghinda tested explanations of code: with the full ASD-STE100 instruction, Claude missed 22 of 47 facts against a control, a 46.8% loss; with a loose "Simple Technical English" instruction it missed 4 of 47. Strict style can delete meaning. That result is why this narration uses Karpathy's softened "80%".
So the narration follows the spirit and a checker enforces a subset. My lint flags sentences over 20 words, passive and progressive verb forms, crowded beats, semicolons, and common complex words that have plain alternatives. It is a heuristic, not a conformance checker. By the lint's own count, the final script has 90 sentences, all passing, a mean of 7.0 words, a maximum of 20, and 633 words in total. Short sentences also help subtitles and translation. The evidence there is mixed, though. Older studies found gains for non-native readers, modern machine translation benefits less, and I found no published study of STE with LLM translation or with Spanish. Using a written-text standard for speech is my extrapolation.
The pipeline at a glance
Facts before frames
Before any animation, five research agents ran in parallel: Karpathy's thread and the reactions (one of them pulled 1,008 quote posts and 1,408 replies through a public API), the ASD-STE100 standard, animation tooling, voices and their licences, and how Claude Code and Codex actually work. The last one produced the fact sheet the narration is allowed to stand on. It covers the agent loop, permission modes, sandboxes, context windows, compaction, prompt caching, KV caches, prefill and decode, and training. Every quote was machine-checked against a saved copy of its source, with zero mismatches.
The fact sheet still caught me. My first script said "The program asks you before risky actions." Claude Code v2.1.283 and later starts in auto mode, where a second model reviews most actions and blocks the ones it judges risky. My script said cloud runs happen "in a container". Both Claude Code on the web and Codex Cloud now describe isolated machines. I wrote that a tool call "must fit a schema". Strict schema enforcement is optional. I paraphrased the reinforcement-learning claim. The source is specific: in May 2025 OpenAI said codex-1 was trained with RL on real-world coding tasks so it "can iteratively run tests until it receives a passing result." Four lines changed. The pictures had been confident with the old wording. Confidence was the bug.
The script: beats and markup
The narration lives in one JSON file as beats: one to four short sentences each, grouped into scenes. Captions and speech often need different text, so beats carry a small markup:
{"id": "l2-2", "text": "It loads its instructions, your project rules from {CLAUDE.md|Claude dot M D} or {AGENTS.md,|Agents dot M D,} and a list of tools."}
The caption shows CLAUDE.md. The voice says "Claude dot M D". A spoken-only segment like {|[whispers] } carries an audio tag for voices that support tags and disappears automatically for voices that do not. The markup keeps a map from every caption word to the exact characters the voice speaks, which is what makes word-level timing possible later.
The voice: why a free local model won
I built the voice layer for three providers and expected to pay for the best one. ElevenLabs released Eleven v4 on September 28, 2026, and it sits at number one on the Artificial Analysis speech arena with an Elo of 1,321. The API list price is $0.08 per 1,000 characters, about $0.36 for a five-minute narration. One trap matters for anyone publishing: ElevenLabs output from the free tier is non-commercial only and must carry "elevenlabs.io" or "11.ai" in the title. For a professional blog, it is a paid plan or nothing.
Then I listened to Kokoro. Kokoro-82M is an open-weights model under Apache-2.0 that runs on the laptop. Its arena Elo is 1,064, far below v4. Its voice table grades af_heart "A", the only voice with that grade. My verdict after the samples: Kokoro af_heart is great, so there is no need to pay ElevenLabs for this purpose. The whole draft narration, 38 beats plus word alignment, took about 40 seconds on an M4 Max. The ElevenLabs path stays in the code, built to the API documentation, and nothing in this video exercised it.
Two other notes from the voice research. OpenAI deprecated tts-1, tts-1-hd and the dated gpt-4o-mini-tts snapshots on October 1, 2026, with removal on January 6, 2027, so I did not build on them. And every beat is cached by a hash of its provider, voice settings and spoken text, so editing one sentence regenerates only its beat. The four fact corrections cost four short syntheses, not a new recording session.
From a spoken word to a frame
Kokoro returns audio, not timestamps. MLX Whisper (large-v3-turbo) listens to each beat and returns word times. Because the script is known, the aligner is timing text we already wrote rather than discovering it. Those word times map back to characters and then to caption words, so CLAUDE.md inherits the span of the spoken "Claude dot M D". ElevenLabs returns character timestamps directly and skips this step.
Everything lands in one timeline file, and the scenes never keep their own clocks. A scene asks:
const {at} = useCues();
// The frame where the narrator says "tool" in beat l3-2:
<Stamp at={at('l3-2', 'tool')} text="TOOL CALL" color={C.purple} x={1450} y={420} />
When the narration moves, the animation moves with it. That cuts both ways. During the fact pass I changed "The program asks you" to "It can ask you", and the render died, because an animation was still waiting for the word "asks". The pipeline now runs a static cue check before every render: it reads every at('beat', 'word') in the scene code and confirms the word is actually spoken.
The look: a reusable Tokyo Night kit
The animation is React, rendered by Remotion. Opus 5.5 wrote the code in the same session. The reusable library is 1,136 lines of TSX: a synthwave grid with particles, scanlines and grain; a HUD with level progress and timecode; the translucent vanja.io watermark; word-by-word captions; a terminal that types; HUD panels; stamps for DENIED, DONE and CACHE HIT; neon beams that carry packets; level title cards; glitch transitions; token chips; meters; and a slow camera drift so no scene sits still. The cast is two generic characters: a neural orb for the model and a small robot for the harness. This video's scenes add 1,273 lines, and the Python pipeline is 1,393. The colors come from the same design tokens as this site.
Karpathy asked for a 3Blue1Brown style. This is not that style, on purpose. The format bet is the bespoke explainer, not the chalkboard. A frame that matches the site reads as one publication.
Why Remotion, not Manim or HyperFrames
Three engines were serious options. Manim Community Edition 0.21.0 is the 3Blue1Brown lineage, MIT-licensed, excellent for mathematics, and its new Typst support removes the LaTeX dependency. HyperFrames 0.8.116 renders HTML and GSAP animations, is Apache-2.0 and was built for agents, but it is about seven months old and sends telemetry by default. Remotion 4.0.532 is the mature one: it publishes twelve official agent skills and a Claude Code plugin. Its licence says individuals and small companies "are allowed to use Remotion to create videos for free (even commercial)", where small means up to three people, which fits a one-person blog. A three-second test rendered in about ten seconds on this machine. For the record, at least one other builder answered Karpathy with the same stack: "Made with Claude Code, Remotion, and local Kokoro."
Sound, render and QA
The sound is synthesized, not licensed. A small numpy program generates the effects (whoosh, glitch, typing, ping, deny, success, impact, riser, data chirps) and a dark synthwave bed that ducks under the voice. No third-party audio is in the file.
Remotion renders H.264 at 1080p and 30 frames per second. ffmpeg then runs two-pass EBU R128 normalization to β16 LUFS integrated and β1.5 dBTP, with AAC audio and fast start for streaming. A half-resolution preview took 3 minutes 40 seconds. The full-resolution render took 12 minutes 26 seconds for 7,967 frames.
One render taught a lesson that no unit test would. To make rendering faster, I swapped an expensive film-grain filter for a noise tile that jumped to a random position every frame. Rendering got faster, and the file grew to 2.1 GB, because when every pixel changes every frame the encoder has nothing to reuse. Static grain looks almost the same and costs almost nothing. The final master is 210 MB. The file on this page is a 98 MB two-pass encode, because the blog's host caps uploads at 100 MB, and side-by-side crops of the two look the same.
QA is a list, not a mood. It checks the cues, loudness, black and frozen frames, and duration against the five-minute cap. The agent then reads contact sheets every four seconds and full-resolution frames of each scene. That pass caught a code line cut off at "autoc", labels stacked on top of each other, a level card bleeding into the next diagram, and a subtitle that split "pull" from "request". It cannot catch everything: the agent sees stills, not motion. Karpathy named that gap in August 2026, when he noted that models "can't easily audit their work because they aren't able to efficiently and natively perceive videos". Fourteen automated tests cover the markup, alignment, captions, lint, cue checker, cache keys and sound.
Captions in two languages
The English captions are burned into the picture, highlighted word by word, so they work in any player. The pipeline also writes SRT and VTT files that follow the Netflix timed-text norms: at most 42 characters per line and two lines, one sentence per cue when it fits, and phrase-aware splits when it does not. Colombian Spanish subtitles come from a per-beat translation file, timed inside each beat, in Colombian style: "usted", "computador", and the loanwords developers there actually use, such as pull request and prompt. The video above carries the English captions; the Spanish file is ready for a Spanish cut. A Colombian Spanish narration is the next step, and it needs a paid voice, so that decision waits.
One command, three agents
The whole procedure lives in one runbook. A /video command in Claude Code points at it, and generated skills give Codex and Grok the same entry point, without a second copy of the rules. Under the command sits a small CLI:
video.py lint how-coding-agents-work # narration style check
video.py tts how-coding-agents-work # voice, alignment, timeline, captions
video.py cues how-coding-agents-work # every animation cue is actually spoken
video.py render how-coding-agents-work # Remotion + loudness normalization
video.py qa how-coding-agents-work # probes, loudness, frame scans, contact sheetsThe cost for this video: $0 for the voice, $0 for the renderer licence, local compute for everything else, and the subscription that runs the agent.
What I did not prove
- Comprehension. No one has watched this with a quiz afterwards. If the video is clearer than a page, that is my judgment, not a measurement.
- ElevenLabs in practice. The provider is implemented from the documentation and was not used here.
- Motion quality. The agent reviewed still frames. A flicker between two frames can pass that review.
- The STE effect. The lint score measures my own checks, not ASD-STE100 compliance, and the translation evidence is mixed.
- Vendor numbers. Arena rankings, prices and model limits are as published on October 3, 2026.
What stays expensive
Karpathy's phrase for this class of output is "large, custom, discardable software artifacts". The MP4 is the cheap object now. Changing one sentence means one new audio file, a new timeline and a 12-minute render. The 264 checked quotes are the expensive part. Even with a script verifying every one, a second pass still found four stale claims. The video can be regenerated tonight from the timeline. A claim that has left its source cannot be regenerated at all.
That is where the work moves when, in Karpathy's words, it rises "into oversight and understanding". The agent writes, voices, animates and checks the pixels. The human job is the part the pixels hide: deciding what is true, and noticing when a clean surface has not been checked.