A 25-second video that has no video file. It's a program, and you render it by asking it for frames.
Speech, one word at a time, cut into contour sections on a picture tube — standing in a dark room that the picture itself lights, revealed by a camera that pulls back over the take.
git clone https://github.com/RayVelez27/you-can-see-code
cd you-can-see-code && npm install && npm start # → http://localhost:8080Most video is a file. Someone opened an editor, arranged clips on a timeline, and exported an mp4. To change one word you reopen the project, re-export, and hope you still have the fonts.
Programmatic video is video as source code. The composition is a program
whose only input is a timestamp. Hand it t = 9.1 and it draws exactly what
should be on screen 9.1 seconds in. Hand it 750 different values of t and
you have a 25-second clip at 30fps.
That one change buys quite a lot:
- Data goes in, video comes out. The captions in this piece are a JSON file. Swap the JSON, get a different video — no editor, no re-export. Do it for a thousand rows and you have a thousand videos.
- It diffs. A composition is text. It reviews in a pull request, it has a blame history, and a regression is a line you can point at.
- It renders in parallel. Because any frame can be drawn without drawing the others first, a render farm can hand frame 1 to one machine and frame 700 to another and staple the results together.
That last property is the whole game, and it's the one that's easy to break.
The natural way to animate is to step: on each tick, move things a bit
further. position += velocity * dt. Every game loop and half the creative
coding on the internet works this way, and it's completely fine when a human
is watching in real time.
It falls apart the moment you render. A render farm asks four browsers for frames in whatever order they come free. If frame 700 only exists because frames 1–699 ran first, in that process, then each worker produces a different frame 700 and the finished video tears between four versions of itself at every worker boundary.
So the rule is: every value must be a closed form of t. No accumulation,
no internal clock, no unhashed randomness, no waiting on a network. Not "move
it a bit further" but "here is exactly where it is at 9.1 seconds."
const scene = await createScene({ canvas, captions });
scene.seek(9.1); // draws the frame at 9.1s. That's the whole API.seek(t) is pure. The same t always gives the same pixels, in any order,
however many times you ask.
This is one composition, written to that contract. The contract is what makes it portable — anything that can ask for a frame can render it:
| A browser | index.html runs it in real time with a requestAnimationFrame loop |
| A frame-scraper | tools/render.mjs — seek, screenshot, pipe to ffmpeg. ~110 lines |
| Remotion | useCurrentFrame() → seek(frame / fps) |
| A render farm | N workers, frames out of order, results identical |
None of those know about each other. They all just ask for frames.
npm start # → http://localhost:8080Space plays and pauses, arrow keys scrub, click starts audio if you've added any.
npm run verify # prove seek(t) is pure
npm run render -- --out out.mp4 # needs ffmpeg on PATHverify is the interesting one. It seeks ten times forward, then the same ten
in reverse, and asserts the hashes match. The reverse leg is the point:
seeking the same time twice in a row will happily agree while an
order-dependence bug is sitting right there. Only arriving at a frame from a
different direction exposes it. CI gates on this.
The piece is a caption engine. The words are data:
[{ "text": "Every", "start": 0.8, "end": 1.18 }]Replace src/captions.example.json, or pass ?captions=path/to/yours.json.
The duration follows the captions automatically. That's the whole
customisation surface for the common case — a lyric video, a quote card, a
title sequence, an ad variant per product.
To go from real speech to that JSON, point the transcriber at audio you have the rights to:
export DEEPGRAM_API_KEY=... # PowerShell: $env:DEEPGRAM_API_KEY="..."
npm run transcribe -- track.mp3 --seconds 25 --out src/captions.json
npm start # open /?captions=src/captions.jsonThat writes one caption per word, each carrying its real onset and its real end. Nothing is padded, deliberately: the fastest words in ordinary speech run about 0.08s — roughly two and a half frames at 30fps — and stretching one would push the next off its onset, which is the sync the whole piece is built on.
Deepgram is a convenience, not a dependency. Any tool that gives you word timings will do; the JSON shape above is the entire contract.
Drop an mp3 at assets/audio.mp3 and the page will play it, nudged back into
step whenever it drifts past 80ms.
The effect fits type to whichever axis runs out first. A phrase fits to width and rasterises short — thin enough that the contour bands eat it and what's left is a smear. A single word is limited by height, so it fills the frame.
This took two wrong turns to find. One line per caption was unreadable at every angle; two balanced lines was readable; one word is better than both. Only the longest words fit to width at all.
The same logic governs the stack's motion. It sways six degrees either side of flat rather than turning continuously, because the legible angles are narrow: flat-on the bands cut clean horizontal sections, but by twenty degrees the diagonal slices have sheared the words into ribbons that spell nothing. A full turn is right for a toy you type into and useless for captions, which exist to be read.
Five layers, composited in this order, and the order is load-bearing.
1 — the word. Rasterised to an offscreen canvas and blurred into a soft field. The blur is doing the real work: threshold a hard-edged glyph and all 288 bands land on the same outline, whereas a gradient crossed at 288 depths gives 288 different contours. Every word is baked once, at load, into its own render target — the blur ping-pongs between two buffers, exactly the state that would make a frame depend on its predecessor.
2 — the flow field. A simplex vector field behind the slices, giving the frame something to do in the silences and something for the bands to cut across. It samples 4D noise on a circle, because a straight line through 3D noise never returns to where it started and this layer has to loop with everything else.
3 — the slices. 288 instanced quads spread along Z and tilted 45°, each keeping only the fragments where the field crosses a threshold. They travel: the stack slides along its own axis, wrapped to exactly one band spacing, so band n arrives precisely where band n−1 was and the flow has no seam. Measured over two frames 0.08s apart, 12.6% of the picture changes while the word itself stands still. The stack is graded front-to-back, because 288 identical bands read as one flat overlay — fading the far ones lets the front rank read as a surface and the rest as depth behind it.
4 — the tube. Barrel-curved glass, scanlines, a roll bar sweeping the frame, phosphor duotone, vignette. The flutter is deliberately close to the frame rate — 36 Hz against 30 fps — so it folds down to a slow beat rather than reading as flicker. That beat is what sells the picture as unstable rather than merely textured.
5 — the room. The tube's picture is used twice: as the face of a monitor, and as the projection map of the spotlight in front of it. That second use is what sells the whole thing — the desk and keyboard are lit by the picture, so the light in the room changes when the word changes, and the word is legible upside-down across the floor. A lamp would have lit the same objects and none of it would have read as a display.
Field and type must be composited into one image before the tube shader runs. A flow field sitting behind the canvas in the page would stay perfectly flat and untinted while the type curved away from it, and the two would visibly not be the same picture.
Every cycle is a whole number over the take — contour flow, line-weight breath, sway, the noise circle, the roll bar and flutter — so nothing has to be faded out at the seam. Only the camera reveal deliberately does not return, because a reveal that loops is not a reveal.
Written down because they outlive this repo. If you are building anything frame-by-frame in a browser, these are the ones that will bite.
Alpha belongs in the colour, never in opacity. An element below opacity 1
is promoted to its own compositing layer and one at 1 is painted directly, so
animating opacity moves a thing in and out of a layer — and a layer
re-rasterises differently depending on what the compositor did before it. Two
frames that computed bit-identical values still differed across 120 pixels
until the alpha moved into rgba().
Seek inside a requestAnimationFrame, and wait for the next one. A WebGL
draw issued from an ordinary task may never reach the compositor, so a
screenshot returns the previous composite. Every captured frame comes back
identical while the canvas is perfectly correct.
A host-driven page must not also drive itself. If your page has a playback
loop it has to stop calling seek() when an external driver takes over, or it
stamps on every frame the driver asked for one tick later. Same symptom as
above, entirely different cause. That's why index.html takes ?paused=1.
Pin your pixel ratio. devicePixelRatio sizes render targets, so left
alone the output becomes a property of the machine that drew it — a retina
laptop and a render worker produce visibly different pictures.
Match your background to your floor. A ground plane ends at a horizon partway up the frame: above it you see the scene background, below it unlit floor. Two different blacks meeting on a hard line reads as banding rather than as a room.
Don't tone-map an already-graded picture. ACES is right for a scene holding real highlights and wrong here — the brightest thing in the room is a screen that already decided its own phosphor curve, and ACES desaturates highlights by design, bleaching the duotone to plain white.
Bloom only what should glow. At a threshold of 0.12 the flow field's own hairlines were over it, so the bloom smeared the whole screen and lifted the blacks — which costs the contrast and the colour, since a duotone's dark end only shows where the picture is actually dark.
src/scene.js the engine — createScene() → seek(t)
src/captions.example.json the built-in demo line
index.html the demo, also the GitHub Pages entry
tools/transcribe.mjs audio → captions via Deepgram
tools/render.mjs frames → mp4
tools/verify.mjs determinism gate
tools/serve.mjs static server
Three external dependencies at runtime, all from a CDN via importmap:
three.js and two of its addon passes. No build step — src/scene.js is the
file that runs.
This started as a template in a private video-template library, where compositions are HTML files a render farm seeks through. It was extracted into a standalone engine because the interesting part — the composition is a pure function of time — has nothing to do with that particular host, and the constraint is worth more than the piece.
MIT for the code. See LICENSE and ATTRIBUTION.md — it builds on three other people's sketches, two of which reached me without an author credit and are marked as such rather than guessed at.
No audio and no transcript of anyone else's work ships here. The built-in captions were written for this repo. If you point the transcriber at a recording, the words that come out are that recording's — mind whose they are before you publish the result.
