Skip to content
zhanglinghao edited this page Oct 2, 2026 · 3 revisions

English · 中文

Can the agent watch the video?

Not directly. It renders still frames and looks at them: contact sheets for composition, strips of consecutive frames for timing, crops for detail. For sound it reads numbers: transcripts, timestamps, loudness, onsets. Gemini can optionally read a whole video, but at about one frame per second it only catches story-level problems and rough sync. That gap is why the checks exist; see How It Works.

Do I need a GPU?

No. Rendering runs in headless Chrome or in Manim; a GPU makes WebGL faster. On a machine without one, bin/vh doctor tells you whether software WebGL works, and you render with --soft-gl (slower, still deterministic). Local voiceover (Qwen3-TTS) needs Apple Silicon; elsewhere use edge or a cloud voice. Blender, used only for path-traced shots, renders on the CPU or a GPU; on a Mac, Metal was several times faster than the CPU in the smoke tests.

Which agents work?

Claude Code and Codex work directly: the installer registers the skill for both, and AGENTS.md is an identical copy of CLAUDE.md. Any coding agent that reads markdown and runs shell commands can follow the same docs. The showcase films were made with Claude Opus 5.5.

Does it cost money? Which parts use paid APIs?

The pure-code path is free: rendering, local Qwen3-TTS, say, edge, music, sound effects, mixing and QA. Paid or metered: cloud voices (gemini has a free tier; dashscope, elevenlabs), --align gemini (free tier, or about $0.005 per audio minute on the paid tier), generative video, and song services. At gate ① the agent lists anything paid, with an estimate, before spending. Current Gemini prices are in playbook/04.

Can I use the styles commercially? What about the references' licenses?

The repo's original content, style presets included, is MIT, except the files that import Blender's Python API (the swatch renderer's blender_render.py, Blender style scenes such as tabletop-miniature/swatch.py, the intro film's galaxy.py), which are GPL-3.0-or-later as Blender asks of published bpy scripts. A few presets adapt text from an older CC BY 4.0 snapshot of lemo-opuscar and keep that attribution. The presets describe a grammar and never copy characters, logos or shots, but what goes into your own film is up to you: fonts are referenced from your machine, not shipped, so check their licenses, and log every asset's source in NOTES.md. The reference repos in references/repos/ are not part of this repo, and some have no license or forbid commercial use: read them, don't reuse them. Per-repo verdicts are in community-skills.md (section 5) and ACKNOWLEDGMENTS.md. For commercial voice work, avoid fish-speech, F5-TTS weights and ChatTTS, which carry non-commercial licenses. And if you pick the Remotion engine, companies of more than three people need a Remotion license.

Why the human gates?

Because early changes are cheap: fixing a storyboard takes minutes, fixing finished code takes hours. Three stops (concept and outline, storyboard, first draft) put your judgment where it saves the most work.

Do I have to come up with the idea?

No. At gate ① the agent offers 2–3 concept cards: each a one-line idea that makes the form tell the content, with the look and the hook it leads to and one frame. Pick one, swap its style preset or hook in your reply, or bring your own concept and it goes straight into the brief. See How It Works.

Can I skip them?

Yes. Ask for a "quick draft" (effort quick), or say plainly in the chat that you don't want to review. The agent writes your exact words and the date into REVIEW.md, and a fresh-context reviewer stands in for the skipped gates. The agent can't decide to skip on its own.

What is effort?

One switch, quick, standard (default) or studio, that sets the gates, how concepts are shown to you (one frame each, or a 10–20 s sketch each at studio), how deep the self-review goes, how many scored review rounds run, how much sound work is done and what gets delivered. A 30-second film takes roughly 10–30 minutes, 1–2 hours, or 3 hours and more. bin/vh effort prints the rules.

What is director mode?

Effort says how hard the agent checks its own work; director mode says what you decide. You can name any of twelve decisions (concept, spec, outline, style, main character, theme music, voice, script, hook, storyboard, edit rhythm, title and cover) and make each one own (the agent shows options and waits), review (it shows a result and carries on unless you object) or delegate (it decides and writes down why in DECISIONS.md). Something like "quick draft, but I'll pick the hook" works. The three gates still stop at standard and studio; director mode only adds stops. See How It Works.

Should I render in 4K?

Only if you want it, or a studio film is going to a platform that plays 4K (Bilibili, YouTube). The film is still written at 1080p and rendered at twice the size (bin/vh new … --res 4k, or Resolution: 4k in the brief), so text looks the same size on screen; drafts and checks stay at 1080p. The hand-drawn engine has no 4K output. Text size depends on where the film is watched, not on the resolution: see How It Works.

Can it cut my own footage?

There's a type for it, 09, but it is experimental. It cuts a talking head, interview or vlog you recorded yourself: filler words and retakes removed, cuts placed on measured quiet gaps, subtitles and explainer graphics added, a vertical version if you want one. Nothing renders until you approve the edit list. It has only been tuned on synthetic material, so the thresholds are a starting point, and some of the tools its docs mention aren't written yet. Your footage is your face and voice: transcription runs locally by default, and cloud services come only after the agent asks you. See Video Types.

Does anything leave my machine?

Not on the pure-code path. What does: cloud voices, --align gemini (it sends your narration audio to Gemini), generative video, and any model download. One easy leak to know about: with GEMINI_API_KEY or GOOGLE_API_KEY set, hyperframes snapshot sends every frame to Gemini for a description unless you pass --describe false, which the repo's docs and hints always include. For footage of yourself, the agent asks before sending anything out.

How do I add a style?

Break down 1–3 works, write STYLE.md and tokens.json, write the swatch, render and check it, rebuild the gallery, open a pull request. The steps are on Style Library.

Why isn't my render identical each time?

It should be: every frame must be a pure function of t. Look for Math.random(), Date.now(), CSS transitions or @keyframes, or state carried from one frame to the next. Compare lossless PNG frames, not MP4s, because video encoding amplifies tiny GPU differences; pixel-identical, or at least 45 dB PSNR with no visible difference, passes. For HyperFrames, --no-browser-gpu makes the text edges repeatable too. More in Troubleshooting.

Where are my API keys stored?

Nowhere in the repo. The tools read keys from environment variables only; .env* files are git-ignored, and LOCAL.md lists variable names, never values. The Gemini key travels in a request header, never in the URL.

How is Chinese handled: fonts, captions, voice?

The workflow docs are written in Chinese first. Captions wrap CJK text properly (use 11 characters per line for vertical video), Qwen3-TTS has five Chinese voices including Beijing and Sichuan accents, and the hand-drawn engine can write characters in true stroke order. For fonts, HyperFrames needs Noto Sans SC loaded explicitly (a Google Fonts <link> or @font-face), and a canvas needs every character preloaded; see Troubleshooting.

Does it work offline?

Mostly, once installed. Projects render from locally installed engines, and local voice, music, SFX and mixing need no network. What does need it: cloud voices and --align gemini, first-time model downloads, Google Fonts links, and bin/vh hf-init (it fetches HyperFrames with npx).

Where do my videos go?

projects/<date>-<slug>/, which git ignores. Drafts and finals are in out/, check images in out/check/. See Getting Started.

Where do the numbers behind the rules come from?

Some come from measurements, written up as lab notes in docs/research/ (also in English): mix levels, how alike the 28 soundtracks are, why the same code rendered different pixels, how long on-screen text has to stay, the shape of a score, and whether concept-first beats floors only. They are records, not rules, and mostly meter readings rather than listening verdicts. See also How It Works.

Are the workflow docs available in English?

Not yet. The README, CHANGELOG, showcase notes and research notes are in English; CLAUDE.md, the playbook and the type docs are Chinese-first. Agents follow them either way. Translations are welcome.

Clone this wiki locally