diff --git a/.gitignore b/.gitignore index 7bbf24a..d62ce77 100644 --- a/.gitignore +++ b/.gitignore @@ -27,3 +27,9 @@ showcase/04-intro-film/assets/films*.jpg showcase/04-intro-film/assets/films*.png showcase/04-intro-film/blender/out/ showcase/04-intro-film/opening/node_modules + +# videos live in the GitHub release "media" (tools/fetch_media.sh puts them here; tools/media.txt lists them) +showcase/*/media/final.mp4 +showcase/04-intro-film/v3/final.mp4 +styles/gallery.mp4 +*.part diff --git a/AGENTS.md b/AGENTS.md index 533d04a..a77aecb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -148,7 +148,7 @@ OpenVideoHarness/ ├── CONTRIBUTING.md 改本仓库本身时的分支、PR 和推送规则(人和 agent 都适用) ├── .github/ CI(Linux + macOS 跑 tools/ci.sh)和 main 分支的规则集 ├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review -├── tools/ bin/vh 背后的脚本(audio/:tts、captions、beats、music、sfx、mix、qa;sheet.py、readcheck.py、review.py;给人拍板用的图:storyboard.py、rhythm.py、style_compare.py、cover_preview.py、audio/roll.py,共用 vhdraw.py);ci.sh 是仓库自检 +├── tools/ bin/vh 背后的脚本(audio/:tts、captions、beats、music、sfx、mix、qa;sheet.py、readcheck.py、review.py;给人拍板用的图:storyboard.py、rhythm.py、style_compare.py、cover_preview.py、audio/roll.py,共用 vhdraw.py);ci.sh 是仓库自检;fetch_media.sh 从 Release 下载样片视频 ├── skills/open-video-harness/ 轻量 skill:在任何目录把做视频的请求引到本仓库 ├── video-types/ 9 类视频(09 实验中):工作流、审美、禁止项、prompt 增量块、自查重点、案例、社区 skill ├── playbook/ 跨类型的通用知识 @@ -168,7 +168,7 @@ OpenVideoHarness/ ├── templates/ 新项目的文件:BRIEF、STORYBOARD、STYLE、REVIEW、DECISIONS、NOTES、LESSONS、TASTE_CHECKLIST;用到时再复制:SCRIPT、CHARACTER、PACKAGING ├── styles/ 风格库:从名作学来的风格预设(STYLE.md + tokens.json + 真渲的 5 s 样片),_swatch/ 是样片渲染器 ├── cases/ 真实案例拆解 + opus55-gallery(社区作品精选) -├── showcase/ 本仓库自己做的片子(源码 + 成片 + 自评记录) +├── showcase/ 本仓库自己做的片子(源码 + 封面和联系表 + 自评记录;成片在 GitHub Release `media`,`tools/fetch_media.sh` 下载) ├── docs/research/ 研究笔记(中英):量过的几件事、数字,和因此改了什么;是实验记录,不是规则(规则在 playbook/) ├── engines/ ClaudeAnimationBase(内置)+ 其他引擎的安装说明 ├── references/ diff --git a/CHANGELOG.md b/CHANGELOG.md index d6269b8..f953998 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,14 @@ ## Unreleased +**Videos move out of git; the intro film is re-paced to 103 s; the READMEs are rewritten** +- Why: the maintainer found the 150.5 s intro "too slow after the opening: the pause after each line is too long… but not a flash either", asked whether a GitHub repo should hold video files at all, and found parts of the README odd to read. +- Videos live in the GitHub release `media` (the five showcase films, the v3 intro, the style reel, the intro's gate ① animatic), not in git. `tools/media.txt` lists the ones the build scripts use (all but the animatic) with their paths in the repo and their sha256; `tools/fetch_media.sh [word…]` downloads them back to those paths, which are now in `.gitignore`. The installer's shallow clone downloads about 80 MB instead of 215 and takes about 185 MB on disk instead of 420 (measured on this branch). The showcase preview GIFs are gone (the READMEs play clips instead). `tools/ci.sh` fails on video files (except the 31 style swatches) and on any file over 8 MB. CONTRIBUTING and CLAUDE.md say where videos go. +- `showcase/04-intro-film` at 103 s (3090 frames): back to the 87.5 s pacing with 15 short holds where a line would flash by (each line stays long enough to read once in either language, then the camera moves), and the music keeps playing through them instead of dropping into a "held breath". A sync check moved four score accents onto their picture events and fixed a clock mix-up that made the case-study stars light before the camera arrived. Six new 1080p chapters for the READMEs; the full film is `intro-film-1080p.mp4` in the release. +- `README.md` and `README.zh-CN.md` rewritten, each on its own instead of translated line by line: what it is, what it makes, quick start, how to ask, how it works (with the A/B result as it was), types and styles, Blender, sound, effort and who decides, requirements, cost and limits, a docs map. The download table and the China-network mirror settings moved to the wiki (Getting Started, China-Network / 国内网络); `bin/vh doctor` points there. +- Reading time: `templates/TASTE_CHECKLIST.md` #5 and playbook/03 add a short-label rule. A few words in a moving shot (node names, station names, gate captions) get one read: max(1.5 s, CJK/7 + other/20 + 0.8 s); titles, claims and numbers to remember keep the on-screen rule. `tools/readcheck.py` checks it for texts marked `data-read="label"` (`--mode label`, a `lab` tag, a budget line, a CI test). Stretching a hold for reading must not stall the music. +- playbook/03: stretching holds for reading time must not slow the whole film or stall the music; playbook/04: when the picture changes, move the score's accent for that event along with its sound effect. + **The installer downloads less: a shallow clone** - Why: after #51 a full clone is about 550 MB on disk (351 MB to download), 336 MB of it `.git`, because earlier versions of the sample films stay in the history (`showcase/04-intro-film/media/final.mp4` alone is 48 MB, and v3's media moved to `v3/`). - `install.sh` clones with `--depth 1 --single-branch`: about 215 MB to download and 420 MB on disk (measured 2026-10-04 at a881f4f). Re-running it on such a checkout fetches the newest commit with `--depth 1` and moves to it with `git reset --keep`, since there is nothing to fast-forward along: untracked and ignored files (`LOCAL.md`, `projects/`) and edits to files the update doesn't touch stay. Before fetching, it stops if the checkout has commits of your own or isn't on a branch that tracks the repo (a detached HEAD, a branch of your own). When the update would overwrite an edited file, or an untracked one that isn't ignored, it stops with a hint and puts `origin/main` back, so `git status` and the next run see the checkout as it was. A full clone (made by hand, or by an earlier installer) is still updated with `git pull --ff-only`, because a depth-1 fetch would make it shallow. diff --git a/CLAUDE.md b/CLAUDE.md index f7afa84..abef202 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -145,7 +145,7 @@ OpenVideoHarness/ ├── CONTRIBUTING.md 改本仓库本身时的分支、PR 和推送规则(人和 agent 都适用) ├── .github/ CI(Linux + macOS 跑 tools/ci.sh)和 main 分支的规则集 ├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review -├── tools/ bin/vh 背后的脚本(audio/:tts、captions、beats、music、sfx、mix、qa;sheet.py、readcheck.py、review.py;给人拍板用的图:storyboard.py、rhythm.py、style_compare.py、cover_preview.py、audio/roll.py,共用 vhdraw.py);ci.sh 是仓库自检 +├── tools/ bin/vh 背后的脚本(audio/:tts、captions、beats、music、sfx、mix、qa;sheet.py、readcheck.py、review.py;给人拍板用的图:storyboard.py、rhythm.py、style_compare.py、cover_preview.py、audio/roll.py,共用 vhdraw.py);ci.sh 是仓库自检;fetch_media.sh 从 Release 下载样片视频 ├── skills/open-video-harness/ 轻量 skill:在任何目录把做视频的请求引到本仓库 ├── video-types/ 9 类视频(09 实验中):工作流、审美、禁止项、prompt 增量块、自查重点、案例、社区 skill ├── playbook/ 跨类型的通用知识 @@ -165,7 +165,7 @@ OpenVideoHarness/ ├── templates/ 新项目的文件:BRIEF、STORYBOARD、STYLE、REVIEW、DECISIONS、NOTES、LESSONS、TASTE_CHECKLIST;用到时再复制:SCRIPT、CHARACTER、PACKAGING ├── styles/ 风格库:从名作学来的风格预设(STYLE.md + tokens.json + 真渲的 5 s 样片),_swatch/ 是样片渲染器 ├── cases/ 真实案例拆解 + opus55-gallery(社区作品精选) -├── showcase/ 本仓库自己做的片子(源码 + 成片 + 自评记录) +├── showcase/ 本仓库自己做的片子(源码 + 封面和联系表 + 自评记录;成片在 GitHub Release `media`,`tools/fetch_media.sh` 下载) ├── docs/research/ 研究笔记(中英):量过的几件事、数字,和因此改了什么;是实验记录,不是规则(规则在 playbook/) ├── engines/ ClaudeAnimationBase(内置)+ 其他引擎的安装说明 ├── references/ diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index df13829..c856422 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -17,7 +17,7 @@ 1. 只推自己的分支。不推 main,不删分支,不动 tag。 2. 推送前跑 `tools/ci.sh --committed`,全部通过才推。它在一份干净的检出上检查 HEAD,结果和 CI 一致;工作区里没提交的改动不算数。修 bug 要先复现,再证明修好,前后对比写进 PR。 3. PR 一律开成草稿。合并由人决定:人在对话里明确让 agent 合并时,agent 先确认 CI 全绿、自己审过 diff,再用 squash 合并。人让 agent"检查通过后合并"时,agent 审过 diff、把草稿转成正式 PR 后,可以打开 auto-merge(squash),不用守着 CI:必需的检查都通过后由 GitHub 合并,有一项失败就不合。规则集允许管理员在 PR 里"绕过规则合并",这个开关只留给人用,agent 不碰。 -4. 不提交 API key、`LOCAL.md`、`projects/` 和渲染产物。测试时改动了受版本管理的样片(`styles//media/`),推送前要还原。 +4. 不提交 API key、`LOCAL.md`、`projects/` 和渲染产物。测试时改动了受版本管理的样片(`styles//media/`),推送前要还原。视频不进 git(风格样片 `swatch.mp4` 除外):成片、样片集锦这类文件放在 GitHub Release `media`,`tools/media.txt` 登记它们在仓库里的路径、文件名和 sha256,`tools/fetch_media.sh` 按它下载。往这个 Release 里放文件是维护者的事,agent 只在维护者要求时做。换一个文件时用新文件名上传(例如 `intro-film-1080p-v2.mp4`),在同一个 PR 里改 `tools/media.txt` 和引用它的链接,合并以后再删旧文件;不要覆盖 main 还在引用的文件,否则在合并之前,按 main 的校验和下载会失败。README 里要直接播放的短片段走 GitHub 的附件链接(每个不到 10 MB)。`tools/ci.sh` 会拦住视频和超过 8 MB 的文件。 5. 用户能感知到的改动,在 `CHANGELOG.md` 的 Unreleased 里记一笔。改了 `CLAUDE.md`,跑 `bin/vh sync-agents` 重新生成 `AGENTS.md`。 6. 提交信息写清改了什么、为什么。末尾可以带 `Co-Authored-By:` 署名行。 @@ -86,4 +86,6 @@ VH_BASH=/bin/bash tools/ci.sh # macOS:用系统自带的 bash 3.2 跑,M 同一天补打了 `v0.1.0`、`v0.2.0`、`v0.2.1` 三个 Release,说明取自 `CHANGELOG.md` 对应的一节。 +另有一个不是版本的 Release `media`(tag `media`,2026-10-04 由 agent 在维护者同意后建):放样片和风格集锦的视频文件,规则见上面第 4 条。它的 tag 不移动,也不受上面两个规则集保护(规则集只管 `v*`);换文件按第 4 条用新文件名,不覆盖。 + 发新版本的做法:把 `CHANGELOG.md` 的 Unreleased 改成 `## vX.Y.Z — 日期`,经 PR 合并进 main,然后由维护者在合并后的提交上建 Release(`gh release create vX.Y.Z --target <提交> --notes-file <这一节>`)。tag 一旦建好就不再移动;发错了就发一个新的补丁版本。改了这两个文件以后,用 `gh api --method PUT repos/ZLHad/OpenVideoHarness/rulesets/ --input <文件>` 同步线上配置,`` 用 `gh api repos/ZLHad/OpenVideoHarness/rulesets` 查。 diff --git a/README.md b/README.md index 3c1fea4..dc8a4b6 100644 --- a/README.md +++ b/README.md @@ -2,23 +2,21 @@ # OpenVideoHarness -**Coding agents like Claude Code and Codex can make videos by writing programs. This makes them do it reliably.** +**A video workbench for Claude Code and Codex: the agent writes the code that renders the film, and you sign off at three checkpoints.** -Explainers, science shorts, product films, music videos, data stories, paper talks, hand-drawn shorts, meme edits and edits of your own footage: 9 video types (09 experimental), 31 styles, one workflow. +Explainers, science shorts, product films, music videos, data stories, paper talks, hand-drawn shorts, meme edits, and (experimental) edits of footage you shot yourself. **English** · [中文](README.zh-CN.md) · [Wiki](https://github.com/ZLHad/OpenVideoHarness/wiki) ![License: MIT](https://img.shields.io/badge/license-MIT-black) ![Agents](https://img.shields.io/badge/agents-Claude%20Code%20%7C%20Codex-orange) -![Engines](https://img.shields.io/badge/engines-HyperFrames%20%7C%20Remotion%20%7C%20Manim%20%7C%20p5.brush%20%7C%20Blender-blue) -![Voice](https://img.shields.io/badge/voice-Qwen3--TTS%20zh%20%7C%20en-purple) -![Styles](https://img.shields.io/badge/styles-31-green) +![Engines](https://img.shields.io/badge/engines-HyperFrames%20%7C%20Manim%20%7C%20Remotion%20%7C%20p5.brush%20%7C%20Blender-blue) -https://github.com/user-attachments/assets/9d305423-0e5b-4276-a347-17a072993b6f +https://github.com/user-attachments/assets/48923404-eff1-41cd-b303-82c9ede51c1f -▶ Chapter 1 of the intro film (the opening, 19 s, 1080p, with sound). The whole film, 150.5 s, plays in eight 1080p chapters in showcase/04-intro-film; the single file is final.mp4. Its first 15.8 s are path-traced in Blender, driven by code: a galaxy of 1.18 million stars in which every star is a film. One continuous HyperFrames + Three.js take follows. An agent made it by following this repo, and the soundtrack is code too. +▶ The opening of the intro film (25 s, with sound). An agent made the film from this repo's docs alone; the opening is rendered in Blender, driven by code, and the soundtrack is code too. The full 103 s is in showcase/04. @@ -26,55 +24,21 @@ https://github.com/user-attachments/assets/9d305423-0e5b-4276-a347-17a072993b6f ## What this is -AI can already make "a video from one sentence". It doesn't paint the pictures. It writes a program that works out what every frame looks like; a browser (or Manim) renders the frames one by one, and sound is added at the end. The community has made plenty of black-hole explainers, hand-drawn music videos, launch films and math animations this way. +Coding agents can already make a video from one sentence: they write a program that computes every frame, a browser (or Manim) renders the frames, and a soundtrack goes on top. Getting a good one every time is the hard part. The same model can be brilliant once and then lose the pacing, glow everything and invent numbers the next time, because it has no production process to follow and no way to check its own work. -The hard part is making it **reliable**. The same model is stunning one day and a mess the next: rushed pacing, glow everywhere, made-up numbers. The model is smart enough. What it lacks is a craft to follow. +This repo is that process and those checks, written as docs an agent can follow and a CLI, `bin/vh`: -OpenVideoHarness is that craft, written as docs and tools an agent can follow: +- **A workflow for each kind of video.** Ask for a science short or a launch film and the agent reads the workflow for that type: which engine, which steps, what looks good, what is off limits. +- **Three stops for your sign-off.** Direction and outline, storyboard, first draft. Direction gets settled before any code is written, when changing it costs least. For a quick try, say so and it renders straight away. +- **It checks its own work.** An agent can't watch video or hear sound, so it reads rendered frames and audio measurements against a 20-point checklist, and a second agent that didn't make the film scores it. +- **Sound included.** Chinese and English voiceover, bilingual subtitles, a score and sound effects written as code, mixing and an audio check. +- **A style library, and Blender.** 31 styles distilled from well-known films and design, each with a rendered sample, to borrow from rather than copy; and when a shot needs real glass, volumetric light or a million particles, the agent drives Blender with Python. -- **A method for each kind of video.** Say "make a science short" or "make a launch film", and it reads the workflow for that type: which engine, which steps, what looks good, what is off limits. -- **It stops and asks you three times.** At the outline, the storyboard and the first draft, it waits for your go-ahead. Direction gets settled while changes are still cheap. For a quick try, switch to the `quick` level and it just renders. -- **More than one taste.** 31 styles learned from famous work, each with a real rendered sample. They are references a concept can borrow from, not templates. -- **It checks its own work.** An agent can't watch video or hear sound, so it looks at rendered frames, measures the mix, and fixes things until a checklist passes. A reviewer who didn't make the film then scores it. -- **Sound included.** Chinese and English voiceover (a local open-source model), bilingual subtitles, music and sound effects written as code, mixing and a final audio check. -- **Code-driven Blender.** When a shot needs real glass, nebulae and volumetric light, depth of field or a million particles, the agent writes Python that has Blender path-trace the plate, in a sandbox, and joins it with the web engine's pictures. The intro film opens this way; see [below](#code-driven-blender-3d-and-film-grade-effects). +It isn't a new rendering engine. It sits on top of HyperFrames, Manim, Remotion, p5.brush and Blender and tells the agent how to use them well. -It isn't a new rendering engine. It is a layer of "how to do it" on top of engines like HyperFrames, Manim, Remotion, p5.brush and Blender. +## What it makes -## Start in 30 seconds - -```bash -curl -fsSL https://raw.githubusercontent.com/ZLHad/OpenVideoHarness/main/install.sh | bash -``` - -This one command: -- clones the newest version of the repo into `~/OpenVideoHarness`, without its history (about 420 MB; the sample films are in it); -- installs the dependencies of the built-in hand-drawn engine and the style renderer (about 210 MB); -- fetches 30 read-only reference repos (about 195 MB; `--no-refs` skips them); -- registers the `open-video-harness` skill for Claude Code and Codex, so saying "make a video" in any folder leads the agent here. - -That is about 825 MB on disk. The first render and the sound tools download more the first time you use them: [what gets downloaded](#what-gets-downloaded) lists how much and where. New here? The wiki's [Getting Started](https://github.com/ZLHad/OpenVideoHarness/wiki/Getting-Started) page goes from nothing to a first video. - -Then open Claude Code (or Codex) and say what you want: - -```bash -cd ~/OpenVideoHarness && claude -``` - -```text -Make a 30-second vertical science short: why does a low-orbit satellite's signal change pitch? English voiceover, English and Chinese subtitles. -``` - -What happens next: -1. It offers two or three concepts (an idea that makes the form itself tell the content, each with one frame), then a one-screen outline; the style follows from the concept you pick, or borrows from the style library. -2. Then a storyboard and a sheet of keyframes. -3. Then a first draft, along with the two or three things it likes least. - -Only when you say yes does it render the final. - -## See what it makes - -The five films below were made by an agent **reading only this repo's docs**. Every video plays right here (1080p, with sound; the 31-style reel at its native 720p), with the repo's name in a corner; click a title for its folder (brief, storyboard, review notes, fix-up log, full source), a good start for a similar video. Each request is a short quote from a longer brief, and “Request” links to the whole text. Films 00–03 were finished silent, and their sound was fitted afterwards. +An agent made each of these from this repo's docs alone. Click a title for its folder: the request, the storyboard, the review notes and all the source, a ready starting point for a similar film. Each request is a short quote; "Request" links to the full text. @@ -82,375 +46,241 @@ The five films below were made by an agent **reading only this repo's docs**. Ev https://github.com/user-attachments/assets/1f5873bd-a83f-4d5e-96fe-08aff0c3a96c -the original file +Download - + - + - + - + - + - +
02 · Vertical science short (HyperFrames · 24.8 s · 1080×1920)
A LEO satellite's Doppler shift, readable with the sound off.
Request: “为什么低轨卫星的信号会"变调"?——多普勒频移” (Why does a LEO satellite's signal "change pitch"? The Doppler shift.)
Suggested workflow (standard effort):
bin/vh new short leo-doppler --aspect 9:16
sound: bin/vh tts, music, mix … profile=short, mux
Sound: Chinese voiceover (Gemini TTS, voice Aoede), a light score, the satellite beacon made audible; zh/en soft subtitles, off by default (captions are burned in)
02 · Vertical science short (HyperFrames · 24.8 s · 1080×1920)
A low-orbit satellite's Doppler shift, readable with the sound off. Chinese voiceover, Chinese and English subtitles.
Request: "为什么低轨卫星的信号会'变调'?——多普勒频移" (why does a low-orbit satellite's signal change pitch? The Doppler shift)
https://github.com/user-attachments/assets/c2368f76-b157-4a48-9945-8048a515efd4 -the original file +Download 03 · 3b1b-style math explainer (Manim CE · 25 s · 1920×1080)
Sine waves stack into a square wave, with the Gibbs overshoot as the payoff.
Request: “Building a square wave from sine waves”
Suggested workflow (standard effort):
bin/vh new math fourier-square-wave --style dark-math
sound: bin/vh tts, music, mix … profile=explainer, mux
Sound: English voiceover (Gemini TTS, voice Iapetus), a piano score kept 13 LU under the voice, each harmonic sounding its own tone
03 · Math explainer in the 3Blue1Brown style (Manim · 25 s · 1920×1080)
Sine waves stack into a square wave and land on the Gibbs overshoot. English narration; each harmonic sounds its own note.
Request: "Building a square wave from sine waves"
https://github.com/user-attachments/assets/da028240-fcff-4e02-95a0-a0abd3238d07 -the original file +Download 01 · Hand-drawn character short (p5.brush · 12 s · 1920×1080)
A hand-painted pantomime: no text, no voice, three shots.
Request: “Clawd tries to film a falling autumn leaf with a tiny hand-cranked movie camera; the wind keeps snatching the leaf just as Clawd frames it …”
Suggested workflow (standard effort):
bin/vh new handdrawn leaf --style watercolor-pastoral
sound: bin/vh music, sfx, mix … profile=cartoon, mux
Sound: a cartoon score (pizzicato, celesta, flute, xylophone) plus foley that follows every action
01 · Hand-drawn character short (p5.brush · 12 s · 1920×1080)
A three-shot pantomime with no words and no voice; the score and foley follow every move.
Request: "Clawd tries to film a falling autumn leaf with a tiny hand-cranked movie camera; the wind keeps snatching the leaf just as Clawd frames it …"
https://github.com/user-attachments/assets/7b5e2d6a-1683-4a01-8cf8-2b27b5bf0c97 -the original file +Download 00 · Launch short (HyperFrames · 20 s · 1920×1080)
Built from the repo's own terminal, folders and contact sheets.
Request: “Produce the README hero video — a short launch film for OpenVideoHarness itself …”
Suggested workflow (standard effort):
bin/vh new promo launch-film
sound: bin/vh music, sfx, mix … profile=promo, mux
Sound: a minimal electronic score plus foley on every on-screen action (typing, clicks, whooshes on the cuts), no voice
00 · Launch short (HyperFrames · 20 s · 1920×1080)
Built from the repo's own terminal, folders and contact sheets, with a sound for every move on screen.
Request: "Produce the README hero video — a short launch film for OpenVideoHarness itself …"
-https://github.com/user-attachments/assets/e4fa59b8-3941-4fdd-ada4-62b4be4f3a82 +https://github.com/user-attachments/assets/dfa4a89e-ab81-4d2a-abab-5ba3eec57fbe -Chapter 5 of 8 (the self-review loop and the final cut); all eight · the original file +Chapter 3 of 6; all six · Download (297 MB) 04 · Intro film (a galaxy of films, then one take) (Blender + HyperFrames + Three.js · 150.5 s · 1920×1080)
The repo's own product film. The opening is path-traced in Blender: every star is a film, the galaxy collapses into a supernova and is flattened into a sea of films. Then one continuous 3D take through the repo.
Request, at the gates: “玻璃、宇宙、星穹……令人瘫坐眩晕的感觉” (glass, cosmos, a starry sky… the feeling that makes you dizzy), then “或者使用blender?好莱坞大片质感” (or use Blender? Hollywood blockbuster quality)
Suggested workflow (studio effort, --effort studio):
bin/vh new promo intro-film --style monumental-scifi
Blender plates: engines/blender.md; one-take 3D: playbook/08; bin/vh sheet, check
Sound: a code-composed score (a 120 BPM opening, then D minor at 80 BPM) and 124 sound effects, mixed with bin/vh mix … profile=promo
04 · Intro film (Blender + HyperFrames + Three.js · 103 s · 1920×1080)
The repo's own product film. The concept: every star is a film. The opening galaxy is path-traced in Blender, collapses, bursts and is flattened; then one continuous 3D take follows a single request through the whole repo. This chapter shows the three checkpoints and the self-review loop.
Request, at the checkpoints: "玻璃、宇宙、星穹……令人瘫坐眩晕的感觉" (glass, cosmos, a starry sky… dizzying), then "或者使用blender?好莱坞大片质感" (or use Blender? Hollywood blockbuster quality)
https://github.com/user-attachments/assets/7fee4f5a-f081-4b7e-82a6-9be8be9f5b7e -the original file +Download 31 styles, one reel (5-second samples)
The same content in 31 styles learned from famous work.
Borrow from one:
bin/vh new promo launch-film --style cutout-jazz
bin/vh style list shows all 31; details in styles/README.md
Sound: one score per sample (bin/vh music, mix … profile=swatch)
The 31 styles in one reel (about 1.5 s each)
The same content in 31 styles, each with its own music.
bin/vh style list lists them; bin/vh new promo launch-film --style cutout-jazz attaches one as a reference.
-“The original file” under each video links to the MP4 in the repo. Films you make with it are welcome in `showcase/` as a PR. - -
-Similar work from the community, and our breakdowns of it - -| Work | Type | How it was made | Breakdown | -|---|---|---|---| -| [I'm Upping My P(doom)](https://x.com/other__reality/status/2102514581684052169) | Hand-drawn MV · 156 s | p5.brush, 9 chapters drawn by parallel subagents | [cases/mv-pdoom.md](cases/mv-pdoom.md) | -| [Functional Emotions](https://x.com/eudaemonea/status/2102610626321490404) | Painted MV · 372 s | A custom WebGL brush renderer, 7 subagents | [cases/mv-functional-emotions.md](cases/mv-functional-emotions.md) | -| [Claude Pop](https://x.com/donaldjewkes/status/2102801274173587569) | Hybrid MV | Generative video as a base, traced over in code; a 12-hour run | [cases/mv-claude-pop.md](cases/mv-claude-pop.md) | -| [The real physics of Interstellar: black holes](https://x.com/AndyL5cc/status/2104519528873103773) | Science explainer · 143 s | A one-sentence request; one black-hole shader carries the film | [cases/explainer-interstellar-blackhole.md](cases/explainer-interstellar-blackhole.md) | -| [You ask an AI one question: the next 3 seconds](https://v.douyin.com/MONt8dOfuEo/) | Knowledge explainer · 126 s | 3 seconds slowed to 2 minutes; a clock and a slow-motion factor stay on screen | [cases/oneshot-five.md](cases/oneshot-five.md) §1 | -| [我眼中的你 (You, as I see you)](https://v.douyin.com/kTpIVOEMsEY/) | Portrait · 231 s | The author says it came out in one go from a short prompt; Claude narrates its user from his notes, quotes and commits | [cases/oneshot-five.md](cases/oneshot-five.md) §2 | -| [Opus 5.5 introduces itself](https://v.douyin.com/bGfv-xMIKWc/) | Motion graphics · 35 s | Each step of how it was made is drawn with that step's technique; the prompt says not to use installed skills | [cases/oneshot-five.md](cases/oneshot-five.md) §3 | -| [FunTech Showreel 2026](https://x.com/tkm_hmng8/status/2105255710531674358) and [Tesseract for design](https://x.com/trymirage/status/2105314048766033999) | Showreel · 50 s; launch film · 31 s | One mascot through a dozen style worlds; one chrome cube from design system to finished video | [cases/oneshot-five.md](cases/oneshot-five.md) §4–5 | -| [Applore promo](https://x.com/decohack/status/2104502625055949242) | Product film · 15 s | One showreel prompt plus real assets | [cases/promo-applore.md](cases/promo-applore.md) | -| [Austerlitz, 2 December 1805](https://x.com/WinterArc2125/status/2103116235009347650) | 3D history film · 301 s | WebGL2 on real terrain; each shot lasts as long as its narration; sound effects are panned and distanced from the picture | [cases/opus55-gallery.md](cases/opus55-gallery.md) §6 | -| [389 community videos](https://github.com/yihui-dev/awesome-opus5-5-videos) and [a 962-work catalog](https://github.com/zhuyansen/awesome-opus-5.5-video) | Mixed | Prompt statistics, categories, curated picks | [cases/opus55-gallery.md](cases/opus55-gallery.md) | - -
- -## 31 styles, not one taste - -AI-made videos drift toward one look: dark background, glow, glass cards, busy animated UI. Unless you say otherwise, that's where an agent goes. - -So we studied famous work and wrote down 31 styles in [`styles/`](styles/), grouped into film titles, brand and launch, data and explainers, illustration and print, Chinese aesthetics, and retro tech. The sources include: -- **film and title design**: Saul Bass's titles, *Se7en*, *Blade Runner 2049*, Wes Anderson's symmetry, Wong Kar-wai's step-printing, Aardman's and Laika's stop-motion miniatures; -- **design**: the Swiss grid, 3Blue1Brown, New York Times data graphics, Otto Neurath's Isotype pictograms; -- **animation and print**: the halftone dots of *Spider-Verse*, 16-bit Super Nintendo pixel art; -- **Chinese aesthetics**: ink wash (水墨), Dunhuang murals, shadow puppetry and guochao. - -Each style is written as instructions an agent can follow: colors and fonts, composition, how things move, how scenes change, what it sounds like, and which clichés to avoid. +The players show clips under 10 MB each (GitHub's limit; 1080p, the style reel 720p), with the repo's name in a corner; the full files are in the [`media` release](https://github.com/ZLHad/OpenVideoHarness/releases/tag/media), outside git, so a clone doesn't download them. Community films in the same vein, and how they were made, are in [cases/](cases/README.md). Films you make with this repo are welcome in `showcase/`. -**Every style comes with a real 5-second sample rendered by this repo**, each with its own code-written music. All 31 samples show exactly the same content, so the only difference is the style: +## Quick start -The 31 style samples, all showing the same content - -Watch them back to back, each with its own sound, in [`styles/gallery.mp4`](styles/gallery.mp4). To use one: +You need macOS or Linux, Node.js 22 or later, Google Chrome, FFmpeg, Python 3 with [uv](https://github.com/astral-sh/uv), and Claude Code or Codex ([full requirements](#requirements-cost-and-limits)). ```bash -bin/vh style list # see all 31 -bin/vh new promo launch-film --style cutout-jazz # attach a style as a reference (several: --style a,b) +curl -fsSL https://raw.githubusercontent.com/ZLHad/OpenVideoHarness/main/install.sh | bash ``` -We learn the *grammar* of these works; we don't copy them. No original characters, logos or shots, and the prompts never say "in the style of" a person. More in [styles/README.md](styles/README.md). - -## Code-driven Blender: 3D and film-grade effects +It installs the repo into `~/OpenVideoHarness` with its dependencies, fetches 30 read-only reference repos, and registers the `open-video-harness` skill for Claude Code and Codex, so asking for a video from any folder finds it. About 590 MB on disk in all. Then: -Some pictures a web engine can't give you: real glass refraction, nebulae and volumetric light, real depth of field, motion blur on a million particles. For those the agent writes Python that drives Blender. Every frame still depends only on time: numpy computes where each star and each card is at time t, Cycles path-traces the plate, and the plate joins HyperFrames' type and UI in the same film. - -

Moments from the intro film's opening, all rendered in Blender: the Earth inside a glass film card, the pull-back to a galaxy, the supernova, the flattened sea of films

+```bash +cd ~/OpenVideoHarness && claude +``` -- **The intro film opens this way.** A galaxy of 1.18 million stars where every star is a film: the camera pulls back from one glass card to the whole galaxy, spirals down, the galaxy collapses into a supernova and is flattened into a sea of films. 15.8 s at 1080p (475 frames) took 1 h 28 min on an M3 Max. The grid and the terminal that follow are HyperFrames, and the two dissolve into each other through the same camera. Source in [`showcase/04-intro-film/blender/`](showcase/04-intro-film/blender/). -- **[`tabletop-miniature`](styles/tabletop-miniature/)** in the style library is rendered in Blender too. -- **It runs in a sandbox.** The agent's bpy scripts get no network, can write only to their output folder, and never see the API keys in your environment. -- **Count the time first.** Render 3–5 frames to calibrate and put "frames × seconds per frame × 1.5" in the brief; 4K takes about 4× as long as 1080p. Long renders run in chunks and can resume. -- **License.** Files that import bpy are GPL-3.0-or-later (Blender's requirement for published bpy scripts); everything else stays MIT. +```text +Make a 30-second vertical science short: why does a low-orbit satellite's signal change pitch? English voiceover, English and Chinese subtitles. +``` -When Blender is worth it, how to bring it into a film, and what went wrong along the way: [engines/blender.md](engines/blender.md) (partly verified; it marks what has actually been run). +What happens next: +1. It proposes two or three directions, each with one frame, and an outline, and waits for you to pick. +2. Then a storyboard with a keyframe for every shot. +3. Then a first draft, with the two or three things it likes least about it. -## How to ask for a video +Only when you say yes does it render the final. The wiki's [Getting Started](https://github.com/ZLHad/OpenVideoHarness/wiki/Getting-Started) walks from nothing to a first video. -You don't need a long brief. Say **what it's about, who it's for and where it goes**. +## Asking for a video -> **Science short:** 45 s vertical, why GPS has to account for relativity. For TikTok and Shorts, English narration, end on one concrete number. +Say what it's about, who it's for and where it will be shown. You don't need a long brief. -> **Math explainer:** in the style of 3Blue1Brown, how a Fourier series builds a square wave piece by piece. 20 s, readable with the sound off. +> **Science short:** 45 seconds, vertical, why GPS has to account for relativity. For YouTube Shorts, with English narration, ending on one concrete number. -> **Product film:** a 30 s launch film for my app in the ink-wash style. Real screenshots only, rhythmic music, sound effects on the key actions. +> **Math explainer:** in the 3Blue1Brown style, how a Fourier series builds a square wave, 20 seconds, readable with the sound off. -> **Paper talk:** turn the core method of `papers/main.tex` into a 3-minute explainer. Copy the paper's details exactly, walk through each equation, English narration with Chinese subtitles. +> **Product film:** a 30-second launch film for my app in the ink-wash style, real screenshots only, sound effects on the key moves. -> **Music video:** a hand-drawn MV for `audio/song.mp3`. No lyrics on screen, the pictures tell the story, and every chorus goes bigger than the last. +> **Paper talk:** turn the core method of `papers/main.tex` into a 3-minute explainer; copy the paper's details exactly, English narration with Chinese subtitles. -> **Learn from someone else's film:** this video is good; break down how it was made, then make a similar one with this workbench. +> **Quick try:** a quick 15-second draft to see whether a cyber-glitch look suits my game trailer; don't ask me anything. -> **Quick try:** a quick 15 s draft to see whether the cyber-glitch style suits my game trailer. Don't ask me anything. +> **Hands-on:** a 90-second explainer on how satellites avoid collisions, at the studio level. I'll choose the hook, the main character, the theme tune, the title and the cover; you decide the rest. ## How it works -

OpenVideoHarness at a glance: a one-sentence request goes through type and effort selection, three human gates, sound-first code rendering and a self-review loop to a finished film; on the right, what the repository provides

+

A one-line request goes through type routing, three human sign-offs, sound-first code rendering and self-review to a finished film

-1. **Pick the type.** From your one sentence, the agent looks up the routing table in [CLAUDE.md](CLAUDE.md), decides what kind of video this is, and reads that type's workflow. -2. **Gate ①, the concept and the outline.** It offers two or three concepts, each a one-line idea such as "slow the 3 seconds after Enter down to 2 minutes", each with one frame; once you pick one, it hands you the outline. The style follows from the concept: its own, or a preset borrowed from the style library. -3. **Gate ②, the storyboard.** For each shot: what the viewer must understand, in what order, and for how long. Plus a preview sheet with one frame per shot. -4. **Sound first.** It makes the narration or music first and measures exactly when every line and beat lands, so the picture follows the sound. -5. **Write the code and check it.** - - After each section, it lays the rendered frames out on a contact sheet, looks at them, and fixes whatever fails a 20-point checklist. - - It measures the mix for dropouts, clicks and missed cues. - - A reviewer who wasn't involved then scores the whole draft on 8 points (the first is whether the idea holds up); each one has to reach 8. -6. **Gate ③, the first draft.** You watch it, and it tells you the parts it likes least. If you can't say what's wrong, it makes two or three versions of one section for you to choose from. -7. **Wrap up.** It renders the final and writes what it learned back into the docs, so the next film starts better. +1. **Route.** The agent looks the request up in the routing table in [CLAUDE.md](CLAUDE.md) and reads that type's workflow. +2. **Directions and outline (first stop).** Two or three one-sentence ideas, such as "slow the three seconds after Enter down to two minutes", each with one frame. Once you pick, the look follows from that idea, or borrows from the style library. +3. **Storyboard (second stop).** Every shot lists what the viewer must take in, in order, and for how long, with one keyframe per shot. +4. **Sound first.** Voiceover or music comes first; every line and beat is measured, and the picture follows the sound. +5. **Code, then check.** After each scene the agent tiles rendered frames into a contact sheet and fixes it against the checklist; it measures the sound for gaps, clipping and missed cues; then a second agent scores the whole film on eight points. +6. **First draft (third stop).** You see the draft and what the agent likes least. If you can't say what's wrong, it makes two or three versions of one passage for you to pick from. +7. **Wrap up.** The final render, and the lessons go back into the docs. +A few rules never relax ([CLAUDE.md](CLAUDE.md) has them all): every frame depends only on its time, so any frame can be rendered alone, in parallel, at any point; when there is sound, the sound sets the length; storyboard before code; numbers, quotes and paper details are copied from the source, and anything uncertain stays out. -### You choose how hard it works +**How is this different from just asking an agent?** We ran one small comparison ([docs/research/06](docs/research/06-concept-first-ab.md), in Chinese): two one-line requests, each made once with the workflow (at the quick level) and once with only a few floor rules, silent, judged blind. The workflow's films had the fresher ideas, but the reviewer found the floors-only films better made and would have posted those both times; time and tokens were about the same. What it flagged (type too small for phones, slow openings) went into the checks. The process aims to save rework, by settling direction before code and catching problems before the render; that saving hasn't been measured yet. -Not every film deserves the full treatment. One switch controls how much effort goes in, at three levels: +## Types and styles -| | `quick` | `standard` (default) | `studio` | +| Type | Good for | Main engine | Docs | |---|---|---|---| -| For | trying a direction, drafts, casual posts | most real videos | launch films, flagship pieces | -| Stops to ask you | never; it just renders | at the outline, storyboard and first draft | the same three, plus a rendered sketch of each concept and a full-length animatic | -| Checks its own work | a contact sheet of the whole film, one at target-screen size, and a strip of the first 2 s | frames and sound, section by section | plus loop seams, determinism and a full audio check | -| Outside reviewer | none | 1 round | at least 3 rounds, all 8 scores at 8+ | -| A 30 s film takes about | 10–30 min | 1–2 h | 3 h or more | +| 01 Math and science explainers | 3Blue1Brown-style animations of how something works | Manim | [01](video-types/01-math-science-explainer.md) | +| 02 Science shorts | TikTok/Douyin, Bilibili, Xiaohongshu, YouTube Shorts | HyperFrames | [02](video-types/02-knowledge-short.md) | +| 03 Product and launch films | Apps, SaaS, open-source projects, feature demos | HyperFrames | [03](video-types/03-product-promo.md) | +| 04 Lyric and music videos | Animation cut to a song | p5.brush or HyperFrames | [04](video-types/04-lyric-music-video.md) | +| 05 Data stories | Animated charts and numbers | HyperFrames + SVG | [05](video-types/05-data-story.md) | +| 06 Paper explainers | Conference videos, research talks | Manim + HyperFrames | [06](video-types/06-paper-explainer.md) | +| 07 Hand-drawn shorts | Watercolor, whiteboard, paper cut-out, character shorts; Chinese characters written stroke by stroke | p5.brush (bundled) | [07](video-types/07-hand-drawn.md) | +| 08 Meme edits | Brutalist, tech-Twitter quick cuts | HyperFrames | [08](video-types/08-brutalist-meme.md) | +| 09 Edits of your own footage (experimental) | Talking heads and interviews: cut filler, add captions and graphics, make a vertical version | HyperFrames | [09](video-types/09-editing-talking-head.md) | -Just say "quick draft" or "make it studio quality" in your request, or start the project with `bin/vh new promo launch --effort studio`. The floor never drops at any level: every frame depends only on time, facts are copied exactly, no audio dropouts, flash-safe. `bin/vh effort` prints the full rules. +For realistic people or physics, combine generated video with code on top ([playbook/05](playbook/05-hybrid-genvideo.md)); for a story with an arc or anything over three minutes, [playbook/09](playbook/09-narrative.md); for opening hooks, titles and covers, [playbook/10](playbook/10-hooks-and-packaging.md). -### You choose what you decide +**Styles.** Left alone, AI video drifts toward one look: dark background, glow, glass cards. [`styles/`](styles/) holds 31 styles distilled from well-known work, among them Saul Bass title sequences, the Swiss grid, 3Blue1Brown, New York Times graphics, the halftone of *Spider-Verse*, Wes Anderson's symmetry, ink wash, Dunhuang murals, shadow puppetry and guochao. Each describes its colors, type, composition, motion, transitions and sound, and each was rendered as a 5-second sample of the same content: -Effort sets how hard the agent checks its own work; director mode sets what you decide yourself. Each of twelve decisions (concept, spec, outline, script, style, hook, main character, theme music, storyboard, edit rhythm, voice, title and cover) can be yours to **own** (it shows you options and waits), yours to **review** (it shows you the result and carries on unless you object), or **delegated** (it decides and writes down why in `DECISIONS.md`). Every stop is, by default, a local page from `bin/vh review`: at most three decisions on the first screen, each with a recommendation and a one-line reply, then pictures, the animatic and music you can play in the browser. `standard` and `studio` still stop at the concept and outline, storyboard and first draft. +The 31 style samples, all with the same content -> **Deep involvement:** a 90 s explainer on how satellites avoid collisions, studio quality. I'll pick the hook, the main character, the theme melody, and the title and cover; decide the rest yourself. +A style is a reference, not a template: borrow one, mix several, or ignore them all. What's borrowed is the visual grammar, never the original's characters, logos or shots. More in [styles/README.md](styles/README.md) (in Chinese). -### Why this makes it reliable +## 3D and film-grade effects: Blender, driven by code -A few hard rules (full version in [CLAUDE.md](CLAUDE.md)): +Some shots are beyond a web engine: refracting glass, volumetric light, real depth of field, motion blur on a million particles. For those the agent writes Python that drives Blender: numpy places every star and card for each frame, Cycles path-traces it, and the result joins HyperFrames' type and interface in one film. -1. **Each frame depends only on time.** No random numbers, no system clock, so the same moment always produces the same frame. That's what allows parallel rendering, checking any single frame, and re-rendering after a one-line change. -2. **When there's sound, sound sets the timing.** The picture lines up with the audio, not the other way round. -3. **Storyboard before code.** Pacing is where AI video most often fails, so first decide what each shot must get across and for how long. -4. **Check every section.** Look at the frames, measure the sound, go through the checklist, fix what fails. -5. **Copy facts exactly.** Numbers, paper details and quotes come straight from the source; anything uncertain gets noted, not put on screen. +

Moments from the intro film's opening, all rendered in Blender: the Earth in a glass card, the pull-back to a galaxy, the burst, the flattened sea of films

-## 9 video types (09 experimental) +The intro film's opening was made this way: a galaxy of 1.18 million stars, 15.8 s and 475 frames, 1 h 28 min at 1080p on an M3 Max. The scripts run sandboxed, with no network and no access to your API keys, and long renders go in resumable chunks. When Blender is worth it, how to join it to a film and what went wrong along the way: [engines/blender.md](engines/blender.md) (in Chinese). -| # | Type | Main engine | The gist | Doc | -|---|---|---|---|---| -| 01 | Math and science explainers | Manim | One color per idea; geometry before algebra | [01](video-types/01-math-science-explainer.md) | -| 02 | Science shorts (vertical or wide) | HyperFrames | Hook in the first second, something new every 3–5 s | [02](video-types/02-knowledge-short.md) | -| 03 | Product and launch films | HyperFrames | Real UI only, and show the product itself | [03](video-types/03-product-promo.md) | -| 04 | Lyric videos and MVs | p5.brush or HyperFrames | On the beat within 1 frame; each chorus bigger | [04](video-types/04-lyric-music-video.md) | -| 05 | Data stories | HyperFrames + SVG | One chart, one point; every number traceable | [05](video-types/05-data-story.md) | -| 06 | Paper and conference videos | Manim + HyperFrames | Paper details copied exactly; figures redrawn as vectors | [06](video-types/06-paper-explainer.md) | -| 07 | Hand-drawn, watercolor, whiteboard, paper-cut | p5.brush (built in) | Handmade and always moving; can write Chinese in stroke order | [07](video-types/07-hand-drawn.md) | -| 08 | Brutalist, meme and fast-cut edits | HyperFrames | Build a grid, then break it; every joke lands in 1 s | [08](video-types/08-brutalist-meme.md) | -| 09 | Editing your own footage, talking heads (experimental) | HyperFrames | Cut where the audio is quiet, not at ASR word times; nothing renders until you approve the edit list | [09](video-types/09-editing-talking-head.md) | +## Sound -A few more guides: -- **Realistic people or real physics:** bring in a generative video model, then layer code on top ([playbook/05](playbook/05-hybrid-genvideo.md)). -- **Effects, transitions, one-take 3D:** [playbook/08](playbook/08-vfx-and-motion-sources.md). -- **Breaking down someone else's film:** [playbook/07](playbook/07-reverse-engineer.md). -- **A story with rises and falls, or a film of 3 minutes or more:** [playbook/09](playbook/09-narrative.md) (structures, beat sheets, the tension curve, act breaks). -- **Posting to short-video platforms: the opening hook, title and cover:** [playbook/10](playbook/10-hooks-and-packaging.md). -- **Music with chapters, a theme you can hum, and real rises and falls:** [playbook/11](playbook/11-composition.md). -- **Research notes, what we measured and what changed because of it:** [docs/research](docs/research/en/README.md) (mix levels, soundtrack distance, render determinism, on-screen reading time, music form, a concept-first A/B test). +An agent can't hear, so the sound is built to be computed and measured: -## Sound +- **Voiceover** (`bin/vh tts`): local, open-source Qwen3-TTS by default (offline and free; five Chinese voices, two English), with Alibaba Cloud Model Studio, ElevenLabs and Gemini TTS as cloud options (showcases 02 and 03 use Gemini). Every line can get its own delivery and stress, and lines can land on the music's beats. +- **Subtitles** (`bin/vh captions`): write the script as `中文 || English` to get Chinese, English or two-line subtitles, burned in or as switchable tracks. +- **Music** (`bin/vh music`): composed as code, so the same score always renders the same audio, with the time of every beat for the picture to hit; it has Chinese instruments such as bianzhong, guzheng and dizi. For your own track, `bin/vh beats` finds the beats. +- **Sound effects** (`bin/vh sfx`): 21 original, code-synthesized effects placed on the frame where the action happens; most vary slightly from use to use, so repeats don't sound identical. +- **Mixing and checks** (`bin/vh mix`, `bin/vh qa`): the music ducks under narration and the mix lands at −14 LUFS; `qa` checks for silence, dropouts, pumping and clipping, and that every cue lands within one frame. -The agent can't hear, so sound is built to be computed and measured: +More in [playbook/04-audio.md](playbook/04-audio.md) (in Chinese). -| You want | Command | Notes | -|---|---|---| -| Voiceover (zh / en) | `bin/vh tts` | Local open-source **Qwen3-TTS** by default: offline, free; the first run downloads about 2 GB of model and about 750 MB of Python packages. 5 Chinese voices (including Beijing and Sichuan accents), 2 English. Interfaces ready for Alibaba Cloud, ElevenLabs and Gemini 3.8 Flash TTS (very expressive; direct the delivery in one sentence) | -| Narration with feeling and rhythm | `bin/vh tts … --beats` | Direct each line on its own, e.g. `[surprised question, fast, stress "one sentence"]`. With music, every line starts on a beat, and key lines can be pinned to a bar start or the drop. Default delivery per video type, frame-aligned tempos and mix settings are in [playbook/04](playbook/04-audio.md) | -| Bilingual subtitles | `bin/vh captions` | Write the script as `中文 \|\| English` and get Chinese, English and two-line subtitles, which can be packed as switchable tracks | -| Music | `bin/vh music` | Composed in code: the same score always gives the same music, plus the exact time of every beat for the picture to hit. Includes Chinese instruments (bells, guzheng, dizi, big drum) and changing time signatures. Using your own track? `bin/vh beats` finds its beats and drum hits. Chapters, a theme and dynamics: [playbook/11](playbook/11-composition.md) | -| Sound effects | `bin/vh sfx` | 21 original synthesized effects, each placed on the frame where its action happens; a sound on the left of the screen comes from the left, and each event gets its own variation of the sound, so repeated cuts don't all sound alike | -| Mix | `bin/vh mix` | A profile per video type sets voice, music and SFX relative to one anchor: the music rides under each narration line to a target level and the SFX are levelled by class; the whole mix is set to −14 LUFS without flattening a cinematic score | -| Mix check | `bin/vh qa` | Measures the finished mix for gaps, dropouts, pumping and clicks, checks every cue lands within 1 frame, and reports the level hierarchy (`qa mix`) | -| Songs | — | Generate in a web service like Suno and import; interfaces ready for ElevenLabs Music and local song models | +## Effort, and who decides -More in [playbook/04-audio.md](playbook/04-audio.md). +Not every film deserves the full process. Say "quick draft" or "studio quality" in the request, or pass `--effort` when you create a project: -## What you end up with +| | `quick` | `standard` (default) | `studio` | +|---|---|---|---| +| For | trying a direction, drafts | most real videos | launches, flagship pieces | +| Stops for you | none | three | three, plus a short sample of each direction and a full-length animatic | +| Second-agent review | none | one round | at least three, all eight scores at 8 or above | +| A 30-second film takes about | 10–30 min | 1–2 h | 3 h or more | -- An **MP4** ready to post, optionally with Chinese and English subtitle tracks; -- A **project you can re-render**: change a line, get a new version, or switch it to another language or style; -- Every file from the process: brief, storyboard, review notes, fix-up log and asset sources; -- Contact sheets and keyframes for review, or to use as a cover. +At every level: no invented facts, no sudden silence in a film with sound, no rapid flashing, and type no smaller than the floor for the target screen. -## Commands `bin/vh` +Apart from the level, you can name what you want to decide yourself: the hook, the look, the main character, the theme tune, the voice, the script, the storyboard, the title and the cover. For those it offers options and waits; the rest it decides, writes down why in the project's `DECISIONS.md`, and you can overrule it at any time. At each stop it builds a local review page (`bin/vh review`) that opens with only the decisions it needs from you, with the frames, animatic and music playable in the browser. -| Command | What it does | -|---|---| -| `doctor` / `setup` | Check your setup / install dependencies and fetch references | -| `types` / `new [--style ] [--effort ] [--aspect 9:16] [--watch phone\|feed\|desktop] [--res 1080p\|4k]` | List the 9 types / start a new project (aspect, target screen and resolution can be set up front) | -| `effort [quick\|standard\|studio]` | What each effort level does | -| `style list` / `style -
+
- - - - - + + + + +