Skip to content

Getting Started

zhanglinghao edited this page Oct 4, 2026 · 6 revisions

English · 中文

This page takes you from nothing to a first video. It assumes you already use a coding agent (Claude Code or Codex). The complete install notes are in the README.

1. What you need

Needs For Required?
macOS or Linux, git the basics yes
Node.js ≥ 22, Google Chrome (or Chromium) rendering in a headless browser yes
FFmpeg encoding, mixing, checks yes
Python 3 + uv the sound tools (tts, beats, music, sfx, qa, mix … profile=), contact sheets and Manim; packages are fetched on first use into uv's cache, not into a global environment yes, for sound and contact sheets
Apple Silicon local Qwen3-TTS voiceover only for local voice
LaTeX equations in Manim only for math explainers

Rendering costs nothing. Cloud voices and generative video charge on their own terms; see the FAQ.

2. Install

curl -fsSL https://raw.githubusercontent.com/ZLHad/OpenVideoHarness/main/install.sh | bash

This clones the newest version of the repo into ~/OpenVideoHarness (without its history), installs the bundled hand-drawn engine and the style swatch renderer, fetches 30 read-only reference repos, registers the open-video-harness skill for Claude Code (~/.claude/skills) and Codex (~/.agents/skills), creates LOCAL.md and finally runs bin/vh doctor.

Options: --dir <path> installs elsewhere, --no-refs skips the reference repos (run references/fetch.sh later, or references/fetch.sh <name> for one), --no-skill skips the skill, --skill claude|codex|all chooses where to register it.

By hand, with a full clone (about 95 MB to download and 180 MB on disk, history included; the clone to make if you want to contribute):

git clone https://github.com/ZLHad/OpenVideoHarness.git && cd OpenVideoHarness
bin/vh setup            # engine deps, reference repos, swatch renderer, a render smoke test
bin/vh install-skill    # optional: register the skill

Restart Claude Code or Codex after registering the skill. From then on, asking for a video in any folder leads the agent to the harness.

To update later, run the install command again with the same options (the one-line install takes them after bash -s --: curl -fsSL https://raw.githubusercontent.com/ZLHad/OpenVideoHarness/main/install.sh | bash -s -- --no-refs). In the installer's clone it fetches only the newest commit; it keeps LOCAL.md and projects/, and stops rather than overwrite a file you changed. In a full clone you can also git pull, then references/fetch.sh. A full clone made before 2026-10-04 can't: the history was rewritten that day to drop the old videos, and every commit hash changed. Clone again and move LOCAL.md and projects/ over; don't merge or rebase the old clone, which brings the videos back (CHANGELOG).

What gets downloaded

The install itself takes about 590 MB on disk: the clone, about 185 MB (85 MB to download; the showcase videos are in a GitHub release, and tools/fetch_media.sh downloads them when you rebuild a showcase), the Node packages of the bundled engine and the style renderer, about 210 MB, and the reference repos, about 195 MB. The first uses pull in more, quietly.

Sizes are approximate, measured on macOS (Apple Silicon); du -sh will show slightly different numbers.

What When Where Size
This repo's newest commit (the showcase videos are in a GitHub release, not in it) install ~/OpenVideoHarness about 85 MB to download, 185 MB on disk
Node packages of the hand-drawn engine and the style renderer install node_modules inside the repo about 210 MB
30 read-only reference repos install, unless --no-refs references/repos/ about 195 MB
Chrome for HyperFrames (chrome-headless-shell) the first hyperframes render ~/.cache/hyperframes about 100 MB to download, 200 MB on disk
Python packages for bin/vh beats, music, sfx, qa and sheet (librosa, numba, scipy …) the first time you run each uv's cache, ~/.cache/uv about 700 MB in all
node_modules of a HyperFrames project each bin/vh new short, promo, data or meme inside that project about 140 MB each
The local Qwen3-TTS voice: mlx-audio and its packages, then the model the first bin/vh tts with the default local provider (qwen); not needed with another provider ~/.cache/uv and ~/.cache/huggingface about 750 MB and 2 GB

The installer writes inside the repo, plus (unless you pass --no-skill) two small skill files under ~/.claude/skills and ~/.agents/skills, and whatever npm keeps in its own cache. The rows from chrome-headless-shell down are fetched later, without asking and with little output: bin/vh runs uv quietly, and in a non-interactive shell (which is how an agent runs commands) the first render only shows "Checking browser…" while Chrome downloads. HyperFrames also keeps a small config and log in ~/.hyperframes. All of this lives under your home directory, outside the repo, and uv's and Hugging Face's caches are shared with your other tools. On a slow or blocked connection (mainland China, for example), China network has mirror settings for each row.

Skipping the references. They are other people's skills and film sources, cloned shallowly so that an agent can read them. Rendering doesn't depend on them. Install with --no-refs, and when a workflow doc points at a references/repos/<name>/ you don't have, fetch just that one: references/fetch.sh hyperframes (about 30 MB). references/fetch.sh with no argument fetches all 30.

3. Check your setup

bin/vh doctor prints one line per item: ✓ fine, ✗ required and missing, ! optional or worth a look. It checks Node, FFmpeg, Chrome (with a real headless WebGL probe), HyperFrames' own Chrome, Python, uv, LaTeX, git, the engine's dependencies, the reference repos (counted against what references/fetch.sh fetches), the Qwen3-TTS caches, LOCAL.md (it warns while the {…} placeholders are still in it), and whether AGENTS.md still matches CLAUDE.md. It ends with what the first runs will still download. If something repo-local is missing, bin/vh setup installs it. A missing uv is a red ✗, because the sound tools and contact sheets run through it.

LOCAL.md is never committed. It tells the agent about your machine: hardware, installed tools and engines, render speed, which API keys you have (variable names only, never values), where your papers or assets live, your default effort level, your default for director mode and your preferences. You can paste the doctor output into it. The template is LOCAL.example.md.

4. Your first video in three commands

# 1. install (ends with bin/vh doctor)
curl -fsSL https://raw.githubusercontent.com/ZLHad/OpenVideoHarness/main/install.sh | bash
# 2. open your agent in the repo
cd ~/OpenVideoHarness && claude        # or: codex
  1. Say what you want, in one sentence:
Make a 30-second vertical science short: why does a low-orbit satellite's signal change pitch? English voiceover, English and Chinese subtitles.

Say what it's about, who it's for and where it will be posted. The agent first works out a few concepts for it, fills in the rest from the video type's defaults and asks at most three questions that change the result. Where it will be posted matters: it sets where the film is watched (a phone, a phone feed or a computer), and that sets the frame and the smallest text size. More example requests: README.

If you want to decide more than the three gates, say which decisions are yours: "I'll pick the hook, the main character and the theme melody; decide the rest yourself." That is director mode.

You can also scaffold first with bin/vh new short my-first-video (list the types with bin/vh types), then open the agent and say what the video is about. --aspect 16:9|9:16|1:1 sets the frame for short, promo, data and meme projects, --watch phone|feed|desktop where it will be watched (the default follows the type and frame), --res 1080p|4k the output resolution (4K is written at 1080p and rendered at twice the size), and --dir <folder> puts the project outside the harness, in a folder you name. bin/vh <command> -h prints any command's usage without running it.

5. What happens at the three gates

At each gate the agent makes, by default, one local page (bin/vh review writes it to out/review/; it opens from disk, with images, video and audio you can play), posts only the decisions and the page's path in the chat, and stops until you answer. The page starts with at most three decisions, each with a recommendation and a one-line reply, and everything else comes below. Your words are copied into the project's REVIEW.md, and what the agent decided alone goes into DECISIONS.md.

Gate You get You answer
① Concept and outline 2–3 concept cards that are far apart: each a one-line idea, the look it leads to (its own, or borrowed from one or more style presets) and its hook, with one frame; then, for the recommended card, the brief (what, who for, length, frame, where it's watched, sound), a 3–7 part outline, the engine and why, and any paid API or model download with a cost estimate; the spec defaults it chose for you; at most 3 questions pick a card (you can swap its preset or hook in the reply); approve, revise or start over on the outline
② Storyboard One page per section of the outline, 3–6 shots each: a keyframe per shot with its reads and narration, the shots it's least sure of marked red, and that section's stretch of the animatic; optionally a rhythm map approve, or name the shots to change ("S04: …")
③ First draft out/draft.mp4 and a contact sheet; which checklist items failed and how they were fixed; the latest review scores; the 2–3 things it likes least, with times and fixes; the source of every number on screen approve (it renders the final), or send notes

Notes work best with a time and a reason: "0:12–0:15, too much text to read". If all you can say is "it doesn't feel right", say that. The agent renders 2–3 variants of one 10–20 s segment (look-dev) and you pick a direction.

With effort quick there are no gates: it renders straight away and asks at most one question (decisions you name still stop). With studio, each concept card at gate ① comes with a 10–20 s sketch and gate ② with a full-length animatic. See How It Works.

6. Where things land

Each video is a folder under projects/ (ignored by git), created by bin/vh new <type> <slug>:

projects/2026-09-30-leo-doppler/
├── BRIEF.md  STYLE.md  STORYBOARD.md  REVIEW.md  DECISIONS.md  NOTES.md  LESSONS.md  TASTE_CHECKLIST.md
├── audio/        script, voiceover, timelines, captions, music, beat map, SFX events, mix
├── assets/       images, fonts, icons (with source and license)
├── style-refs/   copies of the style presets you borrow from, when there are any
├── out/review/   the review pages (gate-1.html …, the latest also index.html)
├── out/check/    contact sheets, strips, crops, storyboard pages
└── out/          draft and final MP4s

The engine files sit in the same folder: a copy of ClaudeAnimationBase for hand-drawn and MV projects, a HyperFrames scaffold for shorts, promos, data stories and memes. For Manim projects, bin/vh new prints the uv commands to set it up. With --style <preset> (several: --style a,b) a copy of each preset goes into style-refs/ as a reference and is listed on the BRIEF's Style refs line; the film's own STYLE.md still starts from the concept, and what it borrows is one line in DECISIONS.md. When a decision comes up, the agent copies in SCRIPT.md, CHARACTER.md or PACKAGING.md (a lean template each). REVIEW.md keeps your words, DECISIONS.md keeps who decided what and why, and NOTES.md keeps the material list, facts, self-review and the asset log. The concept and the spec (frame, resolution, frame rate, length, where it's watched) are at the top of BRIEF.md.

You end up with an MP4 (optionally with switchable Chinese and English subtitle tracks), a project you can re-render after a one-line change, and every file from the process.

Platform notes

macOS is where the harness was built and tested, on Apple Silicon. Local Qwen3-TTS needs Apple Silicon, and the say voices exist only on macOS. CI also runs bin/vh under the system /bin/bash 3.2 that Macs ship with.

Linux is covered by CI (Ubuntu) and by a round of fixes made while setting up an Ubuntu 24.04 machine. Those fixes are on main but not yet in a tagged release, so run git pull if you installed earlier. Install Chromium with your package manager or npx playwright install chromium, or point CHROME_PATH at a browser. Without a GPU, bin/vh doctor tells you whether software WebGL works; if it does, render with --soft-gl (bin/vh setup retries its smoke test that way). For voice, local Qwen3-TTS and say are not available: use edge (free, unofficial) or a cloud provider, see Sound and Voice.

Windows isn't covered: the requirements list macOS or Linux.

Next

Clone this wiki locally