diff --git a/README.md b/README.md
index 59d60c0..b2d3b60 100644
--- a/README.md
+++ b/README.md
@@ -195,6 +195,8 @@ Read [the agent guide](docs/AGENT_GUIDE.md) before automating the CLI and [the p
- Record the real workflow in Cursorful at 1080p, then caption/render it here. Cursorful remains an operator tool, not a project dependency.
- Generate a local ElevenLabs narration file with `npm run voiceover:elevenlabs -- --script narration.txt --voice-id VOICE_ID`.
+- Build a reusable local ElevenLabs phrase library with `npm run voiceover:library -- --budget 36000`; use `--resume` to continue safely. The library is indexed by voice and phrase category, and its manifests contain no API key.
+- Restore the generated library from the [voice library guide](docs/VOICE_LIBRARY.md) when working from a fresh clone.
- Create opt-in, human-reviewed fal assets with `npm run fal:image-edit -- ... --approved-for-generated-marketing` or `npm run fal:reference-video -- ... --approved-for-generated-marketing`.
- Generated images/video are never eBay source-of-truth/main listing photos or evidence of condition. Full setup and QA details: [AI provider workflows](docs/AI_PROVIDERS.md).
@@ -1361,6 +1363,24 @@ npm run render:clip -- \
--vertical-contain
```
+Plan subject-aware 9:16 framing:
+
+```bash
+npm run portrait:analyze -- \
+ --video "/path/to/clip.mp4" \
+ --out "outputs/clip.framing.json" \
+ --auto
+
+npm run render:clip -- \
+ --video "/path/to/clip.mp4" \
+ --captions "outputs/clip.captions.json" \
+ --out "outputs/clip.subject-framed.mp4" \
+ --vertical \
+ --framing "outputs/clip.framing.json"
+```
+
+`--auto` uses an optional local YOLO/OpenCV adapter when its dependencies and model are available. If they are not available, the command writes a reviewable center-framing fallback. For a quick manual adjustment, use `--center-x 0.32` instead. Framing plans are non-destructive JSON inputs; they do not modify the source video.
+
Useful `render:clip` options:
| Option | Meaning |
@@ -1373,6 +1393,8 @@ Useful `render:clip` options:
| `--vertical` | 1080x1920 cropped fill. |
| `--vertical-contain` | 1080x1920 contained with black bars. |
| `--foreground-video FILE` | Optional transparent foreground/subject layer rendered above captions. |
+| `--framing FILE` | JSON subject-framing plan from `portrait:analyze`. |
+| `--center-x N` | Manual horizontal subject center from `0` to `1`. |
| `--fit cover\|contain` | CSS video fit. Normally controlled by `caption-style.json`. |
| `--position NAME` | `left-hook`, `right-hook`, `lower-left`, `center-bottom`, or `center-impact`. |
| `--combine-ms N` | Caption grouping window. |
diff --git a/assets/clipcaptionai-self-ad/plan.svg b/assets/clipcaptionai-self-ad/plan.svg
new file mode 100644
index 0000000..79f4962
--- /dev/null
+++ b/assets/clipcaptionai-self-ad/plan.svg
@@ -0,0 +1,13 @@
+
diff --git a/assets/clipcaptionai-self-ad/qa.svg b/assets/clipcaptionai-self-ad/qa.svg
new file mode 100644
index 0000000..66f869f
--- /dev/null
+++ b/assets/clipcaptionai-self-ad/qa.svg
@@ -0,0 +1,13 @@
+
diff --git a/assets/clipcaptionai-self-ad/render.svg b/assets/clipcaptionai-self-ad/render.svg
new file mode 100644
index 0000000..64b6767
--- /dev/null
+++ b/assets/clipcaptionai-self-ad/render.svg
@@ -0,0 +1,13 @@
+
diff --git a/docs/AI_PROVIDERS.md b/docs/AI_PROVIDERS.md
index abc89ab..ac6b916 100644
--- a/docs/AI_PROVIDERS.md
+++ b/docs/AI_PROVIDERS.md
@@ -17,6 +17,14 @@ npm run voiceover:elevenlabs -- \
The command writes MP3 audio and a sibling generation manifest containing the voice/model IDs, text hash, response request ID, and audio hash. It never writes the key or narration text into that manifest.
+To spend a bounded character budget building reusable local assets across the configured cloned voice and available premade voices:
+
+```bash
+npm run voiceover:library -- --budget 36000 --resume
+```
+
+This writes `outputs/voiceover/elevenlabs-library/library.json`, one MP3 plus one non-secret manifest per phrase, and retries only safe provider failures. Use `--dry-run` before a large batch; the command checks the live subscription balance and leaves a safety reserve. The checked-in phrase catalog covers hooks, workflow, features, captions, B-roll, quality, and calls to action. Generated audio remains subject to human review for pronunciation, tone, and licensing suitability.
+
For the demo, review the exact narration before generation and make sure it clearly explains how Codex and GPT-5.6 were used.
## fal reviewed marketing assets
diff --git a/docs/ASSET_RECOVERY.md b/docs/ASSET_RECOVERY.md
new file mode 100644
index 0000000..edd29c4
--- /dev/null
+++ b/docs/ASSET_RECOVERY.md
@@ -0,0 +1,36 @@
+# Asset Recovery Bundle
+
+The `asset-recovery-2026-07-23` GitHub Release preserves the verified self-created ClipCaptionAI B-roll cards and the local public-source SFX library.
+
+The B-roll bundle contains:
+
+- `plan.svg`
+- `render.svg`
+- `qa.svg`
+
+The separate `clipcaptionai-sfx-library.tar.gz` release asset contains all 256 local SFX files plus `sfx-library/index.json`.
+
+Restore them from a fresh checkout with:
+
+```bash
+mkdir -p /tmp/clipcaptionai-asset-recovery
+gh release download asset-recovery-2026-07-23 \
+ --repo jongan69/ClipCaptionAI \
+ --pattern 'clipcaptionai-cleared-assets.tar.gz' \
+ --dir /tmp/clipcaptionai-asset-recovery
+tar -xzf /tmp/clipcaptionai-asset-recovery/clipcaptionai-cleared-assets.tar.gz \
+ -C .
+```
+
+To restore the SFX library as well:
+
+```bash
+gh release download asset-recovery-2026-07-23 \
+ --repo jongan69/ClipCaptionAI \
+ --pattern 'clipcaptionai-sfx-library.tar.gz' \
+ --dir /tmp/clipcaptionai-asset-recovery
+tar -xzf /tmp/clipcaptionai-asset-recovery/clipcaptionai-sfx-library.tar.gz \
+ -C .
+```
+
+The SFX files are preserved because they are publicly available local assets, but public availability is not the same as verified commercial-use clearance. Review source terms before publishing a video commercially. The `music-library/` manifest explicitly marks its tracks `review_before_commercial_use`, so music remains excluded. The downloaded YouTube/movie `scene-library/` is intentionally excluded until its rights are reviewed.
diff --git a/docs/VOICE_LIBRARY.md b/docs/VOICE_LIBRARY.md
new file mode 100644
index 0000000..ed7b20a
--- /dev/null
+++ b/docs/VOICE_LIBRARY.md
@@ -0,0 +1,29 @@
+# ElevenLabs Voice Library
+
+The repository contains the generator and manifests for the local ElevenLabs phrase library. The generated MP3 files are distributed as a GitHub Release asset so normal clones stay small.
+
+## Restore the generated library
+
+From a fresh checkout:
+
+```bash
+mkdir -p /tmp/clipcaptionai-voice-library
+gh release download voice-library-2026-07-23 \
+ --repo jongan69/ClipCaptionAI \
+ --pattern 'clipcaptionai-elevenlabs-library.tar.gz' \
+ --dir /tmp/clipcaptionai-voice-library
+tar -xzf /tmp/clipcaptionai-voice-library/clipcaptionai-elevenlabs-library.tar.gz \
+ -C outputs/voiceover
+```
+
+The archive restores `outputs/voiceover/elevenlabs-library/`, including 672 MP3 clips, per-clip generation manifests, and `library.json`. Verify the download before extracting it:
+
+```text
+SHA-256: 3fb5e58e7a6acde17ac81c4c78ddb5e7294dd5a80fc89d4a80b142adc81f3d29
+```
+
+The generated audio is reusable production material, but still requires human review for pronunciation, tone, and suitability. The source generator is resumable and checks the live ElevenLabs balance:
+
+```bash
+npm run voiceover:library -- --resume --budget 36000
+```
diff --git a/examples/clipcaptionai-self-ad-brief.txt b/examples/clipcaptionai-self-ad-brief.txt
new file mode 100644
index 0000000..fe35d59
--- /dev/null
+++ b/examples/clipcaptionai-self-ad-brief.txt
@@ -0,0 +1,8 @@
+Meet ClipCaptionAI.
+Turn a creative brief into a real video run.
+Use your footage, images, captions, B-roll, narration, and AI providers.
+Let your coding model direct the workflow through a reproducible CLI.
+Plan the shots and inspect the run before spending provider credits.
+Render a polished cut from the command line.
+Every render gets a manifest, hashes, output metadata, and technical QA.
+ClipCaptionAI. Prompt it. Render it. Ship the cut.
diff --git a/examples/clipcaptionai-self-ad-voiceover.txt b/examples/clipcaptionai-self-ad-voiceover.txt
new file mode 100644
index 0000000..520d403
--- /dev/null
+++ b/examples/clipcaptionai-self-ad-voiceover.txt
@@ -0,0 +1 @@
+Meet ClipCaptionAI. It turns a creative brief into a real video run. Use your own footage and images, add captions, B-roll, narration, music, or AI-generated assets. Your coding model can direct the workflow through a reproducible command line. Every render records its plan, input hashes, output metadata, and technical quality checks. ClipCaptionAI: prompt it, render it, and ship the cut.
diff --git a/package.json b/package.json
index f4e1bab..0d6d8bb 100644
--- a/package.json
+++ b/package.json
@@ -46,6 +46,7 @@
"logo:verify": "tsx scripts/logo/verify-variant.tsx",
"logo:render": "node scripts/logo/render-all.mjs",
"render:clip": "node scripts/render-clip.mjs",
+ "portrait:analyze": "node scripts/portrait-framing.mjs",
"render:batch": "node scripts/render-batch.mjs",
"rerender:clip": "node scripts/rerender-clip.mjs",
"smart:clips": "node scripts/smart-clips.mjs",
@@ -99,6 +100,7 @@
"sample:props": "node scripts/make-sample-props.mjs",
"video": "node scripts/video.mjs",
"voiceover:elevenlabs": "node scripts/generate-elevenlabs-voiceover.mjs",
+ "voiceover:library": "node scripts/generate-elevenlabs-library.mjs",
"fal:image-edit": "node scripts/fal-image-edit.mjs",
"fal:reference-video": "node scripts/fal-reference-video.mjs",
"typecheck": "tsc --noEmit",
diff --git a/scripts/clipkit.mjs b/scripts/clipkit.mjs
index f52f183..8f8d502 100644
--- a/scripts/clipkit.mjs
+++ b/scripts/clipkit.mjs
@@ -1030,6 +1030,7 @@ Examples:
clipcaptionai review-moments --write --format markdown
clipcaptionai auto-clips --links links.txt --max-clips 6
clipcaptionai broll-captions --links links.txt --max-clips 3
+ clipcaptionai portrait-analyze --video input.mp4 --out framing.json --auto
clipcaptionai caption --video "/path/to/video.mp4"
clipcaptionai rotato render ~/Desktop/demo.rotato --screen-media ~/Desktop/app.mp4 --output outputs/mockups/demo.mp4
clipcaptionai voiceover --script narration.txt --voice-id VOICE_ID
@@ -1050,6 +1051,7 @@ Examples:
configurePassthroughCommand(program, 'review-moments', 'Review why moments were flagged as viral and optionally persist scorecards.', runReviewMoments, ['review']);
configurePassthroughCommand(program, 'auto-clips', 'Download YouTube links, pick viral clips, caption, and render.', runAutoClips, ['auto']);
configurePassthroughCommand(program, 'broll-captions', 'Run the B-roll-heavy labeled workflow.', runBrollCaptions, ['heavy']);
+ configurePassthroughCommand(program, 'portrait-analyze', 'Analyze or plan subject-aware 9:16 framing for a video.', (args) => npmRun('portrait:analyze', args), ['portrait']);
configurePassthroughCommand(program, 'caption', 'Caption any existing video with the current caption style.', (args) => npmRun('caption:auto', args));
configurePassthroughCommand(program, 'enhance', 'Add contextual B-roll and captions to an existing edit.', (args) => npmRun('broll:enhance', args));
configurePassthroughCommand(program, 'broll', 'Find reusable B-roll clips from a text prompt file.', runBroll, ['finder']);
diff --git a/scripts/generate-elevenlabs-library.mjs b/scripts/generate-elevenlabs-library.mjs
new file mode 100644
index 0000000..2e14891
--- /dev/null
+++ b/scripts/generate-elevenlabs-library.mjs
@@ -0,0 +1,187 @@
+#!/usr/bin/env node
+import {createHash} from 'node:crypto';
+import fs from 'node:fs';
+import path from 'node:path';
+import {fileURLToPath} from 'node:url';
+import {ensureDir, loadEnv, outputsRoot, parseArgs} from './lib.mjs';
+
+const scriptName = path.basename(fileURLToPath(import.meta.url));
+const args = parseArgs(process.argv.slice(2));
+loadEnv();
+const usage = `
+Usage:
+ npm run voiceover:library -- --budget 36000
+ npm run voiceover:library -- --resume --budget 36000
+
+Options:
+ --budget N Maximum planned characters. Default: 36000.
+ --reserve N Safety reserve below the account limit. Default: 2000.
+ --out-dir DIR Default: outputs/voiceover/elevenlabs-library.
+ --model ID Default: eleven_multilingual_v2.
+ --output-format FORMAT Default: mp3_44100_128.
+ --resume Skip clips whose audio and manifest already exist.
+ --dry-run Show the planned library and estimated character cost.
+ --max-clips N Generate at most N clips after planning.
+
+Requires ELEVENLABS_API_KEY in .env or the environment. The key is never
+written to output files or printed.
+`;
+
+if (args.help || args.h) {
+ console.log(usage);
+ process.exit(0);
+}
+
+const clean = (value) => String(value ?? '').replace(/\s+/g, ' ').trim();
+const sha256 = (value) => createHash('sha256').update(value).digest('hex');
+const numeric = (value, fallback) => {
+ const parsed = Number(value ?? fallback);
+ if (!Number.isFinite(parsed) || parsed < 0) throw new Error(`Invalid numeric value: ${value}`);
+ return parsed;
+};
+const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
+
+const voices = [
+ ['jon', process.env.ELEVENLABS_VOICE_ID, 'configured cloned voice'],
+ ['bella', 'hpp4J3VqNfWAUOO0d1Us', 'professional bright warm'],
+ ['roger', 'CwhRBWXzGAHq8TQ4Fs17', 'laid-back casual resonant'],
+ ['sarah', 'EXAVITQu4vr4xnSDxMaL', 'mature reassuring confident'],
+ ['laura', 'FGY2WhTYpPnrIDTdsKH5', 'enthusiastic social creator'],
+ ['charlie', 'IKne3meq5aSn9XLyUdCD', 'deep confident energetic'],
+ ['liam', 'TX3LPaxmHKxFdv7VOQHJ', 'energetic social creator'],
+ ['alice', 'Xb7hH8MSUJpSbSDYk0k2', 'clear engaging educator'],
+ ['eric', 'cjVigY5qzO86Huf0OWal', 'smooth trustworthy'],
+ ['george', 'JBFqnCBsd6RMkjVDRZzb', 'warm captivating storyteller'],
+ ['callum', 'N2lVS1w4EtoT3dr4eOWO', 'husky character voice'],
+ ['river', 'SAz9YHcvj6GT2YYXdXww', 'relaxed neutral informative'],
+ ['harry', 'SOYHLrjzK2X1ezoPC6cr', 'fierce character voice'],
+ ['matilda', 'XrExE9yKIg1WjnnlVkGX', 'knowledgeable professional'],
+ ['will', 'bIHbv24MWmeRgasZH58o', 'relaxed optimist'],
+ ['jessica', 'cgSgspJ2msm6clMCkdW9', 'playful bright warm'],
+ ['brian', 'nPczCjzI2devNBz1zQrb', 'deep resonant comforting'],
+ ['daniel', 'onwK4e9ZLuTAKqWW03F9', 'steady broadcaster'],
+ ['lily', 'pFZP5JQG7iQjIQuC4Bku', 'velvety actress'],
+ ['adam', 'pNInz6obpgDQGcFmaJgB', 'dominant firm'],
+ ['bill', 'pqHfZKP75CvOlQylNhV4', 'wise mature balanced'],
+].filter(([, voiceId]) => clean(voiceId));
+
+// These are intentionally complete, reusable spoken assets rather than one long ad.
+// Each phrase can be cut into a hook, explainer, transition, or CTA in a future edit.
+const phrases = [
+ ['hook', 'Meet ClipCaptionAI, the command-line video editor built for creative teams and AI models.'],
+ ['hook', 'Start with a brief, a folder of approved assets, and a clear outcome. ClipCaptionAI turns that direction into a video run.'],
+ ['hook', 'Your next product video should not begin with a blank timeline. It should begin with a plan you can inspect.'],
+ ['hook', 'From idea to finished cut, ClipCaptionAI keeps the creative brief, assets, render, and quality checks connected.'],
+ ['workflow', 'Plan a run before spending provider credits. Review the shots, sources, prompts, framing, and export settings first.'],
+ ['workflow', 'The model can direct the workflow, while the CLI keeps every decision explicit, reproducible, and easy to resume.'],
+ ['workflow', 'A run manifest records what was requested, what was rendered, which files were used, and what passed technical QA.'],
+ ['workflow', 'Use dry-run mode to validate paths, providers, and output settings before an external generation call begins.'],
+ ['workflow', 'Resume an existing run instead of starting over. The manifest is the handoff point between planning, rendering, and review.'],
+ ['feature', 'Bring your own footage, product photos, logos, captions, music, sound effects, and approved B-roll into one composition.'],
+ ['feature', 'Generate clean vertical, horizontal, or contained layouts from versioned configuration instead of editing source code.'],
+ ['feature', 'Choose shot recipes, caption styles, audio presets, and export settings that match the channel you are publishing to.'],
+ ['feature', 'Add narration as a real audio input, mix it above a music bed, and verify that the final file is not silently broken.'],
+ ['feature', 'Use local assets for reliable demos, then add OpenAI, ElevenLabs, or fal generation when the brief calls for it.'],
+ ['feature', 'The renderer stays deterministic, so a model can make creative choices without losing control of the final export.'],
+ ['feature', 'Every successful run produces a final artifact, a manifest, hashes, media metadata, and a machine-readable QA result.'],
+ ['feature', 'The desktop app is optional and thin. The CLI is the production surface that works for people, scripts, and coding agents.'],
+ ['captions', 'Captions are part of the composition, not an afterthought. Keep the message readable, paced, and safe inside the frame.'],
+ ['captions', 'Use a clear headline for the hook, supporting copy for the proof, and a concise call to action at the end.'],
+ ['captions', 'A good caption survives muted playback. A good voiceover adds rhythm, context, and confidence without fighting the visuals.'],
+ ['broll', 'Show the work as it happens: the brief becomes a plan, the plan becomes a render, and the render becomes a checked deliverable.'],
+ ['broll', 'Use interface captures, product details, source footage, and workflow cards to make the benefit visible in seconds.'],
+ ['broll', 'B-roll should prove the product promise. Show inputs, decisions, transformations, and the final result instead of decorative noise.'],
+ ['quality', 'Before you ship, check that the file exists, the duration is valid, the dimensions are correct, the codec is supported, and the audio is present.'],
+ ['quality', 'Technical QA catches black screens, missing audio, wrong framing, broken paths, and incomplete renders before your audience does.'],
+ ['quality', 'A passing manifest is evidence about this artifact and this run. It does not pretend that an unverified provider completed work remotely.'],
+ ['quality', 'Keep secrets in the environment. Keep prompts, model IDs, request IDs, hashes, and QA state in the non-secret manifest.'],
+ ['cta', 'ClipCaptionAI. Prompt it, render it, inspect it, and ship the cut.'],
+ ['cta', 'Turn the next creative brief into a video you can actually review. Try ClipCaptionAI today.'],
+ ['cta', 'Stop losing the story between the prompt and the export. Keep the whole run in one place with ClipCaptionAI.'],
+ ['cta', 'Build once, review clearly, and reuse the assets that work. ClipCaptionAI is your model-facing video production CLI.'],
+ ['cta', 'When the brief is ready, the next step is simple: plan the run, render the cut, and let QA tell you what shipped.'],
+];
+
+const apiKey = clean(process.env.ELEVENLABS_API_KEY);
+const budget = numeric(args.budget, 36000);
+const reserve = numeric(args.reserve, 2000);
+const modelId = clean(args.model ?? 'eleven_multilingual_v2');
+const outputFormat = clean(args['output-format'] ?? 'mp3_44100_128');
+const outDir = path.resolve(args['out-dir'] ?? path.join(outputsRoot, 'voiceover', 'elevenlabs-library'));
+const indexPath = path.join(outDir, 'library.json');
+const resume = args.resume === true;
+
+if (!voices.length) throw new Error('No ElevenLabs voices configured. Set ELEVENLABS_VOICE_ID in .env.');
+const planned = [];
+for (const [voiceKey, voiceId, voiceDescription] of voices) {
+ for (const [index, [category, text]] of phrases.entries()) {
+ const id = `${String(index + 1).padStart(2, '0')}-${category}-${voiceKey}`;
+ const audio = path.join(outDir, voiceKey, `${id}.mp3`);
+ planned.push({id, category, voice_key: voiceKey, voice_id: voiceId, voice_description: voiceDescription, text, text_characters: text.length, audio});
+ }
+}
+const maxClips = args['max-clips'] === undefined ? planned.length : Math.floor(numeric(args['max-clips'], planned.length));
+const selected = planned.slice(0, maxClips);
+const pending = resume
+ ? selected.filter((item) => {
+ const manifestPath = item.audio.replace(/\\.mp3$/i, '.generation.json');
+ return !(fs.existsSync(item.audio) && fs.existsSync(manifestPath));
+ })
+ : selected;
+// eleven_multilingual_v2 currently reports a character cost near 0.5x raw text
+// for this account. Keep a conservative ceil per clip, while recording the
+// authoritative provider cost in each generation manifest.
+const estimated = pending.reduce((sum, item) => sum + Math.ceil(item.text_characters * 0.5), 0);
+if (estimated > budget) throw new Error(`Planned text costs ${estimated} characters, above --budget ${budget}. Reduce --max-clips or increase the budget.`);
+
+if (args['dry-run'] === true) {
+ console.log(JSON.stringify({provider: 'elevenlabs', model_id: modelId, output_format: outputFormat, voices: voices.map(([key, id, description]) => ({key, voice_id: id, description})), clips: selected.length, pending_clips: pending.length, estimated_characters: estimated, budget, reserve, output_directory: outDir, dry_run: true}, null, 2));
+ process.exit(0);
+}
+if (!apiKey) throw new Error('ELEVENLABS_API_KEY is required in .env or the environment.');
+ensureDir(outDir);
+
+const subscriptionResponse = await fetch('https://api.elevenlabs.io/v1/user/subscription', {headers: {'xi-api-key': apiKey}});
+if (!subscriptionResponse.ok) throw new Error(`Could not read ElevenLabs subscription (${subscriptionResponse.status}).`);
+const subscription = await subscriptionResponse.json();
+const remaining = Math.max(0, Number(subscription.character_limit ?? 0) - Number(subscription.character_count ?? 0));
+if (estimated > Math.max(0, remaining - reserve)) {
+ throw new Error(`Planned text costs ${estimated} characters, but only ${remaining} remain after the ${reserve}-character safety reserve.`);
+}
+
+const existing = fs.existsSync(indexPath) ? JSON.parse(fs.readFileSync(indexPath, 'utf8')) : null;
+const entries = new Map((existing?.entries ?? []).map((entry) => [entry.id, entry]));
+const failures = [];
+for (const [position, item] of selected.entries()) {
+ const manifestPath = item.audio.replace(/\.mp3$/i, '.generation.json');
+ if (resume && fs.existsSync(item.audio) && fs.existsSync(manifestPath)) continue;
+ const endpoint = `https://api.elevenlabs.io/v1/text-to-speech/${encodeURIComponent(item.voice_id)}?output_format=${encodeURIComponent(outputFormat)}`;
+ let response;
+ for (let attempt = 1; attempt <= 3; attempt += 1) {
+ response = await fetch(endpoint, {method: 'POST', headers: {'content-type': 'application/json', 'xi-api-key': apiKey}, body: JSON.stringify({text: item.text, model_id: modelId})});
+ if (response.ok || ![408, 409, 429, 500, 502, 503, 504].includes(response.status) || attempt === 3) break;
+ await sleep(1500 * attempt);
+ }
+ if (!response.ok) {
+ const detail = (await response.text()).slice(0, 500);
+ failures.push({id: item.id, status: response.status, detail});
+ console.error(`Failed ${position + 1}/${selected.length}: ${item.id} (${response.status})`);
+ continue;
+ }
+ const audioBuffer = Buffer.from(await response.arrayBuffer());
+ if (!audioBuffer.length) throw new Error(`ElevenLabs returned empty audio for ${item.id}.`);
+ ensureDir(path.dirname(item.audio));
+ fs.writeFileSync(item.audio, audioBuffer);
+ const entry = {...item, model_id: modelId, output_format: outputFormat, text_sha256: sha256(item.text), audio_sha256: sha256(audioBuffer), audio_bytes: audioBuffer.length, character_cost: Number(response.headers.get('character-cost') ?? item.text_characters), request_id: response.headers.get('request-id'), created_at: new Date().toISOString(), manifest: manifestPath};
+ fs.writeFileSync(manifestPath, `${JSON.stringify({provider: 'elevenlabs', script: scriptName, ...entry}, null, 2)}\n`);
+ entries.set(item.id, entry);
+ fs.writeFileSync(indexPath, `${JSON.stringify({provider: 'elevenlabs', model_id: modelId, output_format: outputFormat, generated_at: new Date().toISOString(), entries: [...entries.values()], failures}, null, 2)}\n`);
+ console.error(`Generated ${position + 1}/${selected.length}: ${item.id}`);
+ await sleep(250);
+}
+
+const finalEntries = [...entries.values()];
+const totalTextCharacters = selected.reduce((sum, item) => sum + item.text_characters, 0);
+const totalBilledCharacters = finalEntries.reduce((sum, entry) => sum + Number(entry.character_cost ?? entry.text_characters), 0);
+fs.writeFileSync(indexPath, `${JSON.stringify({provider: 'elevenlabs', model_id: modelId, output_format: outputFormat, generated_at: new Date().toISOString(), planned_clips: selected.length, planned_text_characters: totalTextCharacters, estimated_billable_characters: estimated, generated_clips: finalEntries.length, generated_billable_characters: totalBilledCharacters, failures, entries: finalEntries}, null, 2)}\n`);
+console.log(JSON.stringify({provider: 'elevenlabs', output_directory: outDir, index: indexPath, planned_clips: selected.length, generated_clips: finalEntries.length, planned_text_characters: totalTextCharacters, generated_billable_characters: totalBilledCharacters, failures: failures.length, remaining_before_run: remaining}, null, 2));
diff --git a/scripts/portrait-framing.mjs b/scripts/portrait-framing.mjs
new file mode 100644
index 0000000..a507f42
--- /dev/null
+++ b/scripts/portrait-framing.mjs
@@ -0,0 +1,65 @@
+#!/usr/bin/env node
+import fs from 'node:fs';
+import path from 'node:path';
+import {spawnSync} from 'node:child_process';
+import {ensureDir, parseArgs, probeVideo, requireArg} from './lib.mjs';
+
+const usage = `
+Usage:
+ npm run portrait:analyze -- --video input.mp4 --out framing.json [options]
+
+Options:
+ --auto Try the optional local YOLO/OpenCV detector.
+ --center-x N Manual subject center from 0 to 1. Default: 0.5.
+ --model FILE Local YOLO model for --auto. Default: models/yolov8n.pt.
+`;
+
+const args = parseArgs(process.argv.slice(2));
+if (args.help || args.h) {
+ console.log(usage);
+ process.exit(0);
+}
+
+const video = path.resolve(requireArg(args, 'video', usage));
+const out = path.resolve(requireArg(args, 'out', usage));
+const metadata = probeVideo(video);
+const manualCenter = Number(args['center-x'] ?? 0.5);
+if (!Number.isFinite(manualCenter) || manualCenter < 0 || manualCenter > 1) {
+ throw new Error('--center-x must be a number between 0 and 1.');
+}
+
+const fallback = (source = 'fallback', reason = 'No subject detector was available.') => ({
+ schemaVersion: 1,
+ video,
+ source,
+ strategy: metadata.width / metadata.height > 1 ? 'track' : 'contain',
+ confidence: source === 'manual' ? 1 : 0,
+ reason,
+ keyframes: [
+ {at: 0, centerX: manualCenter, confidence: source === 'manual' ? 1 : 0},
+ {at: 1, centerX: manualCenter, confidence: source === 'manual' ? 1 : 0},
+ ],
+});
+
+let plan = args['center-x'] !== undefined ? fallback('manual', 'Manual subject center.') : null;
+if (args.auto && metadata.width / metadata.height > 1) {
+ const detector = spawnSync('python3', [
+ path.join(path.dirname(new URL(import.meta.url).pathname), 'portrait-framing.py'),
+ '--video', video,
+ '--model', path.resolve(String(args.model ?? 'models/yolov8n.pt')),
+ ], {encoding: 'utf8'});
+ if (detector.status === 0) {
+ try {
+ plan = JSON.parse(detector.stdout);
+ plan.video = video;
+ } catch {
+ plan = fallback('fallback', 'Detector returned invalid JSON.');
+ }
+ } else {
+ plan = fallback('fallback', 'Optional YOLO/OpenCV detector unavailable; using center framing.');
+ }
+}
+plan ??= fallback();
+ensureDir(path.dirname(out));
+fs.writeFileSync(out, `${JSON.stringify(plan, null, 2)}\n`);
+console.log(JSON.stringify({ok: true, out, strategy: plan.strategy, source: plan.source, confidence: plan.confidence}));
diff --git a/scripts/portrait-framing.py b/scripts/portrait-framing.py
new file mode 100644
index 0000000..dde2877
--- /dev/null
+++ b/scripts/portrait-framing.py
@@ -0,0 +1,73 @@
+#!/usr/bin/env python3
+"""Optional local subject framing adapter.
+
+This file deliberately has no required project dependency. Install OpenCV,
+Ultralytics, and a local YOLO model to enable --auto; the Node wrapper falls
+back to a reviewable center plan when they are unavailable.
+"""
+import argparse
+import json
+import sys
+
+
+def main():
+ parser = argparse.ArgumentParser()
+ parser.add_argument('--video', required=True)
+ parser.add_argument('--model', required=True)
+ args = parser.parse_args()
+
+ try:
+ import cv2
+ from ultralytics import YOLO
+ except ImportError as exc:
+ print(f'optional detector unavailable: {exc}', file=sys.stderr)
+ return 2
+
+ try:
+ model = YOLO(args.model)
+ capture = cv2.VideoCapture(args.video)
+ if not capture.isOpened():
+ raise RuntimeError('Could not open video.')
+ frame_count = max(1, int(capture.get(cv2.CAP_PROP_FRAME_COUNT)))
+ fps = capture.get(cv2.CAP_PROP_FPS) or 30
+ keyframes = []
+ for index in range(9):
+ frame_index = int((frame_count - 1) * index / 8)
+ capture.set(cv2.CAP_PROP_POS_FRAMES, frame_index)
+ ok, frame = capture.read()
+ if not ok:
+ continue
+ result = model(frame, verbose=False)[0]
+ candidates = []
+ for box in result.boxes:
+ if int(box.cls[0]) != 0:
+ continue
+ x1, y1, x2, y2 = [float(value) for value in box.xyxy[0]]
+ confidence = float(box.conf[0])
+ area = max(0, x2 - x1) * max(0, y2 - y1)
+ candidates.append((area, (x1 + x2) / 2 / max(1, frame.shape[1]), confidence))
+ if candidates:
+ _, center_x, confidence = max(candidates)
+ keyframes.append({
+ 'at': index / 8,
+ 'centerX': max(0.05, min(0.95, center_x)),
+ 'confidence': round(confidence, 4),
+ })
+ capture.release()
+ if not keyframes:
+ raise RuntimeError('No person detections found.')
+ print(json.dumps({
+ 'schemaVersion': 1,
+ 'source': 'detector',
+ 'strategy': 'track',
+ 'confidence': round(sum(item['confidence'] for item in keyframes) / len(keyframes), 4),
+ 'keyframes': keyframes,
+ }))
+ return 0
+ except Exception as exc:
+ print(f'detector failed: {exc}', file=sys.stderr)
+ return 3
+
+
+if __name__ == '__main__':
+ raise SystemExit(main())
diff --git a/scripts/render-clip.mjs b/scripts/render-clip.mjs
index a062b77..ec7dc7e 100644
--- a/scripts/render-clip.mjs
+++ b/scripts/render-clip.mjs
@@ -34,6 +34,8 @@ Options:
--text-opacity N Caption fill opacity. Default: 0.92.
--frames START-END Optional Remotion frame range for proof renders.
--uppercase Render caption text uppercase.
+ --framing FILE JSON framing plan from portrait:analyze.
+ --center-x N Manual horizontal subject center from 0 to 1.
`;
const args = parseArgs(process.argv.slice(2));
@@ -73,6 +75,25 @@ const highlightedWords = args['highlight-words']
? styleConfig.highlightedWords
: [];
+const readFraming = () => {
+ if (args.framing) {
+ const framingPath = path.resolve(String(args.framing));
+ return JSON.parse(fs.readFileSync(framingPath, 'utf8'));
+ }
+ if (args['center-x'] !== undefined) {
+ const centerX = Number(args['center-x']);
+ if (!Number.isFinite(centerX) || centerX < 0 || centerX > 1) {
+ throw new Error('--center-x must be a number between 0 and 1.');
+ }
+ return {
+ strategy: 'track',
+ source: 'manual',
+ keyframes: [{at: 0, centerX}, {at: 1, centerX}],
+ };
+ }
+ return null;
+};
+
const props = {
videoSrc: videoToSrc(video),
foregroundSrc: foregroundVideo ? videoToSrc(foregroundVideo) : null,
@@ -81,6 +102,7 @@ const props = {
height,
fps,
durationInFrames: Math.max(1, Math.ceil(metadata.durationSeconds * fps)),
+ framing: readFraming(),
style: {
...styleConfig,
position: String(args.position ?? styleConfig.position ?? 'left-hook'),
diff --git a/src/captioned-clip.tsx b/src/captioned-clip.tsx
index d00e48d..76f2944 100644
--- a/src/captioned-clip.tsx
+++ b/src/captioned-clip.tsx
@@ -16,6 +16,7 @@ import type {
CaptionMotionPreset,
CaptionPosition,
CaptionStyle,
+ PortraitFraming,
} from './types';
export const captionedClipDefaultProps: CaptionedClipProps = {
@@ -26,6 +27,7 @@ export const captionedClipDefaultProps: CaptionedClipProps = {
height: 1920,
fps: 30,
durationInFrames: 450,
+ framing: null,
style: {
position: 'left-hook',
fit: 'cover',
@@ -543,6 +545,32 @@ const matrixToCss = (matrix: MotionMatrix) =>
const matrixToSvg = (matrix: MotionMatrix) =>
`matrix(${matrix.a} ${matrix.b} ${matrix.c} ${matrix.d} ${matrix.e} ${matrix.f})`;
+const framingCenterX = (framing: PortraitFraming | null | undefined, frame: number, durationInFrames: number) => {
+ if (!framing || framing.strategy !== 'track' || framing.keyframes.length === 0) {
+ return 50;
+ }
+
+ const progress = durationInFrames <= 1 ? 0 : Math.max(0, Math.min(1, frame / (durationInFrames - 1)));
+ const keyframes = [...framing.keyframes]
+ .filter((keyframe) => Number.isFinite(keyframe.at) && Number.isFinite(keyframe.centerX))
+ .sort((a, b) => a.at - b.at);
+ if (keyframes.length === 0) return 50;
+ if (progress <= keyframes[0].at) return Math.max(0, Math.min(100, keyframes[0].centerX * 100));
+
+ for (let index = 1; index < keyframes.length; index += 1) {
+ const next = keyframes[index];
+ const previous = keyframes[index - 1];
+ if (progress <= next.at) {
+ const span = Math.max(0.0001, next.at - previous.at);
+ const local = (progress - previous.at) / span;
+ const center = previous.centerX + (next.centerX - previous.centerX) * local;
+ return Math.max(0, Math.min(100, center * 100));
+ }
+ }
+
+ return Math.max(0, Math.min(100, keyframes[keyframes.length - 1].centerX * 100));
+};
+
const escapeXml = (value: string) =>
value
.replaceAll('&', '&')
@@ -599,10 +627,11 @@ export const CaptionedClip: React.FC = ({
videoSrc,
foregroundSrc,
captions,
+ framing,
style,
}) => {
const frame = useCurrentFrame();
- const {fps, width, height} = useVideoConfig();
+ const {fps, width, height, durationInFrames} = useVideoConfig();
const currentMs = (frame / fps) * 1000;
const normalFontFamily =
style.normalFontFamily ??
@@ -617,6 +646,7 @@ export const CaptionedClip: React.FC = ({
const videoSource = /^https?:\/\//.test(videoSrc)
? videoSrc
: staticFile(videoSrc);
+ const objectPosition = `${framingCenterX(framing, frame, durationInFrames)}% 50%`;
const foregroundSource =
foregroundSrc && /^https?:\/\//.test(foregroundSrc)
? foregroundSrc
@@ -1079,6 +1109,7 @@ export const CaptionedClip: React.FC = ({
width: '100%',
height: '100%',
objectFit: style.fit,
+ objectPosition,
borderRadius: style.videoBorderRadius ?? 0,
filter: style.videoFilter ?? 'none',
}}
@@ -1111,6 +1142,7 @@ export const CaptionedClip: React.FC = ({
muted
style={{
...maskedVideoBaseStyle,
+ objectPosition,
filter: [style.videoFilter ?? 'none', effectNormalFilter]
.filter((value) => value && value !== 'none')
.join(' '),
@@ -1138,6 +1170,7 @@ export const CaptionedClip: React.FC = ({
muted
style={{
...maskedVideoBaseStyle,
+ objectPosition,
filter: [style.videoFilter ?? 'none', effectHighlightFilter]
.filter((value) => value && value !== 'none')
.join(' '),
@@ -1194,6 +1227,7 @@ export const CaptionedClip: React.FC = ({
width: '100%',
height: '100%',
objectFit: style.fit,
+ objectPosition,
borderRadius: style.videoBorderRadius ?? 0,
pointerEvents: 'none',
}}
diff --git a/src/types.ts b/src/types.ts
index 60cc540..be79f05 100644
--- a/src/types.ts
+++ b/src/types.ts
@@ -24,6 +24,18 @@ export type CaptionMotionKeyframe = {
rotateDeg?: number;
};
+export type PortraitFramingKeyframe = {
+ at: number;
+ centerX: number;
+ confidence?: number;
+};
+
+export type PortraitFraming = {
+ strategy: 'center' | 'track' | 'contain';
+ source?: 'manual' | 'detector' | 'fallback';
+ keyframes: PortraitFramingKeyframe[];
+};
+
export type CaptionStyle = {
position: CaptionPosition;
customPosition?: Record;
@@ -107,5 +119,6 @@ export type CaptionedClipProps = {
height: number;
fps: number;
durationInFrames: number;
+ framing?: PortraitFraming | null;
style: CaptionStyle;
};
diff --git a/tests/cli-smoke.test.mjs b/tests/cli-smoke.test.mjs
index 5ec083d..e8ea957 100644
--- a/tests/cli-smoke.test.mjs
+++ b/tests/cli-smoke.test.mjs
@@ -21,6 +21,7 @@ test('clipkit top-level help renders the polished command hub', () => {
assert.match(result.stdout, /review-moments\|review/);
assert.match(result.stdout, /rotato\|mockup/);
assert.match(result.stdout, /video/);
+ assert.match(result.stdout, /portrait-analyze\|portrait/);
assert.match(result.stdout, /fal-reference-video/);
assert.match(result.stdout, /voiceover\|elevenlabs/);
assert.match(result.stdout, /rerender --clip 03-your-website-is-leaking-money --no-captions/);
@@ -86,6 +87,27 @@ test('bin entry works and exposes help output', () => {
assert.match(result.stdout, /rotato\|mockup/);
});
+test('portrait framing planner writes a manual track plan', () => {
+ const outDir = fs.mkdtempSync(path.join(os.tmpdir(), 'cca-portrait-'));
+ const output = path.join(outDir, 'framing.json');
+ const result = spawnSync('node', ['scripts/portrait-framing.mjs',
+ '--video', path.join(projectRoot, 'public', 'listingos-horizontal-demo-rotato-enhanced-20260718.mp4'),
+ '--out', output,
+ '--center-x', '0.32',
+ ], {cwd: projectRoot, encoding: 'utf8'});
+
+ try {
+ assert.equal(result.status, 0, result.stderr);
+ const plan = JSON.parse(fs.readFileSync(output, 'utf8'));
+ assert.equal(plan.source, 'manual');
+ assert.equal(plan.strategy, 'track');
+ assert.equal(plan.keyframes[0].centerX, 0.32);
+ assert.equal(plan.keyframes[1].centerX, 0.32);
+ } finally {
+ fs.rmSync(outDir, {recursive: true, force: true});
+ }
+});
+
test('moments review helper exposes the standalone report command', () => {
const result = spawnSync('node', ['scripts/review-moments.mjs', '--help'], {
cwd: projectRoot,
diff --git a/tests/fixtures/portrait-captions.json b/tests/fixtures/portrait-captions.json
new file mode 100644
index 0000000..b5ef9a8
--- /dev/null
+++ b/tests/fixtures/portrait-captions.json
@@ -0,0 +1,8 @@
+[
+ {
+ "text": "ClipCaptionAI framing smoke test",
+ "startMs": 0,
+ "endMs": 1200,
+ "timestampMs": 600
+ }
+]