You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
fix(knowledge): let video-digest digest caption-less YouTube videos through ASR #6047
/knowledge:video-digest watch <url> cannot digest a YouTube video that has no captions in any
language, although the skill documents an ASR fallback for exactly that case. Once that path is
unblocked, the ASR transcript is second-class: it skips vocabulary biasing, proper-noun repair, and
frame densification. Observed on watch https://www.youtube.com/watch?v=MN9dGgmLyso (a 2h54m
post-live stream, uploaded 2026-09-30), knowledge 0.17.0; the cited code is unchanged on main at 6f5909646 (knowledge 0.18.0). Gap numbers follow the source item; it has no gap 6.
Caption-less media hard-fails acquisition.extraction/acquisition/acquire.js:342-343
returns failVideo(captionResult.error) when caption selection fails, even in full mode with
the video already downloaded (the retry at :332-340 only re-runs the captions pass), and :374-378 fails again on finalCaption. The run exited 1 after downloading the full video with No English captions found (manual EN → auto EN → auto-translate EN ladder exhausted); the info
JSON had empty subtitles and automatic_captions (live_status: post_live). Everything
downstream already handles a caption-less entry: watch/run-watch.js:284
(primary.caption ? ... : ""), transcript/write-transcript.js:239-243
(captionlessEntryCount, per-entry ASR plan) and transcript/transcript-strategy.js (caption
absent selects asr when available, else degrade and proceed). So the documented behavior, SKILL.md:225-232 ("selected automatically for caption-absent entries", "completes the digest
without a transcript"), is unreachable for YouTube, and --transcript-strategy asr does not help
because acquisition fails first. SKILL.md:87 contradicts it: "Caption ladder ... STOP and
surface if exhausted".
No working install path for the ASR prerequisite.transcript/asr-transcribe.js:23 probes
only python, python3, py on PATH. On a machine whose default Python is uv-managed, pip install --user faster-whisper is refused (PEP 668, "externally managed"). The skill gives
no install command and no way to name an interpreter.
faster-whisper 1.2.1 is incompatible with PyAV 19.0.1 (reported, not reproduced this pass). decode_audio passes metadata_errors= to av.open, which PyAV 19 rejects
(TypeError: open() got an unexpected keyword argument 'metadata_errors'). Pinning av<19
(18.1.0) fixed it.
ASR transcripts do not feed densification.watch/run-watch.js:284-289 builds the cues
passed to orchestrateWatching (:297-303) from the caption VTT only. With an ASR transcript
(145 cues, 78 paragraphs) the run recorded densificationWindows: 0 in run-state/watch.json.
The ASR cues already exist inside writeEnvelopeTranscriptArtifacts
(write-transcript.js:192, :262-275) but are not returned to the caller.
The outcome gate cannot pass a talking-head-only video without breaking the synthesis
contract.extraction/evals/check-watch-outcomes.js:350-369 (session-visual-coverage, session-synthesis-depth) are fail severity and count only promoted frames (promotedTs); visual-gaps.md rows are not consulted for sessions. At 1-4 h the floor is 2 promoted frames per
session. context/quality-gates.md:28 allows a session to have a frame "or gap logged", and context/synthesis-contract.md:9 rejects talking-head-only frames. A webcam interview with no
screen content therefore cannot close without promoting frames the contract rejects. This run,
and the earlier boris-cherny-we-cut-80-of-claude-code-s-qyPCVqFUyDo corpus slice, worked
around it with vision-plan-sanctioned "speaker-identity anchor" frames.
The ASR rung gets no vocabulary and no name repair.write-transcript.js:262-265 calls runAsr({ mediaPath, python }) without initialPrompt, although asr-transcribe.js:112,131
accepts one and asr_transcribe.py:51 passes it to Whisper as initial_prompt. The repair
lexicon is built only when a route can reach captions+repair (write-transcript.js:229-233),
never for YouTube's captions default, and from the description and harvested links only, not
the title. ASR output also skips the proper-noun repair pass (repairedTermCount: 0 at :280).
ASR transcript quality papercuts (measured; improvements). 18.5% of words sit in lowercase,
unpunctuated stretches, starting with the first 93 words (a batched-decode artifact), and a
two-person conversation has no speaker labels.
Live-stream metadata duration is stale (observed, not reproduced). The info JSON reported duration: 10446 (the live length) while the served post-live edit probed at 3995 s. The
pipeline correctly used the probed duration, but nothing records the mismatch, so a slice reader
sees 2:54:06 in metadata and a 66-minute transcript.
Evidence
Verified this pass by reading origin/main at 6f5909646: every path:line above for gaps 1, 2,
4, 5 and 7, including the SKILL.md contradiction (:87 versus :225-232) and the gate docs
(quality-gates.md:28, synthesis-contract.md:9).
Reported by the producing run and not re-run here (installs and live probes were out of scope):
Gap 1: a two-check local patch made the caption failure non-fatal when mode === "full" && artifacts.videoPath and returned caption: null. It unblocked the run but
was never committed; the producing worktree is now clean, so no diff survives. The real fix needs
tests.
Gap 2: the working setup was a dedicated venv: uv venv <dir> --python 3.13, then uv pip install faster-whisper nvidia-cublas-cu12 nvidia-cudnn-cu12 "av<19", with the venv's Scripts dir and each site-packages/nvidia/*/bin dir prepended to PATH for the run. A GPU smoke
test of asr_transcribe.py (tiny model, 30 s of audio) ran in about 3 s on CUDA.
Gap 3: the PyAV TypeError and the av<19 fix.
Gap 7, measured against the video: 60 of 109 named-entity occurrences correct (55%); common names
57 of 61, new product, skill and person names 3 of 48. "pstack" came out as "PSAC", "P stack" or
"Psat" 9 of 11 times although the video title contains "pstack"; "Grok Bot" 0 of 15; "Cursor"
heard as "Coursera"; "Karpathy" 0 of 3. A second model (large-v3-turbo) repeats most of these, so
a re-decode does not fix them.
Gap 8: large-v3 and large-v3-turbo disagree on 5.6% of 986 sampled words (2.7% excluding fillers
and spelling); coverage is complete (11,078 words, 166 wpm, no gap over 60 s), so general wording
is sound.
Environment of the run: Windows 11, Git Bash, yt-dlp 2026.08.19, ffmpeg 9.0.2, faster-whisper
1.2.1, ctranslate2 4.8.2, CUDA 12 wheels (cublas 12.9.2.10, cudnn 9.27.0.42), RTX 2000 Ada 8 GB.
Duplicate search (open and closed): video-digest caption ASR, video-digest in:title, youtube-digest, faster-whisper, talking-head OR session-visual-coverage. Only #5982 is
related (below); none covers these gaps.
Proposed approach
Decisions needed first, each with a recommended default:
D1, caption ladder exhausted (gap 1). Recommended: in watch (full mode with media
downloaded), an exhausted ladder is not fatal; the entry proceeds caption-less and the transcript
strategy picks asr or degrades with a recorded transcriptDegradation. In transcript mode
(no media) it stays a STOP. Rewrite SKILL.md:87 to say exactly that. Alternative: keep the STOP
everywhere and delete the ASR-for-caption-absent claim, which leaves caption-less videos
undigestible.
D2, talking-head-only sessions (gap 5). Recommended: let a visual-gaps.md row naming the
session as talking-head-only satisfy session-visual-coverage and session-synthesis-depth
(matching quality-gates.md:28), and keep synthesis-contract.md:9's rejection. Alternative:
bless "speaker-identity anchor" frames in the synthesis contract, which promotes frames that carry
no information.
Then, by gap:
acquire.js: when mode === "full" && artifacts.videoPath, a failed caption selection returns
success with caption: null instead of failVideo (both the first selection and finalCaption). Tests: acquire with no caption files in full mode (succeeds, caption: null)
and in transcript mode (still fails).
Keep the "never auto-installed" contract and give the prerequisite a documented, working path:
a userConfig option naming the ASR interpreter (probed before the PATH candidates), a documented
venv recipe (the uv commands above) under a suggested location such as the plugin data dir, and os.add_dll_directory for each site-packages/nvidia/*/bin in asr_transcribe.py so CUDA works
without PATH edits. Files: asr-transcribe.js, asr_transcribe.py, SKILL.md Prerequisites, plugins/knowledge/.claude-plugin/plugin.json (userConfig).
Document the av<19 pin in the install recipe, with a note to recheck once faster-whisper
releases past 1.2.1.
Return the cues each transcript was built from (caption or ASR) from writeEnvelopeTranscriptArtifacts, and have run-watch.js pass the primary entry's cues to orchestrateWatching instead of re-parsing the VTT.
Implement D2 in check-watch-outcomes.js and align quality-gates.md and synthesis-contract.md.
Build the lexicon from title + description + harvested links for every strategy
(buildRepairLexicon in proper-noun-repair.js), pass it as initialPrompt to the ASR call,
and run repairCues over ASR cues too.
Optional: re-punctuate or re-decode lowercase stretches; speaker labels are out of scope unless
cheap.
Record metadata duration and probed duration side by side in watch.json and the README when
they differ by more than a threshold.
Acceptance criteria
acquire with no caption files in full mode with a downloaded video returns success with caption: null; in transcript mode it still fails. Both covered by tests.
A watch run with no captions and ASR available produces a transcript with transcriptStrategy: asr; with ASR absent it completes with a recorded transcriptDegradation. Covered by a test with a stubbed runAsr / detection.
SKILL.md no longer says both "STOP" and "proceeds" for an exhausted caption ladder.
An interpreter path can be supplied through a userConfig option, and detection tries it
first (unit test with a stubbed spawn).
SKILL.md Prerequisites contains a copyable faster-whisper install recipe that includes the av<19 pin.
With ASR cues and no caption VTT, orchestrateWatching receives a non-empty cues array
(test).
runAsr receives a non-empty initialPrompt when the title or description contains proper
nouns, and ASR cues pass through repairCues (tests).
A slice whose sessions each carry a talking-head-only visual-gaps.md row passes session-visual-coverage and session-synthesis-depth (or the rule D2 chooses), and the two
context docs state the same rule.
A metadata versus probed duration mismatch is recorded in watch.json.
Never auto-install faster-whisper; ASR stays optional and detected at runtime.
Tests must not run a real Whisper model; stub runAsr and detection.
Context
Source: local handoff item 20261003-040023-knowledge-video-digest-captionless-youtube-asr-path.md
(retired into this issue). Related: #5982 (open; data dir and preflight for the same skill), draft
PR #6004 (fixes #5982 E1), #5803 (closed; earlier pipeline fixes, different defects), and a sibling
papercuts issue drafted from the same run (#6048).
Problem
/knowledge:video-digest watch <url>cannot digest a YouTube video that has no captions in anylanguage, although the skill documents an ASR fallback for exactly that case. Once that path is
unblocked, the ASR transcript is second-class: it skips vocabulary biasing, proper-noun repair, and
frame densification. Observed on
watch https://www.youtube.com/watch?v=MN9dGgmLyso(a 2h54mpost-live stream, uploaded 2026-09-30), knowledge 0.17.0; the cited code is unchanged on main at
6f5909646(knowledge 0.18.0). Gap numbers follow the source item; it has no gap 6.extraction/acquisition/acquire.js:342-343returns
failVideo(captionResult.error)when caption selection fails, even infullmode withthe video already downloaded (the retry at
:332-340only re-runs the captions pass), and:374-378fails again onfinalCaption. The run exited 1 after downloading the full video withNo English captions found (manual EN → auto EN → auto-translate EN ladder exhausted); the infoJSON had empty
subtitlesandautomatic_captions(live_status: post_live). Everythingdownstream already handles a caption-less entry:
watch/run-watch.js:284(
primary.caption ? ... : ""),transcript/write-transcript.js:239-243(
captionlessEntryCount, per-entry ASR plan) andtranscript/transcript-strategy.js(captionabsent selects
asrwhen available, else degrade and proceed). So the documented behavior,SKILL.md:225-232("selected automatically for caption-absent entries", "completes the digestwithout a transcript"), is unreachable for YouTube, and
--transcript-strategy asrdoes not helpbecause acquisition fails first.
SKILL.md:87contradicts it: "Caption ladder ... STOP andsurface if exhausted".
transcript/asr-transcribe.js:23probesonly
python,python3,pyon PATH. On a machine whose default Python is uv-managed,pip install --user faster-whisperis refused (PEP 668, "externally managed"). The skill givesno install command and no way to name an interpreter.
decode_audiopassesmetadata_errors=toav.open, which PyAV 19 rejects(
TypeError: open() got an unexpected keyword argument 'metadata_errors'). Pinningav<19(18.1.0) fixed it.
watch/run-watch.js:284-289builds thecuespassed to
orchestrateWatching(:297-303) from the caption VTT only. With an ASR transcript(145 cues, 78 paragraphs) the run recorded
densificationWindows: 0inrun-state/watch.json.The ASR cues already exist inside
writeEnvelopeTranscriptArtifacts(
write-transcript.js:192,:262-275) but are not returned to the caller.contract.
extraction/evals/check-watch-outcomes.js:350-369(session-visual-coverage,session-synthesis-depth) arefailseverity and count only promoted frames (promotedTs);visual-gaps.mdrows are not consulted for sessions. At 1-4 h the floor is 2 promoted frames persession.
context/quality-gates.md:28allows a session to have a frame "or gap logged", andcontext/synthesis-contract.md:9rejects talking-head-only frames. A webcam interview with noscreen content therefore cannot close without promoting frames the contract rejects. This run,
and the earlier
boris-cherny-we-cut-80-of-claude-code-s-qyPCVqFUyDocorpus slice, workedaround it with vision-plan-sanctioned "speaker-identity anchor" frames.
write-transcript.js:262-265callsrunAsr({ mediaPath, python })withoutinitialPrompt, althoughasr-transcribe.js:112,131accepts one and
asr_transcribe.py:51passes it to Whisper asinitial_prompt. The repairlexicon is built only when a route can reach
captions+repair(write-transcript.js:229-233),never for YouTube's
captionsdefault, and from the description and harvested links only, notthe title. ASR output also skips the proper-noun repair pass (
repairedTermCount: 0at:280).unpunctuated stretches, starting with the first 93 words (a batched-decode artifact), and a
two-person conversation has no speaker labels.
duration: 10446(the live length) while the served post-live edit probed at 3995 s. Thepipeline correctly used the probed duration, but nothing records the mismatch, so a slice reader
sees 2:54:06 in metadata and a 66-minute transcript.
Evidence
Verified this pass by reading
origin/mainat6f5909646: everypath:lineabove for gaps 1, 2,4, 5 and 7, including the SKILL.md contradiction (
:87versus:225-232) and the gate docs(
quality-gates.md:28,synthesis-contract.md:9).Reported by the producing run and not re-run here (installs and live probes were out of scope):
mode === "full" && artifacts.videoPathand returnedcaption: null. It unblocked the run butwas never committed; the producing worktree is now clean, so no diff survives. The real fix needs
tests.
uv venv <dir> --python 3.13, thenuv pip install faster-whisper nvidia-cublas-cu12 nvidia-cudnn-cu12 "av<19", with the venv'sScriptsdir and eachsite-packages/nvidia/*/bindir prepended to PATH for the run. A GPU smoketest of
asr_transcribe.py(tiny model, 30 s of audio) ran in about 3 s on CUDA.TypeErrorand theav<19fix.57 of 61, new product, skill and person names 3 of 48. "pstack" came out as "PSAC", "P stack" or
"Psat" 9 of 11 times although the video title contains "pstack"; "Grok Bot" 0 of 15; "Cursor"
heard as "Coursera"; "Karpathy" 0 of 3. A second model (large-v3-turbo) repeats most of these, so
a re-decode does not fix them.
and spelling); coverage is complete (11,078 words, 166 wpm, no gap over 60 s), so general wording
is sound.
Environment of the run: Windows 11, Git Bash, yt-dlp 2026.08.19, ffmpeg 9.0.2, faster-whisper
1.2.1, ctranslate2 4.8.2, CUDA 12 wheels (cublas 12.9.2.10, cudnn 9.27.0.42), RTX 2000 Ada 8 GB.
Duplicate search (open and closed):
video-digest caption ASR,video-digest in:title,youtube-digest,faster-whisper,talking-head OR session-visual-coverage. Only #5982 isrelated (below); none covers these gaps.
Proposed approach
Decisions needed first, each with a recommended default:
watch(full mode with mediadownloaded), an exhausted ladder is not fatal; the entry proceeds caption-less and the transcript
strategy picks
asror degrades with a recordedtranscriptDegradation. Intranscriptmode(no media) it stays a STOP. Rewrite
SKILL.md:87to say exactly that. Alternative: keep the STOPeverywhere and delete the ASR-for-caption-absent claim, which leaves caption-less videos
undigestible.
visual-gaps.mdrow naming thesession as talking-head-only satisfy
session-visual-coverageandsession-synthesis-depth(matching
quality-gates.md:28), and keepsynthesis-contract.md:9's rejection. Alternative:bless "speaker-identity anchor" frames in the synthesis contract, which promotes frames that carry
no information.
Then, by gap:
acquire.js: whenmode === "full" && artifacts.videoPath, a failed caption selection returnssuccess with
caption: nullinstead offailVideo(both the first selection andfinalCaption). Tests: acquire with no caption files in full mode (succeeds,caption: null)and in transcript mode (still fails).
a userConfig option naming the ASR interpreter (probed before the PATH candidates), a documented
venv recipe (the uv commands above) under a suggested location such as the plugin data dir, and
os.add_dll_directoryfor eachsite-packages/nvidia/*/bininasr_transcribe.pyso CUDA workswithout PATH edits. Files:
asr-transcribe.js,asr_transcribe.py,SKILL.mdPrerequisites,plugins/knowledge/.claude-plugin/plugin.json(userConfig).av<19pin in the install recipe, with a note to recheck once faster-whisperreleases past 1.2.1.
writeEnvelopeTranscriptArtifacts, and haverun-watch.jspass the primary entry's cues toorchestrateWatchinginstead of re-parsing the VTT.check-watch-outcomes.jsand alignquality-gates.mdandsynthesis-contract.md.(
buildRepairLexiconinproper-noun-repair.js), pass it asinitialPromptto the ASR call,and run
repairCuesover ASR cues too.cheap.
watch.jsonand the README whenthey differ by more than a threshold.
Acceptance criteria
acquirewith no caption files infullmode with a downloaded video returns success withcaption: null; intranscriptmode it still fails. Both covered by tests.transcriptStrategy: asr; with ASR absent it completes with a recordedtranscriptDegradation. Covered by a test with a stubbedrunAsr/ detection.SKILL.mdno longer says both "STOP" and "proceeds" for an exhausted caption ladder.first (unit test with a stubbed spawn).
av<19pin.orchestrateWatchingreceives a non-emptycuesarray(test).
runAsrreceives a non-emptyinitialPromptwhen the title or description contains propernouns, and ASR cues pass through
repairCues(tests).visual-gaps.mdrow passessession-visual-coverageandsession-synthesis-depth(or the rule D2 chooses), and the twocontext docs state the same rule.
watch.json.Constraints and gotchas
knowledge: video-digest reads its dependency directory only from an environment variable the Bash tool never carries, so its pipeline and preflight fail before any digest work #5982 E2 shows
${user_config.*}in spoke files arrives unsubstituted when read with the Readtool.
faster-whisper import). The install recipe and interpreter option are new here; coordinate so one
probe serves both.
--data-dir), so SKILL.mdline numbers on main will shift once it merges. Rebase on it if it lands first.
main holds at merge time and add a
plugins/knowledge/CHANGELOG.mdentry. If vendored code underplugins/knowledge/vendor/is touched, the vendor version-bump check applies too(
check-vendor-version-bump.sh --check-bump, as in Video-digest pipeline: fix caption class, scene timestamps, gap bounds, max-gap sampling, paragraph overlap (8a-8e) and mark-phase #5803).runAsrand detection.Context
Source: local handoff item
20261003-040023-knowledge-video-digest-captionless-youtube-asr-path.md(retired into this issue). Related: #5982 (open; data dir and preflight for the same skill), draft
PR #6004 (fixes #5982 E1), #5803 (closed; earlier pipeline fixes, different defects), and a sibling
papercuts issue drafted from the same run (#6048).