Skip to content

fix(knowledge): let video-digest digest caption-less YouTube videos through ASR #6047

Description

@kyle-sexton

Problem

/knowledge:video-digest watch <url> cannot digest a YouTube video that has no captions in any
language, although the skill documents an ASR fallback for exactly that case. Once that path is
unblocked, the ASR transcript is second-class: it skips vocabulary biasing, proper-noun repair, and
frame densification. Observed on watch https://www.youtube.com/watch?v=MN9dGgmLyso (a 2h54m
post-live stream, uploaded 2026-09-30), knowledge 0.17.0; the cited code is unchanged on main at
6f5909646 (knowledge 0.18.0). Gap numbers follow the source item; it has no gap 6.

  1. Caption-less media hard-fails acquisition. extraction/acquisition/acquire.js:342-343
    returns failVideo(captionResult.error) when caption selection fails, even in full mode with
    the video already downloaded (the retry at :332-340 only re-runs the captions pass), and
    :374-378 fails again on finalCaption. The run exited 1 after downloading the full video with
    No English captions found (manual EN → auto EN → auto-translate EN ladder exhausted); the info
    JSON had empty subtitles and automatic_captions (live_status: post_live). Everything
    downstream already handles a caption-less entry: watch/run-watch.js:284
    (primary.caption ? ... : ""), transcript/write-transcript.js:239-243
    (captionlessEntryCount, per-entry ASR plan) and transcript/transcript-strategy.js (caption
    absent selects asr when available, else degrade and proceed). So the documented behavior,
    SKILL.md:225-232 ("selected automatically for caption-absent entries", "completes the digest
    without a transcript"), is unreachable for YouTube, and --transcript-strategy asr does not help
    because acquisition fails first. SKILL.md:87 contradicts it: "Caption ladder ... STOP and
    surface if exhausted".
  2. No working install path for the ASR prerequisite. transcript/asr-transcribe.js:23 probes
    only python, python3, py on PATH. On a machine whose default Python is uv-managed,
    pip install --user faster-whisper is refused (PEP 668, "externally managed"). The skill gives
    no install command and no way to name an interpreter.
  3. faster-whisper 1.2.1 is incompatible with PyAV 19.0.1 (reported, not reproduced this pass).
    decode_audio passes metadata_errors= to av.open, which PyAV 19 rejects
    (TypeError: open() got an unexpected keyword argument 'metadata_errors'). Pinning av<19
    (18.1.0) fixed it.
  4. ASR transcripts do not feed densification. watch/run-watch.js:284-289 builds the cues
    passed to orchestrateWatching (:297-303) from the caption VTT only. With an ASR transcript
    (145 cues, 78 paragraphs) the run recorded densificationWindows: 0 in run-state/watch.json.
    The ASR cues already exist inside writeEnvelopeTranscriptArtifacts
    (write-transcript.js:192, :262-275) but are not returned to the caller.
  5. The outcome gate cannot pass a talking-head-only video without breaking the synthesis
    contract.
    extraction/evals/check-watch-outcomes.js:350-369 (session-visual-coverage,
    session-synthesis-depth) are fail severity and count only promoted frames (promotedTs);
    visual-gaps.md rows are not consulted for sessions. At 1-4 h the floor is 2 promoted frames per
    session. context/quality-gates.md:28 allows a session to have a frame "or gap logged", and
    context/synthesis-contract.md:9 rejects talking-head-only frames. A webcam interview with no
    screen content therefore cannot close without promoting frames the contract rejects. This run,
    and the earlier boris-cherny-we-cut-80-of-claude-code-s-qyPCVqFUyDo corpus slice, worked
    around it with vision-plan-sanctioned "speaker-identity anchor" frames.
  6. The ASR rung gets no vocabulary and no name repair. write-transcript.js:262-265 calls
    runAsr({ mediaPath, python }) without initialPrompt, although asr-transcribe.js:112,131
    accepts one and asr_transcribe.py:51 passes it to Whisper as initial_prompt. The repair
    lexicon is built only when a route can reach captions+repair (write-transcript.js:229-233),
    never for YouTube's captions default, and from the description and harvested links only, not
    the title. ASR output also skips the proper-noun repair pass (repairedTermCount: 0 at :280).
  7. ASR transcript quality papercuts (measured; improvements). 18.5% of words sit in lowercase,
    unpunctuated stretches, starting with the first 93 words (a batched-decode artifact), and a
    two-person conversation has no speaker labels.
  8. Live-stream metadata duration is stale (observed, not reproduced). The info JSON reported
    duration: 10446 (the live length) while the served post-live edit probed at 3995 s. The
    pipeline correctly used the probed duration, but nothing records the mismatch, so a slice reader
    sees 2:54:06 in metadata and a 66-minute transcript.

Evidence

Verified this pass by reading origin/main at 6f5909646: every path:line above for gaps 1, 2,
4, 5 and 7, including the SKILL.md contradiction (:87 versus :225-232) and the gate docs
(quality-gates.md:28, synthesis-contract.md:9).

Reported by the producing run and not re-run here (installs and live probes were out of scope):

  • Gap 1: a two-check local patch made the caption failure non-fatal when
    mode === "full" && artifacts.videoPath and returned caption: null. It unblocked the run but
    was never committed; the producing worktree is now clean, so no diff survives. The real fix needs
    tests.
  • Gap 2: the working setup was a dedicated venv: uv venv <dir> --python 3.13, then
    uv pip install faster-whisper nvidia-cublas-cu12 nvidia-cudnn-cu12 "av<19", with the venv's
    Scripts dir and each site-packages/nvidia/*/bin dir prepended to PATH for the run. A GPU smoke
    test of asr_transcribe.py (tiny model, 30 s of audio) ran in about 3 s on CUDA.
  • Gap 3: the PyAV TypeError and the av<19 fix.
  • Gap 7, measured against the video: 60 of 109 named-entity occurrences correct (55%); common names
    57 of 61, new product, skill and person names 3 of 48. "pstack" came out as "PSAC", "P stack" or
    "Psat" 9 of 11 times although the video title contains "pstack"; "Grok Bot" 0 of 15; "Cursor"
    heard as "Coursera"; "Karpathy" 0 of 3. A second model (large-v3-turbo) repeats most of these, so
    a re-decode does not fix them.
  • Gap 8: large-v3 and large-v3-turbo disagree on 5.6% of 986 sampled words (2.7% excluding fillers
    and spelling); coverage is complete (11,078 words, 166 wpm, no gap over 60 s), so general wording
    is sound.

Environment of the run: Windows 11, Git Bash, yt-dlp 2026.08.19, ffmpeg 9.0.2, faster-whisper
1.2.1, ctranslate2 4.8.2, CUDA 12 wheels (cublas 12.9.2.10, cudnn 9.27.0.42), RTX 2000 Ada 8 GB.

Duplicate search (open and closed): video-digest caption ASR, video-digest in:title,
youtube-digest, faster-whisper, talking-head OR session-visual-coverage. Only #5982 is
related (below); none covers these gaps.

Proposed approach

Decisions needed first, each with a recommended default:

  • D1, caption ladder exhausted (gap 1). Recommended: in watch (full mode with media
    downloaded), an exhausted ladder is not fatal; the entry proceeds caption-less and the transcript
    strategy picks asr or degrades with a recorded transcriptDegradation. In transcript mode
    (no media) it stays a STOP. Rewrite SKILL.md:87 to say exactly that. Alternative: keep the STOP
    everywhere and delete the ASR-for-caption-absent claim, which leaves caption-less videos
    undigestible.
  • D2, talking-head-only sessions (gap 5). Recommended: let a visual-gaps.md row naming the
    session as talking-head-only satisfy session-visual-coverage and session-synthesis-depth
    (matching quality-gates.md:28), and keep synthesis-contract.md:9's rejection. Alternative:
    bless "speaker-identity anchor" frames in the synthesis contract, which promotes frames that carry
    no information.

Then, by gap:

  1. acquire.js: when mode === "full" && artifacts.videoPath, a failed caption selection returns
    success with caption: null instead of failVideo (both the first selection and
    finalCaption). Tests: acquire with no caption files in full mode (succeeds, caption: null)
    and in transcript mode (still fails).
  2. Keep the "never auto-installed" contract and give the prerequisite a documented, working path:
    a userConfig option naming the ASR interpreter (probed before the PATH candidates), a documented
    venv recipe (the uv commands above) under a suggested location such as the plugin data dir, and
    os.add_dll_directory for each site-packages/nvidia/*/bin in asr_transcribe.py so CUDA works
    without PATH edits. Files: asr-transcribe.js, asr_transcribe.py, SKILL.md Prerequisites,
    plugins/knowledge/.claude-plugin/plugin.json (userConfig).
  3. Document the av<19 pin in the install recipe, with a note to recheck once faster-whisper
    releases past 1.2.1.
  4. Return the cues each transcript was built from (caption or ASR) from
    writeEnvelopeTranscriptArtifacts, and have run-watch.js pass the primary entry's cues to
    orchestrateWatching instead of re-parsing the VTT.
  5. Implement D2 in check-watch-outcomes.js and align quality-gates.md and
    synthesis-contract.md.
  6. Build the lexicon from title + description + harvested links for every strategy
    (buildRepairLexicon in proper-noun-repair.js), pass it as initialPrompt to the ASR call,
    and run repairCues over ASR cues too.
  7. Optional: re-punctuate or re-decode lowercase stretches; speaker labels are out of scope unless
    cheap.
  8. Record metadata duration and probed duration side by side in watch.json and the README when
    they differ by more than a threshold.

Acceptance criteria

  • acquire with no caption files in full mode with a downloaded video returns success with
    caption: null; in transcript mode it still fails. Both covered by tests.
  • A watch run with no captions and ASR available produces a transcript with
    transcriptStrategy: asr; with ASR absent it completes with a recorded
    transcriptDegradation. Covered by a test with a stubbed runAsr / detection.
  • SKILL.md no longer says both "STOP" and "proceeds" for an exhausted caption ladder.
  • An interpreter path can be supplied through a userConfig option, and detection tries it
    first (unit test with a stubbed spawn).
  • SKILL.md Prerequisites contains a copyable faster-whisper install recipe that includes the
    av<19 pin.
  • With ASR cues and no caption VTT, orchestrateWatching receives a non-empty cues array
    (test).
  • runAsr receives a non-empty initialPrompt when the title or description contains proper
    nouns, and ASR cues pass through repairCues (tests).
  • A slice whose sessions each carry a talking-head-only visual-gaps.md row passes
    session-visual-coverage and session-synthesis-depth (or the rule D2 chooses), and the two
    context docs state the same rule.
  • A metadata versus probed duration mismatch is recorded in watch.json.
  • The video-digest extraction suite passes.

Constraints and gotchas

Context

Source: local handoff item 20261003-040023-knowledge-video-digest-captionless-youtube-asr-path.md
(retired into this issue). Related: #5982 (open; data dir and preflight for the same skill), draft
PR #6004 (fixes #5982 E1), #5803 (closed; earlier pipeline fixes, different defects), and a sibling
papercuts issue drafted from the same run (#6048).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority: highSignificant impact, or blocks an imminent release; staff this cycle.status: needs-decisionAwaiting a human or maintainer judgment call.work-class: structuralRefactors, migrations, contract changes; cross-cutting and hard to reverse.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions