Skip to content

feat(video): read playlists and channels as a corpus, and serve video over MCP - #12

Merged
maxgfr merged 2 commits into
mainfrom
issue-7-video-corpus
Sep 30, 2026
Merged

maxgfr merged 2 commits into
mainfrom
issue-7-video-corpus

Conversation

@maxgfr

@maxgfr maxgfr commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Last of four PRs for #7.

  • webindex video list <playlist|channel> [--limit n] lists the videos without reading them (a channel's /videos tab when the URL names none; tabs and nested playlists skipped), reads each as its own run two at a time, and writes CORPUS.md + corpus.json naming them V1…Vn in listing order. A video that cannot be read keeps its label with the reason, so the numbering never shifts.
  • video search on a corpus directory labels its hits V1…Vn; passages no longer straddle a chapter start.
  • MCP: webindex_video_fetch, webindex_video_search, webindex_video_frames, webindex_video_list. The three that keep a run are annotated readOnlyHint: false (additive, idempotent); under --public-only / --extract-root / --allow-remote, dir is a directory name inside the video root, never a path. webindex_video_list reports progress per video.
  • references/video.md (linked from SKILL.md, served over MCP), README tables and counts (359 exports, 28 commands, 20 MCP tools); untitled chapters read as "Chapter N".

Checked live: the "100 Seconds of Code" playlist, 3 videos, read in 36 s (two from auto-captions, one through whisper), then searched with V# labels; webindex mcp over stdio: tools/list shows the 20 tools, webindex_video_fetch and webindex_video_search answer.

Closes #7

… over MCP

webindex video list reads the first --limit videos of a playlist or channel
(its /videos tab when none is named), two at a time, each kept as its own
run, and writes CORPUS.md and corpus.json naming them V1…Vn; a video that
cannot be read keeps its label. video search on a corpus labels its hits.

MCP gains webindex_video_fetch, webindex_video_search, webindex_video_frames
and webindex_video_list. The three that keep a run are annotated as writing
(additively, idempotent); under a policy their dir is a name inside the
video root, never a path. references/video.md documents the whole of it.

Closes #7
…ements

- under NO_WRITE a run is neither written nor collected: fetch hands the
  transcript back (a run already on disk included), list and frames refuse
- a kept run records the track it was read from and is reused only for a
  request in that language; meta.json is written last, so a run cut short is
  never reused, and a run or corpus that cannot be written is a reason
- a media download is judged on its exit status before any file, and
  leftover fragments are never taken for the video
- a playlist that lists one video twice reads and labels it once
- frames are staged and swapped in whole; frames and list are annotated
  destructive, since they replace earlier frames and an earlier corpus

Refs #7
@maxgfr
maxgfr merged commit ecfff17 into main Sep 30, 2026
3 checks passed
@github-actions

Copy link
Copy Markdown

🎉 This PR is included in version 1.25.0 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

@maxgfr
maxgfr deleted the issue-7-video-corpus branch September 30, 2026 12:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(video): YouTube video engine (transcript ladder, frames, multi-video, MCP)

1 participant