YouTube Transcript Extractor
VoxScripta is a library-first Go toolkit for acquiring normalized, timestamped transcripts from public YouTube videos. A small CLI is included as a development, testing, and diagnostic harness over the same public API.
Status: caption-first development release. Caption discovery, selection, retrieval, WebVTT normalization, checked audio acquisition, and an opt-in local
whisper.cppadapter are implemented. The adapter has comprehensive offline tests and a recorded whisper.cpp 1.9.2 local-runtime evaluation. Public APIs may change before v1.
Applications that summarize, search, cite, or otherwise process video speech need more than a plain text blob. They need timestamps, language and source information, predictable selection rules, cancellation, and errors they can act on.
YouTube's unofficial extraction surfaces change frequently. Instead of embedding a fragile reimplementation, this project uses yt-dlp for caption discovery and retrieval, then normalizes the result behind an idiomatic Go API. The implemented, explicitly configured local whisper.cpp fallback can cover videos that have no usable captions.
- Accept common YouTube URLs or video IDs.
- Discover creator-provided and automatically generated captions.
- Select tracks deterministically using caller-supplied language preferences.
- Preserve ordered start/end timestamps and caption-source metadata.
- Return one provider-independent transcript model.
- Render plain text through the library and structured JSON through the CLI; WebVTT, SRT, and Markdown renderers are possible later additions.
- Respect context cancellation and deadlines.
- Expose useful errors for missing dependencies, invalid input, and unavailable transcripts.
- Permit custom acquisition providers without forcing them on ordinary users.
Optional provider fallback is explicit. transcript.FallbackProvider invokes
its fallback only when the primary provider reports
ErrTranscriptUnavailable; it does not turn cancellation, invalid input,
missing dependencies, or provider failures into additional work. Transcriber
adapters use this composition without becoming core runtime dependencies. The
public AudioSource and Transcriber contracts are
separate, and SpeechToTextProvider composes them while enforcing optional
duration/file-size limits and closing acquired audio on every post-acquisition
path. Cleanup failures are returned. YTDLPAudioSource is the concrete audio
acquisition adapter. WhisperCPPTranscriber passes verified mono 16 kHz 16-bit
PCM WAV through unchanged and uses FFmpeg for other input before consuming
whisper-cli JSON output.
The public package owns domain types, options, orchestration, and stable behavior. Provider-specific details remain internal or behind narrow interfaces.
Importing Go application Development CLI
| |
+---------- public API ----------+
|
selection/orchestration
|
yt-dlp provider
|
manual or automatic captions
|
parse and normalize
|
timestamped Transcript result
When explicitly configured by a library consumer:
no captions -> YTDLPAudioSource -> checked Audio -> Transcriber -> normalize
The CLI will remain a thin API consumer. Extraction logic will not live in cmd/.
The module is github.com/mpsanders/VoxScripta and its public package name is transcript. The default client uses yt-dlp:
client, err := transcript.New(
transcript.WithYTDLPPath("yt-dlp"),
)
if err != nil {
log.Fatal(err)
}
result, err := client.Get(ctx, "https://www.youtube.com/watch?v=VIDEO_ID", transcript.Options{
Languages: []string{"en-AU", "en"},
AllowAutomatic: true,
})
if err != nil {
log.Fatal(err)
}
fmt.Print(result.Text())The canonical result retains segment timing rather than reducing the transcript to text immediately:
type Segment struct {
Start time.Duration
End time.Duration
Text string
}The returned transcript is owned by the caller. A client may be reused concurrently; custom providers must provide their own concurrency safety.
Provider process errors include at most 2 KiB of normalized stderr. URLs and common credential-bearing values are redacted, and command arguments are never included. Applications should still handle diagnostics as potentially sensitive operational data rather than publishing them verbatim.
WebVTT data can already be parsed independently of a provider:
segments, err := transcript.ParseWebVTT(reader)
if err != nil {
log.Fatal(err)
}The development command is ytextract:
ytextract --language en-AU --language en VIDEO_URL
ytextract --format json VIDEO_ID
ytextract --timeout 30s VIDEO_URL
ytextract --manual-only VIDEO_URL
ytextract --check
ytextract --check --yt-dlp /path/to/yt-dlp
ytextract --whisper-model /path/to/ggml-base.bin VIDEO_URLTranscript data is written to stdout and diagnostics to stderr so the command can be composed with other tools.
The CLI includes automatic captions by default. In the library, automatic captions are explicit: set Options.AllowAutomatic to true. Empty language preferences select the video's reported original language when possible, then deterministically fall back to the first eligible track. Translation is outside the caption-only API and may be added later behind a distinct interface.
CLI JSON uses human-readable Go duration strings such as "0s", "1.25s",
and "2m3s" for segment timestamps. It retains the transcript's video,
language, source, provider, and segment structure.
Supplying --whisper-model explicitly enables captions -> local whisper.cpp.
The --whisper-cli, --ffmpeg, --max-audio-duration (default 2h), and
--max-audio-bytes (default 200 MiB) flags configure that fallback. No model is
inferred or downloaded, and no hosted provider is enabled from ambient
credentials.
--check verifies that the configured yt-dlp executable starts and reports
its version. It performs no video or network acquisition.
The caption provider requires a compatible yt-dlp executable available on PATH or supplied explicitly in configuration.
The library will not silently install external tools. Speech-to-text audio
acquisition uses the same caller-installed yt-dlp. The optional whisper.cpp
adapter requires whisper-cli, a caller-selected GGML model, and FFmpeg unless
the input is verified as compatible PCM WAV.
The speech-to-text composition API is transcriber-neutral. Construct a
YTDLPAudioSource with NewYTDLPAudioSource, then configure it on
SpeechToTextProvider. A zero limit disables that limit; negative limits are
invalid. With a positive duration limit, unknown-duration and live inputs are
rejected before download. MaxBytes asks yt-dlp to reject known oversized
downloads and strictly rejects an oversized completed artifact, but upstream
manifest/fragment downloads are not hard-bounded while in flight. Audio.Format
is a lower-case container/file-extension hint, not a codec guarantee. Direct
YTDLPAudioSource.Acquire callers must close Audio.Data promptly to remove
the temporary artifact. SpeechToTextProvider closes it automatically. Cost
and concurrency controls remain adapter-specific work for providers that incur
cost or need an execution cap; they are not represented by misleading generic
fields.
The current public composition is explicit: create the ordinary caption client,
use it as FallbackProvider.Primary, use a configured SpeechToTextProvider
as Fallback, and install that chain in an outer client. The fallback runs only
for ErrTranscriptUnavailable. A compile-tested offline example is included;
the CLI enables the local chain only when --whisper-model explicitly selects
it.
captions, err := transcript.New(transcript.WithYTDLPPath("yt-dlp"))
if err != nil {
return err
}
speech := transcript.SpeechToTextProvider{
AudioSource: transcript.NewYTDLPAudioSource("yt-dlp"),
Transcriber: myTranscriber,
MaxDuration: 2 * time.Hour,
MaxBytes: 200 << 20,
}
client, err := transcript.New(transcript.WithProvider(transcript.FallbackProvider{
Primary: captions,
Fallback: speech,
}))For each requested language in caller order, selection tries an exact tag, then its base tag, then another regional variant. Manual captions beat automatic captions only within the same preference and match rank. Automatic captions require explicit library permission and are enabled by default in the CLI. Speech-to-text runs only through an explicitly configured fallback provider.
The result reports what was actually selected.
- A transcript cannot be guaranteed for every video. Videos may be private, removed, restricted, inaccessible, silent, or unsupported by upstream tools.
- The project will not bypass DRM, authentication, geographic restrictions, or access controls.
- YouTube and
yt-dlpcan change independently; integration behavior therefore requires a documented compatibility and update policy. - Callers are responsible for complying with applicable terms, copyright, privacy, and data-handling requirements.
- The official YouTube captions API is not a general solution for arbitrary public videos because downloading captions requires appropriate authorization over the video.
Development requires Go 1.25 or 1.26. The normalized core has no external runtime dependency. Caption and audio acquisition require a caller-installed yt-dlp; see runtime dependencies.
The repository Makefile provides the common development commands:
make build
make test
make integration
make whisper-integration
make vet
make staticcheck
make vuln
make hardening
make run ARGS="--version"
make checkRun make help for the complete target list, including formatting, race testing,
module tidying, and cleanup. make check performs the deterministic core checks
expected before submitting a change. make hardening additionally runs the race
detector, Staticcheck, and the network-backed Go vulnerability scan; install the
pinned tool versions documented in CONTRIBUTING.md first.
Network-dependent yt-dlp tests are opt-in so the normal unit-test suite remains deterministic. The ordinary suite includes process-level CLI smoke tests backed by a temporary fake provider executable. Run the live tests explicitly with make integration; this requires yt-dlp on PATH and public network access. make test continues to skip the live suite.
The local whisper.cpp runtime test is separately opt-in. Copy .env.example to
.env, set WHISPER_MODEL and WHISPER_SAMPLE to the local model and compatible
WAV sample paths, then run make whisper-integration. WHISPER_CLI defaults to
whisper-cli and can be overridden when the executable is elsewhere. Windows
MinGW/MSYS2 builds can set WHISPER_RUNTIME_DIR to prepend the matching runtime
DLL directory for this target. The local .env is ignored by Git, and
command-line assignments remain available for one-off overrides. The
corresponding environment variables and sample requirements are documented in
docs/DEPENDENCIES.md. The recorded prototype evidence is in
docs/evaluations/whispercpp-2026-08-13.md.
- Goal describes the north star, scope, and completion criteria.
- Roadmap defines implementation milestones and exit criteria.
- TODO is the actionable backlog.
- Ideas holds possible future work that is not yet committed.
- Live test-video matrix records proposed integration fixtures, provenance, validation, and replacement policy.
- Design conversation records the initial exploration of extraction approaches.
- Responsible-use guidance covers privacy, rights, restricted content, rate limits, retention, and caption accuracy.
- Release checklist and changelog describe the pre-release verification and publication process.
The public API is intentionally not fixed yet. Early contributions should align with the goal and current roadmap, keep provider details isolated, include offline tests for deterministic logic, and update documentation when design decisions change.
VoxScripta is available under the MIT License.
