Skip to content

About

Download videos from 1,700+ sites, record live streams, and turn any video into a transcript for AI chats. A Windows app built on yt-dlp, ffmpeg and Whisper.

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

Media Toolkit

Download videos, record live streams, and turn any video into text you can paste into an AI chat.
Everything runs on your own computer: Windows, macOS or Linux.

Latest release Windows, macOS and Linux MIT license Tests

The Download tab with a video preview and format choices

What it does

Download from 1,700+ sites. YouTube, Instagram, TikTok, X, Reddit, Vimeo, Facebook, Twitch and many more. Pick a quality, save only the audio, cut out a part, grab a whole playlist or just the newest videos of a channel, add subtitles, skip sponsor segments, or convert the result with your graphics card.

Transcripts in seconds. Paste a link and get clean text, ready for any AI chat. If the video has captions you have them in about a second; if not, the speech is transcribed on your computer with Whisper, using an NVIDIA graphics card on Windows when you have one. Drop a local video or audio file to transcribe it too. Save as text, notes with timestamps, subtitles or JSON.

Record live streams. Video and sound together, stop whenever you like and keep everything recorded so far. It can wait for a scheduled stream to start, split long recordings into parts, and survives short network drops.

A transcript with Copy for AI chat The queue with finished and running jobs
Transcripts ready to paste into an AI chat Every job with progress, errors and a way forward
Recording a live stream Settings in the light theme
Live recording with stop and save Light and dark themes follow your system

All features

Downloading

  • Works with 1,700+ sites through yt-dlp: YouTube, Instagram, TikTok, X, Reddit, Vimeo, Facebook, Twitch and more
  • Paste one link or many at once; a live preview shows title, length, best quality and an estimated file size
  • Video or audio only: quality from 480p up to 4K, "Plays on any device" (H.264) mode, or the smallest file
  • Audio as MP3, M4A, Opus, FLAC or WAV; containers MP4, MKV or WebM
  • Download only part of a video (from/to times), or pick an exact format from the full format list
  • Playlists and channels right in the preview: all, first or latest N, or specific items, with filters for length, upload date, title, views and file size, reverse or shuffled order, and "join everything into one file"
  • Quick update for channels: skip videos you already have and stop at the first known one
  • Subtitles (official or automatic, any language), embedded or as files
  • SponsorBlock: cut out or just mark sponsor segments
  • Chapters: keep them, split into one file per chapter, or remove chapters by name
  • Re-encode with your graphics card (NVIDIA, Intel or AMD, only offered when it really works) or processor, and even out the volume
  • Extras: cover art and title/artist tags, description, top comments, thumbnail, a shortcut to the page, technical details
  • File name presets (title, channel, date) or your own pattern
  • Sign-in for private or age-restricted videos: use your browser's sign-in, sign in through a separate window, or paste cookies

Transcripts

  • Uses the video's own captions first, so most transcripts are ready in about a second
  • Otherwise transcribes on your computer with Whisper, on an NVIDIA graphics card (Windows) or the processor, with automatic fallback if one does not work
  • Drop any local video or audio file onto the window to transcribe it
  • Language choice with native names (Persian, Arabic, Chinese, Hindi and 20+ more), or translate to English
  • A clean reader with paragraphs or timestamps, word and token count
  • Copy for AI chat adds the title, channel and link; long transcripts can be copied in parts that fit one message each
  • Save as text, notes with timestamps, SRT, VTT or JSON
  • Recent transcripts stay one click away
  • Speech models from tiny (75 MB) to large (3 GB); the app recommends one for your hardware, checks downloads and repairs damaged ones

Live recording

  • Record any live stream yt-dlp can open: YouTube, Twitch, Kick, TikTok, radio and more
  • Video and sound together, copied as streamed: no quality loss and almost no CPU use
  • Stop whenever you like and keep everything recorded so far; the file is always playable
  • Survives network drops and reconnects on its own
  • Waits for scheduled streams to start, up to a time you choose
  • Stop automatically after N minutes, or start a new file every N minutes
  • Sound-only recording for radio and talk streams

The app

  • Each tab shows its own progress and results; the Queue keeps every job, even across restarts
  • Errors in plain language with a fix button (sign in, retry, repair ffmpeg) and technical details one click away
  • Try again, remove, clear with undo, play, open and show in folder on every job
  • Remembers your choices between launches; settings save themselves
  • Light and dark themes that follow your system, keyboard shortcuts, full right-to-left support
  • Windows installer or portable zip, macOS disk image, Linux AppImage or tarball; updates to site support (yt-dlp) from inside the app
  • Keeps working in the background if you close the window; opening it again returns to the running app

Why Media Toolkit

Media Toolkit Online downloader sites yt-dlp on the command line Typical paid downloaders
Sites supported 1,700+ (yt-dlp) a handful 1,700+ dozens to hundreds
Transcripts for AI chats built in, captions or local Whisper no no rarely
Live stream recording that can be stopped and kept yes no partly (no clean stop and keep) some
Runs on your computer, nothing uploaded yes no, your links go to their servers yes yes
Ads, accounts, tracking none usually ads and trackers none accounts, upsells, limits
Easy to use yes yes needs typing commands yes
Price and license free, open source (MIT) free with ads free, open source paid or limited free tier

In short: it gives you the full power of yt-dlp and ffmpeg without the command line, adds local speech-to-text that turns any video into text for an AI chat, and keeps everything on your own computer.

Install

Get the file for your system from the latest release. None of the downloads are code-signed yet, so the first launch needs one extra click; SHA256SUMS.txt in the release lets you check the file first.

Windows 10 and 11

  1. Run MediaToolkit-Setup-<version>.exe. It installs for your user only, so there is no administrator prompt, and adds Start menu and desktop shortcuts.
  2. If SmartScreen says it "protected your PC", choose More info › Run anyway.

Prefer not to install? The portable zip runs from any folder (a USB stick works) and keeps its settings and models in a data folder next to it.

macOS (Apple silicon, macOS 12 or newer)

  1. Open MediaToolkit-<version>-macos-arm64.dmg and drag Media Toolkit into Applications.
  2. The first time, right-click the app › Open › Open (macOS blocks apps from unidentified developers on a normal double-click). If macOS still refuses, run xattr -dr com.apple.quarantine "/Applications/Media Toolkit.app" in Terminal.

For the app window, Media Toolkit uses Chrome, Edge or Brave if you have one, and otherwise opens in your default browser.

Linux (x86-64)

  • AppImage: chmod +x MediaToolkit-<version>-x86_64.AppImage, then run it. Some distributions need libfuse2 for AppImages (sudo apt install libfuse2 on Ubuntu 22.04 and newer).
  • Tarball: unpack MediaToolkit-<version>-linux-x86_64.tar.gz anywhere and run MediaToolkit inside it.

The folder picker uses zenity or kdialog, which most desktops already have.

What gets downloaded later

Everything the app needs is included: Python, yt-dlp and ffmpeg. Two things are downloaded later, only if you want them:

Download Size When
Speech model 75 MB to 3 GB The first time a video without captions is transcribed
GPU support (NVIDIA cuBLAS) 528 MB Windows only: offered in Settings on PCs with an NVIDIA graphics card

On macOS and Linux, transcription runs on the processor. Videos that have captions are instant everywhere.

Upgrading keeps your settings, models and files. The Windows uninstaller asks before removing the app's data folder, and never deletes your downloads.

First steps

  1. Get a transcript: open the Transcript tab, paste a video link and press Enter. Press Copy for AI chat and paste it into any chatbot.
  2. Download a video: paste a link on the Download tab, check the preview, choose Video or Audio only, press Enter.
  3. Private or age-restricted videos: go to Settings › Sign-in for private videos and let the app use the sign-in from your browser.

When a site stops working, it has usually changed something that yt-dlp has already fixed: Settings › About › Site support checks for an update.

Privacy

There is no account, no tracking and no telemetry. The app talks only to:

  • the sites you paste links from (through yt-dlp), and SponsorBlock when you turn on sponsor skipping;
  • Hugging Face, to download a speech model the first time one is needed;
  • PyPI, to download GPU support or a yt-dlp update when you ask for it;
  • GitHub, when you press Check for updates.

Your files, transcripts and sign-in cookies stay on your computer. The app's window talks to a small server on 127.0.0.1 that refuses requests from websites, other computers and anything without the per-launch key.

Command line

MediaToolkit.exe                       open the app (or return to the running one)
MediaToolkit.exe --port 9000           use a fixed port
MediaToolkit.exe --browser             open in your normal browser instead of the app window
MediaToolkit.exe --server              run without a window; keeps going until stopped
MediaToolkit.exe --diagnose MODEL      write a report on a speech model to diagnose.txt
MediaToolkit.exe --diagnose MODEL --repair   ...and download that model again

--host can make the server reachable from other devices on your network; it then prints an address with an access key and requires it on every request. Only use it on a network you trust.

On macOS the program is /Applications/Media Toolkit.app/Contents/MacOS/MediaToolkit, on Linux MediaToolkit in the unpacked folder (or the AppImage itself).

The app keeps its settings, models and log (Settings › About › Open log file) in:

System Folder
Windows %LOCALAPPDATA%\Media Toolkit
macOS ~/Library/Application Support/Media Toolkit
Linux ~/.local/share/media-toolkit

Building from source

Needs Python 3.11 or newer (3.13 recommended).

Windows:

git clone https://github.com/AnotherAH/media-toolkit.git
cd media-toolkit
setup.bat
Start.bat

setup.bat creates a private environment in .venv, installs the dependencies and downloads ffmpeg into bin\.

macOS and Linux:

git clone https://github.com/AnotherAH/media-toolkit.git
cd media-toolkit
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python tools/fetch_ffmpeg.py
.venv/bin/python run.py

Building the release files: install the build tools with pip install -r requirements-build.txt, then run tools\build.py on Windows (installer and portable zip; needs Inno Setup 6, winget install JRSoftware.InnoSetup) or tools/build_unix.py on macOS (a .dmg) and Linux (a tarball, plus an AppImage when appimagetool is installed).

The results land in dist/. Pushing a v<version> tag builds all of them on GitHub Actions into a draft release. See CONTRIBUTING.md for tests and guidelines.

Responsible use

Media Toolkit is a tool for saving and transcribing media you have the right to use: your own uploads, public-domain and Creative Commons works, content whose license allows it, or personal copies where your local law permits them. Respect copyright and the terms of the sites you use. Do not use it to redistribute other people's work.

Media Toolkit is not affiliated with or endorsed by YouTube, Google, Instagram, Meta, TikTok, X, Twitch or any other site or service it mentions. Their names are used only to describe what the app works with.

Built on

Media Toolkit is a thin layer over excellent open-source projects:

Project What it does here
yt-dlp Every site, format selection, playlists, subtitles, SponsorBlock, live stream links
FFmpeg (yt-dlp builds) Joining, converting, recording and audio decoding
faster-whisper and CTranslate2 Speech recognition
Silero VAD Skipping silence
SponsorBlock Sponsor segment data (CC BY-NC-SA 4.0)
FastAPI and Uvicorn The local server behind the window

The installed app includes the license of every component in THIRD-PARTY-NOTICES.txt (also under Settings › About › Third-party licenses). Screenshots show Blender Foundation open movies (CC BY 3.0, © Blender Foundation | blender.org).

License

MIT. The installer also contains third-party software under its own licenses, including FFmpeg under the GPL; see THIRD-PARTY-NOTICES.txt.

About

Download videos from 1,700+ sites, record live streams, and turn any video into a transcript for AI chats. A Windows app built on yt-dlp, ffmpeg and Whisper.

Topics

Resources

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages