Download videos, record live streams, and turn any video into text you can paste into an AI chat.
Everything runs on your own computer: Windows, macOS or Linux.
Download from 1,700+ sites. YouTube, Instagram, TikTok, X, Reddit, Vimeo, Facebook, Twitch and many more. Pick a quality, save only the audio, cut out a part, grab a whole playlist or just the newest videos of a channel, add subtitles, skip sponsor segments, or convert the result with your graphics card.
Transcripts in seconds. Paste a link and get clean text, ready for any AI chat. If the video has captions you have them in about a second; if not, the speech is transcribed on your computer with Whisper, using an NVIDIA graphics card on Windows when you have one. Drop a local video or audio file to transcribe it too. Save as text, notes with timestamps, subtitles or JSON.
Record live streams. Video and sound together, stop whenever you like and keep everything recorded so far. It can wait for a scheduled stream to start, split long recordings into parts, and survives short network drops.
![]() |
![]() |
| Transcripts ready to paste into an AI chat | Every job with progress, errors and a way forward |
![]() |
![]() |
| Live recording with stop and save | Light and dark themes follow your system |
- Works with 1,700+ sites through yt-dlp: YouTube, Instagram, TikTok, X, Reddit, Vimeo, Facebook, Twitch and more
- Paste one link or many at once; a live preview shows title, length, best quality and an estimated file size
- Video or audio only: quality from 480p up to 4K, "Plays on any device" (H.264) mode, or the smallest file
- Audio as MP3, M4A, Opus, FLAC or WAV; containers MP4, MKV or WebM
- Download only part of a video (from/to times), or pick an exact format from the full format list
- Playlists and channels right in the preview: all, first or latest N, or specific items, with filters for length, upload date, title, views and file size, reverse or shuffled order, and "join everything into one file"
- Quick update for channels: skip videos you already have and stop at the first known one
- Subtitles (official or automatic, any language), embedded or as files
- SponsorBlock: cut out or just mark sponsor segments
- Chapters: keep them, split into one file per chapter, or remove chapters by name
- Re-encode with your graphics card (NVIDIA, Intel or AMD, only offered when it really works) or processor, and even out the volume
- Extras: cover art and title/artist tags, description, top comments, thumbnail, a shortcut to the page, technical details
- File name presets (title, channel, date) or your own pattern
- Sign-in for private or age-restricted videos: use your browser's sign-in, sign in through a separate window, or paste cookies
- Uses the video's own captions first, so most transcripts are ready in about a second
- Otherwise transcribes on your computer with Whisper, on an NVIDIA graphics card (Windows) or the processor, with automatic fallback if one does not work
- Drop any local video or audio file onto the window to transcribe it
- Language choice with native names (Persian, Arabic, Chinese, Hindi and 20+ more), or translate to English
- A clean reader with paragraphs or timestamps, word and token count
- Copy for AI chat adds the title, channel and link; long transcripts can be copied in parts that fit one message each
- Save as text, notes with timestamps, SRT, VTT or JSON
- Recent transcripts stay one click away
- Speech models from tiny (75 MB) to large (3 GB); the app recommends one for your hardware, checks downloads and repairs damaged ones
- Record any live stream yt-dlp can open: YouTube, Twitch, Kick, TikTok, radio and more
- Video and sound together, copied as streamed: no quality loss and almost no CPU use
- Stop whenever you like and keep everything recorded so far; the file is always playable
- Survives network drops and reconnects on its own
- Waits for scheduled streams to start, up to a time you choose
- Stop automatically after N minutes, or start a new file every N minutes
- Sound-only recording for radio and talk streams
- Each tab shows its own progress and results; the Queue keeps every job, even across restarts
- Errors in plain language with a fix button (sign in, retry, repair ffmpeg) and technical details one click away
- Try again, remove, clear with undo, play, open and show in folder on every job
- Remembers your choices between launches; settings save themselves
- Light and dark themes that follow your system, keyboard shortcuts, full right-to-left support
- Windows installer or portable zip, macOS disk image, Linux AppImage or tarball; updates to site support (yt-dlp) from inside the app
- Keeps working in the background if you close the window; opening it again returns to the running app
| Media Toolkit | Online downloader sites | yt-dlp on the command line | Typical paid downloaders | |
|---|---|---|---|---|
| Sites supported | 1,700+ (yt-dlp) | a handful | 1,700+ | dozens to hundreds |
| Transcripts for AI chats | built in, captions or local Whisper | no | no | rarely |
| Live stream recording that can be stopped and kept | yes | no | partly (no clean stop and keep) | some |
| Runs on your computer, nothing uploaded | yes | no, your links go to their servers | yes | yes |
| Ads, accounts, tracking | none | usually ads and trackers | none | accounts, upsells, limits |
| Easy to use | yes | yes | needs typing commands | yes |
| Price and license | free, open source (MIT) | free with ads | free, open source | paid or limited free tier |
In short: it gives you the full power of yt-dlp and ffmpeg without the command line, adds local speech-to-text that turns any video into text for an AI chat, and keeps everything on your own computer.
Get the file for your system from the
latest release.
None of the downloads are code-signed yet, so the first launch needs one extra
click; SHA256SUMS.txt in the release lets you check the file first.
- Run
MediaToolkit-Setup-<version>.exe. It installs for your user only, so there is no administrator prompt, and adds Start menu and desktop shortcuts. - If SmartScreen says it "protected your PC", choose More info › Run anyway.
Prefer not to install? The portable zip runs from any folder (a USB stick
works) and keeps its settings and models in a data folder next to it.
- Open
MediaToolkit-<version>-macos-arm64.dmgand drag Media Toolkit into Applications. - The first time, right-click the app › Open › Open (macOS blocks apps
from unidentified developers on a normal double-click). If macOS still
refuses, run
xattr -dr com.apple.quarantine "/Applications/Media Toolkit.app"in Terminal.
For the app window, Media Toolkit uses Chrome, Edge or Brave if you have one, and otherwise opens in your default browser.
- AppImage:
chmod +x MediaToolkit-<version>-x86_64.AppImage, then run it. Some distributions needlibfuse2for AppImages (sudo apt install libfuse2on Ubuntu 22.04 and newer). - Tarball: unpack
MediaToolkit-<version>-linux-x86_64.tar.gzanywhere and runMediaToolkitinside it.
The folder picker uses zenity or kdialog, which most desktops already have.
Everything the app needs is included: Python, yt-dlp and ffmpeg. Two things are downloaded later, only if you want them:
| Download | Size | When |
|---|---|---|
| Speech model | 75 MB to 3 GB | The first time a video without captions is transcribed |
| GPU support (NVIDIA cuBLAS) | 528 MB | Windows only: offered in Settings on PCs with an NVIDIA graphics card |
On macOS and Linux, transcription runs on the processor. Videos that have captions are instant everywhere.
Upgrading keeps your settings, models and files. The Windows uninstaller asks before removing the app's data folder, and never deletes your downloads.
- Get a transcript: open the Transcript tab, paste a video link and press Enter. Press Copy for AI chat and paste it into any chatbot.
- Download a video: paste a link on the Download tab, check the preview, choose Video or Audio only, press Enter.
- Private or age-restricted videos: go to Settings › Sign-in for private videos and let the app use the sign-in from your browser.
When a site stops working, it has usually changed something that yt-dlp has already fixed: Settings › About › Site support checks for an update.
There is no account, no tracking and no telemetry. The app talks only to:
- the sites you paste links from (through yt-dlp), and SponsorBlock when you turn on sponsor skipping;
- Hugging Face, to download a speech model the first time one is needed;
- PyPI, to download GPU support or a yt-dlp update when you ask for it;
- GitHub, when you press Check for updates.
Your files, transcripts and sign-in cookies stay on your computer. The app's
window talks to a small server on 127.0.0.1 that refuses requests from
websites, other computers and anything without the per-launch key.
MediaToolkit.exe open the app (or return to the running one)
MediaToolkit.exe --port 9000 use a fixed port
MediaToolkit.exe --browser open in your normal browser instead of the app window
MediaToolkit.exe --server run without a window; keeps going until stopped
MediaToolkit.exe --diagnose MODEL write a report on a speech model to diagnose.txt
MediaToolkit.exe --diagnose MODEL --repair ...and download that model again
--host can make the server reachable from other devices on your network; it
then prints an address with an access key and requires it on every request.
Only use it on a network you trust.
On macOS the program is /Applications/Media Toolkit.app/Contents/MacOS/MediaToolkit,
on Linux MediaToolkit in the unpacked folder (or the AppImage itself).
The app keeps its settings, models and log (Settings › About › Open log file) in:
| System | Folder |
|---|---|
| Windows | %LOCALAPPDATA%\Media Toolkit |
| macOS | ~/Library/Application Support/Media Toolkit |
| Linux | ~/.local/share/media-toolkit |
Needs Python 3.11 or newer (3.13 recommended).
Windows:
git clone https://github.com/AnotherAH/media-toolkit.git
cd media-toolkit
setup.bat
Start.batsetup.bat creates a private environment in .venv, installs the
dependencies and downloads ffmpeg into bin\.
macOS and Linux:
git clone https://github.com/AnotherAH/media-toolkit.git
cd media-toolkit
python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python tools/fetch_ffmpeg.py
.venv/bin/python run.pyBuilding the release files: install the build tools with
pip install -r requirements-build.txt, then run tools\build.py on Windows
(installer and portable zip; needs Inno Setup 6,
winget install JRSoftware.InnoSetup) or tools/build_unix.py on macOS (a .dmg)
and Linux (a tarball, plus an AppImage when appimagetool is installed).
The results land in dist/. Pushing a v<version> tag builds all of them
on GitHub Actions into a draft release. See CONTRIBUTING.md
for tests and guidelines.
Media Toolkit is a tool for saving and transcribing media you have the right to use: your own uploads, public-domain and Creative Commons works, content whose license allows it, or personal copies where your local law permits them. Respect copyright and the terms of the sites you use. Do not use it to redistribute other people's work.
Media Toolkit is not affiliated with or endorsed by YouTube, Google, Instagram, Meta, TikTok, X, Twitch or any other site or service it mentions. Their names are used only to describe what the app works with.
Media Toolkit is a thin layer over excellent open-source projects:
| Project | What it does here |
|---|---|
| yt-dlp | Every site, format selection, playlists, subtitles, SponsorBlock, live stream links |
| FFmpeg (yt-dlp builds) | Joining, converting, recording and audio decoding |
| faster-whisper and CTranslate2 | Speech recognition |
| Silero VAD | Skipping silence |
| SponsorBlock | Sponsor segment data (CC BY-NC-SA 4.0) |
| FastAPI and Uvicorn | The local server behind the window |
The installed app includes the license of every component in
THIRD-PARTY-NOTICES.txt (also under Settings › About › Third-party
licenses). Screenshots show Blender Foundation open movies
(CC BY 3.0, © Blender
Foundation | blender.org).
MIT. The installer also contains third-party software under its
own licenses, including FFmpeg under the GPL; see THIRD-PARTY-NOTICES.txt.




