Drop in a video file, get a text transcript. Uses OpenAI's Whisper model (or Faster-Whisper for 4x speed) to convert speech to text. Works with multiple languages and can tell apart different speakers in a conversation.
Outputs both .txt and .docx files with timestamps and speaker labels.
This project uses Template Method and Strategy design patterns for a clean, extensible architecture:
- Template Method: Common transcription pipeline (load → preprocess → transcribe → postprocess)
- Strategy: Model-specific inference logic (Whisper, Faster-Whisper, Google Cloud)
See REFACTORING_GUIDE.md and DESIGN_PATTERNS_EXPLAINED.md for details.
pip3 install -r requirements.txtThen use the transcription system programmatically:
from transcriber import Transcriber
from transcription_strategies import WhisperStrategy
# Create strategy
strategy = WhisperStrategy(model_size="base", language="tr")
# Create transcriber
transcriber = Transcriber(strategy=strategy)
# Process video
transcriber.process("video.mp4", output_dir="./output")- Transcribe video/audio files to text
- Multiple transcription backends:
- Standard Whisper (OpenAI)
- Faster-Whisper (CTranslate2) - up to 4x faster with lower memory
- Google Cloud Speech-to-Text API
- Multiple language support
- Speaker separation (local, Hugging Face, or Google Cloud)
- Export to TXT and DOCX formats
- Clean, extensible architecture using design patterns
The system is designed to be used programmatically. See example_usage.py for complete examples.
from transcriber import Transcriber
from transcription_strategies import WhisperStrategy
strategy = WhisperStrategy(
model_size="base", # "base", "small", "medium", "large"
fast_mode=False, # Set to True for faster transcription
language="tr", # Language code or None for auto-detect
device="auto" # "auto", "cpu", "cuda", "mps"
)
transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4", output_dir="./output")Faster-Whisper is a reimplementation of Whisper using CTranslate2, offering:
- ⚡ Up to 4x faster transcription
- 💾 Lower memory usage
- 🎯 Same accuracy as standard Whisper
from transcriber import Transcriber
from transcription_strategies import FasterWhisperStrategy
strategy = FasterWhisperStrategy(
model_size="large-v3", # "tiny", "base", "small", "medium", "large-v2", "large-v3"
fast_mode=False,
language="tr",
compute_type="auto", # "auto", "int8", "float16", "float32"
device="auto" # "auto", "cpu", "cuda"
)
transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4", output_dir="./output")from transcriber import Transcriber
from transcription_strategies import GoogleCloudStrategy
strategy = GoogleCloudStrategy(
credentials_path="path/to/credentials.json",
language_code="tr-TR", # Language code with region
speaker_count=2 # Number of speakers for diarization
)
transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4", output_dir="./output")# Process long videos in chunks (10 minutes per chunk)
transcriber = Transcriber(
strategy=strategy,
chunk_duration=600 # 600 seconds = 10 minutes
)
transcriber.process("long_video.mp4", output_dir="./output")See example_usage.py for more examples.
Both transcription and speaker diarization support GPU acceleration. Set the device parameter when creating strategies:
# NVIDIA GPU
strategy = WhisperStrategy(model_size="base", device="cuda")
# Apple Silicon (M1/M2/M3)
strategy = WhisperStrategy(model_size="base", device="mps")
# CPU (always available)
strategy = WhisperStrategy(model_size="base", device="cpu")
# Auto-detect (default)
strategy = WhisperStrategy(model_size="base", device="auto")Note: GPU acceleration significantly speeds up both Whisper transcription AND pyannote speaker diarization.
The system is designed to be easily extensible. To add a new transcription model:
- Create a new strategy class implementing
TranscriptionStrategy:
from transcription_base import TranscriptionStrategy
class MyModelStrategy(TranscriptionStrategy):
def transcribe(self, audio_path, ...):
# Your transcription logic
return {'text': ..., 'segments': [...]}
def get_speaker_diarization(self, audio_path, transcription):
# Your diarization logic
return [(start, end, speaker), ...]- Use it:
from transcriber import Transcriber
strategy = MyModelStrategy()
transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4")See REFACTORING_GUIDE.md for detailed instructions.
transcription-tr/
├── transcription_base.py # Base classes (Template Method)
├── transcription_strategies.py # Model strategies (Strategy pattern)
├── transcriber.py # Unified interface
├── example_usage.py # Usage examples
├── REFACTORING_GUIDE.md # Architecture guide
└── DESIGN_PATTERNS_EXPLAINED.md # Design patterns explanation
See requirements.txt for full dependencies. Key packages:
openai-whisperorfaster-whisper(for transcription)moviepy(for video/audio processing)python-docx(for DOCX output)pyannote.audio(for speaker diarization, optional)resemblyzer(for local speaker diarization, optional)google-cloud-speech(for Google Cloud, optional)
[Your license here]