Skip to content

Repository files navigation

Video Transcription

Drop in a video file, get a text transcript. Uses OpenAI's Whisper model (or Faster-Whisper for 4x speed) to convert speech to text. Works with multiple languages and can tell apart different speakers in a conversation.

Outputs both .txt and .docx files with timestamps and speaker labels.

Architecture

This project uses Template Method and Strategy design patterns for a clean, extensible architecture:

  • Template Method: Common transcription pipeline (load → preprocess → transcribe → postprocess)
  • Strategy: Model-specific inference logic (Whisper, Faster-Whisper, Google Cloud)

See REFACTORING_GUIDE.md and DESIGN_PATTERNS_EXPLAINED.md for details.

Quick Start

pip3 install -r requirements.txt

Then use the transcription system programmatically:

from transcriber import Transcriber
from transcription_strategies import WhisperStrategy

# Create strategy
strategy = WhisperStrategy(model_size="base", language="tr")

# Create transcriber
transcriber = Transcriber(strategy=strategy)

# Process video
transcriber.process("video.mp4", output_dir="./output")

Features

  • Transcribe video/audio files to text
  • Multiple transcription backends:
    • Standard Whisper (OpenAI)
    • Faster-Whisper (CTranslate2) - up to 4x faster with lower memory
    • Google Cloud Speech-to-Text API
  • Multiple language support
  • Speaker separation (local, Hugging Face, or Google Cloud)
  • Export to TXT and DOCX formats
  • Clean, extensible architecture using design patterns

Usage

The system is designed to be used programmatically. See example_usage.py for complete examples.

Using Whisper

from transcriber import Transcriber
from transcription_strategies import WhisperStrategy

strategy = WhisperStrategy(
    model_size="base",      # "base", "small", "medium", "large"
    fast_mode=False,         # Set to True for faster transcription
    language="tr",           # Language code or None for auto-detect
    device="auto"            # "auto", "cpu", "cuda", "mps"
)

transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4", output_dir="./output")

Using Faster-Whisper

Faster-Whisper is a reimplementation of Whisper using CTranslate2, offering:

  • ⚡ Up to 4x faster transcription
  • 💾 Lower memory usage
  • 🎯 Same accuracy as standard Whisper
from transcriber import Transcriber
from transcription_strategies import FasterWhisperStrategy

strategy = FasterWhisperStrategy(
    model_size="large-v3",   # "tiny", "base", "small", "medium", "large-v2", "large-v3"
    fast_mode=False,
    language="tr",
    compute_type="auto",     # "auto", "int8", "float16", "float32"
    device="auto"            # "auto", "cpu", "cuda"
)

transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4", output_dir="./output")

Using Google Cloud Speech-to-Text

from transcriber import Transcriber
from transcription_strategies import GoogleCloudStrategy

strategy = GoogleCloudStrategy(
    credentials_path="path/to/credentials.json",
    language_code="tr-TR",  # Language code with region
    speaker_count=2         # Number of speakers for diarization
)

transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4", output_dir="./output")

With Chunking for Long Videos

# Process long videos in chunks (10 minutes per chunk)
transcriber = Transcriber(
    strategy=strategy,
    chunk_duration=600  # 600 seconds = 10 minutes
)
transcriber.process("long_video.mp4", output_dir="./output")

See example_usage.py for more examples.

GPU Acceleration

Both transcription and speaker diarization support GPU acceleration. Set the device parameter when creating strategies:

# NVIDIA GPU
strategy = WhisperStrategy(model_size="base", device="cuda")

# Apple Silicon (M1/M2/M3)
strategy = WhisperStrategy(model_size="base", device="mps")

# CPU (always available)
strategy = WhisperStrategy(model_size="base", device="cpu")

# Auto-detect (default)
strategy = WhisperStrategy(model_size="base", device="auto")

Note: GPU acceleration significantly speeds up both Whisper transcription AND pyannote speaker diarization.

Adding New Models

The system is designed to be easily extensible. To add a new transcription model:

  1. Create a new strategy class implementing TranscriptionStrategy:
from transcription_base import TranscriptionStrategy

class MyModelStrategy(TranscriptionStrategy):
    def transcribe(self, audio_path, ...):
        # Your transcription logic
        return {'text': ..., 'segments': [...]}
    
    def get_speaker_diarization(self, audio_path, transcription):
        # Your diarization logic
        return [(start, end, speaker), ...]
  1. Use it:
from transcriber import Transcriber

strategy = MyModelStrategy()
transcriber = Transcriber(strategy=strategy)
transcriber.process("video.mp4")

See REFACTORING_GUIDE.md for detailed instructions.

Project Structure

transcription-tr/
├── transcription_base.py          # Base classes (Template Method)
├── transcription_strategies.py    # Model strategies (Strategy pattern)
├── transcriber.py                 # Unified interface
├── example_usage.py               # Usage examples
├── REFACTORING_GUIDE.md           # Architecture guide
└── DESIGN_PATTERNS_EXPLAINED.md   # Design patterns explanation

Requirements

See requirements.txt for full dependencies. Key packages:

  • openai-whisper or faster-whisper (for transcription)
  • moviepy (for video/audio processing)
  • python-docx (for DOCX output)
  • pyannote.audio (for speaker diarization, optional)
  • resemblyzer (for local speaker diarization, optional)
  • google-cloud-speech (for Google Cloud, optional)

License

[Your license here]

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages