Skip to content

Repository files navigation

Voice Authentication System

πŸŽ™οΈ Voice Authentication System

Your Voice. Your Identity. Zero Compromise.

Register once. Embed forever. Authenticate any voice in 2 lines of Python.

A local-first voice-authentication platform that turns a handful of voice samples into a personalised, installable Python SDK (.whl) β€” powered by a SpeechBrain ECAPA-TDNN speaker-embedding model. No audio ever leaves your infrastructure.


License: MIT DOI CI Tests Code style: Ruff PRs Welcome Made with love

Python FastAPI PyTorch SpeechBrain Docker JWT


🎬 Live walkthrough

Product walkthrough β€” landing, onboarding, live processing pipeline, and SDK download

πŸ“– Table of contents


🎯 Overview

Voice Authentication System verifies who is speaking by comparing a live voice sample against an enrolled voiceprint β€” entirely on your own machine or server.

  1. A user records or uploads 5–20 short voice clips.
  2. The backend cleans each clip, extracts a 192-dimensional embedding, and averages them into one master voiceprint.
  3. That voiceprint is baked into a generated Python package and compiled into a .whl.
  4. Anyone can then authenticate a voice in two lines of Python β€” offline.
Mode Endpoint Produces Use case
Individual POST /api/process authenticate(audio) β†’ bool 1:1 login / identity check
Company / Team POST /api/process_company identify(audio) β†’ "Name" 1:N speaker classification

πŸ“Έ Screenshots

Landing page
Landing β€” WebGL audio-wave hero
Authentication
Auth β€” split-panel login / sign-up
Account type
Account type β€” Individual vs Team / Company
Voice recording
Recording β€” record or upload voice samples
Company upload
Company upload β€” bulk team enrollment
Download SDK
Download β€” grab your generated .whl
Live voice verifier
Live verifier β€” test a voice sample in the browser
Ready to download
Ready state β€” "Your voice is ready"
Processing pipeline with live MFCC heatmap
Processing β€” live pipeline panel with the MFCC heatmap and 192-dim embedding rendered in the browser

✨ Key features

  • πŸ”’ Local-first & private β€” audio never leaves your infrastructure; embeddings and wheels stay on disk.
  • 🧠 State-of-the-art embeddings β€” SpeechBrain ECAPA-TDNN (VoxCeleb) 192-dim speaker vectors.
  • πŸ“¦ Generates a real SDK β€” outputs an installable audioauth .whl with the voiceprint baked in.
  • πŸ‘₯ Two modes β€” individual 1:1 authentication and company-wide 1:N speaker identification.
  • πŸ›‘οΈ Hardened security β€” bcrypt password hashing, JWT bearer auth, CORS allow-list, rate limiting.
  • 🎨 Premium frontend β€” WebGL shader hero, live MFCC/embedding visualizations, in-browser recorder.
  • 🐳 Production-ready β€” Dockerfile, docker-compose, GitHub Actions CI, pytest suite, ruff linting.
Full page-by-page feature tour (click to expand)
Page File Highlights
Landing index.html WebGL hero shader + bloom, How-It-Works flow, "2 lines of auth" code showcase, bento feature grid, pricing, social-proof marquee
Auth auth.html Split glass layout, Login ⇄ Sign-Up tabs, password show/hide, 4-segment strength meter
Account Type onboarding-type.html 4-step stepper, Individual vs Team cards with animated selection
Voice Recording onboarding-record.html Web Audio API + MediaRecorder, 240px mic circle with live level, 10s countdown, drag-and-drop upload
Company Upload onboarding-company.html Spreadsheet-style bulk table, CSV import modal, per-person status badges
Processing onboarding-processing.html Live pipeline panel + MFCC heatmap, waveform/energy, pitch contour, 192-dim embedding bars, UMAP scatter
Download download.html Ready-state hero, authenticated .whl download, live in-browser voice verification demo

πŸ—οΈ Architecture

flowchart LR
    subgraph Browser["🌐 Frontend (static)"]
        UI[HTML + Tailwind + Three.js]
        REC[Web Audio recorder]
    end

    subgraph API["βš™οΈ FastAPI backend"]
        AUTH[JWT auth + bcrypt]
        ROUTES[REST routes]
        PIPE[Voice pipeline]
        BUILD[.whl builder]
    end

    subgraph Store["πŸ’Ύ Local storage"]
        CSV[(login.csv<br/>hashed)]
        NPY[(embedding.npy)]
        WHL[(generated .whl)]
    end

    MODEL[["🧠 ECAPA-TDNN<br/>(SpeechBrain / VoxCeleb)"]]

    UI -->|register / login| AUTH
    REC -->|upload clips| ROUTES
    ROUTES --> PIPE
    PIPE --> MODEL
    PIPE --> NPY
    PIPE --> BUILD --> WHL
    AUTH --> CSV
    WHL -->|download| UI
Loading

πŸ”„ How it works

Enrollment β†’ SDK generation

sequenceDiagram
    actor U as User
    participant FE as Frontend
    participant API as FastAPI
    participant ML as ECAPA-TDNN
    participant FS as Storage
    U->>FE: Record / upload 5–20 clips
    FE->>API: POST /api/process (Bearer JWT)
    API->>API: preprocess + validate each clip
    API->>ML: encode clips β†’ 192-dim vectors
    ML-->>API: embeddings
    API->>API: average β†’ master voiceprint
    API->>FS: save embedding.npy
    API->>API: inject voiceprint β†’ build .whl
    API->>FS: save .whl
    API-->>FE: { build_id, accepted_files }
    FE->>API: GET /api/download (Bearer JWT)
    API-->>U: audioauth-1.0.0.whl
Loading

Verification

sequenceDiagram
    actor U as User
    participant FE as Frontend
    participant API as FastAPI
    participant ML as ECAPA-TDNN
    U->>FE: Provide a test sample
    FE->>API: POST /api/verify (Bearer JWT)
    API->>ML: encode sample β†’ 192-dim vector
    API->>API: cosine similarity vs stored voiceprint
    API-->>FE: { confidence, label, matched }
    FE-->>U: βœ… Strong / 🟑 Partial / ❌ No match
Loading

πŸ”¬ The pipeline

# Stage Module Input β†’ Output
1 Preprocess pipeline/preprocess.py audio file β†’ 16 kHz mono float32, trimmed/padded to 3 s, peak-normalised (rejects clips < 2 s)
2 Embed pipeline/embedding.py audio array β†’ 192-dim ECAPA-TDNN vector
3 Average pipeline/averaging.py N embeddings β†’ 1 master voiceprint (mean)
4 Inject pipeline/injector.py master vector β†’ baked into core.py template
5 Build pipeline/builder.py package folder β†’ installable .whl

Match thresholds (cosine similarity, used in the generated SDK):

Cosine score Verdict
β‰₯ 0.75 βœ… Strong match
0.60 – 0.75 🟑 Partial match
< 0.60 ❌ Rejected

🧰 Tech stack & skills

Backend & API

Python FastAPI Uvicorn Pydantic

Machine Learning & Audio

PyTorch SpeechBrain NumPy librosa Hugging Face

Frontend

HTML5 JavaScript Tailwind CSS Three.js ECharts

Security & Auth

JWT bcrypt

DevOps & Tooling

Docker GitHub Actions pytest Ruff Git

Layer Technology Purpose
API FastAPI, Uvicorn Async REST API + ASGI server
Auth PyJWT, passlib + bcrypt Bearer tokens, password hashing
ML / audio SpeechBrain (ECAPA-TDNN), PyTorch, torchaudio, librosa, soundfile Speaker embeddings + audio decoding
Config pydantic-settings Type-safe, env-driven config
Packaging PyPA build Compiles the per-user SDK .whl
Frontend HTML, Tailwind (CDN), Three.js, ECharts, Web Audio API Animated UI, live visualizations, recording
Dev / CI pytest, ruff, Docker, GitHub Actions Tests, linting, containers, automation

πŸ“Š Project stats

Metric Value
Total files 58
Lines of code (py + js + html + css) ~5,700
Python 1,724 lines Β· 24 files
JavaScript 999 lines Β· 5 files
HTML 2,653 lines Β· 8 files
CSS 353 lines
REST API endpoints 8
Pipeline stages 5
Frontend pages 7
Automated tests 17 βœ…
Embedding dimensionality 192
Model ECAPA-TDNN (VoxCeleb)

πŸ†• What's new in v2

This release hardens the original prototype into a production-ready project.

Area v1 (prototype) v2 (this release)
Passwords ❌ stored in plaintext βœ… bcrypt-hashed (work factor 12)
Authentication ❌ email in request body βœ… JWT bearer tokens
CORS ⚠️ wildcard * βœ… env-driven allow-list
Secrets ⚠️ hard-coded defaults βœ… environment variables
Uploads ❌ unbounded βœ… size + type limits
Brute-force ❌ none βœ… rate-limited auth
Tests ❌ none βœ… 17 automated tests
CI/CD ❌ none βœ… GitHub Actions
Containers ❌ none βœ… Dockerfile + compose
Docs ⚠️ basic βœ… this README + CONTRIBUTING + SECURITY

πŸš€ Quickstart

Option A β€” Docker (recommended)

git clone https://github.com/asadullahirshad3/Voice-Authentication-System.git
cd Voice-Authentication-System
cp .env.example .env          # then edit JWT_SECRET
docker compose up --build

Open http://localhost:8000.

Option B β€” Local Python

git clone https://github.com/asadullahirshad3/Voice-Authentication-System.git
cd Voice-Authentication-System

python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -r Backend/requirements.txt

cp .env.example .env          # set a strong JWT_SECRET
cd Backend
uvicorn app.main:app --reload

The ECAPA-TDNN model (~80 MB) downloads automatically from Hugging Face on the first /api/process call and is cached locally afterwards.


βš™οΈ Configuration keys

All settings come from environment variables (or a .env file). Full list in .env.example.

Key Default Description
JWT_SECRET (placeholder) Set a long random value in production. Signs access tokens.
JWT_ALGORITHM HS256 JWT signing algorithm
ACCESS_TOKEN_EXPIRE_MINUTES 1440 Token lifetime (minutes)
CORS_ORIGINS http://localhost:8000,... Comma-separated browser origin allow-list
MIN_FILES / MAX_FILES 5 / 20 Enrollment clip bounds
MAX_UPLOAD_MB 25 Per-file upload cap
EMAIL_ENABLED false Toggle the welcome email
SMTP_HOST smtp.gmail.com SMTP server host
SMTP_PORT 587 SMTP port
SMTP_USER / SMTP_PASSWORD (empty) SMTP credentials (if email enabled)
DATA_DIR / WORKSPACES_DIR ./data / ./workspaces Runtime storage locations

Generate a secret:

python -c "import secrets; print(secrets.token_urlsafe(64))"

πŸ“‘ API reference

Method Route Auth Description
POST /api/register β€” Create account β†’ returns access_token
POST /api/login β€” Authenticate β†’ returns access_token
POST /api/process πŸ”’ Enroll one voice (5–20 clips) β†’ builds .whl
POST /api/process_company πŸ”’ Enroll a team (PersonName/clip.wav) β†’ builds .whl
POST /api/verify πŸ”’ Score a sample vs the enrolled voiceprint
GET /api/status πŸ”’ Whether the .whl is ready
GET /api/download πŸ”’ Download the generated .whl
GET /api/health β€” Liveness probe

πŸ”’ = requires Authorization: Bearer <token>. Interactive Swagger docs at /docs.

Example: register β†’ enroll β†’ download (click to expand)
# 1. Register (returns an access_token)
TOKEN=$(curl -s -X POST http://localhost:8000/api/register \
  -F name="Asadullah" -F email="me@example.com" -F password="supersecret1" \
  | python -c "import sys,json; print(json.load(sys.stdin)['access_token'])")

# 2. Enroll 5 voice clips
curl -X POST http://localhost:8000/api/process \
  -H "Authorization: Bearer $TOKEN" \
  -F files=@clip1.wav -F files=@clip2.wav -F files=@clip3.wav \
  -F files=@clip4.wav -F files=@clip5.wav

# 3. Download your SDK
curl -X GET http://localhost:8000/api/download \
  -H "Authorization: Bearer $TOKEN" -o audioauth.whl

🐍 Using your generated SDK

pip install audioauth-1.0.0-py3-none-any.whl

Individual authentication:

from audioauth import authenticate

if authenticate("login_attempt.wav"):
    print("βœ… Access granted")
else:
    print("❌ Access denied")

Team identification:

from audioauth import identify

speaker = identify("meeting_clip.wav")   # -> "Alice" | "Bob" | "Unknown"
print(f"Speaking now: {speaker}")

πŸ—ƒοΈ Data model

The user store (data/login.csv) is deliberately simple β€” swap database.py for SQLite/PostgreSQL in production without touching the rest of the app.

erDiagram
    USER {
        string name
        string email PK
        string password_hash "bcrypt"
        string username
        string registered_at
        string whl_path
    }
    VOICEPRINT {
        string username FK
        blob   embedding_npy "192-dim float32"
    }
    WHEEL {
        string username FK
        file   whl "generated SDK"
    }
    USER ||--|| VOICEPRINT : "has"
    USER ||--o| WHEEL : "builds"
Loading
Column Type Notes
name string Display name
email string Primary key / login identity
password_hash string bcrypt hash (never plaintext)
username string Slugified folder name
registered_at string UTC timestamp
whl_path string Path to the built SDK

πŸ“ Project structure

Voice-Authentication-System/
β”œβ”€β”€ Backend/
β”‚   β”œβ”€β”€ app/
β”‚   β”‚   β”œβ”€β”€ main.py            # FastAPI routes
β”‚   β”‚   β”œβ”€β”€ config.py          # env-driven settings
β”‚   β”‚   β”œβ”€β”€ security.py        # bcrypt + JWT
β”‚   β”‚   β”œβ”€β”€ database.py        # CSV user store (hashed)
β”‚   β”‚   β”œβ”€β”€ email_service.py   # optional welcome email
β”‚   β”‚   β”œβ”€β”€ rate_limit.py      # auth brute-force guard
β”‚   β”‚   └── pipeline/          # preprocess β†’ embed β†’ average β†’ inject β†’ build
β”‚   β”œβ”€β”€ template/              # the SDK skeleton baked per-user
β”‚   β”œβ”€β”€ tests/                 # pytest suite (17 tests)
β”‚   └── requirements.txt
β”œβ”€β”€ Frontend/                  # static multi-page UI (7 pages)
β”œβ”€β”€ Docs/Screenshots/          # README imagery
β”œβ”€β”€ .github/workflows/ci.yml   # lint + test on every push
β”œβ”€β”€ Dockerfile
β”œβ”€β”€ docker-compose.yml
β”œβ”€β”€ .env.example
└── LICENSE

πŸ§ͺ Testing & quality

pip install -r Backend/requirements-dev.txt
pytest                 # 17 unit + API tests (pipeline mocked, no models needed)
ruff check backend     # lint
ruff format backend    # format

CI runs all three on every push and pull request via GitHub Actions.

βœ… Verified end-to-end: the full generation pipeline has been run with the real libraries β€” audio preprocessing β†’ voiceprint injection β†’ .whl compilation β†’ installing and importing the generated audioauth SDK (authenticate() present, 192-dim MASTER_EMBEDDING baked in) β€” for both the individual and company build paths.


πŸ” Security

  • Passwords hashed with bcrypt (work factor 12) β€” never stored in plaintext.
  • JWT bearer authentication on every sensitive endpoint; the acting user is resolved from the signed token, not from client-supplied fields.
  • CORS restricted to an explicit env-driven allow-list.
  • Per-file upload size caps and strict extension checks.
  • Rate-limited auth endpoints to blunt brute-force attempts.
  • Secrets read from the environment, never hard-coded.

See SECURITY.md for the hardening checklist and disclosure policy.


πŸ—ΊοΈ Roadmap

  • Swap the CSV store for SQLite/PostgreSQL
  • Liveness / anti-spoofing detection
  • Refresh-token rotation
  • Multi-language SDK targets (JS/Rust bindings)
  • Admin dashboard for enrolled profiles

🏷️ Keywords

voice-authentication Β· speaker-recognition Β· speaker-verification Β· ecapa-tdnn Β· speechbrain Β· voice-biometrics Β· biometrics Β· audio-processing Β· audio-classification Β· deep-learning Β· machine-learning Β· pytorch Β· fastapi Β· python Β· jwt Β· bcrypt Β· rest-api Β· docker Β· github-actions Β· pytest Β· tailwindcss Β· three-js Β· webgl Β· voice-embeddings Β· cosine-similarity Β· on-device-ai Β· privacy-first


πŸ“‘ Citation

If you use this project in your work, please cite it. A machine-readable CITATION.cff is included, so GitHub shows a "Cite this repository" button on the repo sidebar (APA/BibTeX export).

@software{irshad_voice_authentication_system_2026,
  author    = {Irshad, Asadullah},
  title     = {Voice Authentication System},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.22135813},
  url       = {https://doi.org/10.5281/zenodo.22135813},
  license   = {MIT}
}

The archived release is on Zenodo with DOI 10.5281/zenodo.22135813.


πŸ“š More docs


πŸ“ License & author

Released under the MIT License Β© 2026 Asadullah Irshad.

⭐ If you find this project useful, please consider giving it a star!

Made with ❀️ and Python

About

AI-powered voice biometric authentication using SpeechBrain and ECAPA-TDNN for secure speaker verification.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages