Register once. Embed forever. Authenticate any voice in 2 lines of Python.
A local-first voice-authentication platform that turns a handful of voice samples into a
personalised, installable Python SDK (.whl) β powered by a SpeechBrain ECAPA-TDNN
speaker-embedding model. No audio ever leaves your infrastructure.
- Overview
- Screenshots
- Key features
- Architecture
- How it works
- The pipeline
- Tech stack & skills
- Project stats
- What's new in v2
- Quickstart
- Configuration keys
- API reference
- Using your generated SDK
- Data model
- Project structure
- Testing & quality
- Security
- Roadmap
- Keywords
- License & author
Voice Authentication System verifies who is speaking by comparing a live voice sample against an enrolled voiceprint β entirely on your own machine or server.
- A user records or uploads 5β20 short voice clips.
- The backend cleans each clip, extracts a 192-dimensional embedding, and averages them into one master voiceprint.
- That voiceprint is baked into a generated Python package and compiled into a
.whl. - Anyone can then authenticate a voice in two lines of Python β offline.
| Mode | Endpoint | Produces | Use case |
|---|---|---|---|
| Individual | POST /api/process |
authenticate(audio) β bool |
1:1 login / identity check |
| Company / Team | POST /api/process_company |
identify(audio) β "Name" |
1:N speaker classification |
- π Local-first & private β audio never leaves your infrastructure; embeddings and wheels stay on disk.
- π§ State-of-the-art embeddings β SpeechBrain ECAPA-TDNN (VoxCeleb) 192-dim speaker vectors.
- π¦ Generates a real SDK β outputs an installable
audioauth.whlwith the voiceprint baked in. - π₯ Two modes β individual 1:1 authentication and company-wide 1:N speaker identification.
- π‘οΈ Hardened security β bcrypt password hashing, JWT bearer auth, CORS allow-list, rate limiting.
- π¨ Premium frontend β WebGL shader hero, live MFCC/embedding visualizations, in-browser recorder.
- π³ Production-ready β Dockerfile, docker-compose, GitHub Actions CI, pytest suite, ruff linting.
Full page-by-page feature tour (click to expand)
| Page | File | Highlights |
|---|---|---|
| Landing | index.html |
WebGL hero shader + bloom, How-It-Works flow, "2 lines of auth" code showcase, bento feature grid, pricing, social-proof marquee |
| Auth | auth.html |
Split glass layout, Login β Sign-Up tabs, password show/hide, 4-segment strength meter |
| Account Type | onboarding-type.html |
4-step stepper, Individual vs Team cards with animated selection |
| Voice Recording | onboarding-record.html |
Web Audio API + MediaRecorder, 240px mic circle with live level, 10s countdown, drag-and-drop upload |
| Company Upload | onboarding-company.html |
Spreadsheet-style bulk table, CSV import modal, per-person status badges |
| Processing | onboarding-processing.html |
Live pipeline panel + MFCC heatmap, waveform/energy, pitch contour, 192-dim embedding bars, UMAP scatter |
| Download | download.html |
Ready-state hero, authenticated .whl download, live in-browser voice verification demo |
flowchart LR
subgraph Browser["π Frontend (static)"]
UI[HTML + Tailwind + Three.js]
REC[Web Audio recorder]
end
subgraph API["βοΈ FastAPI backend"]
AUTH[JWT auth + bcrypt]
ROUTES[REST routes]
PIPE[Voice pipeline]
BUILD[.whl builder]
end
subgraph Store["πΎ Local storage"]
CSV[(login.csv<br/>hashed)]
NPY[(embedding.npy)]
WHL[(generated .whl)]
end
MODEL[["π§ ECAPA-TDNN<br/>(SpeechBrain / VoxCeleb)"]]
UI -->|register / login| AUTH
REC -->|upload clips| ROUTES
ROUTES --> PIPE
PIPE --> MODEL
PIPE --> NPY
PIPE --> BUILD --> WHL
AUTH --> CSV
WHL -->|download| UI
Enrollment β SDK generation
sequenceDiagram
actor U as User
participant FE as Frontend
participant API as FastAPI
participant ML as ECAPA-TDNN
participant FS as Storage
U->>FE: Record / upload 5β20 clips
FE->>API: POST /api/process (Bearer JWT)
API->>API: preprocess + validate each clip
API->>ML: encode clips β 192-dim vectors
ML-->>API: embeddings
API->>API: average β master voiceprint
API->>FS: save embedding.npy
API->>API: inject voiceprint β build .whl
API->>FS: save .whl
API-->>FE: { build_id, accepted_files }
FE->>API: GET /api/download (Bearer JWT)
API-->>U: audioauth-1.0.0.whl
Verification
sequenceDiagram
actor U as User
participant FE as Frontend
participant API as FastAPI
participant ML as ECAPA-TDNN
U->>FE: Provide a test sample
FE->>API: POST /api/verify (Bearer JWT)
API->>ML: encode sample β 192-dim vector
API->>API: cosine similarity vs stored voiceprint
API-->>FE: { confidence, label, matched }
FE-->>U: β
Strong / π‘ Partial / β No match
| # | Stage | Module | Input β Output |
|---|---|---|---|
| 1 | Preprocess | pipeline/preprocess.py |
audio file β 16 kHz mono float32, trimmed/padded to 3 s, peak-normalised (rejects clips < 2 s) |
| 2 | Embed | pipeline/embedding.py |
audio array β 192-dim ECAPA-TDNN vector |
| 3 | Average | pipeline/averaging.py |
N embeddings β 1 master voiceprint (mean) |
| 4 | Inject | pipeline/injector.py |
master vector β baked into core.py template |
| 5 | Build | pipeline/builder.py |
package folder β installable .whl |
Match thresholds (cosine similarity, used in the generated SDK):
| Cosine score | Verdict |
|---|---|
| β₯ 0.75 | β Strong match |
| 0.60 β 0.75 | π‘ Partial match |
| < 0.60 | β Rejected |
| Layer | Technology | Purpose |
|---|---|---|
| API | FastAPI, Uvicorn | Async REST API + ASGI server |
| Auth | PyJWT, passlib + bcrypt | Bearer tokens, password hashing |
| ML / audio | SpeechBrain (ECAPA-TDNN), PyTorch, torchaudio, librosa, soundfile | Speaker embeddings + audio decoding |
| Config | pydantic-settings | Type-safe, env-driven config |
| Packaging | PyPA build |
Compiles the per-user SDK .whl |
| Frontend | HTML, Tailwind (CDN), Three.js, ECharts, Web Audio API | Animated UI, live visualizations, recording |
| Dev / CI | pytest, ruff, Docker, GitHub Actions | Tests, linting, containers, automation |
| Metric | Value |
|---|---|
| Total files | 58 |
| Lines of code (py + js + html + css) | ~5,700 |
| Python | 1,724 lines Β· 24 files |
| JavaScript | 999 lines Β· 5 files |
| HTML | 2,653 lines Β· 8 files |
| CSS | 353 lines |
| REST API endpoints | 8 |
| Pipeline stages | 5 |
| Frontend pages | 7 |
| Automated tests | 17 β |
| Embedding dimensionality | 192 |
| Model | ECAPA-TDNN (VoxCeleb) |
This release hardens the original prototype into a production-ready project.
| Area | v1 (prototype) | v2 (this release) |
|---|---|---|
| Passwords | β stored in plaintext | β bcrypt-hashed (work factor 12) |
| Authentication | β email in request body | β JWT bearer tokens |
| CORS | * |
β env-driven allow-list |
| Secrets | β environment variables | |
| Uploads | β unbounded | β size + type limits |
| Brute-force | β none | β rate-limited auth |
| Tests | β none | β 17 automated tests |
| CI/CD | β none | β GitHub Actions |
| Containers | β none | β Dockerfile + compose |
| Docs | β this README + CONTRIBUTING + SECURITY |
git clone https://github.com/asadullahirshad3/Voice-Authentication-System.git
cd Voice-Authentication-System
cp .env.example .env # then edit JWT_SECRET
docker compose up --buildOpen http://localhost:8000.
git clone https://github.com/asadullahirshad3/Voice-Authentication-System.git
cd Voice-Authentication-System
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r Backend/requirements.txt
cp .env.example .env # set a strong JWT_SECRET
cd Backend
uvicorn app.main:app --reloadThe ECAPA-TDNN model (~80 MB) downloads automatically from Hugging Face on the first
/api/processcall and is cached locally afterwards.
All settings come from environment variables (or a .env file). Full list in .env.example.
| Key | Default | Description |
|---|---|---|
JWT_SECRET |
(placeholder) | Set a long random value in production. Signs access tokens. |
JWT_ALGORITHM |
HS256 |
JWT signing algorithm |
ACCESS_TOKEN_EXPIRE_MINUTES |
1440 |
Token lifetime (minutes) |
CORS_ORIGINS |
http://localhost:8000,... |
Comma-separated browser origin allow-list |
MIN_FILES / MAX_FILES |
5 / 20 |
Enrollment clip bounds |
MAX_UPLOAD_MB |
25 |
Per-file upload cap |
EMAIL_ENABLED |
false |
Toggle the welcome email |
SMTP_HOST |
smtp.gmail.com |
SMTP server host |
SMTP_PORT |
587 |
SMTP port |
SMTP_USER / SMTP_PASSWORD |
(empty) | SMTP credentials (if email enabled) |
DATA_DIR / WORKSPACES_DIR |
./data / ./workspaces |
Runtime storage locations |
Generate a secret:
python -c "import secrets; print(secrets.token_urlsafe(64))"| Method | Route | Auth | Description |
|---|---|---|---|
POST |
/api/register |
β | Create account β returns access_token |
POST |
/api/login |
β | Authenticate β returns access_token |
POST |
/api/process |
π | Enroll one voice (5β20 clips) β builds .whl |
POST |
/api/process_company |
π | Enroll a team (PersonName/clip.wav) β builds .whl |
POST |
/api/verify |
π | Score a sample vs the enrolled voiceprint |
GET |
/api/status |
π | Whether the .whl is ready |
GET |
/api/download |
π | Download the generated .whl |
GET |
/api/health |
β | Liveness probe |
π = requires Authorization: Bearer <token>. Interactive Swagger docs at /docs.
Example: register β enroll β download (click to expand)
# 1. Register (returns an access_token)
TOKEN=$(curl -s -X POST http://localhost:8000/api/register \
-F name="Asadullah" -F email="me@example.com" -F password="supersecret1" \
| python -c "import sys,json; print(json.load(sys.stdin)['access_token'])")
# 2. Enroll 5 voice clips
curl -X POST http://localhost:8000/api/process \
-H "Authorization: Bearer $TOKEN" \
-F files=@clip1.wav -F files=@clip2.wav -F files=@clip3.wav \
-F files=@clip4.wav -F files=@clip5.wav
# 3. Download your SDK
curl -X GET http://localhost:8000/api/download \
-H "Authorization: Bearer $TOKEN" -o audioauth.whlpip install audioauth-1.0.0-py3-none-any.whlIndividual authentication:
from audioauth import authenticate
if authenticate("login_attempt.wav"):
print("β
Access granted")
else:
print("β Access denied")Team identification:
from audioauth import identify
speaker = identify("meeting_clip.wav") # -> "Alice" | "Bob" | "Unknown"
print(f"Speaking now: {speaker}")The user store (data/login.csv) is deliberately simple β swap database.py for
SQLite/PostgreSQL in production without touching the rest of the app.
erDiagram
USER {
string name
string email PK
string password_hash "bcrypt"
string username
string registered_at
string whl_path
}
VOICEPRINT {
string username FK
blob embedding_npy "192-dim float32"
}
WHEEL {
string username FK
file whl "generated SDK"
}
USER ||--|| VOICEPRINT : "has"
USER ||--o| WHEEL : "builds"
| Column | Type | Notes |
|---|---|---|
name |
string | Display name |
email |
string | Primary key / login identity |
password_hash |
string | bcrypt hash (never plaintext) |
username |
string | Slugified folder name |
registered_at |
string | UTC timestamp |
whl_path |
string | Path to the built SDK |
Voice-Authentication-System/
βββ Backend/
β βββ app/
β β βββ main.py # FastAPI routes
β β βββ config.py # env-driven settings
β β βββ security.py # bcrypt + JWT
β β βββ database.py # CSV user store (hashed)
β β βββ email_service.py # optional welcome email
β β βββ rate_limit.py # auth brute-force guard
β β βββ pipeline/ # preprocess β embed β average β inject β build
β βββ template/ # the SDK skeleton baked per-user
β βββ tests/ # pytest suite (17 tests)
β βββ requirements.txt
βββ Frontend/ # static multi-page UI (7 pages)
βββ Docs/Screenshots/ # README imagery
βββ .github/workflows/ci.yml # lint + test on every push
βββ Dockerfile
βββ docker-compose.yml
βββ .env.example
βββ LICENSE
pip install -r Backend/requirements-dev.txt
pytest # 17 unit + API tests (pipeline mocked, no models needed)
ruff check backend # lint
ruff format backend # formatCI runs all three on every push and pull request via GitHub Actions.
β
Verified end-to-end: the full generation pipeline has been run with the real
libraries β audio preprocessing β voiceprint injection β .whl compilation β
installing and importing the generated audioauth SDK (authenticate() present, 192-dim
MASTER_EMBEDDING baked in) β for both the individual and company build paths.
- Passwords hashed with bcrypt (work factor 12) β never stored in plaintext.
- JWT bearer authentication on every sensitive endpoint; the acting user is resolved from the signed token, not from client-supplied fields.
- CORS restricted to an explicit env-driven allow-list.
- Per-file upload size caps and strict extension checks.
- Rate-limited auth endpoints to blunt brute-force attempts.
- Secrets read from the environment, never hard-coded.
See SECURITY.md for the hardening checklist and disclosure policy.
- Swap the CSV store for SQLite/PostgreSQL
- Liveness / anti-spoofing detection
- Refresh-token rotation
- Multi-language SDK targets (JS/Rust bindings)
- Admin dashboard for enrolled profiles
voice-authentication Β· speaker-recognition Β· speaker-verification Β· ecapa-tdnn Β·
speechbrain Β· voice-biometrics Β· biometrics Β· audio-processing Β· audio-classification Β·
deep-learning Β· machine-learning Β· pytorch Β· fastapi Β· python Β· jwt Β· bcrypt Β·
rest-api Β· docker Β· github-actions Β· pytest Β· tailwindcss Β· three-js Β· webgl Β·
voice-embeddings Β· cosine-similarity Β· on-device-ai Β· privacy-first
If you use this project in your work, please cite it. A machine-readable
CITATION.cff is included, so GitHub shows a "Cite this repository"
button on the repo sidebar (APA/BibTeX export).
@software{irshad_voice_authentication_system_2026,
author = {Irshad, Asadullah},
title = {Voice Authentication System},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.22135813},
url = {https://doi.org/10.5281/zenodo.22135813},
license = {MIT}
}The archived release is on Zenodo with DOI
10.5281/zenodo.22135813.
- CHANGELOG.md β version history (v1 β v2)
- CONTRIBUTING.md β how to contribute
- SECURITY.md β security policy & hardening checklist
- PUBLISHING.md β step-by-step GitHub publishing guide
Released under the MIT License Β© 2026 Asadullah Irshad.
β If you find this project useful, please consider giving it a star!
Made with β€οΈ and Python






