- Overview
- Live Product Preview
- The Core Idea β Why No Training?
- The AI Model β ECAPA-TDNN Deep Dive
- System Architecture
- Backend β Deep Code Analysis
- Frontend β Page-by-Page Walkthrough
- What the User Gets β The .whl Package
- Complete User Journey
- Industry Use Cases
- API Reference
- Project Structure
- Local Setup & Installation
- Configuration
- Security Considerations
- Roadmap
- Author
Audio Classification Hub is a production-grade, full-stack Voice Authentication Platform that
converts spoken voice samples into a permanent, portable identity vector β a voiceprint β and
packages it into a downloadable Python .whl file that any developer can install and use offline in 2 lines of code.
This is not a cloud-locked SaaS. The intelligence ships with the user.
| Traditional Voice Auth | Audio Classification Hub |
|---|---|
| Requires cloud API calls on every verification | Fully offline after .whl install |
| Vendor lock-in, per-call billing | One-time registration, zero recurring cost |
| Data sent to third-party servers | Voiceprint stays on-premise |
| SDK tied to platform version | Pure Python wheel, works anywhere |
| Weeks to integrate | 2 lines of Python |
The landing page features a WebGL shader canvas rendering animated chromatic voice waveforms built with Three.js. Every section uses glassmorphism, scroll-reveal animations, and a dark-mode design system.
| Page | What You See |
|---|---|
Landing (index.html) |
Animated waveform hero, 3-step explainer, bento feature grid, marquee social proof |
Auth (auth.html) |
Split-screen login/signup with live password strength indicator |
Type Select (onboarding-type.html) |
Individual vs. Team/Company card selector with animated ping rings |
Voice Recorder (onboarding-record.html) |
Live browser microphone recorder + drag-and-drop upload zone (5β20 samples) |
Processing (onboarding-processing.html) |
Real-time pipeline: MFCC heatmap, waveform, pitch contour, 192-dim embedding bars, UMAP projection |
Download (download.html) |
SDK download, code snippets, voice verification playground |
This is the most important architectural decision in the project.
Most voice recognition tutorials tell you to:
- Collect thousands of hours of labeled speech data
- Train a classifier from scratch (weeks of GPU time)
- Re-train whenever you add a new user
- Deploy a heavy model that classifies into fixed categories
This approach cannot scale for personal authentication. If you trained on 1,000 users and a new user joins, you'd need to re-train the entire network.
We use a concept from metric learning:
Pretrained Model (ECAPA-TDNN)
β
Maps any voice β 192-dimensional vector space
β
Voices from the SAME person β vectors that are CLOSE together
Voices from DIFFERENT people β vectors that are FAR apart
The model was pre-trained on VoxCeleb β 2,000+ speakers, 1M+ utterances from YouTube. It learned the universal geometry of human voice space. We do not train anything new. We just:
- Encode the user's 5β20 voice samples into 192-dim vectors
- Average them into a single master voiceprint
- At verification time, compute cosine similarity between master and live sample
- If similarity > threshold β authenticated
This is the same principle powering FaceID.
Benefits:
- β Zero training time β new user registration takes ~60 seconds
- β No GPU required at inference β runs on any CPU
- β Adding users does not affect accuracy for others
- β
Model is compact and ships inside the
.whl
| Property | Value |
|---|---|
| Architecture | ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation β Time Delay Neural Network) |
| Pre-trained on | VoxCeleb 1 & 2 (2,000+ speakers, 1M+ utterances) |
| Output | 192-dimensional L2-normalized embedding vector |
| Input format | 16kHz mono float32 numpy array |
| Inference device | CPU (no GPU needed) |
| Source | HuggingFace Hub speechbrain/spkrec-ecapa-voxceleb |
| Framework | SpeechBrain + PyTorch |
Raw Audio (16kHz PCM)
β
Frame-level Feature Extraction
(Filter banks / MFCCs at each time step)
β
TDNN Layers with Dilation
(Captures short + long-range temporal dependencies)
β
SE-Res2Block (Squeeze-and-Excitation + Residual)
(Channel attention β emphasizes discriminative voice features)
β
Multi-scale Feature Aggregation (MFA)
(Combines features from ALL TDNN layers)
β
Attentive Statistics Pooling
(Collapses variable-length sequence β fixed representation)
β
Fully Connected Layer
β
192-dim L2-normalized Vector β THE VOICEPRINT
The 192-dim space is specifically tuned for speaker discrimination:
- Each dimension captures abstract acoustic properties (vocal tract shape, pitch patterns, speaking style)
- L2 normalization means all vectors lie on a unit hypersphere β cosine similarity equals dot product
- Empirically the sweet spot between expressiveness and computational cost
# Raw cosine similarity in [-1.0, +1.0]
cosine = dot(master, test_emb) / (norm(master) * norm(test_emb))
# Display confidence: mapped to 0-100% with +20 boost for UX clarity
# cosine 0.20 β 40% | 0.45 β 65% | 0.62 β 82% | 0.80 β 100%
confidence = min(100, max(0, (cosine * 100) + 20))| Raw Cosine | Display % | Label | Match? |
|---|---|---|---|
| >= 0.62 | >= 82% | π’ Strong Match | β Yes |
| >= 0.45 | >= 65% | π‘ Partial Match | β Yes |
| >= 0.25 | >= 45% | π Weak Match | β No |
| < 0.25 | < 45% | π΄ No Match | β No |
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β AUDIO CLASSIFICATION HUB β
ββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββ€
β FRONTEND β BACKEND β
β (HTML + JS + CSS) β (FastAPI + Python) β
β β β
β index.html β βββββββββββββββββββββββββββββββββββββ β
β auth.html ββββββββββββΌββΊβ FastAPI (Uvicorn :8000) β β
β onboarding pages β βββββββββββββββββ¬ββββββββββββββββββββ β
β download.html β β β
β β ββββββββββββΌβββββββββββ β
β js/api.js β β ML PIPELINE β β
β js/recorder.js β β 1. preprocess.py β β
β js/processing.js β β 2. embedding.py β β
β js/hero-shader.js β β (ECAPA-TDNN) β β
β css/style.css β β 3. averaging.py β β
β β β 4. injector.py β β
β β β 5. builder.py β β
ββββββββββββββββββββββββββββ€ ββββββββββββ¬ββββββββββββ β
β β β
β βββββββββββββββββΌβββββββββββββββββββββ β
β β FILE SYSTEM DB β β
β β DataBase/login.csv (registry) β β
β β DataBase/<user>/embedding.npy β β
β β workspaces/<user>_<id>/dist/*.whl β β
β βββββββββββββββββββββββββββββββββββββββ β
β βββββββββββββββββββββββββββββββββββββββ β
β β SMTP Email (.whl attached) β β
β βββββββββββββββββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββ
The backend is built with FastAPI, chosen for its async-first design, automatic OpenAPI docs
at /docs, Pydantic validation for multipart forms, and BackgroundTasks for non-blocking email delivery.
The entry point registers 7 REST API endpoints and mounts the frontend static assets at the
same paths the HTML expects (/css/, /js/), so no separate web server is needed.
# Frontend HTML + Backend API on same origin β no CORS complexity
app.mount("/css", StaticFiles(directory=str(css_dir)), name="css")
app.mount("/js", StaticFiles(directory=str(js_dir)), name="js")Global error handling ensures all exceptions return clean JSON β the JS safeJson() parser always succeeds.
| Method | Endpoint | Purpose |
|---|---|---|
POST |
/api/register |
Create new user account |
POST |
/api/login |
Validate credentials |
POST |
/api/process |
Upload voice samples β run full ML pipeline β build .whl |
POST |
/api/process_company |
Multi-person bulk upload β multi-embedding .whl |
GET |
/api/download |
Stream the generated .whl file |
GET |
/api/status |
Poll whether .whl build is ready |
POST |
/api/verify |
Verify a voice sample against stored voiceprint |
GET |
/ and page routes |
Serve all frontend HTML pages |
Every uploaded audio file goes through a 3-backend loading cascade:
torchaudio (soundfile β sox_io) βββΊ succeeds for WAV/FLAC/OGG
β fail
soundfile + resampy βββΊ succeeds for WAV/FLAC
β fail
librosa (requires ffmpeg) βββΊ catches MP4/WebM/M4A
After loading, audio is:
- Resampled to exactly 16,000 Hz (ECAPA-TDNN requirement)
- Mixed down to mono
- Validated β rejected if < 2 seconds
- Padded or trimmed to exactly 3 seconds (48,000 samples)
- Peak-normalized to [-1.0, +1.0]
SR = 16000 # ECAPA-TDNN requires exactly 16kHz
DURATION = 3 # seconds window used for embedding
MIN_SEC = 2.0 # reject files shorter than thisThe ECAPA-TDNN model is loaded once into a module-level global _model (lazy singleton)
to avoid reloading ~100MB on every request.
Windows Compatibility Fix: On Windows without Developer Mode, pathlib.Path.symlink_to()
raises [WinError 1314]. SpeechBrain calls this during model caching. The module monkey-patches it:
# Patch BEFORE any SpeechBrain import
_orig_symlink_to = Path.symlink_to
def _safe_symlink_to(self, target, target_is_directory=False):
try:
_orig_symlink_to(self, target, target_is_directory)
except OSError:
shutil.copy2(str(target_p), str(self)) # fall back to file copy
Path.symlink_to = _safe_symlink_toInference:
def get_embedding(audio: np.ndarray) -> np.ndarray:
tensor = torch.FloatTensor(audio).unsqueeze(0) # (1, N)
with torch.no_grad():
emb = model.encode_batch(tensor) # (1, 1, 192)
return emb.squeeze().numpy().astype(np.float32) # (192,)Multiple samples produce slightly different vectors (phrasing, noise). Averaging builds a centroid that is more robust:
def build_master(embeddings: list) -> np.ndarray:
stacked = np.stack(embeddings, axis=0) # (N, 192)
master = np.mean(stacked, axis=0) # (192,)
return master.astype(np.float32)Rule of thumb: More samples = more accurate centroid. 5 is minimum; 20 is recommended maximum.
This is what makes the .whl work without any server calls. The user's 192-float embedding
is literally baked into Python source code:
# core_template.py contains: EMBEDDING = {{EMBEDDING}}
# The injector replaces the placeholder with actual numbers:
embedding_str = repr(master.tolist()) # "[0.023, -0.114, ...]"
final_code = template_code.replace("{{EMBEDDING}}", embedding_str)When the .whl is imported, no database, no network, no model download is needed.
Company Mode injects a dict of {person_name: [192-float-list]} instead, enabling
1:N speaker identification from a single package.
subprocess.run([sys.executable, "-m", "build"], cwd=str(build_dir))Uses Python build module (PEP 517/518). Each build gets a UUID-isolated workspace so concurrent
builds never interfere. Workspace lives outside the server directory to prevent uvicorn --reload
from triggering on generated files.
Uses a flat-file CSV approach β simple, portable, zero-dependency:
DataBase/
βββ login.csv # name, email, password, username, registered_at, whl_path
βββ mohit_jadav/
β βββ voices/ # raw uploaded audio files
β βββ embedding.npy # 192-dim master voiceprint
βββ harsh_jadav/
βββ voices/
βββ embedding.npy
create_user()β creates user row AND theDataBase/<username>/folder atomicallysave_whl_path()β updates whl_path column after a successful buildget_whl_path()β lookup used by/api/downloadand/api/status
After every successful pipeline run, a non-blocking background task sends a welcome email with the .whl attached:
background_tasks.add_task(
send_welcome_email,
name=user["name"],
email=email,
whl_path=whl_path,
) # Returns API response immediately β email sends asynchronouslyEmail failures are caught and logged but never crash the user-facing request.
The frontend is a pure HTML + CSS + JavaScript SPA (no React, no Vue) served directly by FastAPI.
Tech Stack:
- Tailwind CSS (CDN) for utility classes
- Three.js for WebGL voice waveform shaders
- Apache ECharts for processing visualization charts
- Font Awesome 6 for iconography
- Google Fonts β Space Grotesk, Inter, JetBrains Mono
- Custom
css/style.cssβ glassmorphism design system, CSS tokens, micro-animations
| Element | Implementation |
|---|---|
| Hero Canvas | Three.js WebGL shader β animated chromatic waveform sinusoids |
| Scroll Reveal | IntersectionObserver with staggered data-delay |
| 3-Step How It Works | Animated ping rings, waveform bars via @keyframes wfPulse |
| Code Window | Multi-tab snippet (Python / FastAPI / Node / cURL) |
| Bento Grid | 4-column masonry feature cards with radial gradient glows |
| Social Proof Marquee | CSS infinite scroll strip |
| CTA | Start Free β auth.html / View Docs β download.html |
- Split-screen β left: animated canvas + tagline, right: auth card
- Tab toggle β Login / Sign Up with animated active state
- Password strength meter β 4-segment bar (Too weak β Strong) via real-time regex
- Session management β
Session.save()tosessionStorage; redirects based onwhl_ready
- 4-step progress stepper with animated connectors (Step 1 active)
- Individual card β
onboarding-record.html(single voiceprint) - Team/Company card β
onboarding-company.html(multi-user folder upload) - Card selection triggers animated bounce-in checkmark
Record Mode (WebRTC):
MediaRecorder APIβ browser microphone asaudio/webmblobs- Live countdown timer with
requestAnimationFrame - 3 concentric animated pulse rings during recording
- Playback bar + Re-record + Next sample controls
- Counter:
0 / 20 capturedin real time
Upload Mode (Drag & Drop):
dragover/dropevents with visual border feedback<input type="file" multiple accept=".wav,.mp3,.flac,.ogg,.m4a">- Validates 5β20 files before enabling Process button
Left Panel β Pipeline Steps:
- 5 animated steps: Loading Audio β Preprocessing β MFCC Extraction β ECapa Embedding β Packaging .whl
- CSS
transform: translateX(4px)on active step - Real elapsed timer + gradient progress bar
linear-gradient(90deg, #6366F1, #00D4FF)
Right Panel β ECharts Visualizations:
| Chart | What It Shows |
|---|---|
| MFCC Heatmap | 13 Mel-Frequency Cepstral Coefficient tracks across time |
| Waveform + Energy Envelope | Amplitude over time with energy overlay |
| Pitch Contour (F0) | Fundamental frequency curve across the utterance |
| 192-dim Embedding Bars | All 192 embedding dimensions visualized |
| UMAP Scatter Plot | 2D projection: user voiceprint vs. population cluster |
Backend Polling (every 3 seconds, max 6 minutes):
setInterval(async () => {
const data = await apiStatus(user.email);
if (data.whl_ready) {
clearInterval(pollTimer);
forceComplete(); // snap UI to 100%
showDownloadCTA(); // no auto-redirect β user decides
}
}, 3000);- Download button β
GET /api/download?email=<email> - Multi-language code snippets (Python, FastAPI, Node.js, cURL)
- Live Voice Verify playground β record test sample, see animated confidence result card
After completing registration, the user receives:
audioauth-1.0.0-py3-none-any.whl
audioauth/
βββ __init__.py
βββ core.py β YOUR VOICEPRINT BAKED IN (192-float Python list)
βββ ...
core.py contains your embedding as a static Python constant:
EMBEDDING = [0.0234, -0.1142, 0.4471, 0.0891, -0.2310, ...] # 192 floatspip install audioauth-1.0.0-py3-none-any.whl# 2 lines. That is all.
from audioauth import VoiceAuth
auth = VoiceAuth()
result = auth.verify("sample.wav") # Returns: True / FalseAdvanced usage:
details = auth.verify_detailed("voice.wav")
print(details["confidence"]) # 0.0 to 100.0
print(details["label"]) # "Strong Match" / "Partial Match" / ...
print(details["matched"]) # True / False
print(details["cosine_score"]) # raw cosine similarity
# Company / multi-user mode
auth = VoiceAuth(mode="company")
speaker = auth.identify("voice.wav") # returns matched person nameFastAPI integration:
from fastapi import FastAPI, UploadFile
from audioauth import VoiceAuth
app = FastAPI()
auth = VoiceAuth()
@app.post("/check-voice")
async def check_voice(file: UploadFile):
contents = await file.read()
return {"authenticated": auth.verify_bytes(contents)}- β No SpeechBrain model weights β embedding already extracted, model not needed
- β No network calls β runs 100% offline
- β No cloud dependency β your voiceprint, your server
- β No GPU requirement β pure NumPy cosine similarity at ~50ms latency
REGISTRATION FLOW
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β 1. LANDING PAGE β
β Learns about product β Clicks "Start Free" β
β β β
β 2. AUTH PAGE β
β Sign Up β POST /api/register β
β β Redirected to Type Selection β
β β β
β 3. TYPE SELECTION β
β Choose: Individual OR Team/Company β
β β β
β 4. VOICE RECORDING β
β Record 5β20 samples OR drag-and-drop files β
β β Click "Process Voice" β
β β β
β 5. PIPELINE (~60 seconds, async) β
β POST /api/process β
β βββ Stage 1: clean_audio() x N files β
β βββ Stage 2: get_embedding() β N x [192] β
β βββ Stage 3: build_master() β [192] β
β βββ Stage 4: inject_embedding() β core.py β
β βββ Stage 5: build_whl() β .whl β
β β β
β 6. PROCESSING PAGE (polls /api/status every 3s) β
β MFCC, waveform, pitch, embedding, UMAP charts β
β β Done β Download .whl CTA appears β
β β β
β 7. DOWNLOAD PAGE β
β GET /api/download β .whl streamed to browser β
β + Welcome email with .whl attached β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
DAILY AUTHENTICATION FLOW
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β pip install audioauth-1.0.0-py3-none-any.whl β
β from audioauth import VoiceAuth β
β auth = VoiceAuth() β
β result = auth.verify("voice.wav") β
β β β
β 1. Load [192] from core.py (static constant) β
β 2. Load audio β 16kHz numpy array β
β 3. Compute cosine similarity β
β 4. Return True/False + confidence β
β β
β ~50ms latency Β· 0 network calls Β· 0 GPU β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Audio Classification Hub targets both B2C (developers, prosumers) and B2B (enterprises, SaaS).
- Phone banking authentication β replace KBA with voice biometrics
- IVR gating β authenticate callers before agent transfer
- Anti-fraud β detect account takeover by verifying caller identity
- PCI-DSS voice channel compliance
- Telehealth patient verification before prescription refills
- EHR access control β voice-lock records for specific clinicians
- HIPAA-compliant (voiceprint on-premise, zero cloud exposure)
- Remote work time-tracking β employees clock in by voice
- Secure facility access β voice-gated physical security
- Executive document DLP enforcement by speaker verification
- Passive agent-side verification as call connects
- VIP customer detection β priority queue routing
- Callback authentication before resuming sensitive sessions
- Online exam authentication throughout the session
- Voice-based lecture attendance check-in
- Adaptive learning gates β unlock advanced content for verified learners
- Password manager voice unlock β biometric second factor
- Smart home voice commands β prevent spoofing by non-members
- Parental controls β voice-gated content restrictions
- Ship the
.whlinside your product installer - Wrap
/api/verifyas a microservice sidecar for any service - IoT / Edge devices β offline voice auth for Raspberry Pi, embedded systems
Request (multipart/form-data): name, email, password (min 6 chars)
{ "status": "ok", "message": "Account created", "username": "mohit_jadav" }Request: email, password
{ "status": "ok", "name": "Mohit Jadav", "email": "...", "username": "mohit_jadav", "whl_ready": false }Request: email, files[] β 5 to 20 audio files (.wav .mp3 .flac .ogg .m4a)
{
"status": "ok",
"message": "Voice embedding created and .whl built successfully",
"build_id": "a3f91c2b",
"accepted_files": 8,
"skipped_files": 0
}Bulk multi-person enrollment. Files structured as PersonName/filename.wav (5β20 per person).
{ "status": "ok", "whl_ready": true, "whl_filename": "audioauth-1.0.0-py3-none-any.whl" }Response: application/octet-stream β the .whl binary.
Request: email, file (single audio)
{
"status": "ok",
"confidence": 78.4,
"cosine_score": 0.5840,
"label": "Partial Match",
"matched": true,
"color": "amber",
"icon": "fa-circle-half-stroke"
}Full_Working/
β
βββ π Backend/
β βββ π your_server/
β βββ main.py β FastAPI app, all routes, static serving
β βββ database.py β CSV read/write, user management
β βββ login.py β auth validation, welcome email via SMTP
β βββ requirements.txt
β βββ .env.example β SMTP config template
β βββ π pipeline/
β β βββ preprocess.py β audio loading & cleaning (3-backend cascade)
β β βββ embedding.py β ECAPA-TDNN inference + HuggingFace model load
β β βββ averaging.py β N embeddings β 1 master voiceprint
β β βββ injector.py β bake embedding into Python source code
β β βββ builder.py β python -m build β .whl package
β βββ π pretrained_models/
β β βββ spkrec-ecapa-voxceleb/ β ECAPA-TDNN weights (auto-downloaded)
β βββ π template/ β .whl template with {{EMBEDDING}} placeholder
β
βββ π Fronted/
β βββ index.html β Landing page
β βββ auth.html β Login / Sign Up
β βββ onboarding-type.html β Individual vs. Company selector
β βββ onboarding-record.html β Voice recorder + file upload
β βββ onboarding-company.html β Bulk multi-person upload
β βββ onboarding-processing.html β Pipeline visualizer + polling
β βββ download.html β SDK download + verify playground
β βββ π css/style.css β Full design system (glassmorphism, tokens, animations)
β βββ π js/
β βββ api.js β HTTP client + Session management
β βββ main.js β Navbar, footer, scroll-reveal
β βββ recorder.js β MediaRecorder API, WebRTC capture
β βββ processing.js β ECharts: MFCC, waveform, UMAP, embedding
β βββ hero-shader.js β Three.js WebGL waveform shader
β
βββ π DataBase/
β βββ login.csv β User registry (auto-created)
β βββ π <username>/
β βββ π voices/ β Raw uploaded audio samples
β βββ embedding.npy β 192-dim master voiceprint
β
βββ π workspaces/
β βββ π <username>_<build_id>/
β βββ π build/dist/
β βββ audioauth-*.whl
β
βββ π Screenshots/
βββ Screenshot 2026-08-01 190641.png
| Requirement | Version |
|---|---|
| Python | 3.10+ |
| pip | latest |
| ffmpeg | optional (for .webm/.mp4 support) |
| SMTP access | Gmail App Password recommended |
git clone https://github.com/mohitjadav/audio-classification-hub.git
cd audio-classification-hub/Full_Workingcd Backend/your_server
python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # macOS / Linux
pip install -r requirements.txtcp .env.example .envSMTP_HOST=smtp.gmail.com
SMTP_PORT=587
SMTP_USER=your_email@gmail.com
SMTP_PASSWORD=your_gmail_app_passwordGmail App Password: Google Account β Security β 2-Step Verification β App Passwords β Generate
uvicorn main:app --reload --host 0.0.0.0 --port 8000First startup downloads the ECAPA-TDNN model (~100MB):
[EMBEDDING] Path.symlink_to patched β fallback to copy on Windows.
[EMBEDDING] Downloading model to: pretrained_models/spkrec-ecapa-voxceleb
[EMBEDDING] Download complete.
[EMBEDDING] ECAPA-TDNN loaded successfully.
INFO: Uvicorn running on http://0.0.0.0:8000
Subsequent starts use the cached model:
[EMBEDDING] Using cached model at: pretrained_models/spkrec-ecapa-voxceleb
Navigate to http://localhost:8000
| Variable | Default | Description |
|---|---|---|
SMTP_HOST |
smtp.gmail.com |
SMTP server hostname |
SMTP_PORT |
587 |
SMTP port (TLS) |
SMTP_USER |
β | Sender email address |
SMTP_PASSWORD |
β | Gmail App Password |
Audio constants (pipeline/preprocess.py):
| Constant | Default | Description |
|---|---|---|
SR |
16000 |
Target sample rate (Hz) |
DURATION |
3 |
Seconds of audio per embedding |
MIN_SEC |
2.0 |
Minimum audio length accepted |
Confidence thresholds (main.py β /api/verify):
| Label | Threshold | matched |
|---|---|---|
| Strong Match | >= 82% | True |
| Partial Match | >= 65% | True |
| Weak Match | >= 45% | False |
| No Match | < 45% | False |
| Area | Current State | Production Recommendation |
|---|---|---|
| Password Storage | Plain text in CSV | Hash with bcrypt or argon2-cffi |
| Session | sessionStorage (browser) |
JWT tokens with python-jose |
| Database | CSV flat file | PostgreSQL with SQLAlchemy |
| CORS | allow_origins=["*"] |
Restrict to your domain |
| File Validation | Extension + size checks | Also validate magic bytes |
| SMTP | Credentials in .env |
Managed email (SendGrid, Postmark) |
| Rate Limiting | None | Add slowapi for /api/process |
| HTTPS | None (local dev) | Nginx reverse proxy + Let's Encrypt |
- PostgreSQL + SQLAlchemy β replace CSV database
- JWT Authentication β stateless session tokens
- Password hashing β bcrypt
- Multi-tenancy β org-level API keys for B2B
- Anti-spoofing β liveness detection (replay attack prevention)
- Noise augmentation β improve robustness during registration
- Dashboard β admin panel with user management, verification logs
- Docker β
Dockerfile+docker-compose.ymlfor one-command deployment - HTTPS / Nginx β production deployment guide
- Model quantization β INT8 ECAPA-TDNN for faster CPU inference
- Webhook support β POST verification result to developer callback URL
Full-Stack AI Engineer Β Β·Β Deep Learning Β Β·Β FastAPI Β Β·Β Voice Biometrics
"I don't just build models β I build systems that put intelligence in developers' hands without cloud dependency."
Core Stack: Python Β Β·Β FastAPI Β Β·Β PyTorch Β Β·Β SpeechBrain Β Β·Β NumPy Β Β·Β HTML/CSS/JS Β Β·Β TailwindCSS Β Β·Β Three.js
Domains: Machine Learning Β Β·Β Deep Learning Β Β·Β Audio AI Β Β·Β Web Development Β Β·Β Voice Biometrics Β Β·Β API Design
Built with passion for making AI accessible.
Audio Classification Hub Β βΒ Β© 2026 Mohit Jadav. MIT License.
