Skip to content

Latest commit

Β 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Audio Classification Hub



πŸŽ™οΈ Audio Classification Hub

Register once. Embed forever. Authenticate any voice in 2 lines of Python.


Python FastAPI SpeechBrain PyTorch License Status


B2C Β· B2B Β· Enterprise Β |Β  Voice Biometric Authentication as a Service Β |Β  Zero Cloud Lock-In


πŸ“– Table of Contents


🌟 Overview

Audio Classification Hub is a production-grade, full-stack Voice Authentication Platform that converts spoken voice samples into a permanent, portable identity vector β€” a voiceprint β€” and packages it into a downloadable Python .whl file that any developer can install and use offline in 2 lines of code.

This is not a cloud-locked SaaS. The intelligence ships with the user.

What Makes It Different

Traditional Voice Auth Audio Classification Hub
Requires cloud API calls on every verification Fully offline after .whl install
Vendor lock-in, per-call billing One-time registration, zero recurring cost
Data sent to third-party servers Voiceprint stays on-premise
SDK tied to platform version Pure Python wheel, works anywhere
Weeks to integrate 2 lines of Python

πŸ–₯️ Live Product Preview

The landing page features a WebGL shader canvas rendering animated chromatic voice waveforms built with Three.js. Every section uses glassmorphism, scroll-reveal animations, and a dark-mode design system.

Page What You See
Landing (index.html) Animated waveform hero, 3-step explainer, bento feature grid, marquee social proof
Auth (auth.html) Split-screen login/signup with live password strength indicator
Type Select (onboarding-type.html) Individual vs. Team/Company card selector with animated ping rings
Voice Recorder (onboarding-record.html) Live browser microphone recorder + drag-and-drop upload zone (5–20 samples)
Processing (onboarding-processing.html) Real-time pipeline: MFCC heatmap, waveform, pitch contour, 192-dim embedding bars, UMAP projection
Download (download.html) SDK download, code snippets, voice verification playground

πŸ’‘ The Core Idea β€” Why No Training?

This is the most important architectural decision in the project.

Traditional Deep Learning Approach ❌

Most voice recognition tutorials tell you to:

  1. Collect thousands of hours of labeled speech data
  2. Train a classifier from scratch (weeks of GPU time)
  3. Re-train whenever you add a new user
  4. Deploy a heavy model that classifies into fixed categories

This approach cannot scale for personal authentication. If you trained on 1,000 users and a new user joins, you'd need to re-train the entire network.

Our Approach β€” Pretrained Embedding + Cosine Similarity βœ…

We use a concept from metric learning:

Pretrained Model (ECAPA-TDNN)
        ↓
Maps any voice β†’ 192-dimensional vector space
        ↓
Voices from the SAME person  β†’ vectors that are CLOSE together
Voices from DIFFERENT people β†’ vectors that are FAR apart

The model was pre-trained on VoxCeleb β€” 2,000+ speakers, 1M+ utterances from YouTube. It learned the universal geometry of human voice space. We do not train anything new. We just:

  1. Encode the user's 5–20 voice samples into 192-dim vectors
  2. Average them into a single master voiceprint
  3. At verification time, compute cosine similarity between master and live sample
  4. If similarity > threshold β†’ authenticated

This is the same principle powering FaceID.

Benefits:

  • βœ… Zero training time β€” new user registration takes ~60 seconds
  • βœ… No GPU required at inference β€” runs on any CPU
  • βœ… Adding users does not affect accuracy for others
  • βœ… Model is compact and ships inside the .whl

🧠 The AI Model β€” ECAPA-TDNN Deep Dive

Model: speechbrain/spkrec-ecapa-voxceleb

Property Value
Architecture ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation β€” Time Delay Neural Network)
Pre-trained on VoxCeleb 1 & 2 (2,000+ speakers, 1M+ utterances)
Output 192-dimensional L2-normalized embedding vector
Input format 16kHz mono float32 numpy array
Inference device CPU (no GPU needed)
Source HuggingFace Hub speechbrain/spkrec-ecapa-voxceleb
Framework SpeechBrain + PyTorch

ECAPA-TDNN Architecture Flow

Raw Audio (16kHz PCM)
        ↓
  Frame-level Feature Extraction
  (Filter banks / MFCCs at each time step)
        ↓
  TDNN Layers with Dilation
  (Captures short + long-range temporal dependencies)
        ↓
  SE-Res2Block (Squeeze-and-Excitation + Residual)
  (Channel attention β€” emphasizes discriminative voice features)
        ↓
  Multi-scale Feature Aggregation (MFA)
  (Combines features from ALL TDNN layers)
        ↓
  Attentive Statistics Pooling
  (Collapses variable-length sequence β†’ fixed representation)
        ↓
  Fully Connected Layer
        ↓
  192-dim L2-normalized Vector  ← THE VOICEPRINT

Why 192 Dimensions?

The 192-dim space is specifically tuned for speaker discrimination:

  • Each dimension captures abstract acoustic properties (vocal tract shape, pitch patterns, speaking style)
  • L2 normalization means all vectors lie on a unit hypersphere β€” cosine similarity equals dot product
  • Empirically the sweet spot between expressiveness and computational cost

Cosine Similarity β†’ Confidence Mapping

# Raw cosine similarity in [-1.0, +1.0]
cosine = dot(master, test_emb) / (norm(master) * norm(test_emb))

# Display confidence: mapped to 0-100% with +20 boost for UX clarity
# cosine 0.20 β†’ 40%  |  0.45 β†’ 65%  |  0.62 β†’ 82%  |  0.80 β†’ 100%
confidence = min(100, max(0, (cosine * 100) + 20))
Raw Cosine Display % Label Match?
>= 0.62 >= 82% 🟒 Strong Match βœ… Yes
>= 0.45 >= 65% 🟑 Partial Match βœ… Yes
>= 0.25 >= 45% 🟠 Weak Match ❌ No
< 0.25 < 45% πŸ”΄ No Match ❌ No

πŸ—οΈ System Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                    AUDIO CLASSIFICATION HUB                         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚       FRONTEND           β”‚              BACKEND                     β”‚
β”‚   (HTML + JS + CSS)      β”‚         (FastAPI + Python)               β”‚
β”‚                          β”‚                                          β”‚
β”‚  index.html              β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”‚
β”‚  auth.html    ───────────┼─►│      FastAPI (Uvicorn :8000)      β”‚  β”‚
β”‚  onboarding pages        β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β”‚
β”‚  download.html           β”‚                  β”‚                       β”‚
β”‚                          β”‚       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”           β”‚
β”‚  js/api.js               β”‚       β”‚     ML PIPELINE      β”‚           β”‚
β”‚  js/recorder.js          β”‚       β”‚  1. preprocess.py    β”‚           β”‚
β”‚  js/processing.js        β”‚       β”‚  2. embedding.py     β”‚           β”‚
β”‚  js/hero-shader.js       β”‚       β”‚     (ECAPA-TDNN)     β”‚           β”‚
β”‚  css/style.css           β”‚       β”‚  3. averaging.py     β”‚           β”‚
β”‚                          β”‚       β”‚  4. injector.py      β”‚           β”‚
β”‚                          β”‚       β”‚  5. builder.py       β”‚           β”‚
└───────────────────────────       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜           β”‚
                           β”‚                  β”‚                       β”‚
                           β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
                           β”‚  β”‚         FILE SYSTEM DB              β”‚ β”‚
                           β”‚  β”‚  DataBase/login.csv  (registry)     β”‚ β”‚
                           β”‚  β”‚  DataBase/<user>/embedding.npy      β”‚ β”‚
                           β”‚  β”‚  workspaces/<user>_<id>/dist/*.whl  β”‚ β”‚
                           β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
                           β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
                           β”‚  β”‚  SMTP Email (.whl attached)          β”‚ β”‚
                           β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
                           β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

βš™οΈ Backend β€” Deep Code Analysis

The backend is built with FastAPI, chosen for its async-first design, automatic OpenAPI docs at /docs, Pydantic validation for multipart forms, and BackgroundTasks for non-blocking email delivery.

FastAPI Application (main.py)

The entry point registers 7 REST API endpoints and mounts the frontend static assets at the same paths the HTML expects (/css/, /js/), so no separate web server is needed.

# Frontend HTML + Backend API on same origin β€” no CORS complexity
app.mount("/css", StaticFiles(directory=str(css_dir)), name="css")
app.mount("/js",  StaticFiles(directory=str(js_dir)),  name="js")

Global error handling ensures all exceptions return clean JSON β€” the JS safeJson() parser always succeeds.

API Endpoints

Method Endpoint Purpose
POST /api/register Create new user account
POST /api/login Validate credentials
POST /api/process Upload voice samples β†’ run full ML pipeline β†’ build .whl
POST /api/process_company Multi-person bulk upload β†’ multi-embedding .whl
GET /api/download Stream the generated .whl file
GET /api/status Poll whether .whl build is ready
POST /api/verify Verify a voice sample against stored voiceprint
GET / and page routes Serve all frontend HTML pages

ML Pipeline β€” 5-Stage Processing

Stage 1 β€” pipeline/preprocess.py β€” Audio Cleaning

Every uploaded audio file goes through a 3-backend loading cascade:

torchaudio (soundfile β†’ sox_io)  ──► succeeds for WAV/FLAC/OGG
         ↓ fail
soundfile + resampy               ──► succeeds for WAV/FLAC
         ↓ fail
librosa (requires ffmpeg)         ──► catches MP4/WebM/M4A

After loading, audio is:

  • Resampled to exactly 16,000 Hz (ECAPA-TDNN requirement)
  • Mixed down to mono
  • Validated β€” rejected if < 2 seconds
  • Padded or trimmed to exactly 3 seconds (48,000 samples)
  • Peak-normalized to [-1.0, +1.0]
SR       = 16000   # ECAPA-TDNN requires exactly 16kHz
DURATION = 3       # seconds window used for embedding
MIN_SEC  = 2.0     # reject files shorter than this

Stage 2 β€” pipeline/embedding.py β€” Voice Vector Extraction

The ECAPA-TDNN model is loaded once into a module-level global _model (lazy singleton) to avoid reloading ~100MB on every request.

Windows Compatibility Fix: On Windows without Developer Mode, pathlib.Path.symlink_to() raises [WinError 1314]. SpeechBrain calls this during model caching. The module monkey-patches it:

# Patch BEFORE any SpeechBrain import
_orig_symlink_to = Path.symlink_to

def _safe_symlink_to(self, target, target_is_directory=False):
    try:
        _orig_symlink_to(self, target, target_is_directory)
    except OSError:
        shutil.copy2(str(target_p), str(self))  # fall back to file copy

Path.symlink_to = _safe_symlink_to

Inference:

def get_embedding(audio: np.ndarray) -> np.ndarray:
    tensor = torch.FloatTensor(audio).unsqueeze(0)   # (1, N)
    with torch.no_grad():
        emb = model.encode_batch(tensor)             # (1, 1, 192)
    return emb.squeeze().numpy().astype(np.float32)  # (192,)

Stage 3 β€” pipeline/averaging.py β€” Master Voiceprint

Multiple samples produce slightly different vectors (phrasing, noise). Averaging builds a centroid that is more robust:

def build_master(embeddings: list) -> np.ndarray:
    stacked = np.stack(embeddings, axis=0)   # (N, 192)
    master  = np.mean(stacked, axis=0)       # (192,)
    return master.astype(np.float32)

Rule of thumb: More samples = more accurate centroid. 5 is minimum; 20 is recommended maximum.

Stage 4 β€” pipeline/injector.py β€” Embedding Baking

This is what makes the .whl work without any server calls. The user's 192-float embedding is literally baked into Python source code:

# core_template.py contains:  EMBEDDING = {{EMBEDDING}}
# The injector replaces the placeholder with actual numbers:
embedding_str = repr(master.tolist())              # "[0.023, -0.114, ...]"
final_code = template_code.replace("{{EMBEDDING}}", embedding_str)

When the .whl is imported, no database, no network, no model download is needed.

Company Mode injects a dict of {person_name: [192-float-list]} instead, enabling 1:N speaker identification from a single package.

Stage 5 β€” pipeline/builder.py β€” Wheel Packaging

subprocess.run([sys.executable, "-m", "build"], cwd=str(build_dir))

Uses Python build module (PEP 517/518). Each build gets a UUID-isolated workspace so concurrent builds never interfere. Workspace lives outside the server directory to prevent uvicorn --reload from triggering on generated files.


Database Layer (database.py)

Uses a flat-file CSV approach β€” simple, portable, zero-dependency:

DataBase/
β”œβ”€β”€ login.csv                  # name, email, password, username, registered_at, whl_path
β”œβ”€β”€ mohit_jadav/
β”‚   β”œβ”€β”€ voices/                # raw uploaded audio files
β”‚   └── embedding.npy          # 192-dim master voiceprint
└── harsh_jadav/
    β”œβ”€β”€ voices/
    └── embedding.npy
  • create_user() β€” creates user row AND the DataBase/<username>/ folder atomically
  • save_whl_path() β€” updates whl_path column after a successful build
  • get_whl_path() β€” lookup used by /api/download and /api/status

Email Delivery System (login.py)

After every successful pipeline run, a non-blocking background task sends a welcome email with the .whl attached:

background_tasks.add_task(
    send_welcome_email,
    name=user["name"],
    email=email,
    whl_path=whl_path,
)  # Returns API response immediately β€” email sends asynchronously

Email failures are caught and logged but never crash the user-facing request.


🎨 Frontend β€” Page-by-Page Walkthrough

The frontend is a pure HTML + CSS + JavaScript SPA (no React, no Vue) served directly by FastAPI.

Tech Stack:

  • Tailwind CSS (CDN) for utility classes
  • Three.js for WebGL voice waveform shaders
  • Apache ECharts for processing visualization charts
  • Font Awesome 6 for iconography
  • Google Fonts β€” Space Grotesk, Inter, JetBrains Mono
  • Custom css/style.css β€” glassmorphism design system, CSS tokens, micro-animations

Page 1 β€” index.html β€” Landing Page

Element Implementation
Hero Canvas Three.js WebGL shader β€” animated chromatic waveform sinusoids
Scroll Reveal IntersectionObserver with staggered data-delay
3-Step How It Works Animated ping rings, waveform bars via @keyframes wfPulse
Code Window Multi-tab snippet (Python / FastAPI / Node / cURL)
Bento Grid 4-column masonry feature cards with radial gradient glows
Social Proof Marquee CSS infinite scroll strip
CTA Start Free β†’ auth.html / View Docs β†’ download.html

Page 2 β€” auth.html β€” Sign In / Sign Up

  • Split-screen β€” left: animated canvas + tagline, right: auth card
  • Tab toggle β€” Login / Sign Up with animated active state
  • Password strength meter β€” 4-segment bar (Too weak β†’ Strong) via real-time regex
  • Session management β€” Session.save() to sessionStorage; redirects based on whl_ready

Page 3 β€” onboarding-type.html β€” Account Type Selection

  • 4-step progress stepper with animated connectors (Step 1 active)
  • Individual card β†’ onboarding-record.html (single voiceprint)
  • Team/Company card β†’ onboarding-company.html (multi-user folder upload)
  • Card selection triggers animated bounce-in checkmark

Page 4 β€” onboarding-record.html β€” Voice Capture

Record Mode (WebRTC):

  • MediaRecorder API β€” browser microphone as audio/webm blobs
  • Live countdown timer with requestAnimationFrame
  • 3 concentric animated pulse rings during recording
  • Playback bar + Re-record + Next sample controls
  • Counter: 0 / 20 captured in real time

Upload Mode (Drag & Drop):

  • dragover / drop events with visual border feedback
  • <input type="file" multiple accept=".wav,.mp3,.flac,.ogg,.m4a">
  • Validates 5–20 files before enabling Process button

Page 5 β€” onboarding-processing.html β€” Live Pipeline Visualizer

Left Panel β€” Pipeline Steps:

  • 5 animated steps: Loading Audio β†’ Preprocessing β†’ MFCC Extraction β†’ ECapa Embedding β†’ Packaging .whl
  • CSS transform: translateX(4px) on active step
  • Real elapsed timer + gradient progress bar linear-gradient(90deg, #6366F1, #00D4FF)

Right Panel β€” ECharts Visualizations:

Chart What It Shows
MFCC Heatmap 13 Mel-Frequency Cepstral Coefficient tracks across time
Waveform + Energy Envelope Amplitude over time with energy overlay
Pitch Contour (F0) Fundamental frequency curve across the utterance
192-dim Embedding Bars All 192 embedding dimensions visualized
UMAP Scatter Plot 2D projection: user voiceprint vs. population cluster

Backend Polling (every 3 seconds, max 6 minutes):

setInterval(async () => {
    const data = await apiStatus(user.email);
    if (data.whl_ready) {
        clearInterval(pollTimer);
        forceComplete();   // snap UI to 100%
        showDownloadCTA(); // no auto-redirect β€” user decides
    }
}, 3000);

Page 6 β€” download.html β€” SDK Download & Verify

  • Download button β†’ GET /api/download?email=<email>
  • Multi-language code snippets (Python, FastAPI, Node.js, cURL)
  • Live Voice Verify playground β€” record test sample, see animated confidence result card

πŸ“¦ What the User Gets β€” The .whl Package

After completing registration, the user receives:

audioauth-1.0.0-py3-none-any.whl

Package Contents

audioauth/
β”œβ”€β”€ __init__.py
β”œβ”€β”€ core.py          ← YOUR VOICEPRINT BAKED IN (192-float Python list)
└── ...

core.py contains your embedding as a static Python constant:

EMBEDDING = [0.0234, -0.1142, 0.4471, 0.0891, -0.2310, ...]  # 192 floats

How Developers Use It

pip install audioauth-1.0.0-py3-none-any.whl
# 2 lines. That is all.
from audioauth import VoiceAuth

auth = VoiceAuth()
result = auth.verify("sample.wav")  # Returns: True / False

Advanced usage:

details = auth.verify_detailed("voice.wav")
print(details["confidence"])    # 0.0 to 100.0
print(details["label"])         # "Strong Match" / "Partial Match" / ...
print(details["matched"])       # True / False
print(details["cosine_score"])  # raw cosine similarity

# Company / multi-user mode
auth = VoiceAuth(mode="company")
speaker = auth.identify("voice.wav")   # returns matched person name

FastAPI integration:

from fastapi import FastAPI, UploadFile
from audioauth import VoiceAuth

app = FastAPI()
auth = VoiceAuth()

@app.post("/check-voice")
async def check_voice(file: UploadFile):
    contents = await file.read()
    return {"authenticated": auth.verify_bytes(contents)}

What Is NOT in the Package

  • ❌ No SpeechBrain model weights β€” embedding already extracted, model not needed
  • ❌ No network calls β€” runs 100% offline
  • ❌ No cloud dependency β€” your voiceprint, your server
  • ❌ No GPU requirement β€” pure NumPy cosine similarity at ~50ms latency

πŸ”„ Complete User Journey

                    REGISTRATION FLOW
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚   1. LANDING PAGE                                    β”‚
    β”‚      Learns about product β†’ Clicks "Start Free"      β”‚
    β”‚                    ↓                                 β”‚
    β”‚   2. AUTH PAGE                                       β”‚
    β”‚      Sign Up β†’ POST /api/register                    β”‚
    β”‚      β†’ Redirected to Type Selection                  β”‚
    β”‚                    ↓                                 β”‚
    β”‚   3. TYPE SELECTION                                   β”‚
    β”‚      Choose: Individual OR Team/Company              β”‚
    β”‚                    ↓                                 β”‚
    β”‚   4. VOICE RECORDING                                 β”‚
    β”‚      Record 5–20 samples OR drag-and-drop files      β”‚
    β”‚      β†’ Click "Process Voice"                         β”‚
    β”‚                    ↓                                 β”‚
    β”‚   5. PIPELINE (~60 seconds, async)                   β”‚
    β”‚      POST /api/process                               β”‚
    β”‚      β”œβ”€β”€ Stage 1: clean_audio() x N files            β”‚
    β”‚      β”œβ”€β”€ Stage 2: get_embedding() β†’ N x [192]        β”‚
    β”‚      β”œβ”€β”€ Stage 3: build_master() β†’ [192]             β”‚
    β”‚      β”œβ”€β”€ Stage 4: inject_embedding() β†’ core.py       β”‚
    β”‚      └── Stage 5: build_whl() β†’ .whl                 β”‚
    β”‚                    ↓                                 β”‚
    β”‚   6. PROCESSING PAGE (polls /api/status every 3s)    β”‚
    β”‚      MFCC, waveform, pitch, embedding, UMAP charts   β”‚
    β”‚      β†’ Done β†’ Download .whl CTA appears              β”‚
    β”‚                    ↓                                 β”‚
    β”‚   7. DOWNLOAD PAGE                                   β”‚
    β”‚      GET /api/download β†’ .whl streamed to browser    β”‚
    β”‚      + Welcome email with .whl attached              β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

                  DAILY AUTHENTICATION FLOW
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
    β”‚   pip install audioauth-1.0.0-py3-none-any.whl       β”‚
    β”‚   from audioauth import VoiceAuth                    β”‚
    β”‚   auth = VoiceAuth()                                 β”‚
    β”‚   result = auth.verify("voice.wav")                  β”‚
    β”‚                    ↓                                 β”‚
    β”‚   1. Load [192] from core.py (static constant)       β”‚
    β”‚   2. Load audio β†’ 16kHz numpy array                  β”‚
    β”‚   3. Compute cosine similarity                       β”‚
    β”‚   4. Return True/False + confidence                  β”‚
    β”‚                                                      β”‚
    β”‚   ~50ms latency Β· 0 network calls Β· 0 GPU            β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

🏒 Industry Use Cases β€” Where It Fits

Audio Classification Hub targets both B2C (developers, prosumers) and B2B (enterprises, SaaS).

🏦 Banking & Financial Services

  • Phone banking authentication β€” replace KBA with voice biometrics
  • IVR gating β€” authenticate callers before agent transfer
  • Anti-fraud β€” detect account takeover by verifying caller identity
  • PCI-DSS voice channel compliance

πŸ₯ Healthcare

  • Telehealth patient verification before prescription refills
  • EHR access control β€” voice-lock records for specific clinicians
  • HIPAA-compliant (voiceprint on-premise, zero cloud exposure)

🏒 Enterprise HR & Access Control

  • Remote work time-tracking β€” employees clock in by voice
  • Secure facility access β€” voice-gated physical security
  • Executive document DLP enforcement by speaker verification

πŸ“ž Call Centers & Customer Experience

  • Passive agent-side verification as call connects
  • VIP customer detection β†’ priority queue routing
  • Callback authentication before resuming sensitive sessions

πŸŽ“ EdTech & Proctoring

  • Online exam authentication throughout the session
  • Voice-based lecture attendance check-in
  • Adaptive learning gates β€” unlock advanced content for verified learners

πŸ” Consumer Apps (B2C)

  • Password manager voice unlock β€” biometric second factor
  • Smart home voice commands β€” prevent spoofing by non-members
  • Parental controls β€” voice-gated content restrictions

πŸ—οΈ Developer / Platform (API Economy)

  • Ship the .whl inside your product installer
  • Wrap /api/verify as a microservice sidecar for any service
  • IoT / Edge devices β€” offline voice auth for Raspberry Pi, embedded systems

πŸ“‘ API Reference

POST /api/register

Request (multipart/form-data): name, email, password (min 6 chars)

{ "status": "ok", "message": "Account created", "username": "mohit_jadav" }

POST /api/login

Request: email, password

{ "status": "ok", "name": "Mohit Jadav", "email": "...", "username": "mohit_jadav", "whl_ready": false }

POST /api/process

Request: email, files[] β€” 5 to 20 audio files (.wav .mp3 .flac .ogg .m4a)

{
  "status": "ok",
  "message": "Voice embedding created and .whl built successfully",
  "build_id": "a3f91c2b",
  "accepted_files": 8,
  "skipped_files": 0
}

POST /api/process_company

Bulk multi-person enrollment. Files structured as PersonName/filename.wav (5–20 per person).

GET /api/status?email=<email>

{ "status": "ok", "whl_ready": true, "whl_filename": "audioauth-1.0.0-py3-none-any.whl" }

GET /api/download?email=<email>

Response: application/octet-stream β€” the .whl binary.

POST /api/verify

Request: email, file (single audio)

{
  "status": "ok",
  "confidence": 78.4,
  "cosine_score": 0.5840,
  "label": "Partial Match",
  "matched": true,
  "color": "amber",
  "icon": "fa-circle-half-stroke"
}

πŸ“ Project Structure

Full_Working/
β”‚
β”œβ”€β”€ πŸ“‚ Backend/
β”‚   └── πŸ“‚ your_server/
β”‚       β”œβ”€β”€ main.py               β€” FastAPI app, all routes, static serving
β”‚       β”œβ”€β”€ database.py           β€” CSV read/write, user management
β”‚       β”œβ”€β”€ login.py              β€” auth validation, welcome email via SMTP
β”‚       β”œβ”€β”€ requirements.txt
β”‚       β”œβ”€β”€ .env.example          β€” SMTP config template
β”‚       β”œβ”€β”€ πŸ“‚ pipeline/
β”‚       β”‚   β”œβ”€β”€ preprocess.py     β€” audio loading & cleaning (3-backend cascade)
β”‚       β”‚   β”œβ”€β”€ embedding.py      β€” ECAPA-TDNN inference + HuggingFace model load
β”‚       β”‚   β”œβ”€β”€ averaging.py      β€” N embeddings β†’ 1 master voiceprint
β”‚       β”‚   β”œβ”€β”€ injector.py       β€” bake embedding into Python source code
β”‚       β”‚   └── builder.py        β€” python -m build β†’ .whl package
β”‚       β”œβ”€β”€ πŸ“‚ pretrained_models/
β”‚       β”‚   └── spkrec-ecapa-voxceleb/    β€” ECAPA-TDNN weights (auto-downloaded)
β”‚       └── πŸ“‚ template/          β€” .whl template with {{EMBEDDING}} placeholder
β”‚
β”œβ”€β”€ πŸ“‚ Fronted/
β”‚   β”œβ”€β”€ index.html                β€” Landing page
β”‚   β”œβ”€β”€ auth.html                 β€” Login / Sign Up
β”‚   β”œβ”€β”€ onboarding-type.html      β€” Individual vs. Company selector
β”‚   β”œβ”€β”€ onboarding-record.html    β€” Voice recorder + file upload
β”‚   β”œβ”€β”€ onboarding-company.html   β€” Bulk multi-person upload
β”‚   β”œβ”€β”€ onboarding-processing.html β€” Pipeline visualizer + polling
β”‚   β”œβ”€β”€ download.html             β€” SDK download + verify playground
β”‚   β”œβ”€β”€ πŸ“‚ css/style.css          β€” Full design system (glassmorphism, tokens, animations)
β”‚   └── πŸ“‚ js/
β”‚       β”œβ”€β”€ api.js                β€” HTTP client + Session management
β”‚       β”œβ”€β”€ main.js               β€” Navbar, footer, scroll-reveal
β”‚       β”œβ”€β”€ recorder.js           β€” MediaRecorder API, WebRTC capture
β”‚       β”œβ”€β”€ processing.js         β€” ECharts: MFCC, waveform, UMAP, embedding
β”‚       └── hero-shader.js        β€” Three.js WebGL waveform shader
β”‚
β”œβ”€β”€ πŸ“‚ DataBase/
β”‚   β”œβ”€β”€ login.csv                 β€” User registry (auto-created)
β”‚   └── πŸ“‚ <username>/
β”‚       β”œβ”€β”€ πŸ“‚ voices/            β€” Raw uploaded audio samples
β”‚       └── embedding.npy         β€” 192-dim master voiceprint
β”‚
β”œβ”€β”€ πŸ“‚ workspaces/
β”‚   └── πŸ“‚ <username>_<build_id>/
β”‚       └── πŸ“‚ build/dist/
β”‚           └── audioauth-*.whl
β”‚
└── πŸ“‚ Screenshots/
    └── Screenshot 2026-08-01 190641.png

πŸš€ Local Setup & Installation

Prerequisites

Requirement Version
Python 3.10+
pip latest
ffmpeg optional (for .webm/.mp4 support)
SMTP access Gmail App Password recommended

Step 1 β€” Clone

git clone https://github.com/mohitjadav/audio-classification-hub.git
cd audio-classification-hub/Full_Working

Step 2 β€” Backend Setup

cd Backend/your_server
python -m venv venv
venv\Scripts\activate          # Windows
# source venv/bin/activate     # macOS / Linux
pip install -r requirements.txt

Step 3 β€” Configure Environment

cp .env.example .env
SMTP_HOST=smtp.gmail.com
SMTP_PORT=587
SMTP_USER=your_email@gmail.com
SMTP_PASSWORD=your_gmail_app_password

Gmail App Password: Google Account β†’ Security β†’ 2-Step Verification β†’ App Passwords β†’ Generate

Step 4 β€” Run

uvicorn main:app --reload --host 0.0.0.0 --port 8000

First startup downloads the ECAPA-TDNN model (~100MB):

[EMBEDDING] Path.symlink_to patched β†’ fallback to copy on Windows.
[EMBEDDING] Downloading model to: pretrained_models/spkrec-ecapa-voxceleb
[EMBEDDING] Download complete.
[EMBEDDING] ECAPA-TDNN loaded successfully.
INFO:     Uvicorn running on http://0.0.0.0:8000

Subsequent starts use the cached model:

[EMBEDDING] Using cached model at: pretrained_models/spkrec-ecapa-voxceleb

Step 5 β€” Open

Navigate to http://localhost:8000


βš™οΈ Configuration

Variable Default Description
SMTP_HOST smtp.gmail.com SMTP server hostname
SMTP_PORT 587 SMTP port (TLS)
SMTP_USER β€” Sender email address
SMTP_PASSWORD β€” Gmail App Password

Audio constants (pipeline/preprocess.py):

Constant Default Description
SR 16000 Target sample rate (Hz)
DURATION 3 Seconds of audio per embedding
MIN_SEC 2.0 Minimum audio length accepted

Confidence thresholds (main.py β€” /api/verify):

Label Threshold matched
Strong Match >= 82% True
Partial Match >= 65% True
Weak Match >= 45% False
No Match < 45% False

πŸ”’ Security Considerations

Area Current State Production Recommendation
Password Storage Plain text in CSV Hash with bcrypt or argon2-cffi
Session sessionStorage (browser) JWT tokens with python-jose
Database CSV flat file PostgreSQL with SQLAlchemy
CORS allow_origins=["*"] Restrict to your domain
File Validation Extension + size checks Also validate magic bytes
SMTP Credentials in .env Managed email (SendGrid, Postmark)
Rate Limiting None Add slowapi for /api/process
HTTPS None (local dev) Nginx reverse proxy + Let's Encrypt

πŸ—ΊοΈ Roadmap

  • PostgreSQL + SQLAlchemy β€” replace CSV database
  • JWT Authentication β€” stateless session tokens
  • Password hashing β€” bcrypt
  • Multi-tenancy β€” org-level API keys for B2B
  • Anti-spoofing β€” liveness detection (replay attack prevention)
  • Noise augmentation β€” improve robustness during registration
  • Dashboard β€” admin panel with user management, verification logs
  • Docker β€” Dockerfile + docker-compose.yml for one-command deployment
  • HTTPS / Nginx β€” production deployment guide
  • Model quantization β€” INT8 ECAPA-TDNN for faster CPU inference
  • Webhook support β€” POST verification result to developer callback URL

πŸ‘¨β€πŸ’» Author


Mohit Jadav

Full-Stack AI Engineer Β Β·Β  Deep Learning Β Β·Β  FastAPI Β Β·Β  Voice Biometrics


"I don't just build models β€” I build systems that put intelligence in developers' hands without cloud dependency."


GitHub LinkedIn


Core Stack: Python Β Β·Β  FastAPI Β Β·Β  PyTorch Β Β·Β  SpeechBrain Β Β·Β  NumPy Β Β·Β  HTML/CSS/JS Β Β·Β  TailwindCSS Β Β·Β  Three.js

Domains: Machine Learning Β Β·Β  Deep Learning Β Β·Β  Audio AI Β Β·Β  Web Development Β Β·Β  Voice Biometrics Β Β·Β  API Design



Built with passion for making AI accessible.

Audio Classification Hub Β β€”Β  Β© 2026 Mohit Jadav. MIT License.

About

Production-ready AI voice biometric authentication platform using FastAPI, SpeechBrain ECAPA-TDNN, and Tensorflow. Generate offline Python SDKs (.whl) for secure speaker verification with zero cloud dependency.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages