Asad Kassamali
projects
audio-search — Speaker-attributed search and Q&A over conversational audio: faster-whisper → pyannote diarization → hybrid retrieval → grounded answers with timestamp citations, orchestrated in Dagster. Companion experiments asked whether audio adds speech-act signal beyond text
coding-agent-evals — Do 3B–7B code models entrench on their first solution strategy? Kaplan-Meier survival analysis over 910 approach decisions
Evals-and-LLM-as-Judge — How reliable is LLM-as-judge? Five models, 200 claim-verification items
RAG — Retrieval → evidence selection → grounded answering, optimized for measurable reliability. A cite-or-abstain tiered policy took correctness 0.52 → 0.73 and groundedness 0.72 → 0.99 while raising coverage.
Hillstrom-emails-experiment — Pre-analysis plan, health checks, multiple-comparisons control, decision memo.
Prop99-SDID — Synthetic Control vs. Synthetic DiD with a full placebo and robustness suite.
Also here: Prop 47 synthetic control (near-null, reported as such) · M5 forecasting · reciprocal ranking · RL for TFT (paused)
Currently: writing up the audio-search experiments (leakage audit, label perturbation, prosody placebo) · next: deep RL