Skip to content
View aokassamali's full-sized avatar

Block or report aokassamali

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
aokassamali/README.md

Asad Kassamali

projects

audio-search — Speaker-attributed search and Q&A over conversational audio: faster-whisper → pyannote diarization → hybrid retrieval → grounded answers with timestamp citations, orchestrated in Dagster. Companion experiments asked whether audio adds speech-act signal beyond text

coding-agent-evals — Do 3B–7B code models entrench on their first solution strategy? Kaplan-Meier survival analysis over 910 approach decisions

Evals-and-LLM-as-Judge — How reliable is LLM-as-judge? Five models, 200 claim-verification items

RAG — Retrieval → evidence selection → grounded answering, optimized for measurable reliability. A cite-or-abstain tiered policy took correctness 0.52 → 0.73 and groundedness 0.72 → 0.99 while raising coverage.

Hillstrom-emails-experiment — Pre-analysis plan, health checks, multiple-comparisons control, decision memo.

Prop99-SDID — Synthetic Control vs. Synthetic DiD with a full placebo and robustness suite.

Also here: Prop 47 synthetic control (near-null, reported as such) · M5 forecasting · reciprocal ranking · RL for TFT (paused)

Currently: writing up the audio-search experiments (leakage audit, label perturbation, prosody placebo) · next: deep RL

Pinned Loading

  1. coding-agent-evals coding-agent-evals Public

    Do 3B–7B code models entrench on their first solution strategy? An eval harness and Kaplan-Meier survival analysis over 910 approach decisions.

    Python

  2. audio-search audio-search Public

    Speaker-attributed search and Q&A over conversational audio: faster-whisper, pyannote diarization, hybrid retrieval, and grounded LLM answers with timestamp citations.

    Python

  3. Prop99-SDID-Causal-and-Robustness-Suite Prop99-SDID-Causal-and-Robustness-Suite Public

    Synthetic Control vs. Synthetic DiD on Prop 99, with a full placebo and robustness suite reporting sensitivity ranges rather than point estimates.

    Python

  4. RAG RAG Public

    Retrieval → evidence selection → grounded answering, optimized for measurable reliability: hybrid retrieval, cross-encoder reranking, and a cite-or-abstain tiered policy.

    Python

  5. Evals-and-LLM-as-Judge Evals-and-LLM-as-Judge Public

    How reliable is LLM-as-judge? Inter-model agreement and failure modes across 5 local and frontier models on 200 claim-verification items.

    Python

  6. Hillstrom-emails-experiment Hillstrom-emails-experiment Public

    End-to-end A/B workflow on a 3-arm RCT: pre-analysis plan, health checks, multiple-comparisons control, decision memo, and an uplift extension that honestly underperforms treat-all.

    Python