Interpretability and efficient reasoning in language models. Previously an ML Engineering Intern at Shopify, building and evaluating support-triage models.
Computer Science, University of Waterloo · website · linkedin
| The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes Sole author |
In the studied binary tasks, compliant truth and action labels coincide, leaving the probes unidentified. Complementary evaluation labels give AUROC(action) = 1 − AUROC(truth). Mixed-context fitting reaches 1.000 versus 0.006 conventional AUROC on Gemma-9B, averaged over three seeds. code · arXiv |
| The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning First author · COLM 2026, Efficient Reasoning workshop · Spotlight |
A causal halt direction at layer 18, moved into the weights. Hook-free per-problem self-halt: −24% thinking at held accuracy, five unseen benchmarks, 24 problems, no RL. code · arXiv |
| Deep Learning Model for Invasive Ductal Carcinoma Detection Sole author · IEEE CCECE 2025 |
Deep model for invasive ductal carcinoma detection. repo · IEEE Xplore |
| Human Action Detection using FMCW mmWave Radar First author · CVIS 2024 |
Action recognition off raw radar returns. JCVIS |
| Off-Axis Drift: Internalizing a Halt Direction Needs More Than Its Scalar Projection First author · with Tinuade Adeleke |
Evaluated activation-target training for hook-free early stopping across 1.5B–14B reasoning models, measuring compression, accuracy costs, and cross-domain transfer. Completed a 376-job development comparison of 46 candidate recipes, selecting full-vector and on-axis reconstruction pairs at three shortening targets; held-out validation is pending. |
| RL training dynamics | Dense-checkpoint probing across three historical GRPO seeds under a length penalty: in two of three seeds a first-answer-token concentration readout ended at 0.81–0.85 while held-out accuracy fell 43–45%. A pre-registered four-arm factorial testing how GRPO's std and length normalizations jointly change the penalty's accuracy cost has its baseline arm bit-identical to TRL across 168 cases and is not yet run; the pipeline is calibrated on cloud H200s with a 50-step baseline run. |
| LOB-Engine | A small C++20 matching-engine prototype exploring price-time priority, SPSC queues, and pooled order storage. |
| RAG Tradeoffs | Benchmarks retrieval accuracy and latency across context lengths, chunk sizes, and 10+ LLMs. |
| Firefighter Robot | Autonomous maze-solving flame extinguisher. Set course records on two mazes, with recorded demonstrations. |
| Mr. Nutz | Poker robotics combining card perception, poker logic, and integrated hardware. |


