Building production-shaped ML, LLM, and multi-agent security systems β verified against real, live infrastructure, not left as design docs.
Every project below is deployed against real infrastructure and documents the actual production incidents hit along the way β root cause and fix, not a cleaned-up version of events. See each repo's docs/incidents.md for the full, honest account.
Core stack:
The most current, differentiated work β agentic systems with real defenses, and a dedicated project that attacks those defenses to prove they hold.
Orchestrates specialist agents that call three other projects in this portfolio as real MCP tools, gated by a pre-dispatch authorization layer that decides whether to even attempt a request β not just filtering results after the fact. Includes a live, Claude-verified prompt-injection guardrail and agent-reasoning layer.
Three-tier adversarial testing platform built on Microsoft's real PyRIT framework β OWASP-taxonomy testing, obfuscation attacks, the published Crescendo multi-turn technique, and agent-manipulation testing against a real planner. Found a genuine planner-level vulnerability; confirmed the target's defense-in-depth design fully contained it.
Cell-level security enforcement (Apache Accumulo) proven against a real Spark-fused threat intelligence pipeline. A live, non-simulated proof that a restricted analyst account sees exactly the data it's cleared for β nothing more.
Core language-model engineering: retrieval quality, evaluation, and reproducible fine-tuning.
Hybrid BM25 + vector retrieval with cross-encoder reranking, deployed live. Extended with per-document security classification (U/S/TS) to support the Multi-Agent Platform's Threat Intel Agent, with the full clearance boundary proven against live, ingested data.
LoRA fine-tuning with dataset versioning, an F1-gated model registry, and verified ONNX/PyTorch parity. A real experiment shows dataset size alone doesn't improve model quality β the reason this pipeline gates on F1, not accuracy.
The foundation: deployment, canary rollouts, drift monitoring, and credential-free multi-cloud CI/CD.
JWT-secured FastAPI service serving a fine-tuned Hugging Face model, with a live canary rollout (10%β50%β100%) gated on real-time Application Insights telemetry and automatic rollback.
Credential-free CI/CD (OIDC) deploying one container to Azure Container Apps, Hugging Face Spaces, and an Azure ML managed endpoint via Bicep β zero stored secrets.
Population Stability Index drift detection and a Streamlit dashboard monitoring the classifier deployed above, auto-triggering retraining via cross-repo repository_dispatch when drift crosses threshold.
The foundational projects show what actually determines whether an ML system is trustworthy in production: infrastructure that deploys without a single stored credential, a fine-tuning pipeline that can reject its own output, a service with a real rollback mechanism, monitoring that closes the loop back into retraining, retrieval that can prove what it will and won't surface, and a data platform that enforces classification at the cell level against a real distributed cluster.
The security projects ask a harder question: what happens when an AI agent is the one calling all of this β and can its defenses actually withstand attack? The Multi-Agent Security Platform is built with authorization checked before dispatch, not after. The Red-Teaming Platform then attacks it for real, using the same tooling (PyRIT, Microsoft's published Crescendo technique) the AI security industry uses in 2026 β and found a genuine vulnerability, which the layered defense contained anyway.
Roughly 100 documented incidents across all eight repos, several of which are the projects catching their own mistakes β false positives, flawed test controls, overclaimed results β before trusting them. That self-correcting discipline is the actual point, more than any individual technology choice.
