I build agentic LLM systems: multi-agent orchestration, retrieval (RAG / GraphRAG), and self-hosted inference. Day job: production LangGraph agents served on vLLM and Kubernetes. Based in Kuala Lumpur, Malaysia — open to relocation: Singapore · London · Dublin.
Production stack: LangGraph · vLLM / Ollama · hybrid retrieval (Neo4j, pgvector, ChromaDB, BM25 + FlashRank reranking) · MCP servers · FastAPI · Docker / Kubernetes.
- osint-prober — a LangGraph agent swarm (Planner → Gatherer → Briefing) that turns a vague natural-language prompt into a knowledge graph and an intelligence brief, fully local via Ollama. The novel part: a deterministic ModernBERT NLI cross-encoder acts as a strict entailment gate on every gathered claim, replacing LLM-as-a-judge — trade-offs in ARCHITECTURE.md.
- telegram-retriever — human-in-the-loop retrieval for LangChain: when retrieval confidence is low, the agent escalates to a human over Telegram and folds the reply back into the chain — HITL as a first-class retriever step, not an afterthought. Packaged and published on PyPI.
- sql-butler — agentic middleware that turns a sandboxed SQL database into a self-optimizing data store: ingest and query agents design schema, generate SQL, and build full-text search; a background agent mines the query log to add missing indexes. The contrarian part: retrieval with no vector database — plain SQL, every agent query logged and auditable.
Going deeper into inference engineering: SGLang's RadixAttention against vLLM's PagedAttention for prefix reuse across agent loops, and the Triton layer underneath — kernel fusion (RMSNorm, SwiGLU) and why decode is memory-bandwidth bound rather than compute bound.