Skip to content
View lauchunchee's full-sized avatar
  • Kuala Lumpur, Malaysia

Block or report lauchunchee

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
lauchunchee/README.md

Leon Lau — Senior AI Engineer

I build agentic LLM systems: multi-agent orchestration, retrieval (RAG / GraphRAG), and self-hosted inference. Day job: production LangGraph agents served on vLLM and Kubernetes. Based in Kuala Lumpur, Malaysia — open to relocation: Singapore · London · Dublin.

Production stack: LangGraph · vLLM / Ollama · hybrid retrieval (Neo4j, pgvector, ChromaDB, BM25 + FlashRank reranking) · MCP servers · FastAPI · Docker / Kubernetes.

Selected work

  • osint-prober — a LangGraph agent swarm (Planner → Gatherer → Briefing) that turns a vague natural-language prompt into a knowledge graph and an intelligence brief, fully local via Ollama. The novel part: a deterministic ModernBERT NLI cross-encoder acts as a strict entailment gate on every gathered claim, replacing LLM-as-a-judge — trade-offs in ARCHITECTURE.md.
  • telegram-retriever — human-in-the-loop retrieval for LangChain: when retrieval confidence is low, the agent escalates to a human over Telegram and folds the reply back into the chain — HITL as a first-class retriever step, not an afterthought. Packaged and published on PyPI.
  • sql-butler — agentic middleware that turns a sandboxed SQL database into a self-optimizing data store: ingest and query agents design schema, generate SQL, and build full-text search; a background agent mines the query log to add missing indexes. The contrarian part: retrieval with no vector database — plain SQL, every agent query logged and auditable.

Now

Going deeper into inference engineering: SGLang's RadixAttention against vLLM's PagedAttention for prefix reuse across agent loops, and the Triton layer underneath — kernel fusion (RMSNorm, SwiGLU) and why decode is memory-bandwidth bound rather than compute bound.

Contact

lauchunchee@gmail.com · LinkedIn

Pinned Loading

  1. llm-rss-digest llm-rss-digest Public

    Autonomous LangGraph agent that monitors RSS feeds and writes executive briefings with local LLMs

    Python

  2. osint-prober osint-prober Public

    Autonomous OSINT agent swarm — LangGraph ReAct agents build a knowledge graph and intelligence brief from a name, fully local via Ollama

    Python

  3. policy-qa-bot policy-qa-bot Public

    RAG assistant for insurance policies with clause-level citations — Docling ingestion, FAISS retrieval

    Python

  4. sql-butler sql-butler Public

    Agentic middleware that turns a sandboxed SQL database into a self-optimizing data store — no vector DB, plain SQL

    Python