M.S. Computer Science (AI track) and four years of professional engineering. I partner with enterprise clients to evaluate, deploy and monitor AI systems running in their own infrastructure, and make their AI quality measurable — continuous evaluation loops on task-specific golden datasets, so a quality regression fails a build instead of reaching users.
I specialize in the evaluation, reliability, and deployment of specific AI systems in production. Instead of relying on generic benchmarks, I partner with clients to build task-specific Golden Datasets and set up continuous evaluation loops (CI/CD for LLMs) for their RAG and agentic workflows. I design rubric-based grading pipelines that combine deterministic code-checks with "LLM-as-a-Judge" to measure Relevance, Faithfulness, Correctness, and Coherence. I ship systems with backend rigor: typed APIs, automated tests, eval telemetry, and secure deployment on AWS, Azure, and client-owned on-prem infrastructure.
- Currently: Forward Deployed AI Engineer @ Axitem Software Solution Inc. — Evaluating AI systems, building CI/CD loops for LLMs, and turning vague quality complaints into tracked, reproducible defects.
- Client impact: Accelerated time-to-value · Cut hallucinations via strict evaluation · 40% faster ticket triage · 60% faster queries
- Focused on: rubric-based grading, hybrid retrieval, rank fusion, and continuous eval loops that fail a build when quality regresses
- Reach me: chiranjeevigundu1@gmail.com
Featured: hybrid-rag
Hybrid document retrieval — dense embeddings plus Postgres full-text search, fused by Reciprocal Rank Fusion, with an eval harness that gates quality in CI. One core, three faces: Python library, HTTP API, and an MCP server. MIT, no API key required, runs fully local.
The point of the repo is the measurement. On a committed 22-case golden set:
| mode | MRR | paraphrase | exact identifiers |
|---|---|---|---|
| dense only | 0.898 | 0.917 | 0.917 |
| lexical only | 0.273 | 0.000 | 0.833 |
| hybrid (RRF) | 0.920 | 0.917 | 1.000 |
| hybrid + cross-encoder rerank | 0.888 | 0.856 | 1.000 |
Two things that table is meant to show. The arms have genuinely opposite blind spots — lexical scores 0.000 on every paraphrase query while beating nothing on identifiers, which is the measured premise for fusing them rather than picking one. And the cross-encoder reranker lost: −0.032 MRR overall, −0.061 on paraphrase. It ships disabled by default, with the code path and the measurement both kept, because "we tried it and it did not help on this corpus" is a result.
- Dense arm:
BAAI/bge-base-en-v1.5(768d) via ONNX Runtime, with the query/passage asymmetry bge is trained for - Lexical arm: Postgres
ts_rank_cdover a generatedtsvector— cover density, deliberately not BM25 (no IDF term) - Fusion: Reciprocal Rank Fusion, k=60, ranks only
- Metrics: MRR, recall@k, nDCG@10 — positional, no LLM judge
- Structure-aware chunking keeps tables and code fences atomic
Reproduce the table with python -m ragkit eval --compare. Corpus is small and deliberately committed — 4 documents, 29 chunks — so the numbers are honest about their scope rather than unfalsifiable.
aadyon-assist — Self-hosted AI life-ops platform (MIT). Single-agent tool-calling loop with an explicit step cap and an append-only message log; any action with a real-world side effect becomes a proposal a human approves. Multi-tenant isolation via Postgres Row-Level Security enforced at the database. Six Docker Compose services, 174 PyTest cases, GitHub Actions CI.
llmkit — Apache-2.0. The LLM plumbing aadyon-assist and synapse-storage-system both needed, extracted so the copies stop drifting: tier-based routing over LiteLLM, one traced chat() chokepoint, parsing helpers. Six modules, no mandatory heavy dependency — the core imports with zero optional extras, and CI asserts that.
floci-lab — Deployment platform for the above, in both directions: a local AWS-emulator loop, and AWS CDK that runs two services on ECS Fargate behind one load balancer with a shared RDS Postgres 16 instance. The interesting constraint was cost — zero NAT gateways by design, tasks in public subnets reachable only from the load balancer's security group.
synapse-storage-system — Vision-model document triage for a NAS archive. Classification records how it decided — model, heuristic, or mock — so a filename-based fallback is never counted as an AI classification.
AI / LLM — RAG · Hybrid Retrieval · Rank Fusion (RRF) · Cross-Encoder Reranking · Embeddings · Semantic Search · LLM Evaluation & Eval Sets · MRR / recall@k / nDCG · Grounding & Citation · Tool-Calling Agents · Prompt Engineering · OpenAI API · Anthropic Claude · LangChain · LangGraph · LlamaIndex · Hugging Face · Ollama · LiteLLM · MCP · ONNX Runtime · Guardrails AI · LoRA/QLoRA Fundamentals
Backend & Data — Python · FastAPI · Pydantic · REST APIs · PostgreSQL 16 · pgvector · Postgres Full-Text Search · Row-Level Security · Redis · Pinecone · SQL Profiling (EXPLAIN/ANALYZE) · Indexing & Query Optimization · ETL & Ingestion · JWT Auth · Node.js · React
Cloud & DevOps — AWS (ECS Fargate, RDS, S3, ECR, Secrets Manager, CloudWatch, IAM Identity Center, VPC) · AWS CDK (TypeScript) · Infrastructure as Code · Azure · Docker · GitHub Actions · CI/CD · PyTest · yoyo migrations
Languages — Python · Java · SQL · JavaScript · TypeScript · Bash · PowerShell
M.S. Computer Science (AI Track), Saint Louis University · B.Tech ECE, JNTU Kakinada


