A local Retrieval-Augmented Generation (RAG) API that answers questions using information from a personal profile document — instead of letting an LLM guess. Built with FastAPI, ChromaDB, and Ollama, and extended into a multi-user AI directory with per-user data isolation.
When a question is submitted to /ask, the API:
- Retrieves the most relevant chunks from a vector store using semantic search
- Augments the prompt with that retrieved context
- Generates a grounded answer using a local LLM
The same pipeline is extended with a POST /documents endpoint so multiple
users can each submit their own profile — and /ask can be scoped to a
single user via metadata filtering, so one person's questions never surface
another person's data.
| Tool | Role |
|---|---|
| FastAPI | Serves the /ask and /documents endpoints, tested via Swagger UI |
| ChromaDB | Vector store — stores profile chunks as embeddings and runs semantic search |
| Ollama | Runs two local models: nomic-embed-text (embeddings) and qwen2.5:0.5b (generation) |
rag-api-fastapi/
├── app/
│ ├── main.py # FastAPI app: /ask and /documents endpoints
│ ├── build_knowledge_base.py # Chunk profile.txt, embed, and store in ChromaDB
│ └── profile.txt # Source knowledge document
├── docs/
│ ├── architecture/
│ │ └── rag_api_architecture.png
│ ├── governance/ # Project Charter, RAID Log, RACI, Quality Gate, UAT
│ └── screenshots/ # Evidence from the actual run
├── requirements.txt
└── README.md
- Manual RAG walkthrough — before writing any API code, retrieval and
generation were tested independently to confirm the pipeline logic
(
docs/screenshots/01_manual_rag_demo.png). - Build the knowledge base —
profile.txtis chunked, embedded withnomic-embed-text, and stored in a ChromaDB collection (docs/screenshots/02_knowledge_base_build.png). /askendpoint — retrieves the most relevant chunks for a question, augments the prompt, and generates an answer withqwen2.5:0.5b(docs/screenshots/03_ask_endpoint_swagger.png)./documentsendpoint — stores a new user's profile with auser_nametag in the vector metadata (docs/screenshots/04_documents_endpoint_multiuser.png).- Multi-user filtering —
/ask?user=Xfilters ChromaDB retrieval byuser_name, verified by confirming a second user's query only ever returns that user's data (docs/screenshots/05_multiuser_filter_verification.png).
# 1. Install Ollama and pull the models
ollama pull nomic-embed-text
ollama pull qwen2.5:0.5b
# 2. Set up the Python environment
python -m venv venv
source venv/bin/activate # or venv\Scripts\activate on Windows
pip install -r requirements.txt
# 3. Build the knowledge base
cd app
python build_knowledge_base.py
# 4. Start the API
uvicorn main:app --reload
# 5. Test it
# Open http://127.0.0.1:8000/docs for Swagger UI, or:
curl "http://127.0.0.1:8000/ask?question=What%20is%20my%20name%3F"GET /ask?question=What is my name?
{
"question": "What is my name?",
"answer": "Your name is Jaswant Singh.",
"context_used": [
"My name is Jaswant Singh.",
"I'm currently learning about cloud computing, AI, and DevOps.",
"For fun, I enjoy sci-fi movies, working out, and eating good food."
]
}
POST /documents
{
"user_name": "Jordan",
"content": "My name is Jordan. I work as a data analyst..."
}
GET /ask?question=What are their hobbies?&user=Jordan
{
"question": "What are their hobbies?",
"answer": "Their hobbies are rock climbing and playing chess.",
"context_used": [...],
"filtered_by_user": "Jordan"
}
This project follows the same governance pattern as the rest of this portfolio — a lightweight but honest paper trail, not a formality:
| Document | What it covers |
|---|---|
| Project Charter | Objectives, scope, stakeholders, success criteria |
| RAID Log | Risks, assumptions, issues, dependencies — including known gaps like no auth on the API |
| RACI Matrix | Role responsibilities across the delivery |
| Quality Gate Scorecard | Honest pass/fail assessment — including a "Fail" on authentication, called out rather than hidden |
| UAT Checklist | Test cases run against the actual API, with one flagged as not yet covered |
Documented in full in the RAID log and quality gate scorecard:
- No authentication on
/documentsor/ask— anyone with network access can write or read data. Top priority before any real deployment. qwen2.5:0.5bis a small model chosen for local speed; answer fluency is a known trade-off. Swappable for a larger local or hosted model.- Line-based chunking is naive and doesn't handle long paragraphs or sentence boundaries — fine for a short demo profile, not production-grade.
- Unknown-user queries (
userparam with no matching data) aren't formally tested yet — carried forward as an open UAT item.
Built as a hands-on project following NextWork's "Build a RAG API with FastAPI" curriculum, then extended with governance documentation and the multi-user directory as portfolio additions.
