Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

RAG API with FastAPI

A local Retrieval-Augmented Generation (RAG) API that answers questions using information from a personal profile document — instead of letting an LLM guess. Built with FastAPI, ChromaDB, and Ollama, and extended into a multi-user AI directory with per-user data isolation.

RAG API architecture

What this is

When a question is submitted to /ask, the API:

  1. Retrieves the most relevant chunks from a vector store using semantic search
  2. Augments the prompt with that retrieved context
  3. Generates a grounded answer using a local LLM

The same pipeline is extended with a POST /documents endpoint so multiple users can each submit their own profile — and /ask can be scoped to a single user via metadata filtering, so one person's questions never surface another person's data.

Tech stack

Tool Role
FastAPI Serves the /ask and /documents endpoints, tested via Swagger UI
ChromaDB Vector store — stores profile chunks as embeddings and runs semantic search
Ollama Runs two local models: nomic-embed-text (embeddings) and qwen2.5:0.5b (generation)

Project structure

rag-api-fastapi/
├── app/
│   ├── main.py                    # FastAPI app: /ask and /documents endpoints
│   ├── build_knowledge_base.py    # Chunk profile.txt, embed, and store in ChromaDB
│   └── profile.txt                # Source knowledge document
├── docs/
│   ├── architecture/
│   │   └── rag_api_architecture.png
│   ├── governance/                # Project Charter, RAID Log, RACI, Quality Gate, UAT
│   └── screenshots/                # Evidence from the actual run
├── requirements.txt
└── README.md

How it works

  1. Manual RAG walkthrough — before writing any API code, retrieval and generation were tested independently to confirm the pipeline logic (docs/screenshots/01_manual_rag_demo.png).
  2. Build the knowledge base — profile.txt is chunked, embedded with nomic-embed-text, and stored in a ChromaDB collection (docs/screenshots/02_knowledge_base_build.png).
  3. /ask endpoint — retrieves the most relevant chunks for a question, augments the prompt, and generates an answer with qwen2.5:0.5b (docs/screenshots/03_ask_endpoint_swagger.png).
  4. /documents endpoint — stores a new user's profile with a user_name tag in the vector metadata (docs/screenshots/04_documents_endpoint_multiuser.png).
  5. Multi-user filtering — /ask?user=X filters ChromaDB retrieval by user_name, verified by confirming a second user's query only ever returns that user's data (docs/screenshots/05_multiuser_filter_verification.png).

Running it locally

# 1. Install Ollama and pull the models
ollama pull nomic-embed-text
ollama pull qwen2.5:0.5b

# 2. Set up the Python environment
python -m venv venv
source venv/bin/activate   # or venv\Scripts\activate on Windows
pip install -r requirements.txt

# 3. Build the knowledge base
cd app
python build_knowledge_base.py

# 4. Start the API
uvicorn main:app --reload

# 5. Test it
# Open http://127.0.0.1:8000/docs for Swagger UI, or:
curl "http://127.0.0.1:8000/ask?question=What%20is%20my%20name%3F"

Example request/response

GET /ask?question=What is my name?

{
  "question": "What is my name?",
  "answer": "Your name is Jaswant Singh.",
  "context_used": [
    "My name is Jaswant Singh.",
    "I'm currently learning about cloud computing, AI, and DevOps.",
    "For fun, I enjoy sci-fi movies, working out, and eating good food."
  ]
}
POST /documents
{
  "user_name": "Jordan",
  "content": "My name is Jordan. I work as a data analyst..."
}

GET /ask?question=What are their hobbies?&user=Jordan

{
  "question": "What are their hobbies?",
  "answer": "Their hobbies are rock climbing and playing chess.",
  "context_used": [...],
  "filtered_by_user": "Jordan"
}

Governance

This project follows the same governance pattern as the rest of this portfolio — a lightweight but honest paper trail, not a formality:

Document What it covers
Project Charter Objectives, scope, stakeholders, success criteria
RAID Log Risks, assumptions, issues, dependencies — including known gaps like no auth on the API
RACI Matrix Role responsibilities across the delivery
Quality Gate Scorecard Honest pass/fail assessment — including a "Fail" on authentication, called out rather than hidden
UAT Checklist Test cases run against the actual API, with one flagged as not yet covered

Known gaps and roadmap

Documented in full in the RAID log and quality gate scorecard:

  • No authentication on /documents or /ask — anyone with network access can write or read data. Top priority before any real deployment.
  • qwen2.5:0.5b is a small model chosen for local speed; answer fluency is a known trade-off. Swappable for a larger local or hosted model.
  • Line-based chunking is naive and doesn't handle long paragraphs or sentence boundaries — fine for a short demo profile, not production-grade.
  • Unknown-user queries (user param with no matching data) aren't formally tested yet — carried forward as an open UAT item.

Origin

Built as a hands-on project following NextWork's "Build a RAG API with FastAPI" curriculum, then extended with governance documentation and the multi-user directory as portfolio additions.

About

A local Retrieval Augmented Generation API that answers questions from a personal profile document using FastAPI, ChromaDB, and Ollama, extended into a multi user directory with per user data isolation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages