Retrieval-Augmented Generation (RAG) answers questions from your documents: find the relevant passages first, then ask the model to answer using only those passages, with citations.
This project implements the same RAG app from scratch and with a managed service, behind
one CLI and one UI. The scratch engine comes with two chunking strategies, so you get three
engines to compare: scratch (markdown chunking), scratch-semantic (LLM chunking), and
file-search.
scratch/ |
file_search/ |
|
|---|---|---|
| Chunking | Ours: markdown-aware (scratch) or semantic by gemini-3.5-flash-lite (scratch-semantic) |
Google: whitespace chunks (size configurable) |
| Embeddings | gemini-embedding-2, called by us |
gemini-embedding-2, called by Google at upload |
| Vector store | ChromaDB (open source, embedded, on local disk) | Managed File Search store |
| Retrieval | Cosine similarity, top-k, score threshold | Server-side, top-k |
| Generation | gemini-3.8-flash with a prompt we build |
gemini-3.8-flash with the FileSearch tool |
| Citations | [n] markers the model writes, from numbered passages |
Grounding metadata → [n] markers we insert |
| You can see | Every chunk, score, and the exact prompt | Retrieved chunks and grounding only |
Build them all, compare the answers side by side, and you'll understand what a managed RAG service does for you, and what it hides.
Requires Python 3.10+ and a Gemini API key (get one in AI Studio).
git clone https://github.com/ModernPath/rag-example.git
cd rag-example
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt # needs google-genai >= 2.27
cp .env.example .env # then set GEMINI_API_KEY in .env
python rag.py index --engine all # index sample_docs/ with all three engines
python rag.py ask "How many vacation days do I get after four years?" --engine all
python ui/app.py # http://localhost:5020 — compare side by sideThe sample corpus in sample_docs/ is a fictional company's handbook, travel policy,
IT security policy, and product FAQ (tidy markdown), plus an office guide copied from an
intranet web page (plain text, no headings, menu and footer included) and a benefits page
written in Finnish. Each has specific facts (hotel caps, SLA percentages, deadlines), so
you can check whether an answer is correct.
INDEX (once, and again when documents change)
documents ──► chunk (titles in the base language) ──► embed ──► store
│
ASK (every question) ▼
question ──► translate to the base language ──► embed ──► retrieve top-k ──► prompt with numbered passages ──► LLM ──► answer + [n] citations
(replies in the question's language)
scratch/chunking.pysplits markdown at headings and remembers the heading path (Travel and Expense Policy > Hotels). It packs whole paragraphs into chunks of ≤1200 chars and repeats the last paragraph in the next chunk (overlap), so facts near a boundary keep their context.scratch/embedder.pyturns text into 768-dim vectors withgemini-embedding-2. Read the module docstring: the new model has three traps (see Gotchas).scratch/vector_store.pystores the vectors in ChromaDB running embedded (chromadb.PersistentClient): no server, everything persists indata/scratch/. We pass our own embeddings (embedding_function=None) and use cosine space. Chroma returns cosine distance, so similarity =1 - distance. Each chunk's metadata (document, heading path, document hash) also tells the engine which document versions are already indexed.scratch/engine.pyties it together: incremental indexing,retrieve()(no LLM), andask(), which builds the grounded prompt and flags which sources were cited. The chunker is a parameter:ScratchEngine()uses markdown chunking,semantic_engine()the semantic chunker, each with its own index (data/scratch/,data/scratch-semantic/).
Inspect retrieval without the LLM. This is the most useful debugging tool in RAG:
python rag.py search "Can I use a USB stick?"
# [1] score=0.674 it-security.md — IT Security Policy > Devices
# [2] score=0.617 travel-and-expenses.md — Travel and Expense Policy > Taxis ...Markdown chunking needs headings. A web page copied as text has none: one long paragraph covers office access, visitors, and meeting rooms, with a menu above and a footer below. Size-based cuts then land mid-topic, and every chunk gets the same title (the file name).
The semantic chunker numbers the document's sentences and asks gemini-3.5-flash-lite for
JSON: where each chunk starts, a descriptive title, and whether it is boilerplate.
{"chunks": [{"start": 1, "title": "Navigation", "skip": true},
{"start": 4, "title": "Helsinki office guide - Location and access badge", "skip": false},
{"start": 9, "title": "Helsinki office guide - Visitor policy", "skip": false}, ...]}We cut the original text at those sentences. The model never rewrites content (no paraphrased or invented facts), its output is tiny (fast, cheap), and the boundaries are sanitized so nothing is lost whatever it returns. Boilerplate chunks are dropped.
python rag.py chunks helsinki-office-guide.txt # markdown chunker
python rag.py chunks helsinki-office-guide.txt --engine scratch-semantic # semantic chunker
python rag.py search "Can I bring a visitor?" --engine scratch # 0.737, titled "helsinki-office-guide.txt"
python rag.py search "Can I bring a visitor?" --engine scratch-semantic # 0.824, "Helsinki office guide - Visitor policy"On the tidy markdown docs both strategies cut at roughly the same places. The trade-off: one LLM call per document at index time (~3 s per document, run in parallel) and boundaries that can vary a little between runs.
Embeddings are multilingual, but not symmetric: a Finnish question against English passages
scores lower than the same question in English, so thresholds and top-k drift per language.
And the semantic chunker writes titles that are embedded with the text. So the index has one
base language (RAG_BASE_LANGUAGE, default en), used on both sides:
- Index: the semantic chunker writes every chunk title in the base language, even for
tyosuhde-edut.md, which is in Finnish. Chunk text is never translated: answers quote the source as written. - Ask: one small structured-output call (
gemini-3.5-flash-lite) detects the question's language and translates it into the base language. Retrieval and the grounded prompt use the translation; the system prompt tells the model to reply in the language the question was asked in. If the question is already in the base language, the original wording is kept.
python rag.py ask "Montako lomapäivää saan neljän vuoden jälkeen?" --engine all
# Translated (fi → en): How many vacation days do I get after four years?
# ── scratch (2.1s) ──── Neljän vuoden jälkeen saat 30 lomapäivää. [1]
python rag.py search "How much is the bicycle benefit per year?" --engine scratch-semantic
# [1] score=0.71 tyosuhde-edut.md — Employee benefits - Company bicycle <- English title, Finnish textThe CLI and UI translate once and hand the same Query to every engine, so the comparison
stays fair and you pay for one translation. RAG_TRANSLATE_QUERIES=0 skips the call when
you know your questions are already in the base language. Changing the base language changes
the semantic chunk titles, so reset and re-index scratch-semantic afterwards.
store = client.file_search_stores.create(config={"display_name": "...", "embedding_model": "models/gemini-embedding-2"})
operation = client.file_search_stores.upload_to_file_search_store(
file="sample_docs/it-security.md",
file_search_store_name=store.name,
config={"display_name": "it-security.md",
"custom_metadata": [{"key": "source", "string_value": "it-security.md"}],
"chunking_config": {"white_space_config": {"max_tokens_per_chunk": 300, "max_overlap_tokens": 40}}},
)
while not operation.done: # indexing is asynchronous
time.sleep(2)
operation = client.operations.get(operation)
response = client.models.generate_content(
model="gemini-3.8-flash",
contents="Can I use a USB stick?",
config=types.GenerateContentConfig(
tools=[types.Tool(file_search=types.FileSearch(file_search_store_names=[store.name], top_k=5))]),
)
response.candidates[0].grounding_metadata # retrieved chunks + which sentence cites which chunkWhat you still own with the managed service:
- Keeping the store in sync.
data/file_search/state.jsonmaps each file's content hash to its remote document, soindexre-uploads only changed files and deletes removed ones. - Chunking parameters. Chunk size is the single biggest quality lever in every engine.
- Citations.
insert_citations()turnsgrounding_supportsinto[n]markers.
The Google docs also show the newer Interactions API. It is marked experimental in
the SDK, so this example uses generate_content. The equivalent call:
interaction = client.interactions.create(
model="gemini-3.8-flash", input="Can I use a USB stick?",
tools=[{"type": "file_search", "file_search_store_names": [store.name]}],
)python rag.py index [--engine scratch|scratch-semantic|file-search|all] [--docs FOLDER]
python rag.py ask "question" [--engine ...] [--top-k 5] [--no-sources] # any language
python rag.py search "question" [--engine scratch|scratch-semantic] [--top-k 5] # chunks + scores, no answer LLM call
python rag.py chunks FILE [--engine scratch|scratch-semantic] # preview chunking, no index
python rag.py status [--engine ...]
python rag.py reset --engine ... # file-search: deletes the remote storeUse your own documents: python rag.py index --engine all --docs ~/my-notes (.md and
.txt). File Search also accepts PDF, DOCX, and code files; extending the scratch engine
to PDFs is a good exercise.
| Variable | Default | Purpose |
|---|---|---|
RAG_MODEL |
gemini-3.8-flash |
Generation model |
RAG_EMBEDDING_MODEL |
gemini-embedding-2 |
Embedding model (scratch + new File Search stores) |
RAG_CHUNKING_MODEL |
gemini-3.5-flash-lite |
Model that picks chunk boundaries for scratch-semantic |
RAG_DATA_DIR |
./data |
Where indexes and store state live |
RAG_BASE_LANGUAGE |
en |
ISO 639-1 code of the index: semantic chunk titles and translated questions |
RAG_TRANSLATION_MODEL |
gemini-3.5-flash-lite |
Detects the question's language and translates it |
RAG_TRANSLATE_QUERIES |
1 |
0 skips translation (questions are assumed to be in the base language) |
pytest # offline: fake embedder, LLM, and File Search client
RAG_LIVE_TESTS=1 pytest tests/test_live.py # real API; creates and deletes a temporary storegemini-embedding-2aggregates lists.embed_content(contents=["a", "b"])returns one embedding for both texts combined. To batch, wrap each text in its owntypes.Content. Older code written forgemini-embedding-001silently breaks here.task_typeis ignored bygemini-embedding-2. Put the task in the text instead:task: search result | query: …for questions,title: … | text: …for documents.- Similarity thresholds are model-specific. With
gemini-embedding-2at 768 dims, relevant passages here score ~0.65–0.85 and unrelated questions ~0.5–0.57, so the scratch engine uses 0.6. Recalibrate withrag.py searchfor your corpus. - Grounding offsets are UTF-8 bytes, not characters. With text like "10:00–14:00" (an en dash is 3 bytes), inserting citations by string index puts them in the wrong place.
- google-genai 2.27+ is needed for
embedding_modelwhen creating a File Search store. SDK 2.x also logs an "automatic function calling" warning on everygenerate_contentunless you setautomatic_function_calling=AutomaticFunctionCallingConfig(disable=True). - Keep the
genai.Clientreferenced while using it. The SDK closes its HTTP connection when the client is garbage-collected, so chained one-liners can fail with "client has been closed". - Rate limits. A
429 RESOURCE_EXHAUSTEDcan still surface after the SDK's retries. The CLI and UI report it as an error instead of crashing. - Changing the chunker doesn't re-index. Sync compares document hashes, so after
editing chunking code or parameters run
rag.py reset --engine scratch(orscratch-semantic) and index again. Changing the embedding model is detected automatically.
- Tune chunking. Re-index with
max_chars=400and then3000. Which questions get better or worse? Why do small chunks lose context and big chunks dilute similarity? Then comparescratchandscratch-semanticon questions about the office guide. - Contextual chunks. Extend the semantic chunker's JSON with a one-sentence
contextper chunk ("This section of the travel policy covers…") and embed it with the chunk. Does retrieval improve for short chunks? - Hybrid search. Add a keyword score (e.g. BM25 or simple term overlap) to the cosine
score in
ChromaVectorStore.search. Try a query with an exact product name or number. - Metadata filtering. Pass
metadata_filter='source="it-security.md"'inFileSearch, and add the equivalentdocument=filter to the scratch engine. - Multi-hop questions. Ask "I'm flying 7 hours to London. Which class and what hotel
budget?" The answer needs two sections. Inspect
top_kand the retrieved sources. - Make it agent memory. Wrap
ScratchEngine.retrieve()as a function tool for a Gemini agent (automatic function calling with a Python callable), so the agent decides when to look things up and cites what it found. - Cross-language retrieval. Run
rag.py search "bicycle benefit" --engine scratchand--engine scratch-semantic: the markdown chunker keeps the Finnish heading as the title, the semantic chunker wrote an English one. Then setRAG_TRANSLATE_QUERIES=0and ask the Finnish question again. How much does the score drop without translation? TryRAG_BASE_LANGUAGE=fi(reset and re-indexscratch-semanticfirst). - Swap the vector store. Point Chroma at a server (
chromadb.HttpClient, run withchroma run --path ./chroma-data), or implementChromaVectorStore's six methods on Qdrant, LanceDB, or pgvector, without changingengine.py.
rag-example/
├── AGENTS.md # rules for extending this project (engines, chunkers, stores, tests)
├── rag.py # CLI for all engines
├── rag_common.py # config, Document loading, sync planning, Answer/Source types
├── translation.py # base language: detect + translate questions, title language for chunks
├── scratch/ # chunking.py | semantic_chunking.py → embedder.py → vector_store.py → engine.py
├── file_search/engine.py # store sync, FileSearch tool, grounding → citations
├── ui/app.py # Flask side-by-side comparison (port 5020)
├── sample_docs/ # fictional company documents (markdown + one web page as text + one in Finnish)
├── data/ # ChromaDB files + File Search state (gitignored)
├── .env.example # copy to .env and set GEMINI_API_KEY
├── requirements.txt
└── tests/ # offline unit tests + opt-in live tests