100% local AI research pipeline. No API keys. No cloud. No data leakage.
A three-agent research swarm that ingests your local documents (FOIA dumps, PDFs, text files), cross-references them against live web data, and produces an opinionated intelligence dossier -- all running on your own hardware.
Built with CrewAI + Ollama + DuckDuckGo.
Three specialized AI agents work sequentially, each building on the previous agent's output:
LOCAL DOCUMENTS LIVE WEB
| |
[ Archivist ] [ Investigator ]
Extracts facts, Verifies claims,
names, dates, finds gaps, checks
redactions, gaps news, cites sources
| |
+---------------+---------------+
|
[ Synthesizer ]
Compares both sides,
calls out discrepancies,
writes the final dossier
|
INTELLIGENCE DOSSIER
The Archivist -- Reads every document in your local folder using RAG (Retrieval-Augmented Generation). Extracts names, dates, organizations, claims, and flags redactions or suspicious omissions. Reports only what the documents contain. Never fabricates.
The Investigator -- Takes the Archivist's findings and hits DuckDuckGo to verify claims, fill gaps, and find recent news. Classifies every finding as CORROBORATES, CONTRADICTS, ADDS CONTEXT, or UNVERIFIABLE. Cites all sources.
The Synthesizer -- Receives both reports and writes a structured intelligence dossier with an executive summary, evidence matrix, discrepancy report, and a definitive assessment. No hedging.
- Python 3.10+
- Ollama running locally (install)
- GPU recommended (tested on RTX 3080 Ti, works on CPU but slower)
cd crewai
pip install -r requirements.txtollama pull llama3.3 # LLM (swap for any model you prefer)
ollama pull nomic-embed-text # Embedding model for RAGDrop PDFs and/or text files into the foia_dump/ directory:
foia_dump/
uap_report_2023.pdf
senate_hearing_transcript.txt
classified_memo_redacted.pdf
python research_swarm.py "UAP FOIA disclosure analysis"The dossier gets saved to output/dossier_<topic>_<timestamp>.md.
python research_swarm.py "TOPIC" [OPTIONS]
| Flag | Description | Default |
|---|---|---|
"topic" |
Research topic or question (required) | -- |
-d, --docs |
Path to document directory | ./foia_dump |
-o, --output |
Output directory for dossiers | ./output |
-m, --model |
Ollama LLM model name | llama3.3 |
--embed-model |
Ollama embedding model for RAG | nomic-embed-text |
--no-web |
Skip web research (document analysis only) | disabled |
# Basic usage
python research_swarm.py "Pentagon UAP program timeline"
# Use a different model
python research_swarm.py "Bell System monopoly history" -m mistral
# Point to a different document folder
python research_swarm.py "FOIA redaction patterns" -d ./my_documents
# Document analysis only -- no web searches
python research_swarm.py "Classified memo analysis" --no-web
# Everything custom
python research_swarm.py "Corporate lobbying records" \
-d ./lobbying_docs \
-o ./reports \
-m qwen2.5:32b \
--embed-model mxbai-embed-largeAll defaults can also be set via environment variables:
| Variable | Description |
|---|---|
OLLAMA_MODEL |
Default LLM model |
OLLAMA_URL |
Ollama server URL (default: http://localhost:11434) |
EMBED_MODEL |
Default embedding model |
DOC_DIR |
Default document directory |
OUTPUT_DIR |
Default output directory |
The final dossier is a markdown file with seven sections:
- Executive Summary -- The real story in 3-5 sentences
- Key Findings -- Numbered, each with a confidence level (HIGH/MEDIUM/LOW)
- Evidence Matrix -- Document evidence vs. web evidence, side by side
- Discrepancy Report -- Every contradiction between docs and public info
- What They're Not Telling You -- Deliberate omissions and redaction analysis
- Assessment -- Definitive, opinionated conclusion
- Leads for Further Investigation -- Actionable next steps
Before burning GPU cycles, the script verifies:
- Ollama is running and reachable
- The LLM model is pulled
- The embedding model is pulled
- The document directory exists and contains files
- The output directory exists (creates it if not)
If anything fails, you get a clear error message telling you exactly what to fix.
Why local embeddings matter: The PDFSearchTool from crewai-tools defaults to OpenAI embeddings if you don't configure it. That means your documents would be sent to OpenAI's servers for vectorization -- defeating the entire point of running locally. This project explicitly configures Ollama's nomic-embed-text as the embedding provider. Nothing leaves your machine.
One PDFSearchTool per file: The PDFSearchTool expects individual file paths, not directories. The script discovers all PDFs in your document folder and creates a dedicated RAG tool for each one, giving the Archivist focused search capabilities per document.
Task context chaining: Each task passes its output to the next via CrewAI's context parameter. The Investigator sees the Archivist's full report. The Synthesizer sees both. This is how the agents actually build on each other's work rather than operating in isolation.
| Use Case | Model | VRAM |
|---|---|---|
| Best quality | llama3.3 |
~16 GB |
| Good balance | qwen2.5:14b |
~10 GB |
| Lower VRAM | mistral |
~5 GB |
| Fast iteration | llama3.2:3b |
~3 GB |
| Embeddings | nomic-embed-text |
~300 MB |
crewai/
research_swarm.py # The entire pipeline -- single file, no framework bloat
requirements.txt # Three dependencies
foia_dump/ # Your documents go here
output/ # Dossiers land here
MIT