Academic Project
Course: Information Retrieval (IR)
Institution: Minia National University
Term: Summer 2026 Semester
Nova IR System is an end-to-end Information Retrieval engine and interactive benchmarking suite built from the ground up on the standardized CISI benchmark collection.
The system implements and evaluates two foundational IR paradigms:
- Vector Space Model (VSM): TF-IDF representation with Cosine Similarity.
- Probabilistic Retrieval Model: Okapi BM25 with an Inverted Index for sub-millisecond query execution.
The project features a modern, responsive web dashboard built with Streamlit, enabling real-time search with keyword highlight snippets, live Cranfield evaluation metrics (Precision@K, Recall@K, MAP), and a full dataset explorer.
- 🧹 Robust NLP Preprocessing Pipeline: Tokenization, case folding, punctuation stripping, NLTK stop-word removal, and Porter Stemmer morphological reduction.
- ⚡ High-Performance Inverted Indexing: Fast term-to-posting map reducing search complexity from linear corpus scans to sub-millisecond lookups.
- 🔍 Dual Retrieval Algorithms:
- TF-IDF: Sublinear term frequency scaling and Cosine Similarity vector ranking.
-
BM25Okapi: Probabilistic ranking with term frequency saturation (
$k_1 = 1.5$ ) and document length normalization ($b = 0.75$ ).
- 📊 Cranfield Evaluation Benchmark: Automated evaluation over ground-truth relevance judgments (
CISI.REL) computing MAP, Precision@10, and Recall@10 with side-by-side metric delta comparisons and interactive charts. - 🎨 Modern Streamlit Web UI:
-
Search Hub: Real-time query execution, latency indicators, ranked cards with metadata badges, and stem-based keyword highlighting (
<mark>tags). - Evaluation Dashboard: One-click benchmark runner with summary metric cards, model comparison tables, and query-level breakdown.
- Dataset Explorer: Paginated search and inspection of all 1,460 documents by ID, title, author, or keywords.
- Quick Test Queries: One-click selection of sample CISI queries directly from the sidebar.
-
Search Hub: Real-time query execution, latency indicators, ranked cards with metadata badges, and stem-based keyword highlighting (
Evaluated across the 76 ground-truth labeled CISI queries (Cutoff @
| Metric | TF-IDF (Vector Space) | BM25 (Probabilistic) | Winner & Improvement |
|---|---|---|---|
| Mean Average Precision (MAP) | 0.2113 |
0.2335 |
BM25 (+10.5%) 🚀 |
| Mean Precision@10 | 0.3092 |
0.3724 |
BM25 (+20.4%) 🚀 |
| Mean Recall@10 | 0.1172 |
0.1492 |
BM25 (+27.3%) 🚀 |
-
Term Frequency Saturation (
$k_1=1.5$ ): Limits the diminishing returns of repeated terms, avoiding term-spamming distortion. -
Document Length Normalization (
$b=0.75$ ): Normalizes document scores against the average corpus length ($\text{avgdl}$ ), eliminating bias towards overly verbose documents.
nova-ir-system/
├── app.py # Streamlit web application & user interface
├── preprocessing.py # NLTK tokenization, stop-words, Porter stemmer & highlighting
├── ir_engine.py # TF-IDF (Cosine Similarity) & BM25Okapi (Inverted Index) engines
├── evaluator.py # Cranfield evaluation suite (Precision@K, Recall@K, MAP)
├── cisi_parser.py # CISI dataset parser (.I, .T, .A, .W dot-tags)
├── requirements.txt # Python project dependencies
├── CODE_EXPLANATION.md # Detailed mathematical and code explanation
├── PRESENTATION_SCRIPT.md # Presentation script for faculty presentation
└── data/
├── CISI.ALL # Raw CISI documents collection (1,460 docs)
├── CISI.QRY # Raw CISI queries collection (112 queries)
├── CISI.REL # Raw CISI relevance mappings (3,114 pairs)
├── documents.json # Parsed documents
├── queries.json # Parsed queries
└── relevance.json # Structured relevance ground-truth
- Python 3.10, 3.11, or 3.12 installed.
git clone https://github.com/your-username/nova-ir-system.git
cd nova-ir-system- Windows (PowerShell):
python -m venv venv .\venv\Scripts\Activate.ps1 - macOS / Linux:
python3 -m venv venv source venv/bin/activate
pip install -r requirements.txtpython cisi_parser.pystreamlit run app.pyOpen your browser at http://localhost:8501.
The CISI (Centre for Inventions and Scientific Information) dataset is a classic Information Retrieval benchmark collection consisting of:
- 1,460 Documents: Abstracts and bibliographic metadata in library science and information retrieval.
- 112 Queries: Information requests.
- 76 Labeled Queries: Ground-truth binary relevance assessments (
CISI.REL).
- Institution: Minia National University (MNU)
- Faculty: Faculty of Computers and Artificial Intelligence
- Course: Information Retrieval (IR)
- Academic Term: Summer 2026
- Project Team Members — Minia National University
- Supervised for the Information Retrieval Course (Summer 2026).