Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

⚡ Nova IR System — CISI Search Engine & Evaluation Suite

Python 3.12 Streamlit scikit-learn NLTK Dataset

Academic Project
Course: Information Retrieval (IR)
Institution: Minia National University
Term: Summer 2026 Semester


📌 Project Overview

Nova IR System is an end-to-end Information Retrieval engine and interactive benchmarking suite built from the ground up on the standardized CISI benchmark collection.

The system implements and evaluates two foundational IR paradigms:

  1. Vector Space Model (VSM): TF-IDF representation with Cosine Similarity.
  2. Probabilistic Retrieval Model: Okapi BM25 with an Inverted Index for sub-millisecond query execution.

The project features a modern, responsive web dashboard built with Streamlit, enabling real-time search with keyword highlight snippets, live Cranfield evaluation metrics (Precision@K, Recall@K, MAP), and a full dataset explorer.


✨ Key Features

  • 🧹 Robust NLP Preprocessing Pipeline: Tokenization, case folding, punctuation stripping, NLTK stop-word removal, and Porter Stemmer morphological reduction.
  • ⚡ High-Performance Inverted Indexing: Fast term-to-posting map reducing search complexity from linear corpus scans to sub-millisecond lookups.
  • 🔍 Dual Retrieval Algorithms:
    • TF-IDF: Sublinear term frequency scaling and Cosine Similarity vector ranking.
    • BM25Okapi: Probabilistic ranking with term frequency saturation ($k_1 = 1.5$) and document length normalization ($b = 0.75$).
  • 📊 Cranfield Evaluation Benchmark: Automated evaluation over ground-truth relevance judgments (CISI.REL) computing MAP, Precision@10, and Recall@10 with side-by-side metric delta comparisons and interactive charts.
  • 🎨 Modern Streamlit Web UI:
    • Search Hub: Real-time query execution, latency indicators, ranked cards with metadata badges, and stem-based keyword highlighting (<mark> tags).
    • Evaluation Dashboard: One-click benchmark runner with summary metric cards, model comparison tables, and query-level breakdown.
    • Dataset Explorer: Paginated search and inspection of all 1,460 documents by ID, title, author, or keywords.
    • Quick Test Queries: One-click selection of sample CISI queries directly from the sidebar.

🏆 Benchmark Evaluation Results

Evaluated across the 76 ground-truth labeled CISI queries (Cutoff @ $K=10$):

Metric TF-IDF (Vector Space) BM25 (Probabilistic) Winner & Improvement
Mean Average Precision (MAP) 0.2113 0.2335 BM25 (+10.5%) 🚀
Mean Precision@10 0.3092 0.3724 BM25 (+20.4%) 🚀
Mean Recall@10 0.1172 0.1492 BM25 (+27.3%) 🚀

Why BM25 Outperforms TF-IDF:

  1. Term Frequency Saturation ($k_1=1.5$): Limits the diminishing returns of repeated terms, avoiding term-spamming distortion.
  2. Document Length Normalization ($b=0.75$): Normalizes document scores against the average corpus length ($\text{avgdl}$), eliminating bias towards overly verbose documents.

📁 Repository Structure

nova-ir-system/
├── app.py                 # Streamlit web application & user interface
├── preprocessing.py       # NLTK tokenization, stop-words, Porter stemmer & highlighting
├── ir_engine.py           # TF-IDF (Cosine Similarity) & BM25Okapi (Inverted Index) engines
├── evaluator.py           # Cranfield evaluation suite (Precision@K, Recall@K, MAP)
├── cisi_parser.py         # CISI dataset parser (.I, .T, .A, .W dot-tags)
├── requirements.txt       # Python project dependencies
├── CODE_EXPLANATION.md    # Detailed mathematical and code explanation
├── PRESENTATION_SCRIPT.md # Presentation script for faculty presentation
└── data/
    ├── CISI.ALL           # Raw CISI documents collection (1,460 docs)
    ├── CISI.QRY           # Raw CISI queries collection (112 queries)
    ├── CISI.REL           # Raw CISI relevance mappings (3,114 pairs)
    ├── documents.json     # Parsed documents
    ├── queries.json       # Parsed queries
    └── relevance.json     # Structured relevance ground-truth

🚀 Getting Started

1. Prerequisites

  • Python 3.10, 3.11, or 3.12 installed.

2. Clone the Repository

git clone https://github.com/your-username/nova-ir-system.git
cd nova-ir-system

3. Create & Activate Virtual Environment

  • Windows (PowerShell):
    python -m venv venv
    .\venv\Scripts\Activate.ps1
  • macOS / Linux:
    python3 -m venv venv
    source venv/bin/activate

4. Install Dependencies

pip install -r requirements.txt

5. Parse the Dataset (If not already generated)

python cisi_parser.py

6. Launch the Streamlit App

streamlit run app.py

Open your browser at http://localhost:8501.


📚 Dataset Information

The CISI (Centre for Inventions and Scientific Information) dataset is a classic Information Retrieval benchmark collection consisting of:

  • 1,460 Documents: Abstracts and bibliographic metadata in library science and information retrieval.
  • 112 Queries: Information requests.
  • 76 Labeled Queries: Ground-truth binary relevance assessments (CISI.REL).

🎓 Academic Context

  • Institution: Minia National University (MNU)
  • Faculty: Faculty of Computers and Artificial Intelligence
  • Course: Information Retrieval (IR)
  • Academic Term: Summer 2026

👥 Contributors

  • Project Team Members — Minia National University
  • Supervised for the Information Retrieval Course (Summer 2026).

About

An Information Retrieval system & Streamlit search engine benchmarking TF-IDF vs BM25 with Inverted Indexing on the CISI dataset. Course project at Minia National University (Summer 2026).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages