A tool facilitating matching columns across tabular datasets. It also serves as an experiment suite for state-of-the-art schema matching methods.
-
Updated
Oct 4, 2026 - Python
A tool facilitating matching columns across tabular datasets. It also serves as an experiment suite for state-of-the-art schema matching methods.
Valentine scalable deployment for VLDB demo
Adversarial paper red-teamer + dataset opportunity scout — one chat, two LangGraph flows. Calibrated PASS/REVISE/FAIL verdicts, measured against real peer reviews.
Deterministic key and join discovery for structured datasets
Master thesis: Holistic Schema Matching at Scale
Your dataset discovery and curation buddy.
JCDL 2025 Paper "Multi-Disciplinary Dataset Discovery from Citation-Verified Literature Contexts" which matching research questions to cited datasets.
An AI-maintained knowledge base that maps economics research ideas to the best empirical datasets and executable acquisition routes.
Code, data, and dataset browser for an LREC 2026 study of language dataset visibility in low-resource multilingual NLP, comparing catalogue counts with citation-traced datasets.
Search-first dataset discovery platform with DuckDB, FastAPI, background workers, and a reproducible local demo path.
Evidence-led agentic dataset discovery for Hugging Face
LangGraph multi-agent system that crawls a site, finds datasets, then downloads them or extracts them with LLM-generated parsers — with feedback/retry loops, RAG memory, and a Markdown report. Runs on any OpenAI-compatible endpoint (vLLM).
Independent educational implementation of a Graph RAG pipeline for explainable dataset discovery.
Discover underexplored biomedical datasets through transparent, deterministic scoring. A scientific instrument for finding GEO, SRA, Zenodo, ENA, HCA, Expression Atlas, and Open Targets datasets that deserve a second look — local-first, BYOK, fully auditable.
Evidence-aware dataset intelligence and recommendation system for finding, evaluating, ranking, and diversifying public datasets for machine learning tasks.
Safety-first Python CLI for discovering open-data resources, deterministically classifying inspected content, and explicitly loading verified datasets into DuckDB. v0.1.0 has been released. Working hard towards v0.2.0.
To associate your repository with the dataset-discovery topic, visit your repo's landing page and select "manage topics."