Skip to content

Latest commit

 

History

27 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

TIME — temporal reasoning in real-world scenarios

NeurIPS 2025 Spotlight✨ TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios

Shaohang Wei, Wei Li, Feifan Song, Wen Luo,
Tianyi Zhuang, Haochen Tan, Zhijiang Guo, Houfeng Wang

Peking University  ·  Noah's Ark Lab
Contact: shaohang@stu.pku.edu.cn

Peking University        Huawei Noah's Ark Lab

Accepted to the Datasets & Benchmarks track.

Paper: arXiv Project website Code: GitHub TIME dataset on Hugging Face TIME-Lite dataset on Hugging Face NeurIPS 2025 Spotlight

Project Website ↗  ·  Overview  ·  Dataset  ·  Results  ·  Getting started  ·  Citation

Overview

TIME is a benchmark for temporal reasoning in real-world scenarios. It contains 38,522 QA pairs across 3 levels and 11 fine-grained tasks, organized into TIME-Wiki, TIME-News, and TIME-Dial. These settings capture three challenges: intensive temporal information, fast-changing event dynamics, and complex temporal dependencies in social interactions.

We evaluate reasoning and non-reasoning models across these scenarios and tasks, and study how test-time scaling affects temporal reasoning. TIME-Lite provides a human-annotated subset of 943 QA pairs for standardized evaluation.

TIME overview: three levels of temporal reasoning across Wiki, News, and Dial

Dataset

The complete benchmark and its human-annotated subset share the same three scenario groups. For a compact evaluation, start with TIME-Lite.

ScenarioTIMETIME-Lite
Wiki13,848322
News19,958299
Dial4,716322
All scenarios38,522943

Number of QA pairs. TIME is the complete benchmark; TIME-Lite is the high-quality, human-annotated subset.

Detailed statistics by task and scenario
TaskTIMETIME-Lite
WikiNewsDialTotalWikiNewsDialTotal
All tasks1384819958471638522322299322943
Ext.1261021914803003060
Loc.12991800447354630303090
Comp.11261800450337624302478
D.C.11511800450340128302886
O.C.12991800450354930303090
E.R.12871800450353730303090
O.R.12881800450353830303090
R.R.12871800450353730303090
C.T.12631800450351330303090
T.L.13003758450550830293089
C.F.12871800450353730303090

Task abbreviations: Ext. (Extract), Loc. (Localization), Comp. (Computation), D.C. (Duration Compare), O.C. (Order Compare); E.R. (Explicit Reasoning), O.R. (Order Reasoning), R.R. (Relative Reasoning); C.T. (Co-temporality), T.L. (Timeline), C.F. (Counterfactual).

Construction pipeline

TIME dataset construction pipeline

Evaluation results

The following radar charts compare model performance on the three TIME-Lite sub-datasets.

TIME-Lite-Wiki TIME-Lite-News TIME-Lite-Dial
TIME-Lite-Wiki results TIME-Lite-News results TIME-Lite-Dial results

Click any chart to view its full-resolution labels and legend.

Getting started

1. Set up the repository

Install Git LFS and clone this repository:

git lfs install
git clone https://github.com/sylvain-wei/TIME.git
cd TIME
pip install -r evaluation/requirements.txt

2. Download a dataset

TIME-Lite — recommended for a compact evaluation:

bash scripts/download_data_time_lite.sh

TIME — complete benchmark:

bash scripts/download_data_time.sh

The datasets are also available directly on Hugging Face: TIME · TIME-Lite.

3. Configure and run evaluation

Set model and dataset_path in the corresponding evaluation script for your local setup.

The provided scripts need local adaptation: check the download archive locations and evaluation arguments. The prompt templates referenced by evaluation/eval.py are not included in this repository.

TIME-Lite:

bash scripts/eval_time_lite.sh

TIME:

bash scripts/eval_time.sh

Citation

If you find this work helpful, please consider starring this repository, giving the TIME dataset on Hugging Face an upvote, and citing our paper.

GitHub stars

@article{wei2025time,
  title={TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios},
  author={Wei, Shaohang and Li, Wei and Song, Feifan and Luo, Wen and Zhuang, Tianyi and Tan, Haochen and Guo, Zhijiang and Wang, Houfeng},
  journal={arXiv preprint arXiv:2505.12891},
  year={2025}
}

About

[NeurIPS 2025 (Spotlight✨)] TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenario

Topics

Resources

Stars

33 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages