NeurIPS 2025 Spotlight✨ TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
Shaohang Wei, Wei Li, Feifan Song, Wen Luo,
Tianyi Zhuang, Haochen Tan, Zhijiang Guo, Houfeng Wang
Peking University · Noah's Ark Lab
Contact: shaohang@stu.pku.edu.cn
Accepted to the Datasets & Benchmarks track.
Project Website ↗ · Overview · Dataset · Results · Getting started · Citation
TIME is a benchmark for temporal reasoning in real-world scenarios. It contains 38,522 QA pairs across 3 levels and 11 fine-grained tasks, organized into TIME-Wiki, TIME-News, and TIME-Dial. These settings capture three challenges: intensive temporal information, fast-changing event dynamics, and complex temporal dependencies in social interactions.
We evaluate reasoning and non-reasoning models across these scenarios and tasks, and study how test-time scaling affects temporal reasoning. TIME-Lite provides a human-annotated subset of 943 QA pairs for standardized evaluation.
The complete benchmark and its human-annotated subset share the same three scenario groups. For a compact evaluation, start with TIME-Lite.
| Scenario | TIME | TIME-Lite |
|---|---|---|
| Wiki | 13,848 | 322 |
| News | 19,958 | 299 |
| Dial | 4,716 | 322 |
| All scenarios | 38,522 | 943 |
Number of QA pairs. TIME is the complete benchmark; TIME-Lite is the high-quality, human-annotated subset.
Detailed statistics by task and scenario
| Task | TIME | TIME-Lite | ||||||
|---|---|---|---|---|---|---|---|---|
| Wiki | News | Dial | Total | Wiki | News | Dial | Total | |
| All tasks | 13848 | 19958 | 4716 | 38522 | 322 | 299 | 322 | 943 |
| Ext. | 1261 | 0 | 219 | 1480 | 30 | 0 | 30 | 60 |
| Loc. | 1299 | 1800 | 447 | 3546 | 30 | 30 | 30 | 90 |
| Comp. | 1126 | 1800 | 450 | 3376 | 24 | 30 | 24 | 78 |
| D.C. | 1151 | 1800 | 450 | 3401 | 28 | 30 | 28 | 86 |
| O.C. | 1299 | 1800 | 450 | 3549 | 30 | 30 | 30 | 90 |
| E.R. | 1287 | 1800 | 450 | 3537 | 30 | 30 | 30 | 90 |
| O.R. | 1288 | 1800 | 450 | 3538 | 30 | 30 | 30 | 90 |
| R.R. | 1287 | 1800 | 450 | 3537 | 30 | 30 | 30 | 90 |
| C.T. | 1263 | 1800 | 450 | 3513 | 30 | 30 | 30 | 90 |
| T.L. | 1300 | 3758 | 450 | 5508 | 30 | 29 | 30 | 89 |
| C.F. | 1287 | 1800 | 450 | 3537 | 30 | 30 | 30 | 90 |
Task abbreviations: Ext. (Extract), Loc. (Localization), Comp. (Computation), D.C. (Duration Compare), O.C. (Order Compare); E.R. (Explicit Reasoning), O.R. (Order Reasoning), R.R. (Relative Reasoning); C.T. (Co-temporality), T.L. (Timeline), C.F. (Counterfactual).
The following radar charts compare model performance on the three TIME-Lite sub-datasets.
| TIME-Lite-Wiki | TIME-Lite-News | TIME-Lite-Dial |
|---|---|---|
![]() |
![]() |
![]() |
Click any chart to view its full-resolution labels and legend.
Install Git LFS and clone this repository:
git lfs install
git clone https://github.com/sylvain-wei/TIME.git
cd TIME
pip install -r evaluation/requirements.txtTIME-Lite — recommended for a compact evaluation:
bash scripts/download_data_time_lite.shTIME — complete benchmark:
bash scripts/download_data_time.shThe datasets are also available directly on Hugging Face: TIME · TIME-Lite.
Set model and dataset_path in the corresponding evaluation script for your local setup.
The provided scripts need local adaptation: check the download archive locations and evaluation arguments. The prompt templates referenced by evaluation/eval.py are not included in this repository.
TIME-Lite:
bash scripts/eval_time_lite.shTIME:
bash scripts/eval_time.shIf you find this work helpful, please consider starring this repository, giving the TIME dataset on Hugging Face an upvote, and citing our paper.
@article{wei2025time,
title={TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios},
author={Wei, Shaohang and Li, Wei and Song, Feifan and Luo, Wen and Zhuang, Tianyi and Tan, Haochen and Guo, Zhijiang and Wang, Houfeng},
journal={arXiv preprint arXiv:2505.12891},
year={2025}
}




