REVIEW 13 cited by
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have showcased remarkable reasoning capabilities, yet they remain susceptible to errors, particularly in temporal reasoning tasks involving complex temporal logic. Existing research has explored LLM performance on temporal reasoning using diverse datasets and benchmarks. However, these studies often rely on real-world data that LLMs may have encountered during pre-training or employ anonymization techniques that can inadvertently introduce factual inconsistencies. In this work, we address these limitations by introducing novel synthetic datasets specifically designed to assess LLM temporal reasoning abilities in various scenarios. The diversity of question types across these datasets enables systematic investigation into the impact of the problem structure, size, question type, fact order, and other factors on LLM performance. Our findings provide valuable insights into the strengths and weaknesses of current LLMs in temporal reasoning tasks. To foster further research in this area, we are open-sourcing the datasets and evaluation framework used in our experiments: https://huggingface.co/datasets/baharef/ToT.
Forward citations
Cited by 13 Pith papers
-
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.
-
Temporal Preference Concepts and their Functions in a Large Language Model
Temporal preference in Qwen3-4B-Instruct-2507 localizes to layers 17–35 (especially L24 attention), has curved residual-stream geometry, is behaviorally unstable, and can be bidirectionally steered.
-
Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models
Temporal Attractor Steering resolves 29-57% of parametric temporal conflicts in open-weight LLMs while preserving 85-99% accuracy on non-conflict queries.
-
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
On QEDBench, frontier LLM judges over-score university math proofs by up to +0.36 on average relative to human experts, while some solver models fail badly on discrete-combinatorial problems.
-
TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
In TempoBench's formally verified temporal-causality tests, frontier LLMs score 65.6% F1 on normal causal attribution but 7.5% on hard instances, while trace simulation stays far higher.
-
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
TSAIA, a new benchmark, tests eight LLMs on 1,054 multi-step time series tasks and finds they cannot reliably complete the required workflows.
-
Hatevolution: What Static Benchmarks Don't Tell Us
Static hate speech benchmarks rank models differently from time-sensitive evaluations, with correlation coefficients near zero or negative, so high benchmark scores do not guarantee robustness to language change.
-
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
A private, contamination-resistant benchmark of 250 olympiad-level programming problems shows top reasoning models reaching about 36% solve rates, far above conventional models.
-
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
New temporal benchmarks show LLMs struggle with outdated facts, and a structured knowledge-organization memory improves accuracy over ICL and RAG.
-
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
On GeoTemp, the best open model answers only 56% of two-city time questions and 33% when an hour shift is added, despite near-perfect scores on pure time arithmetic.
-
Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
With chain-of-thought prompting and text inputs, GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet reach or exceed human-level set-shifting on the WCST, but not with visual inputs or direct answers.
-
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.
-
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.
Discussion (0). Sign in to comment.