Pith. sign in

Paper Citation Record · LEDGER

Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.09170.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09170 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:39:31.908208Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.787818Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 21db5bb0-2e97-4efa-a790-8c6953049514 · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.451997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.451997Z digest=sha256:dac795e41f79d012ccd126d98a9748b3e86b4178ed66cc4b7a24a7e67c4ff2af

Observation 71a21248-58c0-4b15-b128-e53497781a7f · inbound

DateLogicQA: Benchmarking Temporal Biases in Large Language Models cites this paper.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.958984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.958984Z digest=sha256:32fbac818ad2a96e0dd6995561b9df8f51b846b7e8e91d4c8839d1068c1a6006

Observation 0c5c3514-b454-4ea0-ace4-e3758bafacfb · inbound

Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning cites this paper.

Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:06:05.211951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:06:05.211951Z digest=sha256:beaa39adf44a50142ab8aa08425873cda12fd8999755b1a5e1a4015de1077f4a

Observation 3caea6ea-dc99-4f52-895b-8cb16f54b20c · inbound

ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events cites this paper.

ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:13.128986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:03:13.128986Z digest=sha256:6c5d1f7feecd91b769ffb83c32cbb345ba92bbc2f4067d67067ee34605661900

Observation a2c53ff8-b438-4eed-9498-3c184a07db16 · inbound

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! cites this paper.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.171743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.171743Z digest=sha256:c96a6e526c7e45aa3b0837d6ce3a33743d1e5eb3de3dfb7e268a0c893bad9938

Observation 701b9690-42cd-4c21-a48b-c459aed6ca9f · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.276239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:41ac2b8dc9b38937068c9cbf371eec4695b22dca834a64fc9e49f5d0a0861267

Observation a59d186c-7fc3-48c5-9aa1-90e8e9f54f9b · inbound

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents cites this paper.

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:12.799685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:12.799685Z digest=sha256:9b648ad0b49c2f3a11639feefcd727149c8f3d17c43be8c88366586857ddc7ce

Observation 936d4ef5-3030-4be6-aaef-a88066aa067a · inbound

ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models cites this paper.

ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:06.234436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:06.234436Z digest=sha256:43bdb15fcfe67b6c3553066080ed11a4f20043c1b2b05b43ce43fa867cb25923

Observation d0eaca58-7d6f-46f1-aefc-b450aff8d67c · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.936216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:31.936216Z digest=sha256:e95428816b45e0c276b71d4754f326d208bb403235762a4e3ecb2018e9c53b5e

Observation a7b3db6e-b99f-4219-b7f5-47935a16ef79 · inbound

Around the World in 24 Hours: Probing LLM Knowledge of Time and Place cites this paper.

Around the World in 24 Hours: Probing LLM Knowledge of Time and Place Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:00.683628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:55:00.683628Z digest=sha256:59ee8956471d040a2ac6a8470225dedb713c07505f630967aa50010eb728350f

Observation ef337db1-38f1-45d4-9584-580608afbe32 · inbound

Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs cites this paper.

Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:48.795828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:48.795828Z digest=sha256:50964bb67204ed1e188f8e73e2719af4803bf49a4942beac975fcb208e83b48b

Observation 257705c7-de3a-4bd4-beb8-34177dabd9b4 · inbound

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics cites this paper.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.361435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.361435Z digest=sha256:2b3309c6579b1a318c60641edfe6276bb060d05256679a07521cf8859dee82b4

Observation 008eed55-694a-4f47-b6c7-f3108658c80e · inbound

Hatevolution: What Static Benchmarks Don't Tell Us cites this paper.

Hatevolution: What Static Benchmarks Don't Tell Us Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:09.583331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:09.583331Z digest=sha256:f85c17f647f1dbe7770a060b6a02578e9145826356e9ffadfbb265aba7a81bbb

Observation ae106fff-8865-4154-b1c8-e276fe35ec00 · inbound

Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications cites this paper.

Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:06.179002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:06.179002Z digest=sha256:0be1c5c2dcf57ff590234f070d45b694b1069fbe191db4a77c10b7d3bf089266

Observation 6fff8a26-24ac-41e9-92b6-df2e3a05366c · inbound

When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference cites this paper.

When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:12:58.671594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:12:58.671594Z digest=sha256:5655d521aa58f66fa6486059c57a016ad5b8c44890b3eed698c7f44c91e16da8

Observation c67c6b83-db5a-435f-892a-17f3e41efe8e · inbound

TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation cites this paper.

TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:39.776350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:39.776350Z digest=sha256:991a63ac42ee0fc60dd10da5860f8ccf74422be64b008b0a7feda7291592718f

Observation 526f11c0-2670-4579-98da-bf6a8271bf3f · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.280536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.280536Z digest=sha256:c4643d35322f089ee98201c66e09d56f06ae3a26e0534f3ae4da35c646b4bbd8

Observation 37ce2429-69dd-49b2-8572-ac186b57fc03 · inbound

UserGPT Technical Report cites this paper.

UserGPT Technical Report Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.733342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T03:10:51.555653Z digest=sha256:5dd6f6388f344d0023926ed46f8673cc622d3b23e4ea59b769b4070679fa52fd

Observation 934c5d4d-7a78-4616-8dd7-b5f02c674d6a · inbound

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi cites this paper.

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.589100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T11:18:08.326304Z digest=sha256:e60b0754ae406afc44b3e1e7b2e4e41be76abaf006d6545b04ac8d936102063b

Observation 94b808d1-1c5e-4460-b603-3ca3912f676e · inbound

DateSAT: A Framework for Solving Date and Period Constraints cites this paper.

DateSAT: A Framework for Solving Date and Period Constraints Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:02.905459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T23:40:46.386350Z digest=sha256:4aa91352bec96d9e9b6a6f9cd888ef4910e1530669f31f7e68c685c22f4e7254

Observation b9a4bcb1-7ead-46b3-a648-b8d1b0ebdf04 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:05:47.163039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:378fcbbe9405d3db1bec7071afca03d07a85b29331a06778529dcae9b04e765a

Observation c330128e-aa3c-4114-86bb-0ddb63591bdc · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:19233e852ae8f80468883e57ab207cf05ebc0d3a95fd891f581d291317d67762

Observation eac11a36-2567-455d-8e75-e941472b1e9c · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.660552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:d61ab57bb1de04992bb9ade1ff0f0cb01d3aedeafec0ec14d5ff2f3f68e67d8c

Observation e658e87e-c84b-450e-bac7-6044229ab2c6 · inbound

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models cites this paper.

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.789947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T17:35:09.053339Z digest=sha256:6c2d9ee437da4c6b46ea2ccfa8f34bd4cf34b003f7a4eaba143ea75b40b16b9c

Observation 56df4af4-6623-441c-9983-a6bfa62726ee · inbound

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models cites this paper.

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:06.378923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:48:06.378923Z digest=sha256:77b478bff52e0f998528dad28fbe5b28b897e42060ed5da67a24a5d102c6d937

Observation 7d0781ab-3bc1-4bfd-80c5-66fde80ecc12 · inbound

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs cites this paper.

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:15:45.637923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T23:50:48.584391Z digest=sha256:d0dbae74bda11091af20e73f1f94f2211f9bc16309ba45af4627c5e1bfb9e4ca

Observation 4b3b03a4-708c-4392-b423-7b3f4c5ed25c · inbound

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding cites this paper.

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:39:31.908208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:39:31.908208Z digest=sha256:8ff3f4b4a652dbed2ae344568a3f6187002f2516492f30916d638ff2a9ba7498