Pith. sign in

Paper Citation Record · LEDGER

TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2311.17667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.17667 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:18.581968Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a5a4d92e-344e-4c89-a34b-07fb45d3282a · inbound

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models cites this paper.

VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:07:09.430884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:07:09.430884Z digest=sha256:36e8add9b7ba32adfa31eb46ca520b64fe2ad05b7eee1b2109dd84d2f7048665

Observation 47293d90-b7b2-4877-9aac-faaa4aa22773 · inbound

Rethinking Emotion Annotations in the Era of Large Language Models cites this paper.

Rethinking Emotion Annotations in the Era of Large Language Models TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T18:31:19.938121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:31:19.938121Z digest=sha256:3db10577da6cdcdb01face81d083ac3a4bb7fadb0e1fa44029af1e43be53105c

Observation f62b4e9b-bc55-4f96-8d39-28c78f2dd1b1 · inbound

ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events cites this paper.

ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:13.120901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:03:13.120901Z digest=sha256:b12da3d65660c5ea5d4c8ab118e42561b81fa65b344d139aa6177ad415f1d22f

Observation 0275c38d-0d90-4f62-a50d-b3102b686a91 · inbound

TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References cites this paper.

TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:18.581968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:18.581968Z digest=sha256:f4a35bfca1347feb151c126013c469dfb9fab8dd00f860ef7cf5cbb03706a51d

Observation f97b287e-e644-473f-930c-e8b0ae22ed8c · inbound

ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning cites this paper.

ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:37.284558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:37.284558Z digest=sha256:d96b2eda6dfd3406e3a8689e5cd3ea9b5e0445b2c811d8f081447764bef4ab52

Observation 678faa0b-1548-43f3-ba7b-a663094c0c93 · inbound

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs cites this paper.

Time-R1: Towards Comprehensive Temporal Reasoning in LLMs TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:00:12.943458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:00:12.943458Z digest=sha256:3d5e85b23a2f19b7380affc89da51ceae67bda7f0fba8237c906d02abf4b03c4

Observation a3e9f2f7-2aa0-4acc-8083-40f2d7188618 · inbound

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents cites this paper.

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:12.611434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:12.611434Z digest=sha256:790cfc0e7c2aeb5a12db60cd051e90d939532c6abe218c908784daaf0bfc1055

Observation 0da581eb-c4ba-4a0f-8fb7-430121e8e573 · inbound

LTLZinc: a Benchmarking Framework for Continual Learning and Neuro-Symbolic Temporal Reasoning cites this paper.

LTLZinc: a Benchmarking Framework for Continual Learning and Neuro-Symbolic Temporal Reasoning TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:03.602018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:03.602018Z digest=sha256:8850cdb290a075e8eea955a8acc94f5e979704e360d7989c38ff5aaf61f782b1

Observation 88a368e2-29cc-4909-b946-228946b9edf1 · inbound

AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models cites this paper.

AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:56:17.506210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T03:45:47.145667Z digest=sha256:10e86dfa69426f66377bdb70c616bc7bb37d64f55d530da285c2e417c3049687

Observation 36ed7d88-5a66-4a35-ab7f-5dba122824aa · inbound

All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs cites this paper.

All Circuits Lead to Rome: Rethinking Functional Anisotropy in Circuit and Sheaf Discovery for LLMs TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 171

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:52:57.505992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-14T20:52:47.074893Z digest=sha256:c682b7902f8c7548a4ff7dfeb85bfe7ae593b911542a755aa04979efbb5c4352

Observation 90b7ed50-5e3d-436f-95e2-807e66059b66 · inbound

ClinQueryAgent: A Conversational Agent for Population Health Management cites this paper.

ClinQueryAgent: A Conversational Agent for Population Health Management TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:33:56.170844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T01:31:07.031424Z digest=sha256:716b541ca4556f8353afa81c903595e16c3ce7c91fc9ecb038b337a0185478e1

Observation f66a79a5-6e20-4dd4-b220-68a79bf86bbe · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.280787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:9975cea84379c74775ad56aa4eaebb0d196810aa1b6606effc48e52b62e272f1

Observation 90141d3e-756f-4758-ab28-ae74d8cec817 · inbound

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation cites this paper.

MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T21:55:17.374149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T21:55:17.374149Z digest=sha256:4905b8b994736fdca1423991aa6fc3776bf64cf99f593ee5f07cf5b3c058019a