Pith. sign in

REVIEW 3 cited by

Back to the Future: Towards Explainable Temporal Reasoning with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01074 v2 pith:H26TCN2U submitted 2023-10-02 cs.CL cs.AI

classification cs.CLcs.AI
keywords temporalreasoningexplainablefuturepredictiontaskabilityevent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Temporal reasoning is a crucial NLP task, providing a nuanced understanding of time-sensitive contexts within textual data. Although recent advancements in LLMs have demonstrated their potential in temporal reasoning, the predominant focus has been on tasks such as temporal expression and temporal relation extraction. These tasks are primarily designed for the extraction of direct and past temporal cues and to engage in simple reasoning processes. A significant gap remains when considering complex reasoning tasks such as event forecasting, which requires multi-step temporal reasoning on events and prediction on the future timestamp. Another notable limitation of existing methods is their incapability to provide an illustration of their reasoning process, hindering explainability. In this paper, we introduce the first task of explainable temporal reasoning, to predict an event's occurrence at a future timestamp based on context which requires multiple reasoning over multiple events, and subsequently provide a clear explanation for their prediction. Our task offers a comprehensive evaluation of both the LLMs' complex temporal reasoning ability, the future event prediction ability, and explainability-a critical attribute for AI applications. To support this task, we present the first multi-source instruction-tuning dataset of explainable temporal reasoning (ExpTime) with 26k derived from the temporal knowledge graph datasets and their temporal reasoning paths, using a novel knowledge-graph-instructed-generation strategy. Based on the dataset, we propose the first open-source LLM series TimeLlaMA based on the foundation LlaMA2, with the ability of instruction following for explainable temporal reasoning. We compare the performance of our method and a variety of LLMs, where our method achieves the state-of-the-art performance of temporal prediction and explanation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Analyzing the Role of Context in Forecasting with Large Language Models

    cs.CL 2025-01 reject novelty 6.0 of 10

    A 614-question Metaculus benchmark shows news context improves LLM binary-forecast accuracy by 2 to 6 points and few-shot examples slightly hurt, but the setup risks leaking outcomes into the news articles.

  2. ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events

    cs.LG 2025-01 conditional novelty 6.0 of 10

    ChronoSense evaluates LLMs on all 13 Allen interval relations and temporal arithmetic, finding weak, inconsistent performance and signs of memorization across seven models.

  3. Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A survey of crowd-based future event prediction from text, plus a new eight-component data model for representing individual forecast statements.

Pith tools