Pith. sign in

REVIEW 5 cited by

TimeSeriesExam: A time series understanding exam

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14752 v1 pith:ORMNM27O submitted 2024-10-18 cs.AI cs.CL

classification cs.AIcs.CL
keywords seriestimemodelsabilityllmstimeseriesexamunderstandunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have recently demonstrated a remarkable ability to model time series data. These capabilities can be partly explained if LLMs understand basic time series concepts. However, our knowledge of what these models understand about time series data remains relatively limited. To address this gap, we introduce TimeSeriesExam, a configurable and scalable multiple-choice question exam designed to assess LLMs across five core time series understanding categories: pattern recognition, noise understanding, similarity analysis, anomaly detection, and causality analysis. TimeSeriesExam comprises of over 700 questions, procedurally generated using 104 carefully curated templates and iteratively refined to balance difficulty and their ability to discriminate good from bad models. We test 7 state-of-the-art LLMs on the TimeSeriesExam and provide the first comprehensive evaluation of their time series understanding abilities. Our results suggest that closed-source models such as GPT-4 and Gemini understand simple time series concepts significantly better than their open-source counterparts, while all models struggle with complex concepts such as causality analysis. We believe that the ability to programatically generate questions is fundamental to assessing and improving LLM's ability to understand and reason about time series data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HEARTS: Benchmarking LLM Reasoning on Health Time Series

    cs.LG 2026-02 conditional novelty 7.0 of 10

    A 110-task benchmark across 20 health signal modalities shows current LLMs underperform specialized models and depend on simple heuristics rather than robust time-series reasoning.

  2. TS-Haystack: A Multi-Task Retrieval Benchmark for Long-Context Time-Series Reasoning

    cs.LG 2026-02 unverdicted novelty 7.0 of 10

    TS-Haystack benchmark shows time-series language models degrade sharply on long contexts while an agentic retrieval system using classifier tools matches or beats them on 9 of 10 tasks.

  3. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.

  4. Towards Interpretable Time Series Foundation Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    After fine-tuning on 180 synthetic mean-reverting series annotated by a large multimodal model, small Qwen models can describe trend direction, noise intensity, and extremum location in natural language.

  5. ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large-Scale Multitask Dataset

    cs.CL 2025-06 conditional novelty 4.0 of 10

    ITFormer aligns time-series encoder outputs with a frozen large language model via lightweight instruction tokens, and EngineMT-QA provides a four-task aero-engine benchmark for temporal-textual QA.

Pith tools