Pith. sign in

REVIEW 5 cited by

MIRAI: Evaluating LLM Agents for Event Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.01231 v1 pith:3NBG74IY submitted 2024-07-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords agentseventsforecastinginternationalbenchmarkmiraiapisassessing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in Large Language Models (LLMs) have empowered LLM agents to autonomously collect world information, over which to conduct reasoning to solve complex problems. Given this capability, increasing interests have been put into employing LLM agents for predicting international events, which can influence decision-making and shape policy development on an international scale. Despite such a growing interest, there is a lack of a rigorous benchmark of LLM agents' forecasting capability and reliability. To address this gap, we introduce MIRAI, a novel benchmark designed to systematically evaluate LLM agents as temporal forecasters in the context of international events. Our benchmark features an agentic environment with tools for accessing an extensive database of historical, structured events and textual news articles. We refine the GDELT event database with careful cleaning and parsing to curate a series of relational prediction tasks with varying forecasting horizons, assessing LLM agents' abilities from short-term to long-term forecasting. We further implement APIs to enable LLM agents to utilize different tools via a code-based interface. In summary, MIRAI comprehensively evaluates the agents' capabilities in three dimensions: 1) autonomously source and integrate critical information from large global databases; 2) write codes using domain-specific APIs and libraries for tool-use; and 3) jointly reason over historical knowledge from diverse formats and time to accurately predict future events. Through comprehensive benchmarking, we aim to establish a reliable framework for assessing the capabilities of LLM agents in forecasting international events, thereby contributing to the development of more accurate and trustworthy models for international relation analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters

    cs.CL 2026-07 reject novelty 7.0 of 10

    When forecasters are barred from reading post-cutoff text, retrieval still improves Brier score on 8 of 9 LLMs, but only on markets Reddit had discussed beforehand; on speculative topics retrieval makes forecasts worse.

  2. Robust Human-AI Complementarity under Uncertainty

    cs.LG 2026-07 accept novelty 6.5 of 10

    Negative human-AI error correlation is required for robust complementarity under uncertainty about AI quality; current LLMs exhibit positive correlations on forecasting tasks.

  3. FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

    cs.CL 2026-01 conditional novelty 6.0 of 10

    FutureOmni, a 919-video, 1,034-question audio-visual future-forecasting benchmark, shows top MLLMs reach only 64.8% accuracy, and OFF tuning improves open models.

  4. Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A position paper advocating large-scale training of event forecasting LLMs, with proposals for label selection, counterfactual training data, auxiliary rewards, and multi-source datasets.

  5. Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Spatio-temporal foundation models are organized into a pipeline of data harmonization, model design, training, and adaptation, with a data property taxonomy for model selection.

Pith tools