Pith. sign in

Paper Citation Record · LEDGER

Pitfalls in Evaluating Language Model Forecasters

As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 10 inbound Pith citation observations for arXiv:2506.00723.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00723 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:13.238164Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:31:04.385535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.596036Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be73dd28-81e0-42a8-8dc8-e01252230a94 · outbound

This paper cites Who predicted 2022?, 2023.

Pitfalls in Evaluating Language Model Forecasters Who predicted 2022?, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.610428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.622895Z digest=sha256:338b2cfea36e48e43b39b12968f85242bef76a880badcff1d9cf190caef80717

Observation 15730f2b-f060-411c-b62f-2c89ace57b7b · outbound

This paper cites A backtesting protocol in the era of machine learning.

Pitfalls in Evaluating Language Model Forecasters A backtesting protocol in the era of machine learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.485240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.696705Z digest=sha256:d087e0c985959641dd2eb98f47c60f4f74d0c2554e3c34d81fe37a175c0af3ba

Observation ce37d9e5-8aa5-4eec-9649-fe766e973ef4 · outbound

This paper cites The probability of backtest overfitting.

Pitfalls in Evaluating Language Model Forecasters The probability of backtest overfitting

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.360878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.739151Z digest=sha256:c0d452b065f7d86ab2b2612c3f59185119f17be3dabc2cec27e5cb136857010a

Observation c33d3321-55c5-4d0e-8c73-437e02d93db8 · outbound

This paper cites Contra papers claiming superhuman AI forecasting, 2024.

Pitfalls in Evaluating Language Model Forecasters Contra papers claiming superhuman AI forecasting, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.233335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.846511Z digest=sha256:0f8ec01ee0ec52918cfe0ef0fbc59bb235ee5f6761b378c7bf12cbb145f47a38

Observation de7a0767-3adf-40c2-a02e-3b53253eb1c4 · outbound

This paper cites Long-horizon predictability: a cautionary tale.

Pitfalls in Evaluating Language Model Forecasters Long-horizon predictability: a cautionary tale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.071504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.927181Z digest=sha256:1f5397f71a1fb95bf6eaa4ae3782f724c4e3333541625561014b49a689f3d410

Observation 4c3f244a-ed12-4e92-89a0-e1d1596abd4e · outbound

This paper cites AI forecasting bots incoming: comment section, 2024.

Pitfalls in Evaluating Language Model Forecasters AI forecasting bots incoming: comment section, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.940861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.989448Z digest=sha256:c96b29a28b858c48a066540451542497392553fcb1a2a9e9882364eb44e91167

Observation 037507d0-6e59-480a-8909-10b6c301879e · outbound

This paper cites Point-in-time vs.

Pitfalls in Evaluating Language Model Forecasters Point-in-time vs

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.780192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.026963Z digest=sha256:7c36c9f9fa5931e2b304f3e467066374931b1742b05762c8c1bf4f16143e8ced

Observation 81d3b48b-6987-454d-a686-529767dd38d6 · outbound

This paper cites Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle.

Pitfalls in Evaluating Language Model Forecasters Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.072436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.072436Z digest=sha256:6196184426bbb352ee929318c2ac0f6a2d15ccb343d57499393d9f369fe9f084

Observation 1034a219-7dc6-4945-bbc2-7064498400e2 · outbound

This paper cites Polymarket settles a market incorrectly -- again, 2024.

Pitfalls in Evaluating Language Model Forecasters Polymarket settles a market incorrectly -- again, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.689678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.160852Z digest=sha256:168583605101f670eeee8b5f470a299717e3a94279579596551795fbd9f1ae01

Observation 084f463d-3204-4746-9242-a9438e42eb8c · outbound

This paper cites Survivorship bias and mutual fund performance.

Pitfalls in Evaluating Language Model Forecasters Survivorship bias and mutual fund performance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.564852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.248562Z digest=sha256:2f860ab06087863873bf7ba95ce777b1cff29bbf4c8e3e023d28560f71114385

Observation 93bbe428-7ea9-49c9-9544-5ce20d872745 · outbound

This paper cites Knowledge cutoff issues of GPT -4o regarding Phan et al.

Pitfalls in Evaluating Language Model Forecasters Knowledge cutoff issues of GPT -4o regarding Phan et al

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-07T12:03:13.757585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.321182Z digest=sha256:89df4a9565185031e1bf80c1794998dbcde7810b2d718638d500b5bba864878b

Observation 1c467d53-bd57-4b8a-9770-f932f675f4ec · outbound

This paper cites Approaching human-level forecasting with language models, 2024.

Pitfalls in Evaluating Language Model Forecasters Approaching human-level forecasting with language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.416133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.355084Z digest=sha256:7d378f6ad60eddbe3eaf0f3d60b1d7e5c00877754523df7c9074db11ab54779e

Observation 2f2fb106-dc43-41f6-b143-ff665d282ee4 · outbound

This paper cites Introducing the SalemCSPi forecasting tournament, 2022.

Pitfalls in Evaluating Language Model Forecasters Introducing the SalemCSPi forecasting tournament, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.267368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.431995Z digest=sha256:91c87df6a52b60d733c69cde1f1db686caab090cc30d419c42097a0f35b2f46e

Observation 10b7605b-328b-414b-bfb2-c073521dcc7b · outbound

This paper cites The emerging science of machine learning benchmarks.

Pitfalls in Evaluating Language Model Forecasters The emerging science of machine learning benchmarks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.531282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.531282Z digest=sha256:bc46c9e266907d210d4ce1ffd520a14f623022aaa90a7f85a6fc5163b1c63318

Observation 4206fdb4-2ab0-45c9-8bfb-4752d2fa8abb · outbound

This paper cites Reasoning and Tools for Human-Level Forecasting.

Pitfalls in Evaluating Language Model Forecasters Reasoning and Tools for Human-Level Forecasting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.604051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.604051Z digest=sha256:dfaa3aad1fdc36793c276385e4e507e65c9fefc7cd8c3eb43ddbeee4320f8e3d

Observation a23e1240-1b11-4fd8-b1a1-b396f8969ca0 · outbound

This paper cites asgeirtj/system\_prompts\_leaks/claude.txt, 2025.

Pitfalls in Evaluating Language Model Forecasters asgeirtj/system\_prompts\_leaks/claude.txt, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.056698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.709829Z digest=sha256:e0bda805486107c7c9adb2e125fdcaf8bfec74cb5511e2387fbc4b7fb05d03b5

Observation 4261ec35-ba33-4678-b329-acfc6b24422d · outbound

This paper cites ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities.

Pitfalls in Evaluating Language Model Forecasters ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.828930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.828930Z digest=sha256:bb75f0148d3126eb207e88f35a59cd41279d7fa98e79a866abfa5432264e345d

Observation 3af1629b-f9e9-4d1c-893f-d1f57703d2f4 · outbound

This paper cites an unresolved cited work.

Pitfalls in Evaluating Language Model Forecasters Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:03:15.733151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.987478Z digest=sha256:31de6580103d2bce6fa704fa5a49ed637515e5e72bb9b26b45fb01ddbc5d2f94

Observation c3a89c4a-108d-4a58-a55e-a451665968ab · outbound

This paper cites Questionable practices in machine learning.

Pitfalls in Evaluating Language Model Forecasters Questionable practices in machine learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:11.167415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:11.167415Z digest=sha256:bfaecd2bc56cce8b5e158ff6501718be55bdb48ae0ca7614c9cd25f737ab0a46

Observation e4337b8f-8b2c-44c6-b27f-7d7c9406e5ab · outbound

This paper cites Acx2025 tournament, 2025.

Pitfalls in Evaluating Language Model Forecasters Acx2025 tournament, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:15.449049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.289628Z digest=sha256:f8ab5010fef9b36794bdc0a7b3699aefbbdff2186eda3a959d19895a974be020

Observation f1dd5fa7-feed-4d11-8c08-b93fb932308d · outbound

This paper cites Consistency Checks for Language Model Forecasters.

Pitfalls in Evaluating Language Model Forecasters Consistency Checks for Language Model Forecasters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:11.460739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:11.460739Z digest=sha256:95e2564b6c5bfb5590efb927307e8c3baaed5aaa804fdd65ce7f390be919ee91

Observation cb59e2c4-b5c7-4314-91db-9dc4371b1c52 · outbound

This paper cites LLMs are superhuman forecasters, 2024.

Pitfalls in Evaluating Language Model Forecasters LLMs are superhuman forecasters, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:15.261662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.576117Z digest=sha256:e1cac69530ee113e36ed3a78a9b1659a893aaea9959c5399a9f62fcff093d51e

Observation 03c3e534-b96b-4478-a828-ba15118f16e4 · outbound

This paper cites White House planning face-to-face meeting with Biden , Xi , 2023.

Pitfalls in Evaluating Language Model Forecasters White House planning face-to-face meeting with Biden , Xi , 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:15.152038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.683329Z digest=sha256:6c04bfbfdf95fcefedca841a9b52ee93595bc0ee00b825911a47bfb18ab59573

Observation 620e4317-0d9c-4058-8413-8af487b3e1ed · outbound

This paper cites Against calibration, 2023.

Pitfalls in Evaluating Language Model Forecasters Against calibration, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.980940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.840935Z digest=sha256:d2fe008c7aed8d67aa106135651fb12027d8b8ebe27fe930ef9a186f479defca

Observation 293735f7-c0f8-4e09-975c-8e979dfeb9e0 · outbound

This paper cites Elicitation of personal probabilities and expectations.

Pitfalls in Evaluating Language Model Forecasters Elicitation of personal probabilities and expectations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.776963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.012201Z digest=sha256:8a62d2dc3492d2589f6167b6343856f87bec2d63b1655b92d6baf05e66ce6ff5

Observation bdbc0f9b-f1ae-4efe-a513-e64993ed50e7 · outbound

This paper cites Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy.

Pitfalls in Evaluating Language Model Forecasters Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.605476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.184050Z digest=sha256:5ca816be4a35951ac2f41054492261b86e6c070dcca5e4609b26d7634a668320

Observation 648b7389-af1c-4a8e-8eb0-89b27978a1cd · outbound

This paper cites Alignment Problems With Current Forecasting Platforms.

Pitfalls in Evaluating Language Model Forecasters Alignment Problems With Current Forecasting Platforms

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:13.513497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.352939Z digest=sha256:69eb924870f6c42970dbf121d1317d7935c2871da08e280895c465c57e61c048

Observation 25dddb8a-09c6-497c-b130-65a4dcef2246 · outbound

This paper cites Capital asset prices: A theory of market equilibrium under conditions of risk.

Pitfalls in Evaluating Language Model Forecasters Capital asset prices: A theory of market equilibrium under conditions of risk

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.404486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.469848Z digest=sha256:8ed85efcbdafd6ca0b3cfd1b601cc57c10305e3858c77c1b56852b0f62304124

Observation 74a58ae5-6806-4c10-8ac8-76e40365a06d · outbound

This paper cites Risk-adjusted performance of mutual funds.

Pitfalls in Evaluating Language Model Forecasters Risk-adjusted performance of mutual funds

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.286663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.611811Z digest=sha256:28d2b3b716e669b7312db46cf771e97d1303ae08db83ab72ef11f5a2a3aa6003

Observation b5daf215-3ab5-484f-bcd8-4303013fc21f · outbound

This paper cites PROPHET : An inferable future forecasting benchmark with causal intervened likelihood estimation.

Pitfalls in Evaluating Language Model Forecasters PROPHET : An inferable future forecasting benchmark with causal intervened likelihood estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:12.780222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:12.780222Z digest=sha256:c8322dfe132a5d1c154a9cd8a5344fd22db171201102d27f10e7c3e54f920c1f

Observation 5860ebad-17b9-471c-98f0-d79321ead696 · outbound

This paper cites an unresolved cited work.

Pitfalls in Evaluating Language Model Forecasters Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:12.926278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:12.926278Z digest=sha256:63523e87b305fe35089cf7ad87a55f4ff15c5d3fdc22156621da81d72529f23a

Observation 0b6b46fc-3ac1-4d28-be7b-f2bda0e20f4c · outbound

This paper cites Biden , Xi talks in san francisco, 2023.

Pitfalls in Evaluating Language Model Forecasters Biden , Xi talks in san francisco, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.088043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:13.000775Z digest=sha256:88d8c66dbd0e3d02c75653d496cb6deca041d30f570d573c5ca8dc3b73057960

Observation 45783a36-cbed-4005-8c2a-bace7637202f · outbound

This paper cites Haooowang/llm-knowledge-cutoff-dates, 2025.

Pitfalls in Evaluating Language Model Forecasters Haooowang/llm-knowledge-cutoff-dates, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:13.939127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:03:13.072803Z digest=sha256:5b31d6bf1eb66f178481b6b422cc9ba668b49e970645dfd70bc2a8b99ad6e98f

Observation 563017a7-53f1-4f38-8427-44b697a05264 · outbound

This paper cites Continual Learning for Large Language Models: A Survey.

Pitfalls in Evaluating Language Model Forecasters Continual Learning for Large Language Models: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:13.143938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:13.143938Z digest=sha256:e4d41198be84fad38bf7b534ee301ff75381764269c6ccc63765df7a93bed21b

Observation f9828a89-a1a4-4c55-8a63-67ecea5018c4 · outbound

This paper cites Forecasting Future World Events with Neural Networks.

Pitfalls in Evaluating Language Model Forecasters Forecasting Future World Events with Neural Networks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:13.238164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:13.238164Z digest=sha256:43a4e2e7bdde7a0477b5b4abc51c7371f192c0433b1e14adeb1593a6d4b2a9b6

Pith citing papers

Observation ea63d7ae-fdf2-44c3-a6f4-daf17c074fd1 · inbound

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts cites this paper.

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Pitfalls in Evaluating Language Model Forecasters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:31.301050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:22:31.301050Z digest=sha256:b14035cbf8fda46ec11f674270667c7ccae9db26c87ac86a837fad8276161e64

Observation dc9bdf19-a263-4c7d-ace7-a5a1d7382167 · inbound

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs cites this paper.

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Pitfalls in Evaluating Language Model Forecasters

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:06:01.790093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T04:14:17.395227Z digest=sha256:69ffd554013738c9c54158a81e78cd47c9b98312c8a168a06d90d9007d834a04

Observation d7b565c1-5342-4cde-8ccc-302464aeda7e · inbound

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs cites this paper.

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Pitfalls in Evaluating Language Model Forecasters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T19:33:20.775008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T19:33:20.775008Z digest=sha256:38e4db129015c6196c73e28dedb00e0983024deb7c70b97481dc68fca016deb9

Observation 4fddfe23-5ead-4288-9c8b-e0e9468a9566 · inbound

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking cites this paper.

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking Pitfalls in Evaluating Language Model Forecasters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:16.852549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T16:29:51.187292Z digest=sha256:c125a4806c2270a84255ab4ec28b286907ac5c3cc53545460992027a54845e0f

Observation 581b2b62-0eb3-4eb6-8def-1f9932c4e0ac · inbound

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most cites this paper.

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Pitfalls in Evaluating Language Model Forecasters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:24:38.182496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-22T05:24:10.513564Z digest=sha256:891c06f685740e6b0c2a1751e19308bcf53bdb4dc3f187717bb6b764f60a2e95

Observation 1308da45-df6a-4fa8-9541-48f021f3391b · inbound

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most cites this paper.

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Pitfalls in Evaluating Language Model Forecasters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:05:26.173990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-25T06:04:28.277835Z digest=sha256:9d4510f2ca341f59f87c8183f77cc36c1d33b4598bdf93026f1b8d495d86b43d

Observation 96a2d95c-533a-408f-a97b-695b4234adc6 · inbound

Verifiable Rewards for Calibrated Probabilistic Forecasting cites this paper.

Verifiable Rewards for Calibrated Probabilistic Forecasting Pitfalls in Evaluating Language Model Forecasters

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:18.746566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T19:45:23.721172Z digest=sha256:79edd92b5921604f44f94bf6c894cf96be615e1dd89f225477e333bb7ad87ce3

Observation f722227c-1417-4514-a413-5aff1b4cdb29 · inbound

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry cites this paper.

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry Pitfalls in Evaluating Language Model Forecasters

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.597385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T14:40:49.038578Z digest=sha256:861eb83b3520c332880544e4764d8b08a1772f721eb24c489c45065317d30047

Observation 7af55740-3045-4809-8594-c92e94661666 · inbound

Global Merger-Arbitrage Forecasting with Language Models cites this paper.

Global Merger-Arbitrage Forecasting with Language Models Pitfalls in Evaluating Language Model Forecasters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T14:36:29.370385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:36:29.370385Z digest=sha256:cf12d626a6e3e4484de8930ff95d8f13775df316b22d4062b7d867b46faa4603

Observation 7f3bb58b-25dc-4411-b7fa-83f44904762a · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Pitfalls in Evaluating Language Model Forecasters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.385535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.385535Z digest=sha256:99e96b358374d33c82d63c43ce2e02e6972d509db7275829ba3d91865fe21bf3