Pith. sign in

Paper Citation Record · LEDGER

Pitfalls in Evaluating Language Model Forecasters

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 10 inbound Pith citation observations for arXiv:2506.00723.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00723 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:13.238164Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:31:04.385535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.596036Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy22
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be73dd28-81e0-42a8-8dc8-e01252230a94 · outbound

This paper cites Who predicted 2022?, 2023.

Pitfalls in Evaluating Language Model Forecasters Who predicted 2022?, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.610428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.622895Z digest=sha256:87eaa2c92c23605744c83dd7ae53e2276fa35d82be40b3bc262a0d2009e80af9

Observation 15730f2b-f060-411c-b62f-2c89ace57b7b · outbound

This paper cites A backtesting protocol in the era of machine learning.

Pitfalls in Evaluating Language Model Forecasters A backtesting protocol in the era of machine learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.485240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.696705Z digest=sha256:780e4421a1fab696380ec6e860a7a2f28e24b011e0ee6f7e519dd4e36afeb939

Observation ce37d9e5-8aa5-4eec-9649-fe766e973ef4 · outbound

This paper cites The probability of backtest overfitting.

Pitfalls in Evaluating Language Model Forecasters The probability of backtest overfitting

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.360878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.739151Z digest=sha256:f56b9e707c62fc027a86a005bdbc3867224501c2e6fb62408191b5caef0e28a2

Observation c33d3321-55c5-4d0e-8c73-437e02d93db8 · outbound

This paper cites Contra papers claiming superhuman AI forecasting, 2024.

Pitfalls in Evaluating Language Model Forecasters Contra papers claiming superhuman AI forecasting, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.233335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.846511Z digest=sha256:28b1e19824af961cdbafad12872c7ff73a9aeb16deed22313e88717af494d590

Observation de7a0767-3adf-40c2-a02e-3b53253eb1c4 · outbound

This paper cites Long-horizon predictability: a cautionary tale.

Pitfalls in Evaluating Language Model Forecasters Long-horizon predictability: a cautionary tale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:17.071504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.927181Z digest=sha256:6df11c8890373042ad30d48f8d9374b998e0a6c208c628868a3ef7c63bdb7ab7

Observation 4c3f244a-ed12-4e92-89a0-e1d1596abd4e · outbound

This paper cites AI forecasting bots incoming: comment section, 2024.

Pitfalls in Evaluating Language Model Forecasters AI forecasting bots incoming: comment section, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.940861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:09.989448Z digest=sha256:e91b1c496df3d82503d87680b42b8a2b311afe707d985a087269272d17aad42a

Observation 037507d0-6e59-480a-8909-10b6c301879e · outbound

This paper cites Point-in-time vs.

Pitfalls in Evaluating Language Model Forecasters Point-in-time vs

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.780192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.026963Z digest=sha256:447fad6f83f9df87c470a8b52c63c32d0f274f591aa189f64f3aeb78bd5f7110

Observation 81d3b48b-6987-454d-a686-529767dd38d6 · outbound

This paper cites Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle.

Pitfalls in Evaluating Language Model Forecasters Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.072436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.072436Z digest=sha256:f41b8b237ef6addd3cc06b05be07b0b758ddbf01527cf9f9533e898329ff4d2b

Observation 1034a219-7dc6-4945-bbc2-7064498400e2 · outbound

This paper cites Polymarket settles a market incorrectly -- again, 2024.

Pitfalls in Evaluating Language Model Forecasters Polymarket settles a market incorrectly -- again, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.689678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.160852Z digest=sha256:1ad5f252ed1ef3aae0da2f1424c5e46d32778f8187888613604d9cff659a6344

Observation 084f463d-3204-4746-9242-a9438e42eb8c · outbound

This paper cites Survivorship bias and mutual fund performance.

Pitfalls in Evaluating Language Model Forecasters Survivorship bias and mutual fund performance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.564852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.248562Z digest=sha256:e29ffc3abc321558e4a743d193341850427f454d7a3519231e8be8233dffabfc

Observation 93bbe428-7ea9-49c9-9544-5ce20d872745 · outbound

This paper cites Knowledge cutoff issues of GPT -4o regarding Phan et al.

Pitfalls in Evaluating Language Model Forecasters Knowledge cutoff issues of GPT -4o regarding Phan et al

Reference 11

Resolution
verified exact
raw_fallback, observed 2026-08-07T12:03:13.757585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.321182Z digest=sha256:b90584cd87ca0b936c117570a9b39de22cb199cab4b833d91b71ceec6da2e723

Observation 1c467d53-bd57-4b8a-9770-f932f675f4ec · outbound

This paper cites Approaching human-level forecasting with language models, 2024.

Pitfalls in Evaluating Language Model Forecasters Approaching human-level forecasting with language models, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.416133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.355084Z digest=sha256:d39bd5f9ed271c981cce5f045f2a80deb6b5c6e7bee16d8feb8fff63a7e531ce

Observation 2f2fb106-dc43-41f6-b143-ff665d282ee4 · outbound

This paper cites Introducing the SalemCSPi forecasting tournament, 2022.

Pitfalls in Evaluating Language Model Forecasters Introducing the SalemCSPi forecasting tournament, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.267368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.431995Z digest=sha256:2a6d9f4e227eb116aec19ccef6026bd1a923b4720f5faa718c5c912d889fae05

Observation 10b7605b-328b-414b-bfb2-c073521dcc7b · outbound

This paper cites The emerging science of machine learning benchmarks.

Pitfalls in Evaluating Language Model Forecasters The emerging science of machine learning benchmarks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.531282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.531282Z digest=sha256:ed9f1cf29f3eb582b217d6845d14f67479edd27ef2f06595aa54fe9529feee87

Observation 4206fdb4-2ab0-45c9-8bfb-4752d2fa8abb · outbound

This paper cites Reasoning and Tools for Human-Level Forecasting.

Pitfalls in Evaluating Language Model Forecasters Reasoning and Tools for Human-Level Forecasting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.604051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.604051Z digest=sha256:666709715b4a5da14faf755f4943649a4ef94e5263fbc8e308bd2a931d1635d0

Observation a23e1240-1b11-4fd8-b1a1-b396f8969ca0 · outbound

This paper cites asgeirtj/system\_prompts\_leaks/claude.txt, 2025.

Pitfalls in Evaluating Language Model Forecasters asgeirtj/system\_prompts\_leaks/claude.txt, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:16.056698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.709829Z digest=sha256:f244d9f2986c8e41fbe02f741ea96ba80ac50baac5cba458a8bd021f6f9d2f72

Observation 4261ec35-ba33-4678-b329-acfc6b24422d · outbound

This paper cites ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities.

Pitfalls in Evaluating Language Model Forecasters ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:10.828930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:10.828930Z digest=sha256:01ba5672afddf3f83bf2832d080a0438741e97ad5625936ef15ed127913dd76a

Observation 3af1629b-f9e9-4d1c-893f-d1f57703d2f4 · outbound

This paper cites an unresolved cited work.

Pitfalls in Evaluating Language Model Forecasters Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:03:15.733151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:10.987478Z digest=sha256:30758467d906d903a8c169b05db62f77ee3d655cf0f89c14f478e73041b68444

Observation c3a89c4a-108d-4a58-a55e-a451665968ab · outbound

This paper cites Questionable practices in machine learning.

Pitfalls in Evaluating Language Model Forecasters Questionable practices in machine learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:11.167415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:11.167415Z digest=sha256:1a861d46c2d7b2de63ca3d1d68c13cf0ce20d4181b81491808024688a1c8c275

Observation e4337b8f-8b2c-44c6-b27f-7d7c9406e5ab · outbound

This paper cites Acx2025 tournament, 2025.

Pitfalls in Evaluating Language Model Forecasters Acx2025 tournament, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:15.449049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.289628Z digest=sha256:8279fe2719480b47cef24ab74ec30b4bcdb777586a7f937e77b34f3e8f064e4a

Observation f1dd5fa7-feed-4d11-8c08-b93fb932308d · outbound

This paper cites Consistency Checks for Language Model Forecasters.

Pitfalls in Evaluating Language Model Forecasters Consistency Checks for Language Model Forecasters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:11.460739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:11.460739Z digest=sha256:496991e1f4192e6910f60f62e98ed2c6309130af4d6a9c2dbb59428717e46af7

Observation cb59e2c4-b5c7-4314-91db-9dc4371b1c52 · outbound

This paper cites LLMs are superhuman forecasters, 2024.

Pitfalls in Evaluating Language Model Forecasters LLMs are superhuman forecasters, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:15.261662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.576117Z digest=sha256:23b297616fccf996690c8f7017cffb969328dd332e09c4bf379bc9788b2521e1

Observation 03c3e534-b96b-4478-a828-ba15118f16e4 · outbound

This paper cites White House planning face-to-face meeting with Biden , Xi , 2023.

Pitfalls in Evaluating Language Model Forecasters White House planning face-to-face meeting with Biden , Xi , 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:15.152038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.683329Z digest=sha256:d12c2931f31d81e228e1609bc5e39acde720ff66fa3fcc6f5b48fa408135c60f

Observation 620e4317-0d9c-4058-8413-8af487b3e1ed · outbound

This paper cites Against calibration, 2023.

Pitfalls in Evaluating Language Model Forecasters Against calibration, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.980940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:11.840935Z digest=sha256:4de239bfd25736b2863b479a5f319c7cbd9c8b237ef9ccc9822aed1d9fb93ae0

Observation 293735f7-c0f8-4e09-975c-8e979dfeb9e0 · outbound

This paper cites Elicitation of personal probabilities and expectations.

Pitfalls in Evaluating Language Model Forecasters Elicitation of personal probabilities and expectations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.776963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.012201Z digest=sha256:9d6debc90acad38a280cae02ce9fc9d9eb13e5b1b220f435e67dcb0e1766aac4

Observation bdbc0f9b-f1ae-4efe-a513-e64993ed50e7 · outbound

This paper cites Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy.

Pitfalls in Evaluating Language Model Forecasters Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.605476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.184050Z digest=sha256:27b942e709a1e438da82c67de00da0d6d9fec5ca1e99db12aee1f8e22d0cc1e1

Observation 648b7389-af1c-4a8e-8eb0-89b27978a1cd · outbound

This paper cites Alignment Problems With Current Forecasting Platforms.

Pitfalls in Evaluating Language Model Forecasters Alignment Problems With Current Forecasting Platforms

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:13.513497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.352939Z digest=sha256:bd2466ac8717060406c634038013441feaeb279016e925ba212ee9a4d15c9eb7

Observation 25dddb8a-09c6-497c-b130-65a4dcef2246 · outbound

This paper cites Capital asset prices: A theory of market equilibrium under conditions of risk.

Pitfalls in Evaluating Language Model Forecasters Capital asset prices: A theory of market equilibrium under conditions of risk

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.404486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.469848Z digest=sha256:768f02d68afbc3ce977184a6213f632d74d823433fd9809281247415f2185110

Observation 74a58ae5-6806-4c10-8ac8-76e40365a06d · outbound

This paper cites Risk-adjusted performance of mutual funds.

Pitfalls in Evaluating Language Model Forecasters Risk-adjusted performance of mutual funds

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.286663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:12.611811Z digest=sha256:b66e421d7fab39088fc958290be15a974e4508d993ac702956a02e22376fc541

Observation b5daf215-3ab5-484f-bcd8-4303013fc21f · outbound

This paper cites PROPHET : An inferable future forecasting benchmark with causal intervened likelihood estimation.

Pitfalls in Evaluating Language Model Forecasters PROPHET : An inferable future forecasting benchmark with causal intervened likelihood estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:12.780222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:12.780222Z digest=sha256:f1a146cd29aa43215996c7615f949ef02c4c75a00bb755f84b7daf700f7911e9

Observation 5860ebad-17b9-471c-98f0-d79321ead696 · outbound

This paper cites an unresolved cited work.

Pitfalls in Evaluating Language Model Forecasters Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:12.926278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:12.926278Z digest=sha256:b16d4dc9eb58f9a9b77203c1b45f0ad8b573cd40c7beb4375e5336015c8c7f89

Observation 0b6b46fc-3ac1-4d28-be7b-f2bda0e20f4c · outbound

This paper cites Biden , Xi talks in san francisco, 2023.

Pitfalls in Evaluating Language Model Forecasters Biden , Xi talks in san francisco, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:14.088043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:13.000775Z digest=sha256:b3b477b2d65adfb4ac00023c276e0b3a731d214596ee239315a2cc2c0fd5bc93

Observation 45783a36-cbed-4005-8c2a-bace7637202f · outbound

This paper cites Haooowang/llm-knowledge-cutoff-dates, 2025.

Pitfalls in Evaluating Language Model Forecasters Haooowang/llm-knowledge-cutoff-dates, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:13.939127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T12:03:13.072803Z digest=sha256:287d7ff79bc6c69efc33322e41e50db0048151afaa1d5019a5096c83f49944c2

Observation 563017a7-53f1-4f38-8427-44b697a05264 · outbound

This paper cites Continual Learning for Large Language Models: A Survey.

Pitfalls in Evaluating Language Model Forecasters Continual Learning for Large Language Models: A Survey

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:13.143938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:13.143938Z digest=sha256:7a7095628c752eb76ea0477681b034cb5434f3e49f9f5e25cdccb91d17b297c1

Observation f9828a89-a1a4-4c55-8a63-67ecea5018c4 · outbound

This paper cites Forecasting Future World Events with Neural Networks.

Pitfalls in Evaluating Language Model Forecasters Forecasting Future World Events with Neural Networks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:13.238164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:13.238164Z digest=sha256:8c11eb9005974539ce2f8641277d6a3ec032c9e1bb5c2444c32875a407d21a6e

Pith citing papers

Observation ea63d7ae-fdf2-44c3-a6f4-daf17c074fd1 · inbound

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts cites this paper.

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Pitfalls in Evaluating Language Model Forecasters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:31.301050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:22:31.301050Z digest=sha256:84a25050db844b3eae9733e55188821e43403ca5c14e20ea7ae73c82b36e8020

Observation dc9bdf19-a263-4c7d-ace7-a5a1d7382167 · inbound

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs cites this paper.

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Pitfalls in Evaluating Language Model Forecasters

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:06:01.790093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:14:17.395227Z digest=sha256:2a8e5a5d0a75ef5b73377afc64d078437d7466b72df0a4e270501d7998fc1280

Observation d7b565c1-5342-4cde-8ccc-302464aeda7e · inbound

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs cites this paper.

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Pitfalls in Evaluating Language Model Forecasters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T19:33:20.775008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T19:33:20.775008Z digest=sha256:8bfb5ecf37930a0eb1ad39e06a621e80bdc7ae47b411a1fb259aa797354b9e61

Observation 4fddfe23-5ead-4288-9c8b-e0e9468a9566 · inbound

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking cites this paper.

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking Pitfalls in Evaluating Language Model Forecasters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:16.852549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:29:51.187292Z digest=sha256:288f4a8cf3b08da80296e43afdda9286d5396c7d3b6d11a2240e3c1945cf3d67

Observation 581b2b62-0eb3-4eb6-8def-1f9932c4e0ac · inbound

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most cites this paper.

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Pitfalls in Evaluating Language Model Forecasters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:24:38.182496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T05:24:10.513564Z digest=sha256:da206fbc15779f2666d2578b26ce0eb90f07bc4383d207fb707cdb79ef35a25e

Observation 1308da45-df6a-4fa8-9541-48f021f3391b · inbound

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most cites this paper.

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Pitfalls in Evaluating Language Model Forecasters

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:05:26.173990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T06:04:28.277835Z digest=sha256:092fcb528beaa4deea66bf65252b1124064e0173a7fcfa8380081e167bb3e35d

Observation 96a2d95c-533a-408f-a97b-695b4234adc6 · inbound

Verifiable Rewards for Calibrated Probabilistic Forecasting cites this paper.

Verifiable Rewards for Calibrated Probabilistic Forecasting Pitfalls in Evaluating Language Model Forecasters

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:18.746566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T19:45:23.721172Z digest=sha256:e0a276bdc5192af1c997abb85e0835c2ebca7a1f8dcd4ebb63fdd0a7089f83cd

Observation f722227c-1417-4514-a413-5aff1b4cdb29 · inbound

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry cites this paper.

Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry Pitfalls in Evaluating Language Model Forecasters

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.597385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T14:40:49.038578Z digest=sha256:11b4e8cb0f0b868b30bb443017386774568696bfb30ed0429f79f762f16b14be

Observation 7af55740-3045-4809-8594-c92e94661666 · inbound

Global Merger-Arbitrage Forecasting with Language Models cites this paper.

Global Merger-Arbitrage Forecasting with Language Models Pitfalls in Evaluating Language Model Forecasters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T14:36:29.370385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:36:29.370385Z digest=sha256:2839caa5be1fffb21e9c7ea654a1841abf0e205cf11dac84525b7bb1398c852b

Observation 7f3bb58b-25dc-4411-b7fa-83f44904762a · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Pitfalls in Evaluating Language Model Forecasters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.385535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.385535Z digest=sha256:ef275cbce18674a775893c4d3df2e1e103a188ef14b79c7afb49952ddbf63329