Pith. sign in

Paper Citation Record · LEDGER

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

As of 19 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.02985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02985 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:31:04.475518Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa84cace-faa1-4012-bd2f-d475aca4e98c · outbound

This paper cites Look-ahead-bench: A standardized benchmark of look-ahead bias in point-in-time LLMs for finance.arXiv preprint arXiv:2601.13770,.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Look-ahead-bench: A standardized benchmark of look-ahead bias in point-in-time LLMs for finance.arXiv preprint arXiv:2601.13770,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.345601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.345601Z digest=sha256:3d4aff4bc6ac427e09491ee1e884a553620499ad0a36c4322d2a6bc98ff58f2c

Observation d6eff41d-17eb-4136-a1c6-7c3725515f50 · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.033066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.450644Z digest=sha256:1f43ecb79e68dce370c7de15c28ff58a02efd426bde7661d38fe9a800f18c89a

Observation 2578518a-f8ae-4224-b19e-e50c39d6686a · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:04.978959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.469239Z digest=sha256:33c41e650aa9280dd426147f753ab83b2ce8d700857c694996aea4c061b74ff4

Observation b21ca94e-9c29-49a0-85fc-b6201c56ca47 · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.090136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.429286Z digest=sha256:e137f5dff72caadc5f5b8a2d5964a39210c2348f3c3a30d0253b5ca2b45cd25f

Observation 57e002fd-a558-4c7a-acf4-fe54ebfd2a4b · outbound

This paper cites Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.364840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.364840Z digest=sha256:fd21b68f4013f6db54410681b781acc81f69600a25850046cfe8e729816efa31

Observation 88949916-c22c-436b-aa7b-4bae0dce7452 · outbound

This paper cites Table 4 is the map: each theoretical claim of Sections 3 to 5 and the experiment whose headline result carries it.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Table 4 is the map: each theoretical claim of Sections 3 to 5 and the experiment whose headline result carries it

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.072640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.436187Z digest=sha256:ee16b1bb5857820dad17bea03f6e752e64e8e958fe1dc8deb41179e26a0ddb4a

Observation 62b7d53d-17d2-42c7-90eb-eefe20d0813a · outbound

This paper cites Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:31:04.780051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.376212Z digest=sha256:4717095378fd0e7021d5617f073817fb5a453c744a9c459cdc40b6040a60188c

Observation 34e6a539-1707-4c8c-83d1-cb68f3072666 · outbound

This paper cites ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T04:31:04.765917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.379362Z digest=sha256:d99dc2471c9d7eff355671d43bc48fc5571f253dd4dbb2086027e9b55e463b8a

Observation f10255f9-bef7-4875-8408-e3c823662b4e · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Proving Test Set Contamination in Black Box Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.382539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.382539Z digest=sha256:8647b3b8bc4f69223a6b3a126d84fe1006f1726eb6b15253cbbad373c33fb99d

Observation 7f3bb58b-25dc-4411-b7fa-83f44904762a · outbound

This paper cites Pitfalls in Evaluating Language Model Forecasters.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Pitfalls in Evaluating Language Model Forecasters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.385535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.385535Z digest=sha256:2d72956cf219f5bdd87ebd6067efe78f276d8bda28d040460f00d40983f88196

Observation c1155d93-b9f5-4885-bd2a-7d502bc2c26c · outbound

This paper cites A Comprehensive Survey of Contamination Detection Methods in Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.388470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.388470Z digest=sha256:84a295b094b7adac97c36a0c1f3f0a2d4384c4b862def6e81aabfbd6171ba27f

Observation 765eb102-f051-4d11-bad6-a9e8a45ace8c · outbound

This paper cites Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.391835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.391835Z digest=sha256:45a810f88e543e67fb429860cb477e67367abc911deb19c811e7cacf4d60546f

Observation 7d15f498-479f-4ef2-a62c-23e136f7c027 · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.099410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.395288Z digest=sha256:181d725e98253f14d25758153ac007f028ab517cf192775a16dba7c95a5a93df

Observation 853a2f93-9126-4740-8770-3eedf6e9fa71 · outbound

This paper cites Quantifying the effect of test set contamination on generative evaluations.arXiv preprint arXiv:2601.04301,.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Quantifying the effect of test set contamination on generative evaluations.arXiv preprint arXiv:2601.04301,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.398711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.398711Z digest=sha256:df89888ff47f8e952a747d780b60604714e382ee9c42162840ababbae3281b0d

Observation bc60c097-8936-407a-b34d-3e600514dff0 · outbound

This paper cites Colin White, Samuel Dooley, Manley Roberts, Arka Pal, et al.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Colin White, Samuel Dooley, Manley Roberts, Arka Pal, et al

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.405329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.405329Z digest=sha256:c3ec5285c2546256ae3b7d3ce49fcc20a7bebd1303ac1a0251234cc931554a0c

Observation d4342b7b-c2a5-44d1-bc7c-7682006aeb82 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.408456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.408456Z digest=sha256:d86e08112944a3094d702363e5aefae3e4dad6b063d742ee823b93c72b2af7bf

Observation 772e1434-6b8e-4cfa-937d-653c40d82d9d · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Benchmark Data Contamination of Large Language Models: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.411806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.411806Z digest=sha256:023a339e551c33f0cc2e5af84691ad640f6f0ccfdf31c95f34e36ec10b021128

Observation b6d5b39f-7720-483f-a62c-bd18e18337aa · outbound

This paper cites DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.415388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.415388Z digest=sha256:9738157391a45ffb867160ad9ac9b02a26e4045e26f4de143aaa3127d65086ef

Observation 372bad00-77da-4547-981f-6351c5a94418 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.418737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.418737Z digest=sha256:77124d485b02a0a71895aed0aa8f37af24b89a31165ff06488f21986c9c4b5db

Observation 41e2cb85-5083-426e-b99e-3fe4ba1df6b5 · outbound

This paper cites Detecting Data Contamination in LLMs via In-Context Learning.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Detecting Data Contamination in LLMs via In-Context Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.422329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.422329Z digest=sha256:4a932c5646fd955f6d4fc6bcdf0eb298ea60904ed0aa90c4ac1f87094d759e13

Observation cf2f86c7-c6f3-492d-bfa4-9deb7ac76e34 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.425762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.425762Z digest=sha256:d5e01a4fed3da16d3801aeaffd8e2a7d63c36750851d317151136c00a3c06edb

Observation 8b6905b3-5d1a-4114-b422-92d06516544b · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.063406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.439874Z digest=sha256:8f442b755bb5e2e4f04c25d9ad0b3145c38eaa43d066880a34f4a30db06355cf

Observation db180450-b928-4133-9961-150004c49a3a · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.053534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.443472Z digest=sha256:bdeaa5f29272d8909462e82db67e19474cbf605dc90ac20730c0a34f0afdc591

Observation 59330655-316d-4c44-b6a5-75ccfaa33fd2 · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.043282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.447050Z digest=sha256:811c3fd14459392cd64c5e67c178ea2620adb5f37eedb66c172d707c42bed232

Observation 38c1209c-35c2-409b-8cae-1b0c3fa5eb29 · outbound

This paper cites Controls are MiniMax-M3 and Claude-Opus-4.7, whose January 2026 cutoffs leave no leakage discontinuity inside the tested window.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Controls are MiniMax-M3 and Claude-Opus-4.7, whose January 2026 cutoffs leave no leakage discontinuity inside the tested window

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.022778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.453942Z digest=sha256:23bf5889caf9bdee8407584ac016cc299790116cd484a99f7b9a107a6e4df9a4

Observation e5bd1bdb-b5eb-4a63-841f-2cd158db3496 · outbound

This paper cites Eachproblemcarriesitscontest release date, so a model can only have trained on a problem’s solution if the contest occurred before the model’s training cutoff.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Eachproblemcarriesitscontest release date, so a model can only have trained on a problem’s solution if the contest occurred before the model’s training cutoff

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.000807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.461264Z digest=sha256:5079249a655046eba69dc07b736d80dd876a85f8dcbf82da5ee2986f7839f3db

Observation 9e29a8fe-3126-4a30-9921-726385db38a5 · outbound

This paper cites paraphrases.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores paraphrases

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:04.968317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.472308Z digest=sha256:c85877377cddab27240503be8cb54a4d76264c53ee5487b46b2fd449c4bb8ebc

Observation 1686ff19-2009-461f-a0d6-df65bd193856 · outbound

This paper cites Treatment–control PRC excess by domain and pooled, under Platt and isotonic calibration.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Treatment–control PRC excess by domain and pooled, under Platt and isotonic calibration

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:04.955926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.475518Z digest=sha256:d187f06d6e2e55dea4a2c498ae29a93fa02ae19f32c22fec6ea678505e11e8c0

Observation e8e041c5-3ade-46a3-b70c-5f8d4a85de48 · outbound

This paper cites Because the leakage estimand is ajumprather than a level, protocol level effects cancel unless they vary sharply in time.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Because the leakage estimand is ajumprather than a level, protocol level effects cancel unless they vary sharply in time

Reference 910

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.011782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.457610Z digest=sha256:e4f8f71cc72a48ba07cf4b28cfdf4f505b70122db38d338554698514ff9a61c9

Observation 65fe6454-1ca9-49f2-83ea-6f93f1670b22 · outbound

This paper cites Chronologically Consistent Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Chronologically Consistent Large Language Models

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.368666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.368666Z digest=sha256:fcdc71c3ef7faf7730451748e89633dc8e2070f168d71bef798ecb2186a5843c

Observation c79b8ee5-2332-4943-a681-58e9ec2b1c02 · outbound

This paper cites Detecting lookahead bias in LLM forecasts.arXiv preprint arXiv:2512.23847,.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Detecting lookahead bias in LLM forecasts.arXiv preprint arXiv:2512.23847,

Reference 1987

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.361204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.361204Z digest=sha256:6172ef34f8077ad1aca87b50b085ebd504357ee000f8553a0578a1a1df973abb

Observation 5faa8654-e29e-4563-88e5-c5054eb22f46 · outbound

This paper cites The sharp-bounds framing of Theorem 2 follows partial identification (Manski, 2003; Imbens & Manski, 2004).

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores The sharp-bounds framing of Theorem 2 follows partial identification (Manski, 2003; Imbens & Manski, 2004)

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.081622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.432684Z digest=sha256:f714874a51dea9f4cee0c6b893733ac6a39448f81089337c9cfa629c85097a58

Observation c0328c06-8a03-4415-88fd-e754b03f040c · outbound

This paper cites Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.401902Z digest=sha256:8ecc8cb6d44a0156fc5d681b378775428b41ee788ac5c60f37f1301e263f14f2

Observation a0004f2f-ec5b-48d7-ae3c-becedf7c2733 · outbound

This paper cites Time machine GPT.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Time machine GPT

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.108653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.353655Z digest=sha256:bc934c1b5966b4582325cc205ffa54bd9a0e55de20be98436dedd3cf0b764a0a

Observation 5028d306-cf6b-4a91-acca-f6e123629764 · outbound

This paper cites Composition controls have cutoffsafterthe latest problem (Gemini-2.5-Pro, DeepSeek-R1, and the published 2025-cutoff pool).

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Composition controls have cutoffsafterthe latest problem (Gemini-2.5-Pro, DeepSeek-R1, and the published 2025-cutoff pool)

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:04.990275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.465422Z digest=sha256:9df8d689e9ecf9718230bb89cd0d9d290ef94916fc0b7b97918b41448a431ed3

Observation 3b6506ba-f4ae-43db-9bd6-7e0531d18027 · outbound

This paper cites Do Membership Inference Attacks Work on Large Language Models?.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Do Membership Inference Attacks Work on Large Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.357384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.357384Z digest=sha256:579d1fd0e3cffcf3bcdcf46e28c519ed1f9be576b0c2e78a7343c31168e7d170

Observation 8ba30b76-3c39-4208-82c9-d9ee8bfc293e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.372400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.372400Z digest=sha256:e268d578ba97f0b21f0ed3404987a5fbb7fabfd7f73bd534a37d3972c4bf3eed

Observation f3c8b989-f5b5-40e6-8984-7503679df9e1 · outbound

This paper cites Language models are few-shot learners.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Language models are few-shot learners

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.118711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T04:31:04.349752Z digest=sha256:208e557d0d7413c061b1fd8adaa9acc60f7082adbfcb3f78f7c9d0ed0122dca5

Pith citing papers

No inbound Pith citation observations are available.