Pith. sign in

Paper Citation Record · LEDGER

How Benchmark Prediction from Fewer Data Misses the Mark

As of 10 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2506.07673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07673 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:27.021795Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:40:23.654993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.920258Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved38
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26ce0089-b6d4-4d42-b7b8-b406dacf525f · outbound

This paper cites Jordan, and Tijana Zrnic.

How Benchmark Prediction from Fewer Data Misses the Mark Jordan, and Tijana Zrnic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.846205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.564352Z digest=sha256:e7d7c6ef4a4d4403770358b38c82c182ac2c94dd1ff173b721c384a36f329e66

Observation 12e946a6-75f0-4505-a47d-68740937a7ac · outbound

This paper cites PPI++: Efficient Prediction-Powered Inference.

How Benchmark Prediction from Fewer Data Misses the Mark PPI++: Efficient Prediction-Powered Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.571156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.571156Z digest=sha256:4bb0662d2c3eedc5e76b06d36f215b0300586c03edd2d0650e0358471c9e09e8

Observation e8998634-c9ec-4efd-a751-da8bce42a030 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

How Benchmark Prediction from Fewer Data Misses the Mark Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.582523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.582523Z digest=sha256:00f3e618a318b4f90d46e178f5f423fddd721b4d1fdc1366ef939ad7eb106a37

Observation cb62aaaa-d323-4ab9-b403-cb5c3f267220 · outbound

This paper cites The fifth PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The fifth PASCAL recognizing textual entailment challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.595425Z digest=sha256:e972f5e2dde9c313598efc8b04c52e01dd47a1173e3284fc9cccc1de66ad3abe

Observation b456a683-73b7-4e2a-9d30-d37b78bf7543 · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.612857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.612857Z digest=sha256:69baaa5a14819523782bed2e5236cecd05aaa716de1f99fb37a3091d0b01dcca

Observation 4be3af6d-5e00-46bc-b5ee-b7b40d24cfd3 · outbound

This paper cites A singular value thresholding algorithm for matrix completion.

How Benchmark Prediction from Fewer Data Misses the Mark A singular value thresholding algorithm for matrix completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.629844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.629844Z digest=sha256:fc4812688d577494cbc7441600c353b5afa441e6d80af269fdc7af376ea522b0

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:c2788f1ec0672db36a2b9152b75032805d594544f7f8d084a97bb4089d383a1c

Observation 4900b0b7-bbd8-482c-b3ee-442914c09442 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

How Benchmark Prediction from Fewer Data Misses the Mark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.659435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.659435Z digest=sha256:312bc223d7236f4b5f2f2763df6fc9ff3cbc0996d7fa7bc8436131c94fd8c301

Observation 4a51dd97-39e8-48af-ba9d-61a528884761 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Benchmark Prediction from Fewer Data Misses the Mark Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.728273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.728273Z digest=sha256:18862a09ddffdf3254e3de0cd01d2a5998af97d53b9eb3cd2831d7c07f1c4d92

Observation aad63229-20d8-44ad-9299-46d04ceaf4c6 · outbound

This paper cites Computing the testing error without a testing set.

How Benchmark Prediction from Fewer Data Misses the Mark Computing the testing error without a testing set

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.823011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.804792Z digest=sha256:99649569810696b4211ec11ab487032743c7d5235edb7b4702cf8c3d05633b6f

Observation b98538e5-9647-4134-87d6-3619b1fe2cbf · outbound

This paper cites The PASCAL recognising textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The PASCAL recognising textual entailment challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.813136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.845227Z digest=sha256:6e3aa20d6ab14b30388d1759d4f204c2935005c74bf1185d91b334989c639978

Observation d22a3f24-6499-4e4b-86a0-4f7c3537c121 · outbound

This paper cites Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021.

How Benchmark Prediction from Fewer Data Misses the Mark Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.803019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.848483Z digest=sha256:4abd6ee98d9f28b4366ace325bf105f5b1692d0bf5428496a5150595f118929d

Observation 41110b6f-0e2e-4556-bd50-59291e0a0e13 · outbound

This paper cites Automatically constructing a corpus of sentential paraphrases.

How Benchmark Prediction from Fewer Data Misses the Mark Automatically constructing a corpus of sentential paraphrases

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.791643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.851663Z digest=sha256:dbc63f1fba8fd32f605e306675710e4c2d768fa533db7a26bb4f4a42d8739213

Observation 57d1e637-a851-45ce-ac17-aa6b56ab3582 · outbound

This paper cites Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data.

How Benchmark Prediction from Fewer Data Misses the Mark Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.779352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.854801Z digest=sha256:7894ebbae73cdc13f8ff2728cce1a29031bbfd68a3f7b815248e09d5f87e96d8

Observation dce72101-6bdb-446d-9e38-3fdc7196cc92 · outbound

This paper cites Facility location: concepts, models, algo- rithms and case studies.

How Benchmark Prediction from Fewer Data Misses the Mark Facility location: concepts, models, algo- rithms and case studies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.768782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.858336Z digest=sha256:684638ff855cf486b0706948ca4ddc402f3ff29b4b2c0370ee715ef7459efcb5

Observation 3eebed2e-4b99-46e3-95ad-6207737980b7 · outbound

This paper cites Open llm leaderboard v2.

How Benchmark Prediction from Fewer Data Misses the Mark Open llm leaderboard v2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.758038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.861417Z digest=sha256:2684628d13cc9a462e5a2ab8350f3cea5840f5bed86b836d166b4ea8fc11e56a

Observation d1bfc293-0b96-49ef-b95b-6e74709768c6 · outbound

This paper cites Challenges in evaluating AI systems, 2023.

How Benchmark Prediction from Fewer Data Misses the Mark Challenges in evaluating AI systems, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.747484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.864676Z digest=sha256:02537e4fa59c951563e26a9d6652f11395d5170acfb05f68eebfe0ea5863fe55

Observation fef75381-5e12-4c9b-8e76-4763ed271a4e · outbound

This paper cites The third PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The third PASCAL recognizing textual entailment challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.737314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.868362Z digest=sha256:23777aeed8fb7057572c909e1d3f7354bb8c2366f6c30bd9dad9dc4f0107dc05

Observation e92bd0f6-4ca0-4717-9731-1ee70e48c1e3 · outbound

This paper cites An introduction to the augmented inverse propensity weighted estimator.

How Benchmark Prediction from Fewer Data Misses the Mark An introduction to the augmented inverse propensity weighted estimator

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.872011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.872011Z digest=sha256:3de4fd10c2cf473376003c8ce8791d850ca87900d8aa515a490fa713d8f6644f

Observation f75e2726-15af-4092-bb23-e4c5e4ba78b6 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

How Benchmark Prediction from Fewer Data Misses the Mark Great Models Think Alike and this Undermines AI Oversight

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.874878Z digest=sha256:26aff5daaf731ce75d6ebfa2532e45f22e3bd232cd7a3bdea89caede2ad1517e

Observation 011eccdc-8773-4472-a298-376579d64cd4 · outbound

This paper cites A Survey on LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark A Survey on LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.877994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.877994Z digest=sha256:07d0cb8c5fa2f5c1894850c59b9cefa979b1f298ae86d866538d426769c28629

Observation 9a4172e1-1add-4cf7-9d9a-ff19aaaa9c42 · outbound

This paper cites Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N.

How Benchmark Prediction from Fewer Data Misses the Mark Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.719690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.880977Z digest=sha256:85ac80ca677c8c213726016c73d7aa1e69bbf06e5731a5e164520ae344ba2fde

Observation e236eb94-6ade-4c5c-8d7c-372701a630ef · outbound

This paper cites Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings.

How Benchmark Prediction from Fewer Data Misses the Mark Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.482516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.884061Z digest=sha256:e2c3c77716d16bc954326d61221535194ee67201b1dc0e41842b0e1c50d354aa

Observation 9d271dc8-53d5-4f79-a88a-b7ed61a6cf7f · outbound

This paper cites Test-Time Training on Nearest Neighbors for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Test-Time Training on Nearest Neighbors for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.887369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.887369Z digest=sha256:fb95f7302872c5c9d09932a6920e126a5edc38bbc7522abb3ffcd597288c85be

Observation 28058209-5f05-40f2-aee6-3bd86c0af9a6 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Massive Multitask Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.890744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.890744Z digest=sha256:d4e74b9eb61d47389b0fea5bd5d1c86cb71cebc4fbe363220b9a0842c107f598

Observation e3adc9fc-5581-4631-8821-dd64ac57f168 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.894434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.894434Z digest=sha256:68241c549cad01cb642ea51cef27c189d52aef35a0216c5da1cf20b971b17bd1

Observation f0802203-fea5-4bf8-8d0a-3f8e4ffbd548 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

How Benchmark Prediction from Fewer Data Misses the Mark Categorical Reparameterization with Gumbel-Softmax

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.897769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.897769Z digest=sha256:710b0ea83b0c2089c899a5edfa575c21835be0cd3c48905ccd3d0b6a82925ddd

Observation 87a8b5ae-eb96-4c08-9469-491f0d59427c · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

How Benchmark Prediction from Fewer Data Misses the Mark What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.901004Z digest=sha256:8304c2c7a63050dd88f0cad6af106c3e0212288e133b967ddf71270c91c91a71

Observation 2f0551a2-b906-4595-bfc7-fbf5e87d54f4 · outbound

This paper cites Scaling Laws for Neural Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.904300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.904300Z digest=sha256:7a2d47e7399b06839f5bbc3d53d2abc5abc6214306dfdce7fc51bf117532426a

Observation 31b45bc4-f92a-41a9-91b7-4d3ab770d763 · outbound

This paper cites Active testing: Sample- efficient model evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active testing: Sample- efficient model evaluation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.709413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.907237Z digest=sha256:8dc8154b2313926b1703520298e037bf4e788f693990b5810f76b9e15cb3bde7

Observation 0337a2e4-f01d-4d11-bdf5-60f2c848bf25 · outbound

This paper cites Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.405042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.909784Z digest=sha256:58e84aff1c0988b973d70ec21952ea1d2813a970e19cdd02792f16ffc3a2fa49

Observation 7fc985a9-dc69-49a8-8d29-7754f35819ff · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

How Benchmark Prediction from Fewer Data Misses the Mark Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.912846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.912846Z digest=sha256:6bf5828d0555b31d58e641853d5df2b40d74cb2637826c437973241a9c8b48cf

Observation 75fea7a5-cf22-4c32-a251-182db853b17e · outbound

This paper cites Active Evaluation Acquisition for Efficient LLM Benchmarking.

How Benchmark Prediction from Fewer Data Misses the Mark Active Evaluation Acquisition for Efficient LLM Benchmarking

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.915668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.915668Z digest=sha256:b64f1fed84403b61d816ab00dcd588139b40e300b1eac002f23268d9662c97de

Observation 423e05f4-9dde-4337-9584-8a968046cef3 · outbound

This paper cites Manning, Christopher R’e, Diana Acosta-Navas, Drew A.

How Benchmark Prediction from Fewer Data Misses the Mark Manning, Christopher R’e, Diana Acosta-Navas, Drew A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.692361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.919274Z digest=sha256:307c1792a35863ad30ffd5960cfd04001c376ff1bc7d5b07cf85ddbe8db1eda5

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:363ce0deab0e1339b13705fd1ba70c71575cac83e796903e9a3186bce34ee3b9

Observation c16dfc68-05f7-407b-8512-b586c565db80 · outbound

This paper cites Model Similarity Mitigates Test Set Overuse.

How Benchmark Prediction from Fewer Data Misses the Mark Model Similarity Mitigates Test Set Overuse

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.367042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.925314Z digest=sha256:695ad174f09f207a9ceff44565c74793a1cb6eca12c9fb15b12f476146e7c961

Observation ef43aa4b-2f20-4461-a61c-8771e5acea8b · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

How Benchmark Prediction from Fewer Data Misses the Mark Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.682228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.928473Z digest=sha256:1fc0e10fd373a42d9d1cec04c1ac3168a013356a6d3d207958eac98deafa3904

Observation 080a3031-0e61-44d9-a8f4-1d99a863ec85 · outbound

This paper cites How predictable is language model benchmark performance?.

How Benchmark Prediction from Fewer Data Misses the Mark How predictable is language model benchmark performance?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.931350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.931350Z digest=sha256:a4da0f5bcb8c6d719d78f7c4be493d57f845297e324eaa608d2d5149a7aee046

Observation bf06c144-35d8-4081-ae7f-9f0e12da39ea · outbound

This paper cites PredictaBoard: Benchmarking LLM Score Predictability.

How Benchmark Prediction from Fewer Data Misses the Mark PredictaBoard: Benchmarking LLM Score Predictability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.934415Z digest=sha256:f6c0051e5a9024f71f8b7f527e224661a7369639b1ee051cf231e66e6dc03c1d

Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.937721Z digest=sha256:d236ffcc19d3c31f4aecb97be5c95515f7e5b9a561fcee7dca7bfcf8c9635329

Observation 7d31cfa5-4838-48d6-b3ba-5e84408948d3 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

How Benchmark Prediction from Fewer Data Misses the Mark tinyBenchmarks: evaluating LLMs with fewer examples

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.940907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.940907Z digest=sha256:a079c136ab47e09a248e345b8800c567892dfd9dc39bbe5521a662af29371b91

Observation 98b3f156-3235-4eb2-82b4-51d6389c886e · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

How Benchmark Prediction from Fewer Data Misses the Mark Efficient multi-prompt evaluation of LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.944282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.944282Z digest=sha256:28465a1cb408b84e13f1d0ca9b891c018b7031568c1cb41be23934c432508bcb

Observation c2570af8-3b21-4c0a-af04-5be2ba7ee595 · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

How Benchmark Prediction from Fewer Data Misses the Mark SQuAD: 100,000+ questions for machine comprehension of text

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.671738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.947725Z digest=sha256:1a26596e1d4a2104cbd4fdd41834c05611cdb2e5ace165c4221ad9943dab4765

Observation ed27291c-1724-485f-8517-1abc5e842edb · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.950951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.950951Z digest=sha256:29bde69f1eb38a292e2a464868c8b86f1e2fbb5c7a2bf3bba97d7b365d16b8f6

Observation 2f7c9bc9-ee91-4177-a4b8-dd0edfd42832 · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.

How Benchmark Prediction from Fewer Data Misses the Mark Semiparametric efficiency in multivariate regression models with missing data

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.954437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.954437Z digest=sha256:e50771e39b35edd499780e2543b40c3c3a06d05c1ae89031854a6292e984090d

Observation 2738fefe-3532-4e29-96a8-666d66c99118 · outbound

This paper cites Lalor, Robin Jia, and Jordan L.

How Benchmark Prediction from Fewer Data Misses the Mark Lalor, Robin Jia, and Jordan L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.655964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.957403Z digest=sha256:78fef000353650da6014e797d41c3e3f0692bb9c8ea5367025fc18befc054804

Observation b515bf99-dcf7-4650-b11c-3290558b5a87 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

How Benchmark Prediction from Fewer Data Misses the Mark Observational Scaling Laws and the Predictability of Language Model Performance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.960631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.960631Z digest=sha256:2c2a37107c72aebf4daa638fcf6cb4dc830fca5b7b73038d6819a4d7a8fe5abf

Observation b60afbc5-1bd7-4340-ad53-90e5ba059210 · outbound

This paper cites Bernstein, Alexander C.

How Benchmark Prediction from Fewer Data Misses the Mark Bernstein, Alexander C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.646469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.963910Z digest=sha256:494e7d3c1ca669803b849a94b17102c204e1a4664b9c5ab087dd661a2d38e087

Observation 6da416b0-5c7d-4d36-ae7e-bb0990390be9 · outbound

This paper cites Data distillation: A survey.

How Benchmark Prediction from Fewer Data Misses the Mark Data distillation: A survey

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.636832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.966967Z digest=sha256:cb395ab66e9d4f295eca979ea5be422079bb4b2a0f2e101d06f7ff526c598cb8

Observation 1e72658b-55a9-4ae9-a9f7-4698ca13d93c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.970380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.970380Z digest=sha256:4fe81c1185e475f45f2f656d217e84fd3ae02fddd19cac1b2a83d0e3466ee037

Observation 97e17da9-1936-4277-9f66-41108ef856ea · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

How Benchmark Prediction from Fewer Data Misses the Mark Recursive deep models for semantic compositionality over a sentiment treebank

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.627625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.973877Z digest=sha256:e45aca3623770cc8f8b0f48e1493f9631484a1c76b6ceb5bf0a59bdef744fd54

Observation 0ffbcf20-b91c-480a-9863-65da0b63a71c · outbound

This paper cites MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning.

How Benchmark Prediction from Fewer Data Misses the Mark MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.976895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.976895Z digest=sha256:6bfd2470f6b993fa201ecbf1abd59e214cbb68313843edeefc5fd2e8dcf4555a

Observation fad17933-bd87-4d4a-910f-86803c968dd7 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

How Benchmark Prediction from Fewer Data Misses the Mark Test-time training with self-supervision for generalization under distribution shifts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.980129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.980129Z digest=sha256:833f37fa5e2b927cce31ab59f4408e0c037645be6e55cd43656feaffa0b13210

Observation a1446e7b-b1a7-4633-a1a4-c2523c4af7a0 · outbound

This paper cites Le, Ed H.

How Benchmark Prediction from Fewer Data Misses the Mark Le, Ed H

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.610789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.983164Z digest=sha256:75088dd61537280f569d198968ebb0bcd58d8ba8eb85e7b11f8469cf20a2b230

Observation 6a3b525b-d889-45a6-8b7c-05356aa14d5e · outbound

This paper cites Four lectures on probabilistic methods for data science.

How Benchmark Prediction from Fewer Data Misses the Mark Four lectures on probabilistic methods for data science

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.141439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.986668Z digest=sha256:b3fb95bdbc5724224fe31600dce9717e605f3b9c63de48cffdf4c21842570f9e

Observation 5c68714b-37a1-4384-b026-c3e71eb8d677 · outbound

This paper cites Anchor Points: Benchmarking Models with Much Fewer Examples.

How Benchmark Prediction from Fewer Data Misses the Mark Anchor Points: Benchmarking Models with Much Fewer Examples

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.127869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.989961Z digest=sha256:e5566e2056bfafb39d654f5d651e91de9b292483fbab418d39c4159bfa71bb6d

Observation ead71a23-d406-4f36-b537-835d2941fa44 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.601129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:26.993144Z digest=sha256:e3f0b56ce107e165d4b42cc6228d49df12300f5fa4aaa1beb66d8cdcf892fa20

Observation aa39132f-b54c-4a1f-9498-b0cc5d5562c2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.996332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.996332Z digest=sha256:056772e11a7e5b448a06bfc2dfabb9eb4cf75573e70546f85b9cc703221fede4

Observation 1b64fdac-d565-4747-8e86-db0fdbf64704 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark Self-Preference Bias in LLM-as-a-Judge

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.999567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.999567Z digest=sha256:c4cf9db2bbb4c44a6629df7e9cd5c2377c894a69cb4347d96b4344faeadf1163

Observation a11929c3-9425-4a9c-bf3c-74488c36aad4 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.591844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:27.002653Z digest=sha256:770365df7341a22f4bf41ed6516d369163ff46f94addf94601bc1a4b451f126e

Observation 18b48d55-7d95-4d69-8dad-1b3795ef9cab · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

How Benchmark Prediction from Fewer Data Misses the Mark Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.005749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.005749Z digest=sha256:a75adf56e7822d101ae840924958770dcd3b01d68817e4833196f1d8b0244a5b

Observation dff6138b-4a97-4362-b2f6-3655a56c0bba · outbound

This paper cites Automatic evaluation of attribution by large language models.

How Benchmark Prediction from Fewer Data Misses the Mark Automatic evaluation of attribution by large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.582423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:27.008828Z digest=sha256:14a7b001c11489ac8d93dccf4f7d9513600a29ff9aaa3b4d011013cb36e35004

Observation 9cbd5450-d806-4bfd-8b28-5ab478823d84 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Instruction-Following Evaluation for Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.011892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.011892Z digest=sha256:e8bc2c106fcaf29291326a23dca726b26b08fe784f403ea4244997f87849ba77

Observation e8a454e4-155c-44a6-93ef-63ffc29a4e1c · outbound

This paper cites On Speeding Up Language Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark On Speeding Up Language Model Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.015178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.015178Z digest=sha256:98967674916a9372476805fc0088e16ee808a1d5adc24e0e9736666d844e41d5

Observation 8e104514-0e04-4d17-a20c-fc8609e6259e · outbound

This paper cites Probabilistic Bilevel Coreset Selection.

How Benchmark Prediction from Fewer Data Misses the Mark Probabilistic Bilevel Coreset Selection

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.059462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:27.018282Z digest=sha256:56b1187906158fd774c73f0df9d402f112a7c721b5dcc79daf8c227b4c08b7a0

Observation 1e3e299b-9295-4d96-8c3f-399e60260ee5 · outbound

This paper cites How to select datapoints for efficient human evaluation of nlg models?, 2025.

How Benchmark Prediction from Fewer Data Misses the Mark How to select datapoints for efficient human evaluation of nlg models?, 2025

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:34:27.572476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:34:27.021795Z digest=sha256:65ec6002bed3d5e33ee0c3e61e030bdff3f88a08f1fd69575e76fdea4d1bb8ff

Pith citing papers

Observation b27a9196-c3e0-4eb4-9adc-6acd9af2d40c · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees How Benchmark Prediction from Fewer Data Misses the Mark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:98e8894586471593f38d9a0656a97aa3648a6ef075fee5e60fd284f7d69e14f0

Observation 6992b4b1-fe2e-435b-85da-cacb5393970f · inbound

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking cites this paper.

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking How Benchmark Prediction from Fewer Data Misses the Mark

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T05:26:41.529866Z digest=sha256:e7fc5c737b39506c92f3f035b96f0bc99286eeb4f031e96ffaa3164deb62b3ac

Observation 470626fd-38bc-4101-a212-c01eccc4759e · inbound

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences cites this paper.

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences How Benchmark Prediction from Fewer Data Misses the Mark

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T11:24:01.547119Z digest=sha256:35e38d11e0c75a7b802faff339459f1f0a486c9a41ffb9a56a03fbb6b3fa1916

Observation e131494d-8e10-4c20-bf88-bbc35f486f3d · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research How Benchmark Prediction from Fewer Data Misses the Mark

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:1cbe0a77ca3d1a443441f2c0b21ad7e1c06fb21196155ea3b30a69fd735b05d0

Observation cf978fc5-69ca-48b4-a646-c55b8a9574af · inbound

You Don't Need to Run Every Eval cites this paper.

You Don't Need to Run Every Eval How Benchmark Prediction from Fewer Data Misses the Mark

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T08:23:12.145516Z digest=sha256:e2730b21bf0002952bd1ac0e8ff6c530b5ba89174552d5a876487eb7e52ad17f

Observation 497f56ee-452b-4e85-b947-3c7cecea0230 · inbound

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification cites this paper.

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification How Benchmark Prediction from Fewer Data Misses the Mark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T06:40:23.654993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:40:23.654993Z digest=sha256:20a626b5447ac0c539296b12696b9a4963fa21944348ed6422f74bc454fc0708