Pith. sign in

Paper Citation Record · LEDGER

How Benchmark Prediction from Fewer Data Misses the Mark

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 6 inbound Pith citation observations for arXiv:2506.07673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07673 v2

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:27.021795Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:40:23.654993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.920258Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact6
  • verified fuzzy21
  • unresolved38
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26ce0089-b6d4-4d42-b7b8-b406dacf525f · outbound

This paper cites Jordan, and Tijana Zrnic.

How Benchmark Prediction from Fewer Data Misses the Mark Jordan, and Tijana Zrnic

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.846205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.564352Z digest=sha256:454c5ebc53365b1e87085421744222d4957e6440ca1c339d3455bce5dd3036f1

Observation 12e946a6-75f0-4505-a47d-68740937a7ac · outbound

This paper cites PPI++: Efficient Prediction-Powered Inference.

How Benchmark Prediction from Fewer Data Misses the Mark PPI++: Efficient Prediction-Powered Inference

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.571156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.571156Z digest=sha256:45b964e44c3883e4c46423cc564be703a145bb11f3204f67b1107724bf31a8a3

Observation e8998634-c9ec-4efd-a751-da8bce42a030 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

How Benchmark Prediction from Fewer Data Misses the Mark Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.582523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.582523Z digest=sha256:4e3ca0ac1821a612ea586343e4b4806006cfe683f2192863de1b3827e5fa81b5

Observation cb62aaaa-d323-4ab9-b403-cb5c3f267220 · outbound

This paper cites The fifth PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The fifth PASCAL recognizing textual entailment challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.595425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.595425Z digest=sha256:7af1782318e3b281049fa7fe7b1ae2ffaf0320f8bd7c54684d9e3ce8a0d1d860

Observation b456a683-73b7-4e2a-9d30-d37b78bf7543 · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.612857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.612857Z digest=sha256:c6e3f0972d009f54a97f9b3ada3f41217f477f5ed03841302935df7f5da6b820

Observation 4be3af6d-5e00-46bc-b5ee-b7b40d24cfd3 · outbound

This paper cites A singular value thresholding algorithm for matrix completion.

How Benchmark Prediction from Fewer Data Misses the Mark A singular value thresholding algorithm for matrix completion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.629844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.629844Z digest=sha256:4b3cec14c40d2e34dbfbc6fed5778855abc710eb068b4e6060fed147ee75a5f1

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · outbound

This paper cites Humans or LLMs as the Judge? A Study on Judgement Biases.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:8a3529fb0b7bae3feb2a06a5e1f357b9ea2e8a126f7bd072ba70ce00b667ff54

Observation 4900b0b7-bbd8-482c-b3ee-442914c09442 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

How Benchmark Prediction from Fewer Data Misses the Mark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.659435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.659435Z digest=sha256:94b196d3d7d36286ca9723c5a38546a7e5754d81de5b85c0f1be77a0474f987b

Observation 4a51dd97-39e8-48af-ba9d-61a528884761 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Benchmark Prediction from Fewer Data Misses the Mark Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.728273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.728273Z digest=sha256:067782e9a95ad06956b52e7c68fb439844790a9a70af4758a959d872a874525b

Observation aad63229-20d8-44ad-9299-46d04ceaf4c6 · outbound

This paper cites Computing the testing error without a testing set.

How Benchmark Prediction from Fewer Data Misses the Mark Computing the testing error without a testing set

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.823011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.804792Z digest=sha256:52c55972ead59695fb6aa38beceb6c660d6fa7e09768cfc24bb9104778fea88d

Observation b98538e5-9647-4134-87d6-3619b1fe2cbf · outbound

This paper cites The PASCAL recognising textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The PASCAL recognising textual entailment challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.813136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.845227Z digest=sha256:ddf9455aa25e31913a17ef060bb44ecbb5346dd32591b3ef70894a9a8a06e499

Observation d22a3f24-6499-4e4b-86a0-4f7c3537c121 · outbound

This paper cites Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021.

How Benchmark Prediction from Fewer Data Misses the Mark Are labels always necessary for classifier accuracy evaluation? In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15069– 15078, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.803019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.848483Z digest=sha256:28ffd3f44f7fb8a9bcc707c74d75d7be4e64df3c3eda1b601a57f91c4987689a

Observation 41110b6f-0e2e-4556-bd50-59291e0a0e13 · outbound

This paper cites Automatically constructing a corpus of sentential paraphrases.

How Benchmark Prediction from Fewer Data Misses the Mark Automatically constructing a corpus of sentential paraphrases

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.791643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.851663Z digest=sha256:8dcff7cea43a1a262a96215a5ce8090c0e6e80edefd6ab89509ebe3b8722db0f

Observation 57d1e637-a851-45ce-ac17-aa6b56ab3582 · outbound

This paper cites Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data.

How Benchmark Prediction from Fewer Data Misses the Mark Limits to scalable evaluation at the frontier: Llm as judge won’t beat twice the data

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.779352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.854801Z digest=sha256:c6392fed81cbaa34a0e8137df3aa8a72b8922e975d9402abfcb1240f57b73f03

Observation dce72101-6bdb-446d-9e38-3fdc7196cc92 · outbound

This paper cites Facility location: concepts, models, algo- rithms and case studies.

How Benchmark Prediction from Fewer Data Misses the Mark Facility location: concepts, models, algo- rithms and case studies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.768782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.858336Z digest=sha256:492a67cdaaa4d6b6f637e3571e96180c9450754915b514907213126dba52b252

Observation 3eebed2e-4b99-46e3-95ad-6207737980b7 · outbound

This paper cites Open llm leaderboard v2.

How Benchmark Prediction from Fewer Data Misses the Mark Open llm leaderboard v2

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.758038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.861417Z digest=sha256:32e1d1f3e23d14aad58f0447ed11073de2815cf56635ed99388b43eeb307bb4e

Observation d1bfc293-0b96-49ef-b95b-6e74709768c6 · outbound

This paper cites Challenges in evaluating AI systems, 2023.

How Benchmark Prediction from Fewer Data Misses the Mark Challenges in evaluating AI systems, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.747484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.864676Z digest=sha256:b515d4a821b022457778957be8c1b9f5d5548673d0878f7d284cab353f04e339

Observation fef75381-5e12-4c9b-8e76-4763ed271a4e · outbound

This paper cites The third PASCAL recognizing textual entailment challenge.

How Benchmark Prediction from Fewer Data Misses the Mark The third PASCAL recognizing textual entailment challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.737314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.868362Z digest=sha256:111b68b32306e0cf0d9d8240055c6df38c38268513de05447a760e8329a9db0c

Observation e92bd0f6-4ca0-4717-9731-1ee70e48c1e3 · outbound

This paper cites An introduction to the augmented inverse propensity weighted estimator.

How Benchmark Prediction from Fewer Data Misses the Mark An introduction to the augmented inverse propensity weighted estimator

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.872011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.872011Z digest=sha256:869c00af8beb7dcd3078963c0e394d420259f5c1722c67c25a4e86fd55ef40de

Observation f75e2726-15af-4092-bb23-e4c5e4ba78b6 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

How Benchmark Prediction from Fewer Data Misses the Mark Great Models Think Alike and this Undermines AI Oversight

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.874878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.874878Z digest=sha256:19c53fce8ed8c2f4e27fa4e1f7d2db2b631cc20d821a34f431e760fe9f10a20f

Observation 011eccdc-8773-4472-a298-376579d64cd4 · outbound

This paper cites A Survey on LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark A Survey on LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.877994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.877994Z digest=sha256:353eee5b4ddbe72923bfe18b778338a9a9064e10db0d1e3f7900ae3f7e8b3367

Observation 9a4172e1-1add-4cf7-9d9a-ff19aaaa9c42 · outbound

This paper cites Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N.

How Benchmark Prediction from Fewer Data Misses the Mark Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.719690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.880977Z digest=sha256:9f86fe0cc03881ea2394bfef8cd8a5973f2a89dd5e95a0a9681415e5838045a7

Observation e236eb94-6ade-4c5c-8d7c-372701a630ef · outbound

This paper cites Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings.

How Benchmark Prediction from Fewer Data Misses the Mark Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.482516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.884061Z digest=sha256:6334ec71837256a92b11a2d5b3321c382148463b14eb1cd563005411f138ed05

Observation 9d271dc8-53d5-4f79-a88a-b7ed61a6cf7f · outbound

This paper cites Test-Time Training on Nearest Neighbors for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Test-Time Training on Nearest Neighbors for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.887369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.887369Z digest=sha256:acafef9fbee8bc20c90e000b223f393c9cbeb069a7c2bf9a38312c3e614c365f

Observation 28058209-5f05-40f2-aee6-3bd86c0af9a6 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Massive Multitask Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.890744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.890744Z digest=sha256:9a055b97f76b18b4a029326a3b2de26e964c1702ad61c91b82db5c1e2bc7c13f

Observation e3adc9fc-5581-4631-8821-dd64ac57f168 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

How Benchmark Prediction from Fewer Data Misses the Mark Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.894434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.894434Z digest=sha256:b268240c29a02f39efc840964280faaeb4c5fa6fdca1290d596df4d71d443f9c

Observation f0802203-fea5-4bf8-8d0a-3f8e4ffbd548 · outbound

This paper cites Categorical Reparameterization with Gumbel-Softmax.

How Benchmark Prediction from Fewer Data Misses the Mark Categorical Reparameterization with Gumbel-Softmax

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.897769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.897769Z digest=sha256:44f8eec8240432ca2d0c03233f6eeecb8cf0efd4682772c011ec06217d7e1459

Observation 87a8b5ae-eb96-4c08-9469-491f0d59427c · outbound

This paper cites What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams.

How Benchmark Prediction from Fewer Data Misses the Mark What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.901004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.901004Z digest=sha256:e1b90afa262e0fe13abe21d5117ea9129268ead2b4fb3feb69d9799760920c3d

Observation 2f0551a2-b906-4595-bfc7-fbf5e87d54f4 · outbound

This paper cites Scaling Laws for Neural Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling Laws for Neural Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.904300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.904300Z digest=sha256:cceeecd295ba82eca12e491bdfcc4dce6fe9b6d9f0262dfc8a9254e3a43fe69c

Observation 31b45bc4-f92a-41a9-91b7-4d3ab770d763 · outbound

This paper cites Active testing: Sample- efficient model evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active testing: Sample- efficient model evaluation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.709413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.907237Z digest=sha256:2211b14a6f818b785a22838de0f421867d9ef311947457c932eae3ab46c37a0c

Observation 0337a2e4-f01d-4d11-bdf5-60f2c848bf25 · outbound

This paper cites Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model Evaluation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.405042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.909784Z digest=sha256:5b1b1012a18beb99ebea6c42d1a1407444d0813841dc9cc6b84b5e78ec444fd4

Observation 7fc985a9-dc69-49a8-8d29-7754f35819ff · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

How Benchmark Prediction from Fewer Data Misses the Mark Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.912846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.912846Z digest=sha256:80b97df176c3ad13227ecc017e98d7128a2959b5ddb2fe32173bd52709c0bd3b

Observation 75fea7a5-cf22-4c32-a251-182db853b17e · outbound

This paper cites Active Evaluation Acquisition for Efficient LLM Benchmarking.

How Benchmark Prediction from Fewer Data Misses the Mark Active Evaluation Acquisition for Efficient LLM Benchmarking

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.915668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.915668Z digest=sha256:704743b64aed7265eb2d7fda272abfdd23aba22a3d529b06acce608c4d49cbfd

Observation 423e05f4-9dde-4337-9584-8a968046cef3 · outbound

This paper cites Manning, Christopher R’e, Diana Acosta-Navas, Drew A.

How Benchmark Prediction from Fewer Data Misses the Mark Manning, Christopher R’e, Diana Acosta-Navas, Drew A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.692361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.919274Z digest=sha256:b5091c43d1a102a40b1d622e2c964938a56d43ac71858a293e2241cb2f0e5111

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:83dce827d37c3adbcde8bc2e541130730b1ac53e096f36597d52c582e9998197

Observation c16dfc68-05f7-407b-8512-b586c565db80 · outbound

This paper cites Model Similarity Mitigates Test Set Overuse.

How Benchmark Prediction from Fewer Data Misses the Mark Model Similarity Mitigates Test Set Overuse

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.367042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.925314Z digest=sha256:ef248f3c5c886020c68b4839fe7a78ae5b2f5a938be49ca9aa0f1f89f83df161

Observation ef43aa4b-2f20-4461-a61c-8771e5acea8b · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

How Benchmark Prediction from Fewer Data Misses the Mark Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.682228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.928473Z digest=sha256:8f1d98bd5b41094862999bd7db43931bd9f04065b2de5d5e4c2cc6f835eb8cca

Observation 080a3031-0e61-44d9-a8f4-1d99a863ec85 · outbound

This paper cites How predictable is language model benchmark performance?.

How Benchmark Prediction from Fewer Data Misses the Mark How predictable is language model benchmark performance?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.931350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.931350Z digest=sha256:1f2cca93b89ef6219b76fee18d6cc7bc4df0bf63a0901ab3d92f4c59b186f128

Observation bf06c144-35d8-4081-ae7f-9f0e12da39ea · outbound

This paper cites PredictaBoard: Benchmarking LLM Score Predictability.

How Benchmark Prediction from Fewer Data Misses the Mark PredictaBoard: Benchmarking LLM Score Predictability

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.934415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.934415Z digest=sha256:b911f3636523c8ced681da2098dffe8f686f6047056f2981dbead2d370963d8a

Observation 30ef8940-83f3-4a49-8f21-17c2e82ebbe2 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

How Benchmark Prediction from Fewer Data Misses the Mark LLM Evaluators Recognize and Favor Their Own Generations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.937721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.937721Z digest=sha256:0ce694fc20b2c2e3e69483a9802d167c18c1286f6d0682c74fb9a3cb9aa1a1c3

Observation 7d31cfa5-4838-48d6-b3ba-5e84408948d3 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

How Benchmark Prediction from Fewer Data Misses the Mark tinyBenchmarks: evaluating LLMs with fewer examples

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.940907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.940907Z digest=sha256:4def5ef88f5ab9c5c5dca5278e39e21b041019a4f4acd483d0eea835029b40a2

Observation 98b3f156-3235-4eb2-82b4-51d6389c886e · outbound

This paper cites Efficient multi-prompt evaluation of LLMs.

How Benchmark Prediction from Fewer Data Misses the Mark Efficient multi-prompt evaluation of LLMs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.944282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.944282Z digest=sha256:bb05332c0424c1b76924974491d0c81547a49712d32c08bb4ce55e1c583f3520

Observation c2570af8-3b21-4c0a-af04-5be2ba7ee595 · outbound

This paper cites SQuAD: 100,000+ questions for machine comprehension of text.

How Benchmark Prediction from Fewer Data Misses the Mark SQuAD: 100,000+ questions for machine comprehension of text

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.671738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.947725Z digest=sha256:1c9418b95cddace31040335c43895737aff69ffde6f5f953e4c2fd0c2e581db8

Observation ed27291c-1724-485f-8517-1abc5e842edb · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.950951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.950951Z digest=sha256:c049ff6ff3238f2b2a94233c6d6b333609d15f9fe7c7e17a86f6ec3ca1d7c335

Observation 2f7c9bc9-ee91-4177-a4b8-dd0edfd42832 · outbound

This paper cites Semiparametric efficiency in multivariate regression models with missing data.

How Benchmark Prediction from Fewer Data Misses the Mark Semiparametric efficiency in multivariate regression models with missing data

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.954437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.954437Z digest=sha256:fde70623e79de6ec2aaac986a10805c24993daa47298014385643d3a7d32708c

Observation 2738fefe-3532-4e29-96a8-666d66c99118 · outbound

This paper cites Lalor, Robin Jia, and Jordan L.

How Benchmark Prediction from Fewer Data Misses the Mark Lalor, Robin Jia, and Jordan L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.655964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.957403Z digest=sha256:aa6873805593ace6efe47da1a8a65e72c253dd37c5c099a2b92c4ef1dc46c45c

Observation b515bf99-dcf7-4650-b11c-3290558b5a87 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

How Benchmark Prediction from Fewer Data Misses the Mark Observational Scaling Laws and the Predictability of Language Model Performance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.960631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.960631Z digest=sha256:78416ed843ddc0f7d1c1c59bcde87c09afbe2dde561971a039df5f76fafd3f5d

Observation b60afbc5-1bd7-4340-ad53-90e5ba059210 · outbound

This paper cites Bernstein, Alexander C.

How Benchmark Prediction from Fewer Data Misses the Mark Bernstein, Alexander C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.646469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.963910Z digest=sha256:2e49825a4b72d02b0c1e39da5380fa565530931a3866eae7a6aa8e28e084fd1e

Observation 6da416b0-5c7d-4d36-ae7e-bb0990390be9 · outbound

This paper cites Data distillation: A survey.

How Benchmark Prediction from Fewer Data Misses the Mark Data distillation: A survey

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.636832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.966967Z digest=sha256:46ef6c8224fdc602c79a4c874898aa0a0e0f5785f5b04792179b9b3627f07643

Observation 1e72658b-55a9-4ae9-a9f7-4698ca13d93c · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

How Benchmark Prediction from Fewer Data Misses the Mark Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.970380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.970380Z digest=sha256:f03eb1ae8327d625cc03e3fa894abef8ba8b615b04a5e2de4d145cb2f8203ddb

Observation 97e17da9-1936-4277-9f66-41108ef856ea · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

How Benchmark Prediction from Fewer Data Misses the Mark Recursive deep models for semantic compositionality over a sentiment treebank

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.627625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.973877Z digest=sha256:3aae75944077e8601b68f07db7fb08a399b7ec9906b3f45183b2e33eb821baa9

Observation 0ffbcf20-b91c-480a-9863-65da0b63a71c · outbound

This paper cites MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning.

How Benchmark Prediction from Fewer Data Misses the Mark MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.976895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.976895Z digest=sha256:af5d9efbaa6f403192b1384f4ee1996297490ae4d7c8fb87bb72cb1ca50c23a5

Observation fad17933-bd87-4d4a-910f-86803c968dd7 · outbound

This paper cites Test-time training with self-supervision for generalization under distribution shifts.

How Benchmark Prediction from Fewer Data Misses the Mark Test-time training with self-supervision for generalization under distribution shifts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.980129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.980129Z digest=sha256:fcecd0d06b1b36c550fc055b005782c8684344d42f790f89bc7a4b1b2d4f6b19

Observation a1446e7b-b1a7-4633-a1a4-c2523c4af7a0 · outbound

This paper cites Le, Ed H.

How Benchmark Prediction from Fewer Data Misses the Mark Le, Ed H

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.610789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.983164Z digest=sha256:355621419596ae8b034d14de0760b12e4f41c05b8cf755bbcceca5c6ada64faf

Observation 6a3b525b-d889-45a6-8b7c-05356aa14d5e · outbound

This paper cites Four lectures on probabilistic methods for data science.

How Benchmark Prediction from Fewer Data Misses the Mark Four lectures on probabilistic methods for data science

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.141439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.986668Z digest=sha256:d9d12858d418aac64e0fcd61b6f288d3d370db3edbb1dbbc5dbd1452a983f9dc

Observation 5c68714b-37a1-4384-b026-c3e71eb8d677 · outbound

This paper cites Anchor Points: Benchmarking Models with Much Fewer Examples.

How Benchmark Prediction from Fewer Data Misses the Mark Anchor Points: Benchmarking Models with Much Fewer Examples

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.127869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.989961Z digest=sha256:c3829cd6ccca0c5202efa03163092848cf656f9e739412a7d767a78f835097b6

Observation ead71a23-d406-4f36-b537-835d2941fa44 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.601129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:26.993144Z digest=sha256:f7410ba351df9602422955a0bd79f42f2d5097773743574b65a654a42113ceb0

Observation aa39132f-b54c-4a1f-9498-b0cc5d5562c2 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

How Benchmark Prediction from Fewer Data Misses the Mark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.996332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.996332Z digest=sha256:4bc767d9d7f78381874cfb3b9ef4ad14c3e2f65963ef8a4af10a0d2081b30384

Observation 1b64fdac-d565-4747-8e86-db0fdbf64704 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

How Benchmark Prediction from Fewer Data Misses the Mark Self-Preference Bias in LLM-as-a-Judge

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.999567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.999567Z digest=sha256:b57130e7c14b7d365a10aac18876c6f46ac417ec0abd287c7e8d704018502c41

Observation a11929c3-9425-4a9c-bf3c-74488c36aad4 · outbound

This paper cites an unresolved cited work.

How Benchmark Prediction from Fewer Data Misses the Mark Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:34:27.591844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:27.002653Z digest=sha256:9574e28f2440f5a20c6b782d7a998517a3b8d6b2c31d970fd88a2c1a6dd92155

Observation 18b48d55-7d95-4d69-8dad-1b3795ef9cab · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

How Benchmark Prediction from Fewer Data Misses the Mark Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.005749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.005749Z digest=sha256:c70ecc8e1e5691350099b78512ab9d0ae543acf8245f8410c5c32f1f82583f66

Observation dff6138b-4a97-4362-b2f6-3655a56c0bba · outbound

This paper cites Automatic evaluation of attribution by large language models.

How Benchmark Prediction from Fewer Data Misses the Mark Automatic evaluation of attribution by large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:34:27.582423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:27.008828Z digest=sha256:e0af228490f329e3c070eae0aa9088cfdcc30d80a73d9a6618f0414de902310a

Observation 9cbd5450-d806-4bfd-8b28-5ab478823d84 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

How Benchmark Prediction from Fewer Data Misses the Mark Instruction-Following Evaluation for Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.011892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.011892Z digest=sha256:d1b63c38178d077383aac9efa5e351a36681bda395db003bfb2adf1576b7ef13

Observation e8a454e4-155c-44a6-93ef-63ffc29a4e1c · outbound

This paper cites On Speeding Up Language Model Evaluation.

How Benchmark Prediction from Fewer Data Misses the Mark On Speeding Up Language Model Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:27.015178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:27.015178Z digest=sha256:4bc3980b2681947a9158f697a92a3661e4b05220eedd1cc0a252d6479584c778

Observation 8e104514-0e04-4d17-a20c-fc8609e6259e · outbound

This paper cites Probabilistic Bilevel Coreset Selection.

How Benchmark Prediction from Fewer Data Misses the Mark Probabilistic Bilevel Coreset Selection

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:34:27.059462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:27.018282Z digest=sha256:4d8a3131980bd13150e304038e0a4db3b3565152ecbfc1ee108e88f3f9eac38d

Observation 1e3e299b-9295-4d96-8c3f-399e60260ee5 · outbound

This paper cites How to select datapoints for efficient human evaluation of nlg models?, 2025.

How Benchmark Prediction from Fewer Data Misses the Mark How to select datapoints for efficient human evaluation of nlg models?, 2025

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:34:27.572476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:34:27.021795Z digest=sha256:289e0c18584a6939a33c56cfdcdce1bc7de363860ab077e57e5a8966dd52f6b5

Pith citing papers

Observation b27a9196-c3e0-4eb4-9adc-6acd9af2d40c · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees How Benchmark Prediction from Fewer Data Misses the Mark

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:dd8a824975be2741592123b6bee9f11a45114b73f75fcf292148cc0221e13add

Observation 6992b4b1-fe2e-435b-85da-cacb5393970f · inbound

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking cites this paper.

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking How Benchmark Prediction from Fewer Data Misses the Mark

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T05:26:41.529866Z digest=sha256:ca733faffe980b36ff750da7f6f04d4f15a88e66733831691c589149b50bb2c8

Observation 470626fd-38bc-4101-a212-c01eccc4759e · inbound

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences cites this paper.

FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences How Benchmark Prediction from Fewer Data Misses the Mark

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T11:24:01.547119Z digest=sha256:7b9ed6385e17682d61e79c58e9aad9575d309619285ce34e948ba561c32fed70

Observation e131494d-8e10-4c20-bf88-bbc35f486f3d · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research How Benchmark Prediction from Fewer Data Misses the Mark

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:6893514975f9ab76761bc2f29907c8ca2ab301cb25d96e3568df417cb4b4fb86

Observation cf978fc5-69ca-48b4-a646-c55b8a9574af · inbound

You Don't Need to Run Every Eval cites this paper.

You Don't Need to Run Every Eval How Benchmark Prediction from Fewer Data Misses the Mark

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-22T01:23:30.685296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T08:23:12.145516Z digest=sha256:af36646749963eb26c09729e293b5c9553e3173cc39943aab89857eb8782a391

Observation 497f56ee-452b-4e85-b947-3c7cecea0230 · inbound

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification cites this paper.

Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification How Benchmark Prediction from Fewer Data Misses the Mark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T06:40:23.654993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:40:23.654993Z digest=sha256:f0c80faad19a0cb4fb25ccc3d394bee7aa20a7faea0dfdd42d39e3ebf87ace50