Pith. sign in

Paper Citation Record · LEDGER

Quantifying Variance in Evaluation Benchmarks

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2406.10229.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10229 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:07:23.253277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d45ccfe-0234-4aa6-855a-01ac7411d9b5 · inbound

Loss-to-Loss Prediction: Scaling Laws for All Datasets cites this paper.

Loss-to-Loss Prediction: Scaling Laws for All Datasets Quantifying Variance in Evaluation Benchmarks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:23.253277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:07:23.253277Z digest=sha256:833ce86e8156e704bc710014285e6d7240b546732b539987003178a1059a66b9

Observation dbf2f91f-8564-41de-9c90-e3d558a412da · inbound

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models cites this paper.

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models Quantifying Variance in Evaluation Benchmarks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:35:58.924818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:35:58.924818Z digest=sha256:65ecc8f894cadfff427746c7bbc0dee85dd9d3cc6e3a6f4da19df776dc50658d

Observation 5488d2ae-4920-460d-ac76-f9f5a45a5486 · inbound

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic cites this paper.

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic Quantifying Variance in Evaluation Benchmarks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:27.003075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:37:27.003075Z digest=sha256:6d4bd78efa2fbd21359a85878dc021ce5b7751a0d203d32db3b61eadc1213370

Observation 6cb9b301-1ad3-4483-95a2-a675fc773bc9 · inbound

HARP: A challenging human-annotated math reasoning benchmark cites this paper.

HARP: A challenging human-annotated math reasoning benchmark Quantifying Variance in Evaluation Benchmarks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.733078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.733078Z digest=sha256:5cdc3de33224168fc8905b0a76c142b4ee610e8a28ae8000e78365eb3aa15430

Observation 0b1e145f-0b0c-4cbf-938e-5aa28133c018 · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Quantifying Variance in Evaluation Benchmarks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.051526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:518ea1ba9d384037dfe41c50532afe8a6058b9cff6ec5170baff18b7025cf4a9

Observation b0ea9157-82e9-4f2d-b003-8f3f5de0b70d · inbound

A Study of LLMs' Preferences for Libraries and Programming Languages cites this paper.

A Study of LLMs' Preferences for Libraries and Programming Languages Quantifying Variance in Evaluation Benchmarks

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:55:12.391707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-22T22:53:16.951417Z digest=sha256:51604f87008699b994a9c64e1e2e664e17d18282f9bacf56b7088774706d56dc

Observation bc38713e-f5ab-4a85-86d9-645a25d093ce · inbound

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks cites this paper.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Quantifying Variance in Evaluation Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.617692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.617692Z digest=sha256:d6b8321eba88ea203f850b4d00192c1e6457b562aa16cc13d312bb5c75849da0

Observation bef963a3-d7a2-4043-aa23-ec9d4c5e89aa · inbound

Structure-Aware Fill-in-the-Middle Pretraining for Code cites this paper.

Structure-Aware Fill-in-the-Middle Pretraining for Code Quantifying Variance in Evaluation Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:23.000112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:23.000112Z digest=sha256:75d230fcdb93268e39452924761b4dcecf13d6b18ecbf5bdf5734a0c558a2715

Observation 5238fcf6-3d49-42d7-95d7-957763d3bf49 · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Quantifying Variance in Evaluation Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:10.760594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:10.760594Z digest=sha256:4d4c56b4a7c3d41ac7e7e6bb2119735a0a9e66de57313d5d5ab910f5448f6f14

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:637da8b2c3f93da8dc40af9165cb10f1e1047c1750b669ce70d9406edee3667c

Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.311443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.311443Z digest=sha256:d62660b3d1a1f45eb3d62543dd033216617128372b1626a0ef5dbc4c766016fd

Observation e2e17ed5-befc-4d79-9cb8-7275906204c8 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Quantifying Variance in Evaluation Benchmarks

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.439511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.439511Z digest=sha256:cabdf4f2901452a973865b20c82bae11865f10d6d32fce0f2f90c4782b8d418d

Observation 5c0953d6-f1b9-4059-9f27-cdadff721879 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Quantifying Variance in Evaluation Benchmarks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:47.532043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:47.532043Z digest=sha256:4e05fae18208d20fe6ff24cc897f43c86bc50f2a87ea8d66c4d04a71b97f3427

Observation 38ae454e-4f7b-403d-a66c-cc2cee262063 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs Quantifying Variance in Evaluation Benchmarks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.003804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:135db3dcd297164d1227fc6f664ec85196c4401276c4886f95e0538ba9fb13c0

Observation 976900a8-7ad3-4a51-b8f9-ca3797e671d2 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Quantifying Variance in Evaluation Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:55.247281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:55.247281Z digest=sha256:ed725b74aad29d8ef6c87daf743c503bc2cdb8352b4d00fc1097f4a08262c7df

Observation 9f3ce2f6-22e1-4faa-a8a2-1ac5f18decc0 · inbound

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning cites this paper.

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning Quantifying Variance in Evaluation Benchmarks

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:36:11.793707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T08:26:01.717280Z digest=sha256:68aaa19f13a1054bf8a992e821b1f0648fae183483f42edf128e7864ca85a2ce

Observation 507f4142-380b-435a-aab5-d0e2331614df · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models Quantifying Variance in Evaluation Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.592425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:fa20bd59926e7296f89d079f588947a0ff396d99e7796871e1305a8bdbcb3560

Observation 8bbafc1c-e21b-4b90-93a7-af375f2bddfd · inbound

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness cites this paper.

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.039511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-29T12:42:53.018183Z digest=sha256:6b7d1506e49f91378de8e006a0473f7178ad648439ad3c5a20190a5502ee0eb8

Observation 94b16516-8599-4888-a6ea-8900097459f0 · inbound

Resolution Diagnostics for Paired LLM Evaluation cites this paper.

Resolution Diagnostics for Paired LLM Evaluation Quantifying Variance in Evaluation Benchmarks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.254548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T07:15:27.402338Z digest=sha256:14e041e1601bcdc4b31c7f5c775323c0974d09ca000d369a74ceae694a0f2fc3

Observation 30606fe5-3738-4848-a408-cc5b42f6d912 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Quantifying Variance in Evaluation Benchmarks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:36:44.879438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:30c2f30d8b045645b0df56d18893bf63b19d44e904c64ef2bc0024f9b94a1086

Observation 93cf316e-d257-42ee-b052-52a7e1bda135 · inbound

Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty cites this paper.

Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty Quantifying Variance in Evaluation Benchmarks

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:09:00.963089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T23:02:58.564465Z digest=sha256:56668f9125baf2ceb14d1ce3f3456a129e725238c8afd96e6adc8d1a57acf664

Observation 6d6830a2-a2fa-452e-af6b-407fe573777f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.638431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:0e61d5d3c097d1ee6d9edb8cc5254d28d5b1362a14f1f50125a4137aba9f3772

Observation cffd2844-7096-4c59-a7f3-1a2b2b322c4f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.993727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:07b4cca4303c6cf485268cbb4990d49b24ff33d543d1ee2d3deb141e34eef2e0

Observation 7c690aa6-92c3-42d1-886a-5b6d566a38d6 · inbound

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise cites this paper.

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise Quantifying Variance in Evaluation Benchmarks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:58.886380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:29:58.886380Z digest=sha256:64d89c7cd4476f856e4a1261d981f6821def0e4230b0786353c8ce04ed7437e2

Observation 96468dd4-d445-4fba-b111-cd7f773baf37 · inbound

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks cites this paper.

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks Quantifying Variance in Evaluation Benchmarks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:07:19.640776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T14:58:38.688149Z digest=sha256:1aa79727416ac785d91cbe5f8f52410dd48e8a3365d7a2031ed1354af16ad6f1

Observation daf1a889-df9e-4597-bf7c-34b3b488455c · inbound

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally cites this paper.

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally Quantifying Variance in Evaluation Benchmarks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T12:17:16.303388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:17:16.303388Z digest=sha256:679286d6f122f40a7afdf3ae86e5fbec609ea14e570f0a9edfd21692d1de786e

Observation 8c876188-bedd-4059-9116-3271153f8224 · inbound

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks cites this paper.

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks Quantifying Variance in Evaluation Benchmarks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:00:48.305722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:00:48.305722Z digest=sha256:650498fdd05af2c2be949edf3b96eeeac8022b81aad5ebda1b2b49b4313eddc9

Observation c82b8d54-7f25-4e7d-a7c0-596b29880703 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Quantifying Variance in Evaluation Benchmarks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:44.479430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:44.479430Z digest=sha256:7d725299afbd22c36f29681672bb0b71a5d9b979c30303cb745345b288099ae7