Pith. sign in

Paper Citation Record · LEDGER

Quantifying Variance in Evaluation Benchmarks

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2406.10229.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10229 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:07:23.253277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8d45ccfe-0234-4aa6-855a-01ac7411d9b5 · inbound

Loss-to-Loss Prediction: Scaling Laws for All Datasets cites this paper.

Loss-to-Loss Prediction: Scaling Laws for All Datasets Quantifying Variance in Evaluation Benchmarks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:07:23.253277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:07:23.253277Z digest=sha256:833ce86e8156e704bc710014285e6d7240b546732b539987003178a1059a66b9

Observation dbf2f91f-8564-41de-9c90-e3d558a412da · inbound

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models cites this paper.

Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models Quantifying Variance in Evaluation Benchmarks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:35:58.924818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:35:58.924818Z digest=sha256:3843cc417c4548edf618b9224ab046e661536e5bd97b690ec5a0f82416e710be

Observation 5488d2ae-4920-460d-ac76-f9f5a45a5486 · inbound

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic cites this paper.

Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic Quantifying Variance in Evaluation Benchmarks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:37:27.003075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:37:27.003075Z digest=sha256:6d4bd78efa2fbd21359a85878dc021ce5b7751a0d203d32db3b61eadc1213370

Observation 6cb9b301-1ad3-4483-95a2-a675fc773bc9 · inbound

HARP: A challenging human-annotated math reasoning benchmark cites this paper.

HARP: A challenging human-annotated math reasoning benchmark Quantifying Variance in Evaluation Benchmarks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.733078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.733078Z digest=sha256:5cdc3de33224168fc8905b0a76c142b4ee610e8a28ae8000e78365eb3aa15430

Observation 0b1e145f-0b0c-4cbf-938e-5aa28133c018 · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Quantifying Variance in Evaluation Benchmarks

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.051526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:9c322d267638691934d497c15cf0ace455f68d42c6a2fe1f23dee06e7ff5a0f4

Observation b0ea9157-82e9-4f2d-b003-8f3f5de0b70d · inbound

A Study of LLMs' Preferences for Libraries and Programming Languages cites this paper.

A Study of LLMs' Preferences for Libraries and Programming Languages Quantifying Variance in Evaluation Benchmarks

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:55:12.391707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T22:53:16.951417Z digest=sha256:c283199447f22c7986b245b87b118a6006021edd1f86fd1ad4be5c0285c62d6a

Observation bc38713e-f5ab-4a85-86d9-645a25d093ce · inbound

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks cites this paper.

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks Quantifying Variance in Evaluation Benchmarks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:55:56.617692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:55:56.617692Z digest=sha256:d6b8321eba88ea203f850b4d00192c1e6457b562aa16cc13d312bb5c75849da0

Observation bef963a3-d7a2-4043-aa23-ec9d4c5e89aa · inbound

Structure-Aware Fill-in-the-Middle Pretraining for Code cites this paper.

Structure-Aware Fill-in-the-Middle Pretraining for Code Quantifying Variance in Evaluation Benchmarks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:23.000112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:23.000112Z digest=sha256:75d230fcdb93268e39452924761b4dcecf13d6b18ecbf5bdf5734a0c558a2715

Observation 5238fcf6-3d49-42d7-95d7-957763d3bf49 · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Quantifying Variance in Evaluation Benchmarks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:10.760594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:10.760594Z digest=sha256:7a5eb329ce4d3d3219eaa56bf46c33eb68de918fa4756006ae357fee99a3f763

Observation 5705a10f-9f73-4de9-ac11-01fc409cea29 · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.922228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.922228Z digest=sha256:637da8b2c3f93da8dc40af9165cb10f1e1047c1750b669ce70d9406edee3667c

Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.311443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.311443Z digest=sha256:d62660b3d1a1f45eb3d62543dd033216617128372b1626a0ef5dbc4c766016fd

Observation e2e17ed5-befc-4d79-9cb8-7275906204c8 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Quantifying Variance in Evaluation Benchmarks

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.439511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.439511Z digest=sha256:cabdf4f2901452a973865b20c82bae11865f10d6d32fce0f2f90c4782b8d418d

Observation 5c0953d6-f1b9-4059-9f27-cdadff721879 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Quantifying Variance in Evaluation Benchmarks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:47.532043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:47.532043Z digest=sha256:9700894a0da6e1b07cf71f92265894742ef01b784fbff079f9258aa4d9231541

Observation 38ae454e-4f7b-403d-a66c-cc2cee262063 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs Quantifying Variance in Evaluation Benchmarks

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.003804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:5c561b406df473ca7b5af0f0ac971415954cfa7cdcdbeef41b18a5a982e46bb6

Observation 976900a8-7ad3-4a51-b8f9-ca3797e671d2 · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Quantifying Variance in Evaluation Benchmarks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:55.247281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:55.247281Z digest=sha256:ed725b74aad29d8ef6c87daf743c503bc2cdb8352b4d00fc1097f4a08262c7df

Observation 9f3ce2f6-22e1-4faa-a8a2-1ac5f18decc0 · inbound

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning cites this paper.

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning Quantifying Variance in Evaluation Benchmarks

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:36:11.793707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T08:26:01.717280Z digest=sha256:b9835344263989c31f24bc43111b2e156da8f43074b7168f0cf3a9e13574c463

Observation 507f4142-380b-435a-aab5-d0e2331614df · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models Quantifying Variance in Evaluation Benchmarks

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:13.592425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:b58cb5e271e2a959ff1635b506e6b69d9070501eea01c3067f3b42858686dd80

Observation 8bbafc1c-e21b-4b90-93a7-af375f2bddfd · inbound

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness cites this paper.

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness Quantifying Variance in Evaluation Benchmarks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.039511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T12:42:53.018183Z digest=sha256:1c7ff5c916e2f8f0d25749ad2f39897e2fe58caf6c4a516f7768f204d63d86ce

Observation 94b16516-8599-4888-a6ea-8900097459f0 · inbound

Resolution Diagnostics for Paired LLM Evaluation cites this paper.

Resolution Diagnostics for Paired LLM Evaluation Quantifying Variance in Evaluation Benchmarks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.254548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T07:15:27.402338Z digest=sha256:2b32f6f9134eefba1498e3f277651f75cbed3f58d47d97556680fb6977d1197a

Observation 30606fe5-3738-4848-a408-cc5b42f6d912 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Quantifying Variance in Evaluation Benchmarks

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:36:44.879438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:0c167c7ae8f4995fa590ceed050b9ab91978b899ea7574be47ce33955cafeff3

Observation 93cf316e-d257-42ee-b052-52a7e1bda135 · inbound

Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty cites this paper.

Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty Quantifying Variance in Evaluation Benchmarks

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:09:00.963089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T23:02:58.564465Z digest=sha256:cc8978ee685750ff72f2703307fb8d4bc441a715186c78713792c456a0c0ab9b

Observation 6d6830a2-a2fa-452e-af6b-407fe573777f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.638431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:01dbb52c10b5aace55e6d67e5b8448e23f7791561e23e6dc61050577a3d28f6c

Observation cffd2844-7096-4c59-a7f3-1a2b2b322c4f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models Quantifying Variance in Evaluation Benchmarks

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:23.993727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:259aac50761d7155a5acbaca81d2c9efd28c8fe57c6e76363be46c00cee8a127

Observation 7c690aa6-92c3-42d1-886a-5b6d566a38d6 · inbound

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise cites this paper.

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise Quantifying Variance in Evaluation Benchmarks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:58.886380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T02:29:58.886380Z digest=sha256:64d89c7cd4476f856e4a1261d981f6821def0e4230b0786353c8ce04ed7437e2

Observation 96468dd4-d445-4fba-b111-cd7f773baf37 · inbound

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks cites this paper.

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks Quantifying Variance in Evaluation Benchmarks

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:07:19.640776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-10T14:58:38.688149Z digest=sha256:e2352bcfbc3fc9ffae1dda5e6689f61bb686a5ab9212c524f6c6f3622056462c

Observation daf1a889-df9e-4597-bf7c-34b3b488455c · inbound

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally cites this paper.

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally Quantifying Variance in Evaluation Benchmarks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T12:17:16.303388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:17:16.303388Z digest=sha256:679286d6f122f40a7afdf3ae86e5fbec609ea14e570f0a9edfd21692d1de786e

Observation 8c876188-bedd-4059-9116-3271153f8224 · inbound

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks cites this paper.

Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks Quantifying Variance in Evaluation Benchmarks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:00:48.305722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:00:48.305722Z digest=sha256:c3eb4f19318dd250caffb63b767d70ffc1f0546ff8d625c8437b974b07693b65

Observation c82b8d54-7f25-4e7d-a7c0-596b29880703 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Quantifying Variance in Evaluation Benchmarks

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:44.479430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:44.479430Z digest=sha256:d6efd191bd396536bd727daba35c63fc49b2184dc47a3e833daaf45cf10d9b1a