Pith. sign in

Paper Citation Record · LEDGER

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

As of 15 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 5 inbound Pith citation observations for arXiv:2501.04234.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04234 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:43:06.616805Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:30.024509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:47:27.677181Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ee6cec1-6a14-4895-9c83-9fe2550131a8 · outbound

This paper cites GPT-4 Technical Report.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.499332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.499332Z digest=sha256:f89f059fd1dc8ccdd57458a441d1957d4c24b635d2966b7a78d9cb685b28aadc

Observation 6ce775bc-e399-466f-87b3-61d0fddea82c · outbound

This paper cites Bayesian inferences on uncertain ranks and orderings: Application to ranking players and lineups.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Bayesian inferences on uncertain ranks and orderings: Application to ranking players and lineups

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:07.007179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.504544Z digest=sha256:51a020ff61e119bb707f4dd7ec0987a1d8257d16046a999d6dde33e595cb15a8

Observation 4212ce72-189f-44e7-b5af-716038db8c01 · outbound

This paper cites Time for a change: A tutorial for comparing multiple classifiers through Bayesian analysis.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Time for a change: A tutorial for comparing multiple classifiers through Bayesian analysis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.994702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.508818Z digest=sha256:dec967e94e44b0aba202f416d49f4ba11cacfcc58c000e756c85c76599e16211

Observation c040c238-9c07-47ee-85a6-7f3826dee3b5 · outbound

This paper cites Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.982741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.513317Z digest=sha256:6d99688c40913c4b7e0db202e5c176838f0c62789aa4c5b43ab80bafb85abc81

Observation b24df635-f720-4d26-88fc-70487398d3e6 · outbound

This paper cites Accounting for variance in machine learning benchmarks.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Accounting for variance in machine learning benchmarks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.970718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.517895Z digest=sha256:b8d74a760c4650a6e2fcf97b7edd0f500ddf7cb120668b1b82311b4146a41fd1

Observation 5bce9be4-2f6d-4392-92b4-75330a7ec61e · outbound

This paper cites What are the best systems? N ew perspectives on NLP benchmarking.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks What are the best systems? N ew perspectives on NLP benchmarking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.957749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.522408Z digest=sha256:5dacd8f26d6593139f0d66b9712acee4b3015482498e70619a2bda6dfd811e2e

Observation 3c0f5b98-4385-4e21-9fbf-bbdac12fa03b · outbound

This paper cites The Benchmark Lottery.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks The Benchmark Lottery

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.526759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.526759Z digest=sha256:ef2f374cd8f5a1df083027e1db859eb34b0aa182f725ef7fd5707ce007f1cf51

Observation 86d6398b-416a-41b6-96a4-69e0c5ef7c5f · outbound

This paper cites Statistical comparisons of classifiers over multiple data sets.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Statistical comparisons of classifiers over multiple data sets

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.945859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.530628Z digest=sha256:e7324d7abe6c9209a36992de09841908e66a1bf0e1ed3538e5cb6979ba96a53b

Observation 2ae27b53-2cea-4c6a-a380-6436a2d15947 · outbound

This paper cites Bayesian aggregation of order-based rank data.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Bayesian aggregation of order-based rank data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.933812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.534246Z digest=sha256:bfa8b26683159fee81b9de6db1eb6e931c63d64972e8dd0bf4259e8f25ab5ff4

Observation afae105b-2648-43b3-b27e-d32223867a54 · outbound

This paper cites Approximate statistical tests for comparing supervised classification learning algorithms.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Approximate statistical tests for comparing supervised classification learning algorithms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.920637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.538052Z digest=sha256:8f68f5eb82b35fb017d6d7431ceb43bd30b3c9a5e395c581e4b51501531387a2

Observation fa0b86bb-d71f-46a2-9cf1-52fcf4c1d1a3 · outbound

This paper cites Statistical significance testing for natural language processing.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Statistical significance testing for natural language processing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.908376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.541830Z digest=sha256:663b86dba77ed881e87ae172e8f9b52a3bf79eef0e42b124d686b9b3fc122ba3

Observation a03c59fb-ff64-451b-b43c-6be47a494543 · outbound

This paper cites An Introduction to the Bootstrap.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks An Introduction to the Bootstrap

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.545426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.545426Z digest=sha256:42b322fed710f4ded4f73328ef6218d291f186b5b709476eb1ee5fa7ae407e90

Observation 6cf60cdd-e389-4cdb-9b7b-6d4f3e37775b · outbound

This paper cites The graphical presentation of a collection of means.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks The graphical presentation of a collection of means

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.889354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.549358Z digest=sha256:35508d1474d2ccef25ea4512dc0e0962cc3b284a18efe31d7df43b99734c295d

Observation eca38a4f-76de-4952-b108-f88de834458d · outbound

This paper cites League tables and their limitations: S tatistical issues in comparisons of institutional performance.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks League tables and their limitations: S tatistical issues in comparisons of institutional performance

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.878323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.552950Z digest=sha256:0f56a249dfdd6c9e1e693671f45e29ecf9cc04f129f687fcb7be67445ded8ae1

Observation d7da6371-c1dc-4d90-9473-3a247f0731aa · outbound

This paper cites Randomized significance tests in machine translation.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Randomized significance tests in machine translation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.866528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.556318Z digest=sha256:34ff1b5629b22a928e009fa893f771ab15ddf048854510ff001ffbf787f71a0d

Observation 651d2b0e-5c5b-4fd7-8a29-f72b5187211c · outbound

This paper cites Modeling the variability of rankings.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Modeling the variability of rankings

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.853267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.559927Z digest=sha256:b1980831a603bd9b6e5ebb2429f21bd27afe952cccbf36eb09df80f971ee58d3

Observation 2e5a9dd2-22fb-4e09-b1a6-6779713d3a43 · outbound

This paper cites Statistical comparisons of classifiers by generalized stochastic dominance.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Statistical comparisons of classifiers by generalized stochastic dominance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.841019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.563813Z digest=sha256:0aafdc1e78d8eac5b5f2819fe872c34c3f80862ae6f8df45a1c90552214de456

Observation 123283bf-cb14-447c-a32c-e9a9b52625be · outbound

This paper cites Active Bayesian assessment of black-box classifiers.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Active Bayesian assessment of black-box classifiers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.827661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.567510Z digest=sha256:f6b7fdd2dde1ca10cccd628e5b18bdf1f3a3fed2861be117c4aabc47e74f0d71

Observation 0c55c288-6523-4424-abc1-446d5bc5ca62 · outbound

This paper cites Theory of Point Estimation.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Theory of Point Estimation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.814772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.571098Z digest=sha256:e8e3405dda077e43975315b2aa6893d8e85c76726b890f5fd56147845158ac72

Observation acdeee00-44d8-4ef0-b3d4-c74dbc1a7e7b · outbound

This paper cites Bayesian analysis of rank data with covariates and heterogeneous rankers.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Bayesian analysis of rank data with covariates and heterogeneous rankers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.802155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.574833Z digest=sha256:46b3b2ef5ff5943b364dcdf3a88aed9ce3f5d2898bed49dbc02da98add978bba

Observation 000e0b1b-144c-424f-8db9-63accd29a40c · outbound

This paper cites Slice sampling.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Slice sampling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.579923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.579923Z digest=sha256:595a4028319306b2f2ff30c68b908f23033a631bf7321e792c11a401c7985bcd

Observation 8dca4664-9298-47f1-83da-5451cd1b1a9e · outbound

This paper cites Uncertainty in Ranking.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Uncertainty in Ranking

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.583624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.583624Z digest=sha256:6a16775520e0ea481b4004c79d615ec8fea39190b9ffc287b16bd3891d383c66

Observation 730be737-3b13-4fff-97fa-a7ed45edcdc4 · outbound

This paper cites On comparing classifiers: Pitfalls to avoid and a recommended approach.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks On comparing classifiers: Pitfalls to avoid and a recommended approach

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.782703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.587733Z digest=sha256:9576f3031a72e6fca8beb5118615b9a3f5dea92c922c1618589af87df61d97cd

Observation 2790ecde-d15d-4e1c-b6a3-1a9f61e32af1 · outbound

This paper cites an unresolved cited work.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Unresolved cited work

Reference 24

Resolution
verified exact
doi, observed 2026-08-10T21:43:06.651434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.591539Z digest=sha256:cb7995667c9dd9e27a73595a92184f852cecd4d469d1e712ba4c240656e48214

Observation a8292bee-30c6-4278-a8ea-7e7cb8bbfc9e · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.595642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.595642Z digest=sha256:ba6746df69c576f187d7cb2207207cd631ac823f735b420b23ec6ad9d128f1c4

Observation aafa2a5d-bd87-4698-895c-8ba22576d0b3 · outbound

This paper cites Classifier uncertainty: E vidence, potential impact, and probabilistic treatment.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Classifier uncertainty: E vidence, potential impact, and probabilistic treatment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.770303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.600120Z digest=sha256:e453c586751f4526a736247ffd3af2199e61bfcb144d22268f75cf3a27374a6a

Observation 56df0d95-102f-4570-a73d-4924f262a84b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.604550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.604550Z digest=sha256:78a6045e6afa96edac1dda75c014c70a09346cfafaa323093057c68e0cbe64da

Observation 6edff260-ef13-4ad0-9033-af101c7eba67 · outbound

This paper cites Follow the leader (board) with confidence: E stimating p-values from a single test set with item and response variance.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Follow the leader (board) with confidence: E stimating p-values from a single test set with item and response variance

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.757704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.608568Z digest=sha256:891130a58cce5e32ec981d29bc9a72d654cb8e996acf322f4814721215688874

Observation 86801887-c27a-478a-b0c3-9e707c3d3173 · outbound

This paper cites Confidence intervals for population ranks in the presence of ties and near ties.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks Confidence intervals for population ranks in the presence of ties and near ties

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:43:06.744142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T21:43:06.612790Z digest=sha256:44fff820258ae45828c9f1d374717b1c18e843a1a08509e995538019090e15e0

Observation de1b81ab-b83a-4acf-955c-cdf7dd101341 · outbound

This paper cites A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark.

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:43:06.616805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:43:06.616805Z digest=sha256:90256d05850b90556011949cd8b1ff5d7c12b09f9912de86f7a5f9f6917366d0

Pith citing papers

Observation 665aba2b-438d-435f-a8da-e66539ee39eb · inbound

Rethinking LLM Parametric Knowledge as Post-retrieval Confidence for Dynamic Retrieval and Reranking cites this paper.

Rethinking LLM Parametric Knowledge as Post-retrieval Confidence for Dynamic Retrieval and Reranking Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:30.024509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:30.024509Z digest=sha256:3be5b8e1a75100436a2980fe996ff088cc9ce861701ec6b8cc16243110bdb7c9

Observation d087617a-4524-4cd8-a690-bc1ca6097709 · inbound

Unstable Rankings in Bayesian Deep Learning Evaluation cites this paper.

Unstable Rankings in Bayesian Deep Learning Evaluation Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:31:15.186122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T08:34:44.254637Z digest=sha256:41076f668cc53bae4dc1aa320f413cd8aa51d2e568e69a6d69e933a708634c50

Observation 83b60adc-67a0-46a5-a3fc-5402f697f8ad · inbound

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning cites this paper.

A Tale of Two Variances: When Single-Seed Benchmarks Fail in Bayesian Deep Learning Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.788732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T08:26:01.717280Z digest=sha256:79e9cbf33cb2293747c01332af9f1ddecd91920c611c2e4eb92ec61bbe0e1a26

Observation c3a5bc77-6ef5-4ef8-bb6c-936e0b756488 · inbound

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation cites this paper.

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:47:27.679532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:54:22.974336Z digest=sha256:55bcbdcb498321eb92de97527647eb72c9c742a3c1516808e9825ae7e980adab

Observation 3514908c-2504-42d2-9ea3-a515cef37245 · inbound

Quantifying Ranking Uncertainty in LLM Benchmarks cites this paper.

Quantifying Ranking Uncertainty in LLM Benchmarks Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:26.760955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:26.760955Z digest=sha256:ffccd435cd64a0f18a91388631100cf48a80c423818e8c217f365a69eda29d12