Pith. sign in

Paper Citation Record · LEDGER

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

As of 10 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2606.13221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.13221 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:44:53.583327Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved47
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 332fc3a4-d04c-4ade-96cc-a3fca9fbf204 · outbound

This paper cites Chatbot arena: An open platform for evaluating LLMs by human preference.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Chatbot arena: An open platform for evaluating LLMs by human preference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:816a5287a23c81e39c8ce0c32238983015dfddcdd9233e052680a4d1df59365d

Observation 8e044b64-81d6-43ce-bd8e-16729cc11dc6 · outbound

This paper cites compar:ia: The french government’s llm arena to collect french-language human prompts and preference data.arXiv preprint arXiv:2602.06669, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation compar:ia: The french government’s llm arena to collect french-language human prompts and preference data.arXiv preprint arXiv:2602.06669, 2026

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:38:19.228275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:52a510f75772b97433346d43ad8e64b0ae1e551df17e381d8b2aba88ba3d8423

Observation e0cd7a58-f4f9-4677-8693-bf7b028fa164 · outbound

This paper cites Gonzalez, and Ion Stoica.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Gonzalez, and Ion Stoica

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:03d8394c10120b2db195b2dd3020a541e3ab72a37914711ab05ee96baceee165

Observation 26198120-a334-4633-8b87-1f7862e8d8c2 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:3d532caf4b571770aa9d4be37ba98e66cb0b09c32b4aaddc94bca7cf996cc785

Observation d25ab067-2bc7-436e-a5b4-26a286498fd9 · outbound

This paper cites Gonzalez, and Ion Stoica.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Gonzalez, and Ion Stoica

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:44c1499d90ff66651c75407ffa47db7f577a4a26be8c4fe9e8bd8f0a5e291766

Observation 0da20f28-d7e3-4feb-a94e-2faa6510c4b3 · outbound

This paper cites Judging the judges: A systematic study of position bias in LLM-as-a-judge.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Judging the judges: A systematic study of position bias in LLM-as-a-judge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:a0889a8c0673e375c68ebb352e3da82b3c6b62a4df122d21251d0044aeafe9d0

Observation 18dc3a9f-c380-4821-b52a-a260b1e58170 · outbound

This paper cites Bowman, and Shi Feng.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Bowman, and Shi Feng

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:da13a9d79ccd56da7e8bede1421d6e43cdc7c15370b29ded5c38a67cffd58b7d

Observation 0e423bfd-5fc1-4251-967e-e2b69928ed34 · outbound

This paper cites Tuning LLM judge design decisions for 1/1000 of the cost.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Tuning LLM judge design decisions for 1/1000 of the cost

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:0386d94c379146277c0341b52e4a8f94bf3a271d0603017ebbcdf56d44550ef4

Observation cff900a0-2b3d-48f5-9d9c-b2c090b71eea · outbound

This paper cites Investigating non-transitivity in llm-as- a-judge.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Investigating non-transitivity in llm-as- a-judge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:5eff5003a870f458cc39e6384764bda1fa0d9501978ab7521a1d906d1eb516d7

Observation 9a460dd4-141b-42ac-b99b-9fd44ffd2fa6 · outbound

This paper cites Mediocrity is the key for llm as a judge anchor selection, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Mediocrity is the key for llm as a judge anchor selection, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:e8df55ea5879c35449fa4bd2c44942d4098d63a485f90c7abe5c19ba8421bb09

Observation c3240093-8207-4472-9ca3-abe6c4b98f0a · outbound

This paper cites Auto-arena: Automating LLM evaluations with agent peer battles and committee discussions.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Auto-arena: Automating LLM evaluations with agent peer battles and committee discussions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:057df60d36d457f77bc16dc1730b9f6bfe5afec43b28ac20aaaa483d4d068c34

Observation cb73b3e7-dc2c-43de-b475-47805042e3ea · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:b492720b7b4b58171982046e686b5c0f9c50bc4fc90b7485e6469aaa990c404b

Observation 38f0f8a0-fe14-4c8b-bff0-ee588a8ff177 · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 13

Resolution
parse uncertain
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:5612172095b5e75d97e6381fe532307c8007bfedb4bbf8b12b89ac1bddcae5fa

Observation 5e4b3cac-3bdf-47bf-bf18-af8bc6373ecd · outbound

This paper cites Large language models are not fair evaluators.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Large language models are not fair evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:25dc7faca7653430238a651c50297cdf8c32e512545dab3d0170c31a3e6a1c29

Observation e5c3e027-68ba-4ed0-9d57-cff0ac6ad2c0 · outbound

This paper cites Explaining length bias in LLM-based preference evaluations.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Explaining length bias in LLM-based preference evaluations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:8aae85ed2782d5d34a02e325a0d5d1b2286f12e21f312768710be9016e3517d2

Observation e29c6ac0-9407-408b-9239-2740fd2972cf · outbound

This paper cites Self-preference bias in LLM-as-a-judge.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Self-preference bias in LLM-as-a-judge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:607f4ee96cc3bc795a0c00280b30c2e11e86b9113a6bbad4627067c222e85fe3

Observation bf4b7989-3532-4b83-aadd-812e8efe6dcc · outbound

This paper cites Huang, Yunyi Shen, Dennis Wei, and Tamara Broderick.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Huang, Yunyi Shen, Dennis Wei, and Tamara Broderick

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:a53d0833e8b0ea4f563f1839720e1f1e284aca807a5c11488423f082d9550197

Observation 232074ad-a7f8-4052-b479-3fd108fe4092 · outbound

This paper cites Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:33f8cdb2acd96c7d0964ff2a5d4900f0456bb227b452d26e450b81fc0d56e5fe

Observation 16c4037c-fbfe-4f18-a070-303e3438f950 · outbound

This paper cites Bridging human and LLM judgments: Understanding and narrowing the gap.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Bridging human and LLM judgments: Understanding and narrowing the gap

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:20028b8ae61ad0e8a29ebc797f18fc0aa76ca7c4e882a6df0ef01ed5e21544d3

Observation 0366021a-bfcb-4914-a86e-97315da3e86d · outbound

This paper cites Elo uncovered: Robustness and best practices in language model evaluation.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Elo uncovered: Robustness and best practices in language model evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:d654e40b4ca70d6eb3108ca768aa4c66ce154902aa37f328ad3ea6e566fd61f0

Observation 6cd050ff-d8b4-428f-b82a-48d818ba6959 · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:2265c1472408aedcb38acf3562476b27bded81f59cff6828a99e442f7ca65401

Observation 6a7b5359-ac23-443e-87e2-f5233b852812 · outbound

This paper cites am-ELO: A stable framework for arena-based LLM evaluation.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation am-ELO: A stable framework for arena-based LLM evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:fb441cfd50e4d2d4d2b043caf4cd58c1ea117b44f203ba6f98eb6959eebbc2dc

Observation 2d9ca594-51ce-4d0f-9e35-805e2df2d1fd · outbound

This paper cites Beyond bradley-terry models: A general preference model for language model alignment.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond bradley-terry models: A general preference model for language model alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:2dce50724e017eab71efeb60675c61217467a88a7355e914471cc62c1228fbff

Observation 94e2a045-408c-4600-a832-d00ba2f17176 · outbound

This paper cites Nonparamet- ric llm evaluation from preference data, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Nonparamet- ric llm evaluation from preference data, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:e325eb22dc3e15f2940a6974c46c5662c07a003a0b23b972b8cd15823b4b793d

Observation 7f42885f-eeda-46bd-8fb6-775bbd1e4598 · outbound

This paper cites Reward learning from preference with ties, 2024.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Reward learning from preference with ties, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:6b870701063958a1615dc96498bf710339254fd0dce9dddfaf3f0049aecfe476

Observation 48c8e517-8533-441b-ba31-e868e1d70622 · outbound

This paper cites Beyond binary preferences: A principled framework for reward modeling with ordinal feedback, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond binary preferences: A principled framework for reward modeling with ordinal feedback, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:44771127d9e5bfacbbf912c5368bcebd2068b3e4a4b9a3d4866fcd8235d8c13d

Observation c6af74d0-3e86-4514-bedd-defa168391ce · outbound

This paper cites Reward modeling with ordinal feedback: Wisdom of the crowd.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Reward modeling with ordinal feedback: Wisdom of the crowd

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:06c1920d57e6f0c35715a1e5672668721fd77d95dafd0204c10a0176010e8e24

Observation ecfd96d6-2165-4f74-bb7e-84e74a18b423 · outbound

This paper cites Improving LLM-as-a-judge inference with the judgment distribution.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Improving LLM-as-a-judge inference with the judgment distribution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:bd20a97c97adca22c18b2d1dcd606f45e1e93a61a052480ea9175d5f54103703

Observation 5040c1d6-e214-4a40-8d69-0e3d2b2baff1 · outbound

This paper cites Beyond single-point judgment: Distribution alignment for llm-as-a-judge, 2025.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond single-point judgment: Distribution alignment for llm-as-a-judge, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:f7c7862ab0ef990a8339a4e8b78e118e6cf5845e30c48f0293e1183172349dd6

Observation 0e95aabe-4f3a-4e77-bf38-fdeb4dd7545c · outbound

This paper cites Malin, and Yuan Xue.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Malin, and Yuan Xue

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:56cb1091c559afdffb0ad8e16858a4386e45876f22c2c06405af1a43adfc55b5

Observation 8d506c84-a716-4cfe-83f8-7f2b3ba1ef94 · outbound

This paper cites Beyond ordinal preferences: Why alignment needs cardinal human feedback, 2025.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond ordinal preferences: Why alignment needs cardinal human feedback, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:d4e41eea043626e07e4edb2f94185206d545d4e961b21ec7ba56e7b29e1c42b9

Observation 05e12e5c-bf5c-4b34-b1ba-833e28857041 · outbound

This paper cites LLM- rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation LLM- rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:f1bc998189dc1634e69bf43173e62c26462eee3f65eff4c3c69c5838f500ca90

Observation 2ae1ee6f-9199-4ccb-b18e-8b20ac202214 · outbound

This paper cites Quantitative llm judges, 2025.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Quantitative llm judges, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:c3d6d6ab938ecc2f237baccd4ac0f7a998414f46a45dffafa8f90cb08fc2bce6

Observation ce688e7d-e20c-47f6-acac-8fa7220632f1 · outbound

This paper cites Analyzing uncertainty of LLM-as-a-judge: Interval evaluations with conformal prediction.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Analyzing uncertainty of LLM-as-a-judge: Interval evaluations with conformal prediction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:58c0485ee0bb0a2bfbeb375a1bd3e2dd977cbf438107e8da4bebce554ab55575

Observation ee646bfb-db86-47f1-86a3-08c066eebebc · outbound

This paper cites Scope: Selective conformal optimized pairwise llm judging, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Scope: Selective conformal optimized pairwise llm judging, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:deec2ee310664bba210e9a8cfb4ee21bc526a9292f543fdb424e73c1b9ebfb8b

Observation 2ce3422c-5f94-46d0-8248-c00aaff8c46a · outbound

This paper cites Prediction-powered ranking of large language models.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Prediction-powered ranking of large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:4573e4c3f144a6bd400df3011e016f29441f909e78af274261c7315af287d235

Observation 024ef43c-f6b7-4b24-9a86-c86ee5107ad3 · outbound

This paper cites Alex Hofer, Bhuwan Dhingra, Amir Globerson, and William W.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Alex Hofer, Bhuwan Dhingra, Amir Globerson, and William W

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:2b3a33c21fcf5c6a809df0b88149b9289e94a306a65b6d82a4bde45d7c2f1a26

Observation ac81a5fe-a137-44d8-8a39-e35b414e02ad · outbound

This paper cites Adaptive prediction-powered autoeval with reliability and efficiency guarantees.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Adaptive prediction-powered autoeval with reliability and efficiency guarantees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:e6a0d9feb18814f4a0e0360c62c27e4b0940a6b2a1a002e33ca5c40711efd3cb

Observation 1fdc4bee-44d3-48ef-aad8-56296cfc56c9 · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:a96952c69fe263f62dd3e78acca16b21e664e245477b5e099ff71ce47bba00c3

Observation cbad09e6-231c-43f8-8685-c273c9956641 · outbound

This paper cites Rank analysis of incomplete block designs: I.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Rank analysis of incomplete block designs: I

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:fbf0b96d1d9d56dd92b559d5cfaca9708fb09a9fcf5f187ca10b0c5a49470702

Observation f86ae7b3-8398-40af-8b5e-653a270d5c9e · outbound

This paper cites Springer, 2005.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Springer, 2005

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:d4d51e3f32d106f4c8077c722c58e930894d713a00fdbbaeefff0b5e98e6f78e

Observation 9c468e00-21d0-41e0-b189-a235209dda2f · outbound

This paper cites Distribution-free predictive inference for regression.Journal of the American Statistical Associ- ation, 113(523):1094–1111, 2018.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Distribution-free predictive inference for regression.Journal of the American Statistical Associ- ation, 113(523):1094–1111, 2018

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:74b734dad68c7ba0f7ad5b6235284d6f701098dc03f993fc4e02284a06c36492

Observation adc1bcdf-9291-4b32-947e-6561da2d40bb · outbound

This paper cites Normalized nonconformity measures for regression conformal prediction.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Normalized nonconformity measures for regression conformal prediction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:ef857152c152c60e66a01585ad82f25b979ccfed4a0004dea37a3564905a9d8c

Observation bbf38b34-ac84-473c-bd05-06a99be126ff · outbound

This paper cites Are we done with ImageNet?.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Are we done with ImageNet?

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:38:19.231041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:8563c7f0037e79bcf65d01e48fe37a992545c0e7454310c1e08e0efcb330ea75

Observation a4eb6505-7f46-444f-9ebb-79e5551dca55 · outbound

This paper cites Conformal prediction beyond exchangeability.Annals of Statistics, 51(2):816–845, 2023.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Conformal prediction beyond exchangeability.Annals of Statistics, 51(2):816–845, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:c0735583768b25493d16633a2475a0e02a3ee3c200257cfc4e0b8785780eb1fa

Observation 80acbc3c-3085-479a-8706-a0fa841f5d2b · outbound

This paper cites Penalize missing requested parts or deviating from constraints.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Penalize missing requested parts or deviating from constraints

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:48805622e0d88df79dc14b8d85a17d227f0cc5963c34917adc9ee60876d42af7

Observation ca792783-0344-494b-8220-1605466640c7 · outbound

This paper cites Provides useful steps, options, or explanations tailored to the request.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Provides useful steps, options, or explanations tailored to the request

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:55008d4c95be51970a4f0c0d1c3ce545d568f0d75b86fe65b01bfa5971aa9cf3

Observation b9fc972f-2e5f-489e-bffd-aea929457693 · outbound

This paper cites Avoids hallucinations and unwar- ranted specifics.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Avoids hallucinations and unwar- ranted specifics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:3626c34f67e141769852861a2c2aea3fa9daef9f47b62a1c9cf3cb1cc7c20c68

Observation 247fafb6-012d-424b-a4a7-5682aff4c8fd · outbound

This paper cites Addresses all sub-questions and important constraints.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Addresses all sub-questions and important constraints

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:8e7bb1d2ae6dc23d3b9b2e57b424114569e85c94072c07b7f799e9d8d6e0293e

Observation ef5a9767-045d-48fb-b68a-9197c4f81259 · outbound

This paper cites A wins”, “tie.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation A wins”, “tie

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:97a4fa6751d9721096f166a4f81efdddf093214677d029ff552f84fd0f8e6cb8

Pith citing papers

No inbound Pith citation observations are available.