Pith. sign in

Paper Citation Record · LEDGER

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2606.13221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.13221 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T07:44:53.583327Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved47
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 332fc3a4-d04c-4ade-96cc-a3fca9fbf204 · outbound

This paper cites Chatbot arena: An open platform for evaluating LLMs by human preference.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Chatbot arena: An open platform for evaluating LLMs by human preference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:ea310d909d7acc96c93bf46214c319dea03612746e0f0a5fb670078978325912

Observation 8e044b64-81d6-43ce-bd8e-16729cc11dc6 · outbound

This paper cites compar:ia: The french government’s llm arena to collect french-language human prompts and preference data.arXiv preprint arXiv:2602.06669, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation compar:ia: The french government’s llm arena to collect french-language human prompts and preference data.arXiv preprint arXiv:2602.06669, 2026

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:38:19.228275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:a6fb733dd6d7390a7d62e14cb5039a0c096c3a16bc21395a808e2a92bf7e62f8

Observation e0cd7a58-f4f9-4677-8693-bf7b028fa164 · outbound

This paper cites Gonzalez, and Ion Stoica.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Gonzalez, and Ion Stoica

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:b8ba4ecaf24534fc656e6ebc0eba768c9b7947da81168b04174c2be4797a574b

Observation 26198120-a334-4633-8b87-1f7862e8d8c2 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:deb76b27c20a371da69e4f9e22fcdaf9b9102ffcfdd6c067e540e040bd29c0ce

Observation d25ab067-2bc7-436e-a5b4-26a286498fd9 · outbound

This paper cites Gonzalez, and Ion Stoica.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Gonzalez, and Ion Stoica

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:4e4e5f2e7771d1a5ba6802ffa738b7db095a8a989f9ad84a41d6baf03c132297

Observation 0da20f28-d7e3-4feb-a94e-2faa6510c4b3 · outbound

This paper cites Judging the judges: A systematic study of position bias in LLM-as-a-judge.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Judging the judges: A systematic study of position bias in LLM-as-a-judge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:1566be26e67c7e2e6f4c118ca0731910a8041ff9d95ae5f4b37f9aa64df2af5d

Observation 18dc3a9f-c380-4821-b52a-a260b1e58170 · outbound

This paper cites Bowman, and Shi Feng.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Bowman, and Shi Feng

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:36b9a6d17c62b3ac9869646ea8408ee0eaa7d12001138a1cef54c5f9185a7f66

Observation 0e423bfd-5fc1-4251-967e-e2b69928ed34 · outbound

This paper cites Tuning LLM judge design decisions for 1/1000 of the cost.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Tuning LLM judge design decisions for 1/1000 of the cost

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:8c73e49136b5c8eb79ffcd8fb2b3efdfd9088008f9bb803c24661317c26973a6

Observation cff900a0-2b3d-48f5-9d9c-b2c090b71eea · outbound

This paper cites Investigating non-transitivity in llm-as- a-judge.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Investigating non-transitivity in llm-as- a-judge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:54e1c0132044ad8f909cd588fcfe0280d76f9c9cbeb76a29c3ceca7a8a6a0fde

Observation 9a460dd4-141b-42ac-b99b-9fd44ffd2fa6 · outbound

This paper cites Mediocrity is the key for llm as a judge anchor selection, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Mediocrity is the key for llm as a judge anchor selection, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:018cff96c7aa00884ccc1f83955ecbbdae40d1c01cd3011a1c2d5aceaf56ea2c

Observation c3240093-8207-4472-9ca3-abe6c4b98f0a · outbound

This paper cites Auto-arena: Automating LLM evaluations with agent peer battles and committee discussions.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Auto-arena: Automating LLM evaluations with agent peer battles and committee discussions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:227fb2dbc5d4be0c4ef9f6fccbe45b379c75bfcc420b97f697839612f7ab25ef

Observation cb73b3e7-dc2c-43de-b475-47805042e3ea · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:88ae00e5fa4e2b235a1810deb5523b12a33d65c31b69cb2797d477ff7ab41261

Observation 38f0f8a0-fe14-4c8b-bff0-ee588a8ff177 · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 13

Resolution
parse uncertain
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:55e2736cd42a9e11f47cc754e1d580864326856a29f044cd4a43e52013186b2d

Observation 5e4b3cac-3bdf-47bf-bf18-af8bc6373ecd · outbound

This paper cites Large language models are not fair evaluators.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Large language models are not fair evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:69a475e0fdce2b1a17ada78a94bbfbd23163a0c669ffb37ed05370f9cb4af197

Observation e5c3e027-68ba-4ed0-9d57-cff0ac6ad2c0 · outbound

This paper cites Explaining length bias in LLM-based preference evaluations.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Explaining length bias in LLM-based preference evaluations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:28de82f33c0d65d53f74a8f798aa2104c0ad04585aaabbe0e063f73ac913576a

Observation e29c6ac0-9407-408b-9239-2740fd2972cf · outbound

This paper cites Self-preference bias in LLM-as-a-judge.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Self-preference bias in LLM-as-a-judge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:7150ae3f45717e4836d141e1b9235a4c8159a6b30142226106bec68a439a69a0

Observation bf4b7989-3532-4b83-aadd-812e8efe6dcc · outbound

This paper cites Huang, Yunyi Shen, Dennis Wei, and Tamara Broderick.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Huang, Yunyi Shen, Dennis Wei, and Tamara Broderick

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:46fa67317519cb3a59b7f52fcfc1676c9b401b1656018a9bc80da7109cd88c51

Observation 232074ad-a7f8-4052-b479-3fd108fe4092 · outbound

This paper cites Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Smith, Beyza Ermis, Marzieh Fadaee, and Sara Hooker

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:84f70c2341129639eb5f6caaf9a9c3341803163854b81251b3b9cc482145a212

Observation 16c4037c-fbfe-4f18-a070-303e3438f950 · outbound

This paper cites Bridging human and LLM judgments: Understanding and narrowing the gap.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Bridging human and LLM judgments: Understanding and narrowing the gap

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:a6b3fa6d75b79e3232089699029c25a2cea98a6a05c5a624ed3dd4adac4318fd

Observation 0366021a-bfcb-4914-a86e-97315da3e86d · outbound

This paper cites Elo uncovered: Robustness and best practices in language model evaluation.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Elo uncovered: Robustness and best practices in language model evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:21b51f2e3c60197039d6910634f501f5de420d3bcbdf28ece953274e112a48bd

Observation 6cd050ff-d8b4-428f-b82a-48d818ba6959 · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:939efeff2691d2107359551af24c01181ef4e5d38271ac36203eb83d00691870

Observation 6a7b5359-ac23-443e-87e2-f5233b852812 · outbound

This paper cites am-ELO: A stable framework for arena-based LLM evaluation.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation am-ELO: A stable framework for arena-based LLM evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:6a0916202b6221b657764ed92003a7d9572d1e8081bf06c3a4c0fd226ad4793a

Observation 2d9ca594-51ce-4d0f-9e35-805e2df2d1fd · outbound

This paper cites Beyond bradley-terry models: A general preference model for language model alignment.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond bradley-terry models: A general preference model for language model alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:00d26c6df6e7cb3e6f054167f0c271aca0793cbd6bfafb1b75281f0caca2bfe0

Observation 94e2a045-408c-4600-a832-d00ba2f17176 · outbound

This paper cites Nonparamet- ric llm evaluation from preference data, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Nonparamet- ric llm evaluation from preference data, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:c517e1080d3288e3532c9279115985dea216c794b5cfc3d7ddedf4438318efc1

Observation 7f42885f-eeda-46bd-8fb6-775bbd1e4598 · outbound

This paper cites Reward learning from preference with ties, 2024.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Reward learning from preference with ties, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:b1ea63b915387ac4cf6aa22539fae8d3faf9b572ed56dec37645b84310cb13ec

Observation 48c8e517-8533-441b-ba31-e868e1d70622 · outbound

This paper cites Beyond binary preferences: A principled framework for reward modeling with ordinal feedback, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond binary preferences: A principled framework for reward modeling with ordinal feedback, 2026

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:76c13fe7292a86c48b08d6971cef51484c019159e0a49b92a9529d709ff07516

Observation c6af74d0-3e86-4514-bedd-defa168391ce · outbound

This paper cites Reward modeling with ordinal feedback: Wisdom of the crowd.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Reward modeling with ordinal feedback: Wisdom of the crowd

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:9480ed56eba0d9cbddb1548792f37f7f700d26e56f26cc51125c62c7a3d364fe

Observation ecfd96d6-2165-4f74-bb7e-84e74a18b423 · outbound

This paper cites Improving LLM-as-a-judge inference with the judgment distribution.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Improving LLM-as-a-judge inference with the judgment distribution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:d7d2b01eef16b51c8c252ea6335d2191e6c69e402267289662ac84abd892ecc5

Observation 5040c1d6-e214-4a40-8d69-0e3d2b2baff1 · outbound

This paper cites Beyond single-point judgment: Distribution alignment for llm-as-a-judge, 2025.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond single-point judgment: Distribution alignment for llm-as-a-judge, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:e1d835bf39653c14278523892b562dd4712dfac5c077f515c64504253b5c48f3

Observation 0e95aabe-4f3a-4e77-bf38-fdeb4dd7545c · outbound

This paper cites Malin, and Yuan Xue.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Malin, and Yuan Xue

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:fb7d8223bc9421bab62f8230d4c597c9d0b654613459a1da80d6e7eee1dc2759

Observation 8d506c84-a716-4cfe-83f8-7f2b3ba1ef94 · outbound

This paper cites Beyond ordinal preferences: Why alignment needs cardinal human feedback, 2025.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Beyond ordinal preferences: Why alignment needs cardinal human feedback, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:9a7397e0d4c326b4e6847d4258f5899e46c533d61279729ecdcbfb385d3f363f

Observation 05e12e5c-bf5c-4b34-b1ba-833e28857041 · outbound

This paper cites LLM- rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation LLM- rubric: A multidimensional, calibrated approach to automated evaluation of natural language texts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:26995ea26ffe4dda98f933ae95582c4712538bfcde36e205db472151bcc33b51

Observation 2ae1ee6f-9199-4ccb-b18e-8b20ac202214 · outbound

This paper cites Quantitative llm judges, 2025.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Quantitative llm judges, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:b8f1699b628a1687dea8690559a7c121a7426a0fe0624234546a56420308864f

Observation ce688e7d-e20c-47f6-acac-8fa7220632f1 · outbound

This paper cites Analyzing uncertainty of LLM-as-a-judge: Interval evaluations with conformal prediction.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Analyzing uncertainty of LLM-as-a-judge: Interval evaluations with conformal prediction

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:0ac0dfc70e46ec6f67ec6c3085927ef64541fb7772d97624b986240b2d858d80

Observation ee646bfb-db86-47f1-86a3-08c066eebebc · outbound

This paper cites Scope: Selective conformal optimized pairwise llm judging, 2026.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Scope: Selective conformal optimized pairwise llm judging, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:ffc5cdf2f7472f197acdf84f77f58ecae82b6083af2841347f761f8429a7e46d

Observation 2ce3422c-5f94-46d0-8248-c00aaff8c46a · outbound

This paper cites Prediction-powered ranking of large language models.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Prediction-powered ranking of large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:d04db8801029f651943d149f7dcaaba2790e5cb093483fbb3bb0808d613dbafa

Observation 024ef43c-f6b7-4b24-9a86-c86ee5107ad3 · outbound

This paper cites Alex Hofer, Bhuwan Dhingra, Amir Globerson, and William W.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Alex Hofer, Bhuwan Dhingra, Amir Globerson, and William W

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:64af5d0e1764ea0035d22e3771277b929993fe144ba566b2604617dee9e26cdc

Observation ac81a5fe-a137-44d8-8a39-e35b414e02ad · outbound

This paper cites Adaptive prediction-powered autoeval with reliability and efficiency guarantees.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Adaptive prediction-powered autoeval with reliability and efficiency guarantees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:10eafbccb3f1702a1cd2c25e0cab2e58a1643a992ba588b18cb116bdd825cafc

Observation 1fdc4bee-44d3-48ef-aad8-56296cfc56c9 · outbound

This paper cites an unresolved cited work.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:b28f55b8dfa109afb132e4825b243f5479ef0d578c65be901db16ad1df344b50

Observation cbad09e6-231c-43f8-8685-c273c9956641 · outbound

This paper cites Rank analysis of incomplete block designs: I.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Rank analysis of incomplete block designs: I

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:982591a80c562e9d1e116ae805b239838da5b9b20ec23603313f53dffcc03957

Observation f86ae7b3-8398-40af-8b5e-653a270d5c9e · outbound

This paper cites Springer, 2005.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Springer, 2005

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:a873b0c7a1b542dc1f5c4277b1c6ac1cda2d112c5ee7ed5d7c3e864985a5fb97

Observation 9c468e00-21d0-41e0-b189-a235209dda2f · outbound

This paper cites Distribution-free predictive inference for regression.Journal of the American Statistical Associ- ation, 113(523):1094–1111, 2018.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Distribution-free predictive inference for regression.Journal of the American Statistical Associ- ation, 113(523):1094–1111, 2018

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:886f53cae3d290c95ad3589b38df55fed8eca01b9fd235fd32c7baf4222e6272

Observation adc1bcdf-9291-4b32-947e-6561da2d40bb · outbound

This paper cites Normalized nonconformity measures for regression conformal prediction.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Normalized nonconformity measures for regression conformal prediction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:e5ffdcc8d3e0d7326718cce4d9f63fca6b3ef93965ece5e7d9859a44ad6cb67e

Observation bbf38b34-ac84-473c-bd05-06a99be126ff · outbound

This paper cites Are we done with ImageNet?.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Are we done with ImageNet?

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:38:19.231041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:604fa6eeacfec762065754ad910677fb38948eff63be35c1cb78d9393cf0c552

Observation a4eb6505-7f46-444f-9ebb-79e5551dca55 · outbound

This paper cites Conformal prediction beyond exchangeability.Annals of Statistics, 51(2):816–845, 2023.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Conformal prediction beyond exchangeability.Annals of Statistics, 51(2):816–845, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:d766cba5ab5538ed6e83169d4648238c9637a7c917c14a23ab5e01814c9b3a31

Observation 80acbc3c-3085-479a-8706-a0fa841f5d2b · outbound

This paper cites Penalize missing requested parts or deviating from constraints.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Penalize missing requested parts or deviating from constraints

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:6c0fef59d708422e612c977d0f4c1d5a1a2ed42ce15b2673a709473c8e2cabe1

Observation ca792783-0344-494b-8220-1605466640c7 · outbound

This paper cites Provides useful steps, options, or explanations tailored to the request.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Provides useful steps, options, or explanations tailored to the request

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:cd78abb5dfff4b96f2c43ecc715e28649c74f4692117dbcd2c67c0321204327c

Observation b9fc972f-2e5f-489e-bffd-aea929457693 · outbound

This paper cites Avoids hallucinations and unwar- ranted specifics.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Avoids hallucinations and unwar- ranted specifics

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:2586a3627180d43eaf3310053ae935dc540c4b73a0dabca61ff3261cb5744a12

Observation 247fafb6-012d-424b-a4a7-5682aff4c8fd · outbound

This paper cites Addresses all sub-questions and important constraints.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation Addresses all sub-questions and important constraints

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:ccbb471c9d1e3ea2860ca462dd4bbb626a722437849ef02c9788e16fe0c335a5

Observation ef5a9767-045d-48fb-b68a-9197c4f81259 · outbound

This paper cites A wins”, “tie.

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation A wins”, “tie

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T07:44:53.583327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T07:44:53.583327Z digest=sha256:f69f52b50feda32439f2456d3d4d7926aa50e498da6497b9f0d557bc61ff6d54

Pith citing papers

No inbound Pith citation observations are available.