Pith. sign in

Paper Citation Record · LEDGER

Revisiting Uncertainty Estimation and Calibration of Large Language Models

As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 11 inbound Pith citation observations for arXiv:2505.23854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23854 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:42.412963Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:18:20.760827Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved33
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation eeabaf38-434a-4669-9f7d-f022633efb8e · outbound

This paper cites Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:37.551868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:37.551868Z digest=sha256:3b4cbf90a925b253ff953efbbfeca4dfd5cd4e645e6f0b61f2050a271e4e0ac5

Observation 2d07b809-04ab-4715-80f3-5b9e8016ace4 · outbound

This paper cites Adapted large language models can outperform medical experts in clinical text summa- rization.Nature medicine, 30(4):1134–1142, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Adapted large language models can outperform medical experts in clinical text summa- rization.Nature medicine, 30(4):1134–1142, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:45.680203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:37.640897Z digest=sha256:56a055ac9de718bf850fb9c9b472669a61079bfed9ecd5d28d98497b023ff2ea

Observation 8e9b719b-827a-40e3-ba33-cf5c4ec82e35 · outbound

This paper cites MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes.

Revisiting Uncertainty Estimation and Calibration of Large Language Models MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:37.703236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:37.703236Z digest=sha256:7e58c5ce7d21af0290a85c7fe3171ada3d5a3ed593768cb32fdab3db2714d592

Observation ca2020b9-175c-4de5-9c69-cc56cb9257dd · outbound

This paper cites Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects.Authorea Preprints, 1:1–26, 2023.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects.Authorea Preprints, 1:1–26, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:45.466563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:37.785893Z digest=sha256:334b5feb4d649e7b63687510e975c734a135607e5b53819f6e33da00b5f3b767

Observation ef793f4e-92f1-4cf4-8dbc-b79a16f31749 · outbound

This paper cites Large language models in law: A survey.AI Open, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Large language models in law: A survey.AI Open, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:37.854288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:37.854288Z digest=sha256:724e43b3f9d5cbbf6701134245f1f436357ac7712322a3c9ac6275110b36daed

Observation a4e05e20-ae7b-4b17-9782-8e6d68e26011 · outbound

This paper cites Overreliance on ai literature review.Microsoft Research, 339:340, 2022.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Overreliance on ai literature review.Microsoft Research, 339:340, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:45.228040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:37.902306Z digest=sha256:92d32602a85c2993448dd2e76ebac52543b6485fbf99e7066ba19f84df810094

Observation 94622b17-0103-455b-a7d1-8f179f1f1019 · outbound

This paper cites Understanding the eu ai act: Requirements and next steps.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Understanding the eu ai act: Requirements and next steps

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.987647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:37.971286Z digest=sha256:6f51bf8e5bdf12775bd4421ffe877e30b09764c31fb7dd9c56e7b76c9ede0002

Observation f0fd20b6-0dbe-4484-8034-a5adc2f6d468 · outbound

This paper cites Us debt ceiling deal clears major hurdle in congress, 2023.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Us debt ceiling deal clears major hurdle in congress, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.776735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:38.025559Z digest=sha256:53e67c7e1b604e5d3017134290f834c8213ecbd2a8a47ab86b472673bf3b9d98

Observation 3d7996eb-a9ad-476e-b359-4d2f55243ef6 · outbound

This paper cites Shifting attention to relevance: Towards the uncertainty estimation of large language models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Shifting attention to relevance: Towards the uncertainty estimation of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.558433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:38.065319Z digest=sha256:50fd99616c68d3f9d0bbe289718f36b9c0ec7466b912adcba5a458523ae3a910

Observation ef2ee3bb-1ceb-49ee-8c0a-39e3a8832012 · outbound

This paper cites Semantic calibration of llms through the lens of temperature scaling.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Semantic calibration of llms through the lens of temperature scaling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.398144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:38.132286Z digest=sha256:6915fb0f391150f0f414250b0b78bc331814b07a298cb1e7253f6f0bfffb04f1

Observation 7483b66a-e019-446d-9bc9-30d7cf093fe1 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.199784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.199784Z digest=sha256:80efd33bc4765d456d00cdcddd381e1a9bf901b9c72ade88ef68f154fee4b954

Observation 20371349-ad3e-46d9-b448-97107aaabb35 · outbound

This paper cites Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.252917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.252917Z digest=sha256:7ebc2e26543f791371c35dfb5e38c821b38a6a0f6b8e617a9269059f293fd231

Observation d9a07837-1993-4910-bdd7-8388a21795f8 · outbound

This paper cites Mitigating object hallucinations in large vision-language models via attention calibration.arXiv preprint arXiv:2502.01969, 2025.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Mitigating object hallucinations in large vision-language models via attention calibration.arXiv preprint arXiv:2502.01969, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.316495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.316495Z digest=sha256:d09d887e5481781d2b1b44591eaa2fe3875b80fa24f96855f685e37492f07b99

Observation 9d820e3e-20a4-4080-9559-74e1ee58c031 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.390328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.390328Z digest=sha256:21bed7c0545bf9beb7c293deeb4f4a94da31456faba9b981e443501fc92c62df

Observation daf4f791-7a66-4312-9dd6-51f8d5645d6a · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.513072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.513072Z digest=sha256:d0783d675c6cb0dfb49f87ee448429a757b6cb7b256593c5881004759939b0af

Observation 1da2abeb-136f-4094-a287-a324fe04db61 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.624582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.624582Z digest=sha256:6ba40ea54e15c036dd3a901097ed9c1dcb9632b92697ef87fb522cc2ba3c62e2

Observation 3252f1ca-9b54-4095-9185-67ba89f5ccee · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.714851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.714851Z digest=sha256:bbde910102290cb4fa30f9c0c962a6e0be20bc983905b081d6b8adc76524505a

Observation 46d5145c-d84e-4cc8-8d54-ac61ca9238c7 · outbound

This paper cites Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.828049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.828049Z digest=sha256:8c43292cd9826b7e0976e5313c923a648f0fc0d55315f9f3058e38736fc87209

Observation 040bf658-455f-4dcf-a43a-841e691046fb · outbound

This paper cites Perceptions of Linguistic Uncertainty by Language Models and Humans.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Perceptions of Linguistic Uncertainty by Language Models and Humans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.924855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.924855Z digest=sha256:9dbfd43b3f2423748054c570a7cc7d79b73d60ca1fcea357e7f6afe2577c93b3

Observation 561c4e14-e924-45c8-ab52-42a5b1aa0943 · outbound

This paper cites On the Calibration of Large Language Models and Alignment.

Revisiting Uncertainty Estimation and Calibration of Large Language Models On the Calibration of Large Language Models and Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.079405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.079405Z digest=sha256:f8ffe73bcd6f77a5a739b1f46dde946b5db35f0cb793f57b0bc055eb5d29ed66

Observation 43492b8e-c07c-4c8a-a7c6-b8881390c9fd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Revisiting Uncertainty Estimation and Calibration of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.166999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.166999Z digest=sha256:8c4e9d600773decc34956bfc9128b1e1ef449aff8fb7ca8ec6bfeabe2f1d6480

Observation feced0d8-d9af-428f-9ed7-f4f812140968 · outbound

This paper cites Qwen2 and qwen3: The new generation open models.https://qwenlm.github.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Qwen2 and qwen3: The new generation open models.https://qwenlm.github

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.275222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:39.236785Z digest=sha256:17876f116edd5ff9f0e5cf29d7075feccb47c88dd783338c5712012d27f3c278

Observation 96ee5786-bd27-4fbb-ace8-6dcb53a8203a · outbound

This paper cites Llama 4: Multimodal intelligence.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Llama 4: Multimodal intelligence

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.237880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:39.357994Z digest=sha256:fa1336327cec1cc0eb9a6dcc7e64ad915f8838c6f2d301bdbde30ba4a3800efe

Observation a2a320e6-784f-4152-90fa-d40238abbb74 · outbound

This paper cites Introducing o3 and o4 mini.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Introducing o3 and o4 mini

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.141647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:39.442780Z digest=sha256:f54edd0de001e95176aba6da65ea51340be0d9430e7aa838ff8016a2da8385f1

Observation 7ca919f0-c98c-44a9-a08d-1c3439806fef · outbound

This paper cites Grok-3 and the next generation of reasoning models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Grok-3 and the next generation of reasoning models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.048562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:39.522769Z digest=sha256:185dc842b91cd588fcbcf5a4f4e05d3d578ea3226f573bcb7dafa1a53af417d2

Observation d5c417bf-1447-4da1-90d3-fdc1f4bc9285 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension.

Revisiting Uncertainty Estimation and Calibration of Large Language Models TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.709871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.709871Z digest=sha256:08a72d14ce70eb576eed9d82a80199ba7732ddf68359ee4e6ef9f34ccb03fa26

Observation d4fb2b74-c697-4bfd-8579-261b7d2d5daa · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.781726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.781726Z digest=sha256:6a64d06bc3064ab1130c8ccf540989bd930d32a6f73ce9db52f00f5756227d29

Observation 644fd330-d804-43e6-a659-90515ff59a9f · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.839500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.839500Z digest=sha256:643ae73e332fc9569b7399117e8456cf15753b81d821219de39baa5bd58a95cf

Observation 7e671f56-039f-4d6d-9ebf-f699aafa29e0 · outbound

This paper cites Discovering latent knowledge in language models without supervision.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Discovering latent knowledge in language models without supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.954047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:39.912855Z digest=sha256:055af9eee55b797fa17edf10df4e5b79e6786712c1c2523e04cc42a3ac729e69

Observation d07b328d-4fc3-4dde-b7a7-3bc9cf23f8bf · outbound

This paper cites Language Models (Mostly) Know What They Know.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Language Models (Mostly) Know What They Know

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.012125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.012125Z digest=sha256:b97c9cea8bc45ee445e2718e908268f2767cf84b266e1d9fe79a19e97e87f788

Observation 5d544f98-6e7d-473b-832b-d53a9d3d1f1b · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.103165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.103165Z digest=sha256:c0725c9079c60f36f9ed9c4a3917100cf99ca96b92cb09ea0bf2be4e1f8439b6

Observation 98fdfb8b-4d9d-4f32-8da3-7de8b050e27b · outbound

This paper cites On calibration of modern neural networks.

Revisiting Uncertainty Estimation and Calibration of Large Language Models On calibration of modern neural networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.212539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.212539Z digest=sha256:03d19fcb520a0f68ab84ddb4bd5c7d4c2a4e54a0c7a9dce8b627d45ca3734978

Observation dfd90084-708e-4ff0-a687-ead5ff14c5e0 · outbound

This paper cites Benchmarking uncertainty disen- tanglement: Specialized uncertainties for specialized tasks.Advances in neural information processing systems, 37:50972–51038, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Benchmarking uncertainty disen- tanglement: Specialized uncertainties for specialized tasks.Advances in neural information processing systems, 37:50972–51038, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.826331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:40.320314Z digest=sha256:5b3ecc731cba055a1908d30ba066a266ad5a9377c6d3e5acd2a531b59873a30a

Observation edc5dd6f-45f2-44c5-a9fc-b7813b164eed · outbound

This paper cites Measuring calibration in deep learning.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Measuring calibration in deep learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.430239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.430239Z digest=sha256:5fa3c6413615c27d74cec09d64f6d1d06f6741a4e84e9c04240aafeed6edad40

Observation bf21376f-ff4a-48b1-92a2-72b3e96a872f · outbound

This paper cites Predicting good probabilities with supervised learning.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Predicting good probabilities with supervised learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.530832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.530832Z digest=sha256:27835340c031081cf7a7eba13d8fda22d8f09600841b4bc7d6dfac656344bce8

Observation 46cf49c1-3722-4012-8942-0e7e3b5abdae · outbound

This paper cites The Llama 3 Herd of Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models The Llama 3 Herd of Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.568605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.568605Z digest=sha256:29045485ecde69276ead7c23c00f7101259a3686103e244e24e5783dcb736d11

Observation dcc0c5ba-e9ed-4731-8080-4ea7282a70dd · outbound

This paper cites DeepSeek-V3 Technical Report.

Revisiting Uncertainty Estimation and Calibration of Large Language Models DeepSeek-V3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.672286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.672286Z digest=sha256:42866788683713b05daf96ccc77acb7fb4b896cdf83be6159c3619fcf6796738

Observation 28897efc-495f-44b7-bd9a-c3e2e5921e7d · outbound

This paper cites Qwen2.5 Technical Report.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.769319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.769319Z digest=sha256:7edc04a6e867b130fe2faf3a41689809f4618ccc848edb127ae16a92720e7244

Observation dfca067e-37b5-48a8-bf61-e03739bbbe41 · outbound

This paper cites Identifying and mitigating vulnerabilities in llm-integrated applications.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Identifying and mitigating vulnerabilities in llm-integrated applications

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.707808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:40.902069Z digest=sha256:b71854c2c734fda5d0303dc0f3eeaa0e88f7d8759d4a950efe959fe3e4c090d9

Observation b7d3af9e-c9c5-436c-8b93-276255bf9c66 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.026953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.026953Z digest=sha256:56d703271f3afb934a0b7ec686e5bfada975e79badc3a9e93fa85171750c4309

Observation 7e56c294-799c-4035-b00d-968fef65f92d · outbound

This paper cites Introducing gpt-4o: Openai’s new omni model.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Introducing gpt-4o: Openai’s new omni model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.627738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:41.103318Z digest=sha256:28d0e3a9e2947a9b8fa5f41f21a963a69b0b088f4d98885c43e41a61b40390b3

Observation 205b571d-ccb9-4ea5-90d3-c80f83777942 · outbound

This paper cites Gpt-4.1 overview.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Gpt-4.1 overview

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.540243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:41.172123Z digest=sha256:cde24c3f5393ece2db85669218b2d95399abf3070c2663ace054facad8d3f6d2

Observation 512247d6-a942-4999-81d9-86498a28fed7 · outbound

This paper cites Introducing the claude 3 model family.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Introducing the claude 3 model family

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.455340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:41.282128Z digest=sha256:9f0bbfc2297de7cdc6dc39a1fc4c04716dd213acb57b00e7febf98c86f5d6250

Observation 0d888322-b373-4cb2-a7d7-eb0df487fed1 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.387827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.387827Z digest=sha256:1684996a8239e25b5a00f93a4d017954ce5af23d47fcba98fb87c5e23329f5dd

Observation 26d9cfaa-8dd7-44ff-9a27-a25118f1dbd6 · outbound

This paper cites Qwen3 technical report.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Qwen3 technical report

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.338288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:41.466584Z digest=sha256:f4abddf42a043633257159a972cb07a03ba02315cec21510bb0d9d83b2abfbe4

Observation aa3f8ebf-8f15-438f-9205-c886c88f0e47 · outbound

This paper cites The kendall rank correlation coefficient.Encyclopedia of measurement and statistics, 2:508–510, 2007.

Revisiting Uncertainty Estimation and Calibration of Large Language Models The kendall rank correlation coefficient.Encyclopedia of measurement and statistics, 2:508–510, 2007

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.219279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:41.561704Z digest=sha256:3c9bcd9b7b30a7f09d0c21ec87fa0c78978051272133715e4c4605da8ddbefb6

Observation 1b059fc6-3b14-4a6c-acab-2704abb5c733 · outbound

This paper cites i’m not sure, but.

Revisiting Uncertainty Estimation and Calibration of Large Language Models i’m not sure, but

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.108420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:59:41.651489Z digest=sha256:671352872debe283b7c36da518c142b210377651f19332cc2d30823bf62dc2f5

Observation 32c81645-cadb-4dfd-8b68-1c9f47e25d8e · outbound

This paper cites Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.765614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.765614Z digest=sha256:6df29bfda17325881b0f3ddb1813cee8fce0962c36320885d7ca742ba945dce8

Observation dee90a92-4161-4657-a699-700ec5fb9550 · outbound

This paper cites The internal state of an llm knows when it’s lying.

Revisiting Uncertainty Estimation and Calibration of Large Language Models The internal state of an llm knows when it’s lying

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.862601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.862601Z digest=sha256:a2984676a882860bc3c8d715835282b731f8f9ddc4000de1e4273e56faa77f5e

Observation 10675e64-2b09-480d-83f2-e5211afa71ea · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Inference- time intervention: Eliciting truthful answers from a language model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.990181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.990181Z digest=sha256:3460f5ad9cec8c87377cb9bbb80df4c4415ada0fcfb1b5a5c18eb5557a2351f6

Observation 8f7b23fa-2184-4464-a311-2d9a6e5d2f00 · outbound

This paper cites Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:42.079253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.079253Z digest=sha256:ee21a3a07620c7938964e275ce931b3d30a8d9f44679f2ae3303a1e97938d3ae

Observation 8068ad81-a5df-45a1-9833-b1b3eef73a3d · outbound

This paper cites A Survey of Confidence Estimation and Calibration in Large Language Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models A Survey of Confidence Estimation and Calibration in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:42.188284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.188284Z digest=sha256:0eaa4249144bff89bbfb5dd29881e89b135d00933a52dba3b5727bc9668902e6

Observation 914ff920-0b70-419f-ae25-9d0adb44237c · outbound

This paper cites A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice.

Revisiting Uncertainty Estimation and Calibration of Large Language Models A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:42.294102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.294102Z digest=sha256:7cd35bf095be017bff8f39ec9556915d4e89585ff735fe99334c320dea3daada

Observation 772fa1e4-ebcf-417f-9f15-0e801f62ac9d · outbound

This paper cites Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:59:42.412963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.412963Z digest=sha256:8bbc16b69cb46c2d165362d907c6b62a58807c641e9247f62bd89b8b12170ca7

Observation 3747804d-053b-4181-997c-b61cd5e3e967 · outbound

This paper cites an unresolved cited work.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-07T12:59:39.612033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.612033Z digest=sha256:8ba844a41edb6e789a1812003265289bf9f9232f9a25a78180572b8d8b006862

Pith citing papers

Observation 9b3d346c-f242-49e2-a541-b5f46b7c0465 · inbound

Toward Efficient Uncertainty in LLMs through Evidential Knowledge Distillation cites this paper.

Toward Efficient Uncertainty in LLMs through Evidential Knowledge Distillation Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:18:20.760827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:18:20.760827Z digest=sha256:70e1ef20ffc4af793bec032f504dd53298c70613a66fd8e5a77085cb97e4089d

Observation a5ed4c11-5ded-4721-9f57-0600103d5539 · inbound

Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models cites this paper.

Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T20:21:55.864745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:21:55.864745Z digest=sha256:29ad22942ef7e3508f69cd2e8949454ca806d3866d0e3065480076dc438ac132

Observation 82d38082-ecdf-401c-8854-f51fd3962bd9 · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:44:05.672258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:2834277e02dfe39bc87685c34c83f130c56ba50631a3b1a28493c2a226ef4117

Observation 88f41d5f-8b82-49f5-9136-fcab8f089423 · inbound

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment cites this paper.

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:58.469843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T21:13:29.201794Z digest=sha256:fbb835af2bc67818749c6d18c8c8e60ed835deaec904682b9c5cc7d1319f1525

Observation 01f33cb2-5767-4c0b-97d4-68032e0a9035 · inbound

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation cites this paper.

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:11.565452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T18:30:03.367870Z digest=sha256:51b6245ad71f8cad8a1b7ae81565ec5e36c49f582ee0645139bd0adb917cfd1e

Observation 668e8321-cc48-42cd-af59-9e178617f734 · inbound

Hallucinations Undermine Trust; Metacognition is a Way Forward cites this paper.

Hallucinations Undermine Trust; Metacognition is a Way Forward Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.202112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T14:29:54.924293Z digest=sha256:efa056be1a4e997f05f7e3e2b52de707334e1b6d7789c9090a08829d2fd97e61

Observation a9d14390-04ab-4387-9f8e-a6b2c7d95f72 · inbound

Hypothesis generation and updating in large language models cites this paper.

Hypothesis generation and updating in large language models Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-08T21:14:12.533436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T14:43:56.125546Z digest=sha256:8d10030505782f8273285c328d12bf72812e53251391b720cac49329eda6b258

Observation 8b2a0696-382c-455e-9ead-adb525a9ae63 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.021788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:627756671bce347dad2d700187756602c80c9cfeb1a8ff39f4dcef1a9a28a50e

Observation 61d0be8f-0d5f-463c-a19b-f51d597d8b84 · inbound

Retrieval-Augmented Linguistic Calibration cites this paper.

Retrieval-Augmented Linguistic Calibration Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.155542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T05:48:03.641696Z digest=sha256:8cb6c86f2a74c833d8464c72dd3190f8f0154deaa7768aa1627361febf343cb6

Observation 80de674f-289e-40eb-851c-90d35991b82e · inbound

Retrieval-Augmented Linguistic Calibration cites this paper.

Retrieval-Augmented Linguistic Calibration Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.397448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:36:25.106668Z digest=sha256:776da8f54a91f9b287d03185068da00f7c45b39c3397fa361f95c650a770e217

Observation e08bc5bb-8617-45de-bc2e-9a753e99e46e · inbound

Gradient-Guided Reward Optimization for Inference-time Alignment cites this paper.

Gradient-Guided Reward Optimization for Inference-time Alignment Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.645365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T16:43:04.984866Z digest=sha256:7d862b2baeb2988c4d88a4ff096d26d9f421be5876443ef69f7b81ae5746c2c6