Pith. sign in

Paper Citation Record · LEDGER

Revisiting Uncertainty Estimation and Calibration of Large Language Models

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 10 inbound Pith citation observations for arXiv:2505.23854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23854 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:42.412963Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:21:55.864745Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved33
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation eeabaf38-434a-4669-9f7d-f022633efb8e · outbound

This paper cites Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:37.551868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:37.551868Z digest=sha256:dba6d96587f087344690c7ab362f980f590daa4fc137d094096b706cf95db5d7

Observation 2d07b809-04ab-4715-80f3-5b9e8016ace4 · outbound

This paper cites Adapted large language models can outperform medical experts in clinical text summa- rization.Nature medicine, 30(4):1134–1142, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Adapted large language models can outperform medical experts in clinical text summa- rization.Nature medicine, 30(4):1134–1142, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:45.680203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:37.640897Z digest=sha256:e9283848e8696386be58416fa074af079012a10f8390a86545ec918fd84bd3c7

Observation 8e9b719b-827a-40e3-ba33-cf5c4ec82e35 · outbound

This paper cites MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes.

Revisiting Uncertainty Estimation and Calibration of Large Language Models MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:37.703236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:37.703236Z digest=sha256:39c592eefc6bbd9a5f23240c459a9d4325bd075288bbb3db09e00bae1968b7ce

Observation ca2020b9-175c-4de5-9c69-cc56cb9257dd · outbound

This paper cites Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects.Authorea Preprints, 1:1–26, 2023.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Large language models: a comprehensive survey of its applications, challenges, limitations, and future prospects.Authorea Preprints, 1:1–26, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:45.466563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:37.785893Z digest=sha256:1236a1336688bd23e22a901ab08a9ea6d9b22f88916c5d088a96ebf6fb63ecd7

Observation ef793f4e-92f1-4cf4-8dbc-b79a16f31749 · outbound

This paper cites Large language models in law: A survey.AI Open, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Large language models in law: A survey.AI Open, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:37.854288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:37.854288Z digest=sha256:4b42f0d6f3407e69ed1c84c26cf77ca0f2ef2438ae793743f58a6ab893b0e6f4

Observation a4e05e20-ae7b-4b17-9782-8e6d68e26011 · outbound

This paper cites Overreliance on ai literature review.Microsoft Research, 339:340, 2022.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Overreliance on ai literature review.Microsoft Research, 339:340, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:45.228040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:37.902306Z digest=sha256:e2a737a3293bb68ef9c98ccae94a9927606ad5c3507f95eb403aeab4c8590b94

Observation 94622b17-0103-455b-a7d1-8f179f1f1019 · outbound

This paper cites Understanding the eu ai act: Requirements and next steps.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Understanding the eu ai act: Requirements and next steps

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.987647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:37.971286Z digest=sha256:d72d1ba6112f5ecb1062f7de252c80063eb7665297914a100f8c7c4c1e12eb9e

Observation f0fd20b6-0dbe-4484-8034-a5adc2f6d468 · outbound

This paper cites Us debt ceiling deal clears major hurdle in congress, 2023.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Us debt ceiling deal clears major hurdle in congress, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.776735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:38.025559Z digest=sha256:d35d27d73ae7b163d53344c8a76a69101b7f2e67d0aa185b16fe57de45a26c1a

Observation 3d7996eb-a9ad-476e-b359-4d2f55243ef6 · outbound

This paper cites Shifting attention to relevance: Towards the uncertainty estimation of large language models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Shifting attention to relevance: Towards the uncertainty estimation of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.558433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:38.065319Z digest=sha256:5fde2a66decf4c179250fa3d41fc2c1198894119bbdd6d9c5936c826804ecf9d

Observation ef2ee3bb-1ceb-49ee-8c0a-39e3a8832012 · outbound

This paper cites Semantic calibration of llms through the lens of temperature scaling.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Semantic calibration of llms through the lens of temperature scaling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.398144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:38.132286Z digest=sha256:9d97dc60edc2f03e1a175abea353465b1caf5a40b203f4a6d37104574d0f182c

Observation 7483b66a-e019-446d-9bc9-30d7cf093fe1 · outbound

This paper cites Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.199784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.199784Z digest=sha256:269f5304064828bebfa5d6d4f5f6f6e06a3f8cb7e8dd859830347fd5fb25492e

Observation 20371349-ad3e-46d9-b448-97107aaabb35 · outbound

This paper cites Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.252917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.252917Z digest=sha256:c37fc3423568e1c7cfe73d554c9cbf2324e60789ec486906756ba2324c4c96b6

Observation d9a07837-1993-4910-bdd7-8388a21795f8 · outbound

This paper cites Mitigating object hallucinations in large vision-language models via attention calibration.arXiv preprint arXiv:2502.01969, 2025.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Mitigating object hallucinations in large vision-language models via attention calibration.arXiv preprint arXiv:2502.01969, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.316495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.316495Z digest=sha256:715099cdf2f3d24c04162b87e53b3e61b106a1c192be441349bf7f7ce212a026

Observation 9d820e3e-20a4-4080-9559-74e1ee58c031 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.390328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.390328Z digest=sha256:5ec27261f45c87345adfa54cad7dbf32ad8d36325fc5d13938fff529ce387385

Observation daf4f791-7a66-4312-9dd6-51f8d5645d6a · outbound

This paper cites SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.513072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.513072Z digest=sha256:1b5bb003f514c6f9df904f2e07987576b2a700e17cd0f5c1006e48593a66d5e0

Observation 1da2abeb-136f-4094-a287-a324fe04db61 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.624582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.624582Z digest=sha256:73335a46fa2f4909c60ecaeb9b226fc79993e6e82b503fa873c0bfa9f9ebc6c3

Observation 3252f1ca-9b54-4095-9185-67ba89f5ccee · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.714851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.714851Z digest=sha256:9b0e1a9ed73bbe269970d20af5351c526eb94cbf30b90453a1c6cfd8d76e4b9e

Observation 46d5145c-d84e-4cc8-8d54-ac61ca9238c7 · outbound

This paper cites Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.828049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.828049Z digest=sha256:fab2016741619ff56911d5e0e6fb24efd30efcffcb37235e983a865818d10b49

Observation 040bf658-455f-4dcf-a43a-841e691046fb · outbound

This paper cites Perceptions of Linguistic Uncertainty by Language Models and Humans.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Perceptions of Linguistic Uncertainty by Language Models and Humans

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:38.924855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:38.924855Z digest=sha256:44f5fb3f34abf6ff51abc6c7f49e43254dda4685ff5bd8e8ccbbe49f3b1cb89a

Observation 561c4e14-e924-45c8-ab52-42a5b1aa0943 · outbound

This paper cites On the Calibration of Large Language Models and Alignment.

Revisiting Uncertainty Estimation and Calibration of Large Language Models On the Calibration of Large Language Models and Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.079405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.079405Z digest=sha256:6fc1c1c0384208d913781c7f38d1a8d3bf05fb882674bd40940c2e3e8c8f392e

Observation 43492b8e-c07c-4c8a-a7c6-b8881390c9fd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Revisiting Uncertainty Estimation and Calibration of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.166999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.166999Z digest=sha256:5d0c89262ac40f4dbaee3c51c3b3cdb3cd2ce4134742e358ef0021813a2d615c

Observation feced0d8-d9af-428f-9ed7-f4f812140968 · outbound

This paper cites Qwen2 and qwen3: The new generation open models.https://qwenlm.github.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Qwen2 and qwen3: The new generation open models.https://qwenlm.github

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.275222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:39.236785Z digest=sha256:2b347a5ac8396e70ea7f058347a81125e3312e58262ba23c5866843bcb371856

Observation 96ee5786-bd27-4fbb-ace8-6dcb53a8203a · outbound

This paper cites Llama 4: Multimodal intelligence.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Llama 4: Multimodal intelligence

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.237880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:39.357994Z digest=sha256:7734342e2048933aa90e94aa61638a3de7f6a7be4b17ab16c74ce6c8b82043c7

Observation a2a320e6-784f-4152-90fa-d40238abbb74 · outbound

This paper cites Introducing o3 and o4 mini.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Introducing o3 and o4 mini

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.141647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:39.442780Z digest=sha256:8696b925b75f13107d9c02b47049273f203509ae2a656ecad3a0f38e591286fa

Observation 7ca919f0-c98c-44a9-a08d-1c3439806fef · outbound

This paper cites Grok-3 and the next generation of reasoning models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Grok-3 and the next generation of reasoning models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:44.048562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:39.522769Z digest=sha256:891ada9bd61aaa2548a4c98a783debd6d79970e27278d1810b5a9c6512e0cce1

Observation d5c417bf-1447-4da1-90d3-fdc1f4bc9285 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension.

Revisiting Uncertainty Estimation and Calibration of Large Language Models TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.709871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.709871Z digest=sha256:65c5d819c8f41f4bc977040f68c97f41114f9e97e51ccdfe83d66f9b0350eca2

Observation d4fb2b74-c697-4bfd-8579-261b7d2d5daa · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Measuring Massive Multitask Language Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.781726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.781726Z digest=sha256:ab7c2c6b95db5fd95128459f6c922e2802e1c79a06f538a6f5285499fcf37707

Observation 644fd330-d804-43e6-a659-90515ff59a9f · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:39.839500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.839500Z digest=sha256:6993f8f57345dfb070abfd3202f57371ec565cb7874decd6c32d10d45264fcc6

Observation 7e671f56-039f-4d6d-9ebf-f699aafa29e0 · outbound

This paper cites Discovering latent knowledge in language models without supervision.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Discovering latent knowledge in language models without supervision

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.954047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:39.912855Z digest=sha256:f0dbed99ee05f0bf46ac34b1af7696428b4865d8f2a7a1e7d88b39ec9103b617

Observation d07b328d-4fc3-4dde-b7a7-3bc9cf23f8bf · outbound

This paper cites Language Models (Mostly) Know What They Know.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Language Models (Mostly) Know What They Know

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.012125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.012125Z digest=sha256:a049c042281e85c7a5e7ccfa278aea40fb3d98cc8488d5880ceeef2760c5efbb

Observation 5d544f98-6e7d-473b-832b-d53a9d3d1f1b · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Detecting hallucinations in large language models using semantic entropy.Nature, 630(8017):625–630, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.103165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.103165Z digest=sha256:7f93ba4c923c2764a7a81f2baa53c810b32a049416c4a40cc840b19a23334fc5

Observation 98fdfb8b-4d9d-4f32-8da3-7de8b050e27b · outbound

This paper cites On calibration of modern neural networks.

Revisiting Uncertainty Estimation and Calibration of Large Language Models On calibration of modern neural networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.212539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.212539Z digest=sha256:005a3850bcacbe939c2859ea56a24ea68eb7ab245ef4286d1310b7b3f9c7e66e

Observation dfd90084-708e-4ff0-a687-ead5ff14c5e0 · outbound

This paper cites Benchmarking uncertainty disen- tanglement: Specialized uncertainties for specialized tasks.Advances in neural information processing systems, 37:50972–51038, 2024.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Benchmarking uncertainty disen- tanglement: Specialized uncertainties for specialized tasks.Advances in neural information processing systems, 37:50972–51038, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.826331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:40.320314Z digest=sha256:fe92f5ab4e5ffa8d4013a51ab33b4b249f320a4530aed8292a383e9224c30d2e

Observation edc5dd6f-45f2-44c5-a9fc-b7813b164eed · outbound

This paper cites Measuring calibration in deep learning.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Measuring calibration in deep learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.430239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.430239Z digest=sha256:c7dbb6467da5563877c9bcc4c5cb6c2a8d688c5de2b2e16cfe5af654fe3fb38a

Observation bf21376f-ff4a-48b1-92a2-72b3e96a872f · outbound

This paper cites Predicting good probabilities with supervised learning.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Predicting good probabilities with supervised learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.530832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.530832Z digest=sha256:586b069010dfd3090ec0f7f61864f82caa5a18e37917335094a6c8ff7b4b0d4d

Observation 46cf49c1-3722-4012-8942-0e7e3b5abdae · outbound

This paper cites The Llama 3 Herd of Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models The Llama 3 Herd of Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.568605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.568605Z digest=sha256:a3d8e6c164ac16f586041aa98e23d28954c716a070f327d7cc26f884c40a80d1

Observation dcc0c5ba-e9ed-4731-8080-4ea7282a70dd · outbound

This paper cites DeepSeek-V3 Technical Report.

Revisiting Uncertainty Estimation and Calibration of Large Language Models DeepSeek-V3 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.672286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.672286Z digest=sha256:9b447821860ee7e6325f57e8f9703b31b020b9bcf401c7e095bc852fdf6f72d8

Observation 28897efc-495f-44b7-bd9a-c3e2e5921e7d · outbound

This paper cites Qwen2.5 Technical Report.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Qwen2.5 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:40.769319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:40.769319Z digest=sha256:c5f09a2498cd578ad030496088fd4b6119e69bf650fe72bf7398911db6fec0a6

Observation dfca067e-37b5-48a8-bf61-e03739bbbe41 · outbound

This paper cites Identifying and mitigating vulnerabilities in llm-integrated applications.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Identifying and mitigating vulnerabilities in llm-integrated applications

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.707808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:40.902069Z digest=sha256:6c9dfe67af4ef6e3ad436486254ea064b5aeae8e7645e0c85bc357fd053366ab

Observation b7d3af9e-c9c5-436c-8b93-276255bf9c66 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.026953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.026953Z digest=sha256:879243a2112dcbb87a2ca8af7791a68a1d6ece4d5964db59099c0ed5cb809085

Observation 7e56c294-799c-4035-b00d-968fef65f92d · outbound

This paper cites Introducing gpt-4o: Openai’s new omni model.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Introducing gpt-4o: Openai’s new omni model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.627738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:41.103318Z digest=sha256:56a5fff8f3830555c14749a413dfbc409ed97b86c3e13444f7c574e53d37a218

Observation 205b571d-ccb9-4ea5-90d3-c80f83777942 · outbound

This paper cites Gpt-4.1 overview.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Gpt-4.1 overview

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.540243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:41.172123Z digest=sha256:d46659c4ec155ced3dae15b0c56e26e19c12a95b518ea3ae0a36b7f2b53a0ffc

Observation 512247d6-a942-4999-81d9-86498a28fed7 · outbound

This paper cites Introducing the claude 3 model family.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Introducing the claude 3 model family

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.455340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:41.282128Z digest=sha256:eed1c6670faaa3587a1acd715dc5258ccbe0e9c3ab0527ad06b97a10c6cb76c6

Observation 0d888322-b373-4cb2-a7d7-eb0df487fed1 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.387827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.387827Z digest=sha256:316c43fa1f2906b3d04abdc98014c61da9c27c13f9875a3cbd9d10ff3d93fff8

Observation 26d9cfaa-8dd7-44ff-9a27-a25118f1dbd6 · outbound

This paper cites Qwen3 technical report.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Qwen3 technical report

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.338288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:41.466584Z digest=sha256:2805decc11bbac4ac010f5b272463cd7cd4406602abfe64af2f13168e24e8c5d

Observation aa3f8ebf-8f15-438f-9205-c886c88f0e47 · outbound

This paper cites The kendall rank correlation coefficient.Encyclopedia of measurement and statistics, 2:508–510, 2007.

Revisiting Uncertainty Estimation and Calibration of Large Language Models The kendall rank correlation coefficient.Encyclopedia of measurement and statistics, 2:508–510, 2007

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.219279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:41.561704Z digest=sha256:26c223796619d29f58c40ee5bd18c2151772b433b3bc3e50323b6023f5276589

Observation 1b059fc6-3b14-4a6c-acab-2704abb5c733 · outbound

This paper cites i’m not sure, but.

Revisiting Uncertainty Estimation and Calibration of Large Language Models i’m not sure, but

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:43.108420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:59:41.651489Z digest=sha256:421ff98f8134e24bccb15827e9ccdc2a29e63efb2ee02b5f9548ef794a34c8b3

Observation 32c81645-cadb-4dfd-8b68-1c9f47e25d8e · outbound

This paper cites Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.765614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.765614Z digest=sha256:91a901be77746da89315f4be04c58750e8fa92e46485c37d0111d8b5a0b76a8f

Observation dee90a92-4161-4657-a699-700ec5fb9550 · outbound

This paper cites The internal state of an llm knows when it’s lying.

Revisiting Uncertainty Estimation and Calibration of Large Language Models The internal state of an llm knows when it’s lying

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.862601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.862601Z digest=sha256:116af74deb74e972858cd06714c34e8d8df35c2a71f4d3196937bb11ed8da250

Observation 10675e64-2b09-480d-83f2-e5211afa71ea · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Inference- time intervention: Eliciting truthful answers from a language model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:41.990181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:41.990181Z digest=sha256:40b85a207f2a10adf070445c492f803a5bace7c7cc2c7a1c72d6243c698cd3cc

Observation 8f7b23fa-2184-4464-a311-2d9a6e5d2f00 · outbound

This paper cites Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:42.079253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.079253Z digest=sha256:5fda9bfaa39a3bfae392c2c0fca6c87f00aa6a89e2e860993b0b7b506a5ccec9

Observation 8068ad81-a5df-45a1-9833-b1b3eef73a3d · outbound

This paper cites A Survey of Confidence Estimation and Calibration in Large Language Models.

Revisiting Uncertainty Estimation and Calibration of Large Language Models A Survey of Confidence Estimation and Calibration in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:42.188284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.188284Z digest=sha256:00315f976623486178b03765e6f7c35ab051dc8841bee75c0e371a6b8b277260

Observation 914ff920-0b70-419f-ae25-9d0adb44237c · outbound

This paper cites A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice.

Revisiting Uncertainty Estimation and Calibration of Large Language Models A Survey of Uncertainty Estimation in LLMs: Theory Meets Practice

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:42.294102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.294102Z digest=sha256:0491990244d2a2c2dcba946195280cccd73d61bdc94e88accedba74a6f1e81db

Observation 772fa1e4-ebcf-417f-9f15-0e801f62ac9d · outbound

This paper cites Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:59:42.412963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:42.412963Z digest=sha256:22a8f2a533baf7d7356864056d18a31952e2897d2194ec0f04753b860d3c18cc

Observation 3747804d-053b-4181-997c-b61cd5e3e967 · outbound

This paper cites an unresolved cited work.

Revisiting Uncertainty Estimation and Calibration of Large Language Models Unresolved cited work

Reference 2024

Resolution
parse uncertain
no resolver link, observed 2026-08-07T12:59:39.612033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:39.612033Z digest=sha256:b292860dd7d94453ee2e3280d48109c5b37ba589cd1ae8b3f71ec9ec1edc017e

Pith citing papers

Observation a5ed4c11-5ded-4721-9f57-0600103d5539 · inbound

Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models cites this paper.

Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T20:21:55.864745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:21:55.864745Z digest=sha256:2a3c3c8e86d713534fd2dae8eeefdd0d366c7b4f9b2d5722d953a78be6058acc

Observation 82d38082-ecdf-401c-8854-f51fd3962bd9 · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:44:05.672258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:5e6da5b0f251d4329db12b30125c457b7c8347116a30fb53f2d15cde71730752

Observation 88f41d5f-8b82-49f5-9136-fcab8f089423 · inbound

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment cites this paper.

Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:58.469843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:13:29.201794Z digest=sha256:e621ca4c9dfa418c174e6166ede277fd890d46e2927b9cfd94092aedad5f4913

Observation 01f33cb2-5767-4c0b-97d4-68032e0a9035 · inbound

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation cites this paper.

"I Don't Know" -- Towards Appropriate Trust with Certainty-Aware Retrieval Augmented Generation Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:11.565452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:30:03.367870Z digest=sha256:1a9c2cb1794fa778747594fe1e38b88d7c1117bf119ae4d9326a9034d99433fd

Observation 668e8321-cc48-42cd-af59-9e178617f734 · inbound

Hallucinations Undermine Trust; Metacognition is a Way Forward cites this paper.

Hallucinations Undermine Trust; Metacognition is a Way Forward Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:07.202112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T14:29:54.924293Z digest=sha256:ba1c44b54a9955974fe1a4d9f4ee16369a25278a70f4c4ea1d7e34975c3cb537

Observation a9d14390-04ab-4387-9f8e-a6b2c7d95f72 · inbound

Hypothesis generation and updating in large language models cites this paper.

Hypothesis generation and updating in large language models Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-08T21:14:12.533436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T14:43:56.125546Z digest=sha256:493f699e12ef7e96af97490f9a83d81b52cf2c774c8b3c423f8fe8589b5254de

Observation 8b2a0696-382c-455e-9ead-adb525a9ae63 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.021788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:4eae0c0ced29ecdd1c7a5d964fe9769cea6bda6f409256f13132587e14ea5f20

Observation 61d0be8f-0d5f-463c-a19b-f51d597d8b84 · inbound

Retrieval-Augmented Linguistic Calibration cites this paper.

Retrieval-Augmented Linguistic Calibration Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.155542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:48:03.641696Z digest=sha256:16ebcc57f95d3c196ea03ef577342ad1ec8411c3accaf4f82ae103e0dc56f0c7

Observation 80de674f-289e-40eb-851c-90d35991b82e · inbound

Retrieval-Augmented Linguistic Calibration cites this paper.

Retrieval-Augmented Linguistic Calibration Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.397448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:36:25.106668Z digest=sha256:0d48eb8e874ae0ab1835a60084419bcd0114b19b950bc6ba365eb648ee701e05

Observation e08bc5bb-8617-45de-bc2e-9a753e99e46e · inbound

Gradient-Guided Reward Optimization for Inference-time Alignment cites this paper.

Gradient-Guided Reward Optimization for Inference-time Alignment Revisiting Uncertainty Estimation and Calibration of Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.645365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T16:43:04.984866Z digest=sha256:3ab741079e8079dee9d199962c81a1c3861d1d0f25a6ce54f80af805f12d3f32