Pith. sign in

Paper Citation Record · LEDGER

Confidence Calibration in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2605.23909.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23909 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T13:23:25.998584Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T22:46:02.023365Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 3bb1368b-8637-42bd-8fea-4bd2c46109f8 · outbound

This paper cites A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges.

Confidence Calibration in Large Language Models A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:0909eefd37b40666ceb09543e03b25de9e59eb928ba2765c40381e416eb68c11

Observation f17d7937-dc3a-4ff7-acbc-1a3d5286361f · outbound

This paper cites Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova.

Confidence Calibration in Large Language Models Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:28cad7bcf68d3bf46786bd5a1ed30008524a1a9b18838b937718f6af3d797124

Observation ed143514-84d3-438f-9bbe-90b5e7a1ef07 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Confidence Calibration in Large Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:3a8841a51fc2f107197256d86f554ea708d2a4099417444e66586bb78fdf1301

Observation 1a1ef2c0-ac61-409b-a64e-69803af5b3c1 · outbound

This paper cites Do LLMs Implicitly Determine the Suitable Text Difficulty for Users?.

Confidence Calibration in Large Language Models Do LLMs Implicitly Determine the Suitable Text Difficulty for Users?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:7f655a09e4c5ebca15ea7c1f857c91b9b1196b3836a08ab814d827a54e9a93ec

Observation e07311e7-fc2b-46db-b7b9-d8fe1c409e19 · outbound

This paper cites On Calibration of Modern Neural Networks.

Confidence Calibration in Large Language Models On Calibration of Modern Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:e5797b2f74bef944ef34137f184d43e550d30fac20fea019340ad67f3b557b69

Observation bc9285a3-efde-4c96-b9c2-73009a49c621 · outbound

This paper cites InProceedings of the 56th ACM Technical Symposium on Computer Science Education V.

Confidence Calibration in Large Language Models InProceedings of the 56th ACM Technical Symposium on Computer Science Education V

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:3180f0ae11a11b5497312bf975e92bbd4a07fdd213aa8e7e4eb80b839e8d918e

Observation bad2fab0-8ad3-4381-b40a-30cfbb823d6d · outbound

This paper cites Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?.

Confidence Calibration in Large Language Models Can LLMs Estimate Cognitive Complexity of Reading Comprehension Items?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:6e5c7f7485ffb03eb39f2bc9733136b60de7461325186e2f85f5a103394ff838

Observation 2cf8fc05-6147-4117-842d-0113394201a8 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Confidence Calibration in Large Language Models Language Models (Mostly) Know What They Know

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:d98aef22f65615f14db8b85e405efbc638935642e510eb2653b3b6bb222ddb7c

Observation 39db4301-d476-4c17-9ad8-11ad53e2a37c · outbound

This paper cites Why Language Models Hallucinate.

Confidence Calibration in Large Language Models Why Language Models Hallucinate

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:3247c22041fcd7c0593488378c1bd8f65662648d0cb085d13a22544ed70f2b52

Observation c77dc202-54d1-4fe5-b38e-618c19b8e9aa · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Confidence Calibration in Large Language Models Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:e3ac789e015275d8004f6210bb826be7aa0d89a03bbad3c707b9f38dfc42e657

Observation aaed1cd9-d46e-486d-a891-31a5b4da21b8 · outbound

This paper cites Sarah Lichtenstein and Baruch Fischhoff.

Confidence Calibration in Large Language Models Sarah Lichtenstein and Baruch Fischhoff

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:2ea0e9f86385d4be7c90ca9043c2f2ce22216ce49ba82f0b442cfab186e8220d

Observation 8757ac57-c559-4823-82e4-01723cbc4ee2 · outbound

This paper cites an unresolved cited work.

Confidence Calibration in Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:b8881aa5dae173526af129a9418706be2f94a335bc5f2068c6b0141b6f2dcb8e

Observation 0b4ebb65-9591-4c6a-81d9-b03d351a3933 · outbound

This paper cites When are Bayesian model probabilities overconfident?.

Confidence Calibration in Large Language Models When are Bayesian model probabilities overconfident?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:e8db0e0fa8863972a82f18104290b48a284cf0fc852a23f27f5b8264c678ecae

Observation 8e9e26c0-1259-44d9-8d52-67e7696efd72 · outbound

This paper cites Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE).

Confidence Calibration in Large Language Models Understanding Model Calibration -- A gentle introduction and visual exploration of calibration and the expected calibration error (ECE)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:2e6381b7d50ad30f4805e7a64a158c473eae92c5ec6958afa6040ebb32c01cd2

Observation 377f3949-5341-42fb-a312-9654d38ea0d8 · outbound

This paper cites Web page.

Confidence Calibration in Large Language Models Web page

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:df32b0659cffaa98b543dcfa578363b1272f6fd836c63cf285cffb50209a0da0

Observation 78c615ca-479a-4ca0-8ba3-5a52ab9fc33c · outbound

This paper cites GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration.

Confidence Calibration in Large Language Models GRACE: A Granular Benchmark for Evaluating Model Calibration against Human Calibration

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:a8390913da094e1ea59a7654e46f22011828a8f26af508014a50d71dbb224a1f

Observation d31f1d48-005f-4085-8e42-7ad2cafb3e90 · outbound

This paper cites Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christo- pher D.

Confidence Calibration in Large Language Models Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christo- pher D

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:ba1d7cd0deda0c25ccc71a86ef7a39c1c6c66c2666d445db168679c14cbfbe70

Observation c4f7439e-d337-448d-8cf7-f44492eb1078 · outbound

This paper cites Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback.

Confidence Calibration in Large Language Models Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:0e7a3b2265010a7a2f0345a30f4095e28014ae22ae00b0405bd1801bf3ec680e

Observation ef494730-f362-44a7-bf5e-2940ada56aa8 · outbound

This paper cites ArXiv:2506.23464 [cs].

Confidence Calibration in Large Language Models ArXiv:2506.23464 [cs]

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:37515955212962237276710812dd5be923f4f33d66abe0c178fcc2808b122dbe

Observation 77190e65-a272-4f00-8c38-62441e58fe05 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

Confidence Calibration in Large Language Models Crowdsourcing Multiple Choice Science Questions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:3fcfe2210e13a2565c02bc6674674c4e852887105ce7e28148f85900036c1364

Observation 65038dd5-d727-406b-8cf7-2ea4e468fde1 · outbound

This paper cites ArXiv:2505.01997 [cs].

Confidence Calibration in Large Language Models ArXiv:2505.01997 [cs]

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:641f7c042b3626d7905ed68477460509ed4fbcadd61f9d34d64c31c7fb833eb2

Observation 212fface-ed7d-47e1-bc23-d93c823bcc6d · outbound

This paper cites Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs.

Confidence Calibration in Large Language Models Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:127b791e47febbc40b0c139cff298af1b93fdf7978246037abb48f970149efae

Observation 8e44e662-0984-45c6-9ca1-bdd9020e8207 · outbound

This paper cites Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs.

Confidence Calibration in Large Language Models Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:c1ceee1c991180b722e4421c7c608d075e1a5e0024ec99a4ef28374704e30c2a

Observation 572da502-e57c-4723-8c76-364e78efbe62 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

Confidence Calibration in Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:2d6a2168c1ef12bad210a0097a8c99ae003f82a6110161815321863035e5e250

Observation de24ed61-024e-45bf-b294-624b5215f3ca · outbound

This paper cites AR-LSAT: Investigating Analytical Reasoning of Text.

Confidence Calibration in Large Language Models AR-LSAT: Investigating Analytical Reasoning of Text

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:4fd75584d09d871c33261745ed7b3bf7676bf7e097ca2e0bfe8eb439c3290216

Observation 2e5dabe9-470e-4ffd-b9fa-fd600f0cc18e · outbound

This paper cites Reasoning.

Confidence Calibration in Large Language Models Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T13:23:25.998584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:23:25.998584Z digest=sha256:f25770381019f6cf46b5d0af72521c619b510db9a0334d4f5bb496700f2acb2b

Pith citing papers

Observation c08ee2c7-56df-4e05-8858-90b412ecf3b2 · inbound

Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts cites this paper.

Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts Confidence Calibration in Large Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-01T22:48:43.408499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-01T22:46:02.023365Z digest=sha256:c78d3134fdbc3d6778189281218155071a2cc29222764ac997704e945753ad5f