Pith. sign in

Paper Citation Record · LEDGER

Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.13213.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13213 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:36.447956Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:07:23.912546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6b3a3b5c-f38f-44e4-8fee-2de7ebce61bd · inbound

Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning cites this paper.

Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:47:42.669613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T07:45:50.292586Z digest=sha256:31d6632eac77faf02351369222671fdbc91e92e9ebcaeb7568dffc9ae3a13339

Observation 78a6a99f-2285-48aa-817a-140b6cece409 · inbound

Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure cites this paper.

Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:22:38.186542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T06:22:25.496269Z digest=sha256:b68b96450e0c0f91de23af858ee1ea679b1d6f2b1f7e2d8c7d05fd087c848e35

Observation d77c015c-0fb6-4b89-81cc-ca57c5ff6e60 · inbound

Resilient LLM-Empowered Semantic MAC Protocols via Zero-Shot Adaptation and Knowledge Distillation cites this paper.

Resilient LLM-Empowered Semantic MAC Protocols via Zero-Shot Adaptation and Knowledge Distillation Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:36.447956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:36.447956Z digest=sha256:d4f605aae617cd2a3524695572743e5da8a208357daaac0b8ac9cc91ecc23b51

Observation 48e9aed8-a786-4985-8e97-ca104f4f2d9b · inbound

Shapley Uncertainty in Natural Language Generation cites this paper.

Shapley Uncertainty in Natural Language Generation Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:55:57.711866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:55:57.711866Z digest=sha256:619a81a9151f74652d1dff467ba32ece0bb35c7a580e8cc9c54adc72d5ae7f67

Observation 62e24b5d-6bee-4a4e-b606-6396c400275f · inbound

Causal Evidence that Language Models use Confidence to Drive Behavior cites this paper.

Causal Evidence that Language Models use Confidence to Drive Behavior Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:44:05.679600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T09:43:05.524088Z digest=sha256:c113601156c70edb3beeffa5fb40808387c0ad333b6a34deca87bbc84f8d624d

Observation 41208fd2-6bb0-40be-981e-f0c1201f221c · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:37.021662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:efd26b1df9c6ed6f04e0a2ed6e413e581d5392b2ebd41260f55911c1b19f3386

Observation a2506d7f-6ae5-4966-9f74-774be58d8d5c · inbound

Instance-Adaptive Online Multicalibration cites this paper.

Instance-Adaptive Online Multicalibration Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:40.742330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:45:02.146222Z digest=sha256:ff45e551576f2f0460b364409e9a430cf8c6fe0a2d98c5b3951abce1941d23d0

Observation 8682f582-6878-4837-bdb2-f5c8047ba853 · inbound

Instance-Adaptive Online Multicalibration cites this paper.

Instance-Adaptive Online Multicalibration Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:41:21.584990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T09:41:09.689528Z digest=sha256:e09e66e5852b30e8aada7f5396a05a8a9578b454d462a37ec83cb4e258bf4b95

Observation ca86e480-9f27-48d4-84b8-949bc233cd15 · inbound

Task-Aware Calibration: Provably Optimal Decoding in LLMs cites this paper.

Task-Aware Calibration: Provably Optimal Decoding in LLMs Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:56:18.701791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:55:54.933129Z digest=sha256:1e692dff222de79c8fb4022edc5dc2b43c30ec9c1c8c25b7fe2363515535e7eb

Observation fc9ccb72-0e16-4a51-a59c-c0eda646e3c9 · inbound

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts cites this paper.

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:46:18.168776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T08:46:12.102599Z digest=sha256:21573243c8cf181210536f95f7d539e50efa200943a622f77c9388fd195ca8f7

Observation 96734832-56cb-42d8-90c9-2cd72e1a68b6 · inbound

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning cites this paper.

ConSteer-RL: Steering Reasoning Capabilities in Large Language Models via Confidence-Aware Reinforcement Learning Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:07:23.914209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T19:56:09.820812Z digest=sha256:207d5acab4302487f0f579fa2dfedcc2b086d754cd58624f40b3e36305e11fdd