Pith. sign in

Paper Citation Record · LEDGER

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2606.11211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11211 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-04T19:41:53.668529Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:54:49.531461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact7
  • verified fuzzy3
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f87eb00c-b55e-4074-aff3-ed7288f41cb4 · outbound

This paper cites S., and Robbins, H.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models S., and Robbins, H

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.135397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:94e2ede65904a7d5164c70b7ccd0a6a061ed3c2adc0aa1229670663a73adce56

Observation d6739967-4412-4e2a-9d37-28c8c6d44b2f · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.131904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:49cbf71a279b818dc59a9c20d9aa627448ad86196a1c51ffc88fa0ff9ce53d23

Observation 2cde472d-7047-441a-a028-5554742abd39 · outbound

This paper cites and Durrett, G.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models and Durrett, G

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.148597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:aa01ec5f83ef34d00a9a22ffbb93e185ce5eb4bab2daec599d29550655a60e27

Observation 9117c447-94d2-4404-8c11-30177d24e448 · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Adaptive Computation Time for Recurrent Neural Networks

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.032546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:9af811ee05648978f68f3d09039ff4f2f01b3720c0e9379ee5604695f61b7a8d

Observation 58863f1b-a17d-4ea0-aad8-f30999955fcb · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.130285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:0d95b033718f900d59d1ff46550a12b4a61a659263ffab82d0cb73499752c05b

Observation f5f11320-fba6-4092-8c37-1ae056d00f20 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.146722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:36306e3bc15b28aebf2407ef603b2855902055f650383d515227b8024ebc0d8e

Observation 37b82850-abff-4f42-9480-14f87e70ef49 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Language Models (Mostly) Know What They Know

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.026754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:e6d7bc2e0d0bb8a5a3a7dba5e04fe14460dbd67c8c541763b5afa1eef295a5ff

Observation ed7b77ce-7994-4b10-bdae-54e15a80edb4 · outbound

This paper cites S., Reid, M., Matsuo, Y., and Iwasawa, Y.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models S., Reid, M., Matsuo, Y., and Iwasawa, Y

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.153386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:d0886cf1f5aac3b1790150d744e9f4dbe74a4039ab80fd2798884f824257d164

Observation 8313b1c8-b01b-4d4f-84be-d8627593d63a · outbound

This paper cites Let's Verify Step by Step.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Let's Verify Step by Step

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.016854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:19e4077a4e88b094f83b14c265f92cf21a573be9f2db7c50640f2c9f90b07b28

Observation 38b7dd66-97df-498f-ae7c-2816b5a5c5d9 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.140748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:87df7d9037efd3ad4566f433030fde0e4421ab5184a4e5ab291eca4dc6bcc521

Observation 0fe2b956-6397-4cd3-b797-8236d84f9571 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.142893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:065cac12da0acaf276e2a29c9c22ef71eff133ff8499db79929dd83d04b45763

Observation 2ed1e3c5-4b8f-4763-bf78-62e6dae5475e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.019780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:194baa48b5032696a392d647c15815ebcfe6a2eb804850aa77b04e0f0d180713

Observation 3e18dc46-94a8-40de-818b-c45bd05f6cca · outbound

This paper cites Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.029851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:bbb4b97f64e3cf3f3db1a0078bbd37e533d2be8a41956ade454d519c1dde6f59

Observation ed5194e4-7b5e-4a83-a55d-206447994df0 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.144853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:d0f0065514c3ddf4515495ff6fefda60f4f138583cbfd761ecd21759453dbcae

Observation b5230aa5-0d4a-4a49-98b8-714690069b05 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.013847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:1e9db7c28d95e5e3104585540bb7ebf07ea5c6cf3ae0c88b3800fcd15f581834

Observation 2e6fd3ac-b290-4930-9c4e-029cc767df3a · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.136380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:0b1b47b3487fa601352870cd9c5974905dacbda30f421f62763e2cf61cd41f94

Observation 95b704c1-5d80-455c-b195-1aea19123b79 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.138840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:465c84840b1817b046a7367ea40a54576aa6cfe685e8167ed38f86be1fab6caa

Observation a6d37504-b47c-46b1-9600-88db77864a59 · outbound

This paper cites Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:10.023261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:8fe443fcfa2471eae7e914a20cb7dae4e2179b9a72fbacfc195f2fc8461c40d0

Pith citing papers

Observation d8879794-a8c6-45f1-8b95-c0b2815085c5 · inbound

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents cites this paper.

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T17:54:49.531461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:54:49.531461Z digest=sha256:e957fcabda5a5e1bd78f5ac521002ebd4ccff7abef177e07318d5f118ab6d55b