Pith. sign in

Paper Citation Record · LEDGER

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

As of 10 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2606.11211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11211 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-04T19:41:53.668529Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T17:54:49.531461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact7
  • verified fuzzy3
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f87eb00c-b55e-4074-aff3-ed7288f41cb4 · outbound

This paper cites S., and Robbins, H.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models S., and Robbins, H

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.135397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:6fa5a27f20894be3777f3c7b2c1adaf55d2ccf0c1095d1b4e01a81d1e84c0c77

Observation d6739967-4412-4e2a-9d37-28c8c6d44b2f · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.131904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:64e8a873f7b37a39a2d91d6f41498630c0ca276f0ab717d70e328425a1c2fed0

Observation 2cde472d-7047-441a-a028-5554742abd39 · outbound

This paper cites and Durrett, G.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models and Durrett, G

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.148597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:51c655f455480352e5673aed905a49cabdd895b846b34fcb8581cd36fc96e4ee

Observation 9117c447-94d2-4404-8c11-30177d24e448 · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Adaptive Computation Time for Recurrent Neural Networks

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.032546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:524422b7c73a6552409029f77ef294cf7e40a025d13b1f349940006f4668ba70

Observation 58863f1b-a17d-4ea0-aad8-f30999955fcb · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.130285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:c800f7dfe0763ebd0412a54600d760b328eb94e5444e1e2c0a3ebc934c02d7d3

Observation f5f11320-fba6-4092-8c37-1ae056d00f20 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.146722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:f3205f18437d436e8f4aeb4d5aab6ca9406a9ba0741b8094a398fb6b451043b2

Observation 37b82850-abff-4f42-9480-14f87e70ef49 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Language Models (Mostly) Know What They Know

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.026754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:3d83355a12b490d6878df0220a0e28ceb3ec2754762df7273e0b41169c3af95e

Observation ed7b77ce-7994-4b10-bdae-54e15a80edb4 · outbound

This paper cites S., Reid, M., Matsuo, Y., and Iwasawa, Y.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models S., Reid, M., Matsuo, Y., and Iwasawa, Y

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T21:20:09.153386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:d8658f9ebab61c8180c2f30f7eb15c8531ed69d91b2d02bb28d20d2e35735de8

Observation 8313b1c8-b01b-4d4f-84be-d8627593d63a · outbound

This paper cites Let's Verify Step by Step.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Let's Verify Step by Step

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.016854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:4ae88ec48ee663100116ed3025f6f041255ebe58ec8d2b46b359ca4a8526dd9a

Observation 38b7dd66-97df-498f-ae7c-2816b5a5c5d9 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.140748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:48d51623febd0cd06d128eae2634fe0347e182c54e2e4d7c823fc444ad76419c

Observation 0fe2b956-6397-4cd3-b797-8236d84f9571 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.142893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:61026e89bd559083a618c0e9c4699f3a127f24bf3bbe7c085eaa00369bb4c98d

Observation 2ed1e3c5-4b8f-4763-bf78-62e6dae5475e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.019780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:b26abec554dd1e2cb6c5d37621fdbe4c5afa8f88c38e4a55474f482f3fca6727

Observation 3e18dc46-94a8-40de-818b-c45bd05f6cca · outbound

This paper cites Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.029851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:3d787931a54903c818a4846213c5f5deaece01db7ade9aa3296238851f5938ad

Observation ed5194e4-7b5e-4a83-a55d-206447994df0 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.144853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:303cfc8a5883e76f85ea933f7bbbb90c6c88e6df894e8878787e1e4458ed88c8

Observation b5230aa5-0d4a-4a49-98b8-714690069b05 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:10.013847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:5729a79d2ba13b4409d45150728de35b8d073e5a58b7481b1034e1e01f7caf6d

Observation 2e6fd3ac-b290-4930-9c4e-029cc767df3a · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.136380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:39450f5051fb7c482b17c7f93bb21feee6fb7df6aa9d3862ed8618d26def6bfa

Observation 95b704c1-5d80-455c-b195-1aea19123b79 · outbound

This paper cites an unresolved cited work.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-04T21:20:09.138840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:f951bb620a27cb90d7a2bf42851cd44b66b37fd8661e03fa1c81ad2bafb43938

Observation a6d37504-b47c-46b1-9600-88db77864a59 · outbound

This paper cites Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models.

Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:10.023261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T19:41:53.668529Z digest=sha256:e98e69b171c14df43c84953d47db6f88680e55c17291c8007dafd63526619567

Pith citing papers

Observation d8879794-a8c6-45f1-8b95-c0b2815085c5 · inbound

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents cites this paper.

Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T17:54:49.531461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T17:54:49.531461Z digest=sha256:e957fcabda5a5e1bd78f5ac521002ebd4ccff7abef177e07318d5f118ab6d55b