Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:14:41.426953Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 4 inbound Pith citation observations for arXiv:2505.13774.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:14:41.426953Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:40:58.274985Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T11:16:03.303665Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 180ed1c2-402c-4c37-a06f-214f6bc498ca · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e8b36f-5b37-48aa-b2dc-3dc0f300879c · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Claude 3.7 sonnet system card
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34c1dd22-d3b5-4e4d-bc39-676eeb569201 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77de9039-6502-41bf-a255-06d87bedf54d · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Faithfulness Tests for Natural Language Explanations
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d40209e-5118-41dd-a464-adbac0917f02 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b22d84a6-ed08-4bf0-bcdf-323d9f46874d · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Reasoning models don’t always say what they think
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d5c304b5-0ab3-438d-ad9b-0814d8cde3eb · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be611117-1179-47ae-8c47-1f4361c5a698 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Are DeepSeek R1 And Other Reasoning Models More Faithful?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7524a57-ff57-4bbc-8d8c-45020960caec · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff84b833-00c3-4027-86ae-9077de8d2abc · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Are We Done with MMLU?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676e1310-73f4-422b-810d-d2c0efafab99 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04280e3e-e42f-4dd6-96c5-464770b080b9 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d8e926b-da02-4ab7-b2ac-6a6f3058a5bb · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Measuring Massive Multitask Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954a10ba-9621-4d09-996a-cc4a790ad7ab · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ffa26ccd-20ab-45cc-b2ff-6a8cf947a190 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models OpenAI o1 System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe3bdae5-dd59-4cf1-ab0d-2f5e22cdf334 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71115129-531f-45b2-aa7d-81532a1c224b · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc740fec-8a73-49e0-88cc-b5d23f55ed70 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Deepseek-r1 thoughtology: Let’s< think> about llm reasoning.arXiv preprint arXiv:2504.07128, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8be1fde-4913-4d0b-b731-03b5161b69d0 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Openai o3-mini system card, 2025
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23665c23-49a9-46d2-a464-8e8be3187dcd · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Gpqa: A graduate-level google-proof q&a benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1837272f-d26e-473d-a771-f5df8ef09d57 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models On the hardness of faithful chain-of-thought reasoning in large language models, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6170b7d5-2f72-44ad-b10e-eb7df06f0933 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Qwen3, April 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559eca07-c3e8-4c0f-915f-34ebca8f758a · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems, 36:74952–74965, 2023
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08edc20c-6720-4b79-8e12-a0e7bfd0a50c · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Thoughts are all over the place: On the underthinking of o1-like llms, 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae1cc333-b203-444a-85ab-89c5e24c75a0 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Chi, Quoc V Le, and Denny Zhou
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6f3aa6b-0698-4cef-a95f-f9b04bfcf147 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Effectively Controlling Reasoning Models through Thinking Intervention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 421c5a21-ffa4-4c50-857d-3f51eeaaf7cd · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Dynamic early exit in reasoning models, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c17e41-24d1-407d-922d-536833e06d41 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Dissociation of Faithful and Unfaithful Reasoning in LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28e89627-5a55-4984-a927-aef5d81f7385 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models ‘json { “perturbed_option
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6913a7e-b35e-469c-862f-1269869fcdc8 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models EXPLICITLY_CORRECTED
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b6004830-518f-4290-999d-c6890e2b9e1a · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c3a677ba-125c-4c57-9d50-d9484502c829 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models EXPLICITLY_CORRECTED
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a4f29cec-b128-4afd-8e38-b865fa1fc2ea · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 58b87729-4033-4c3a-a41a-ec2482933956 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models EXPLICITLY_CORRECTED
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d0ca7ea5-c00f-46c3-a6a0-be65dcbd3da7 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b170b2c-2e20-41ba-9b3f-fbf88a4689c0 · outbound
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models EXPLICITLY_CORRECTED
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation df660e94-2abd-4754-be99-949a2751cd90 · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7729dbc3-4d62-424b-b6a5-fa630867821b · inbound
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23ad4e73-b7ef-42c6-89ad-399ff2c6e743 · inbound
LLM Reasoning Is Latent, Not the Chain of Thought Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a20b96d2-668e-4fd7-917e-c93fd6d2546b · inbound
Risky Business: Measuring The Faithfulness-Safety Tension Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.