{"id":"1febb274-4d36-4e4b-906f-f844704710ef","arxiv_id":"2505.13774","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across six large reasoning models, counterfactual edits to thinking drafts show that intermediate steps are selectively followed and final answers frequently contradict the draft's stated conclusion.","lead":"The paper tests whether the 'thinking drafts' that reasoning AI models produce before answering actually drive their final answers, by injecting misleading steps and false conclusions. It finds that these models often ignore or override their own drafts, which matters for anyone using the drafts to monitor or control AI behavior.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The CONTINUE-vs-BACKTRACK faithfulness gap may be an artifact of inconsistent classifier rubrics in Appendix B.1.","rationale":"The reader's weakest assumption is that implicit rejection of an inserted counterfactual step is misclassified as unfaithful, which is a plausible construct-validity problem. My stress-test identifies a more specific and more directly load-bearing flaw: the Appendix B.1 classification prompts for shift-mapping interventions define EXPLICITLY_CORRECTED and CONSISTENTLY_FOLLOWED asymmetrically across CONTINUE and BACKTRACK conditions. The same observed behavior—acknowledge the remapping and adopt it—is scored as EXPLICITLY_CORRECTED for CONTINUE (then judged unfaithful if the conclusion shifts) and CONSISTENTLY_FOLLOWED for BACKTRACK (then judged faithful if the conclusion shifts). This alone can manufacture the paper's headline result that BACKTRACK steps are more faithfully integrated than CONTINUE steps. This is an internal inconsistency, not merely a missing human-label validation, so it weakens the paper's main comparative claim. However, the overall central claim that LRMs exhibit selective faithfulness and frequently fail to align with draft conclusions is not fully invalidated: even if the CONTINUE rates are underestimated, the absolute intra-draft faithful rates are around 45-68%, and the Draft-to-Answer results are measured through a separate intervention. The paper should be revised to unify the classification rubric and re-report the affected rate, so the appropriate verdict remains conditional rather than accept; since the reader already returned CONDITIONAL, I mark the verdict as UNCHANGED, with the condition now sharpened to this specific rubric artifact.","tokens_in":15036,"tokens_out":9003,"duration_ms":86197,"concrete_test":"Recompute Table 1's Shift columns using a single, consistent classification rubric for both CONTINUE and BACKTRACK shift-mapping responses. Specifically, relabel all responses that explicitly mention the option remapping and then adopt the shifted conclusion as FOLLOW (or as a separate category) in both conditions, then recompute faithful rates and the BACKTRACK-minus-CONTINUE gap. If the gap changes by more than a few percentage points or changes sign, the A1 finding is an artifact of the asymmetric classifier prompt rather than a property of the models.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline intra-draft result (A1: BACKTRACK steps are more faithfully integrated than CONTINUE steps) is at risk from an internal inconsistency in the behavior classifier, not just from missing human validation. In Appendix B.1, the shift-mapping classification prompt defines the two categories differently across step types. For a CONTINUE shift, EXPLICITLY_CORRECTED is defined as \"the model explicitly detects the discrepancy between the two mappings or reiterates the original mapping,\" with no requirement that the model reject the new mapping. For a BACKTRACK shift, EXPLICITLY_CORRECTED is defined as detecting the discrepancy or reiterating the original mapping \"and doesn't adopt the new mapping,\" while CONSISTENTLY_FOLLOWED explicitly includes \"it recognizes the discrepancy but adopts the new mapping.\" Consequently, the identical model behavior—noticing the option remapping and then adopting the shifted labels—is labeled CORRECTION in the CONTINUE condition and FOLLOW in the BACKTRACK condition. Because the metric δIntra scores CORRECTION against the original conclusion and FOLLOW against the shifted conclusion φ(ANS(T)), the CONTINUE case is counted unfaithful while the BACKTRACK case is counted faithful. This built-in asymmetry inflates the reported BACKTRACK advantage and deflates CONTINUE faithfulness. The paper's Table 1 shift columns, and the aggregated A1 comparison, are therefore not a clean measure of step-type faithfulness; they partly encode a difference in the labeling rubric. The reader's concern about implicit rejection is related but distinct: this is not merely a construct-validity question about what counts as faithful, but an internal inconsistency in how the same response is classified in the two conditions being compared. A reanalysis with a unified rubric could shrink or reverse the BACKTRACK-vs-CONTINUE gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a counterfactual intervention framework for measuring the faithfulness of thinking drafts in large reasoning models (LRMs). It defines two dimensions: Intra-Draft Faithfulness, which tests whether an inserted counterfactual reasoning step (a mistaken CONTINUE or BACKTRACK step) is causally integrated into subsequent reasoning and the draft conclusion, and Draft-to-Answer Faithfulness, which tests whether the final answer depends on the draft and matches its stated conclusion. The framework is applied to six open-source LRMs on GPQA Diamond and an MMLU subset, using thinking drafts from DeepSeek-R1, Qwen3-32B, and the models themselves. The main reported findings are that BACKTRACK steps are integrated more faithfully than CONTINUE steps, explicit corrections yield higher faithfulness than following the inserted step, and the answer stage frequently introduces additional reasoning that can change or contradict the draft conclusion. The authors release code and data for the benchmark.","tokens_in":15272,"tokens_out":4914,"duration_ms":44375,"significance":"If the measurement framework is sound, this paper provides a useful and much-needed benchmark for a property that underpins monitoring and control of reasoning models. The design has notable strengths: the experiments use greedy decoding with temperature 0 for reproducibility, multiple model families and draft sources are covered, and the draft-to-answer comparison between standard and immediate answering (Table 3) is a clean, interpretable way to measure whether the answer stage merely restates the draft. The paper also ships open code and data, which is a concrete asset. However, the intra-draft faithfulness results rest entirely on LLM-based classification rubrics that are not validated against human judgment and that contain a step-type asymmetry in the classifier definitions (Appendix B.1). This asymmetry threatens the headline claim that BACKTRACK steps are more faithfully integrated than CONTINUE steps. The draft-to-answer results are less exposed to this classifier problem and are the more robust part of the contribution.","major_comments":[{"comment":"The shift-mapping classifier uses different definitions for CONTINUE and BACKTRACK. For CONTINUE, EXPLICITLY_CORRECTED is defined as \"the model explicitly detects the discrepancy between the two mappings or reiterate the original mapping,\" with no requirement that the model reject the new mapping. For BACKTRACK, EXPLICITLY_CORRECTED requires that the model \"doesn't adopt the new mapping,\" while CONSISTENTLY_FOLLOWED explicitly includes the case where the model \"recognizes the discrepancy but adopts the new mapping.\" Consequently, the same observable behavior—noticing the remapping and then adopting the shifted labels—is classified as EXPLICITLY_CORRECTED in the CONTINUE condition and as CONSISTENTLY_FOLLOWED in the BACKTRACK condition. Because δIntra in §3 scores CORRECTION against the original conclusion ANS(T) and FOLLOW against the shifted conclusion φ(ANS(T)), this asymmetry systematically increases the measured BACKTRACK faithfulness and decreases the measured CONTINUE faithfulness. The BACKTRACK-vs-CONTINUE gap in Table 1 and finding A1 therefore confound step type with the classifier rubric. The authors should reclassify with a symmetric rubric, or report the ambiguous \"recognizes but adopts\" behavior separately, and re-run the analysis.","section":"Appendix B.1 and §3"},{"comment":"The faithfulness metric δIntra assumes that a model's response to a counterfactual step must be either explicit verbal correction or explicit adoption of the step's logic. Appendix B.3's top example is labeled as an unfaithful step-following case because the model \"does not explicitly mention or correct the mapping but implicitly reverts to the original mapping\" while producing the original answer. This is a case where the final conclusion is faithful to the original draft, yet the framework scores it as unfaithful solely because the behavior is not narrated. If faithful reasoning can occur without explicit narration, the reported absolute faithfulness rates are systematically too low, and the comparison between step types could be distorted. The paper should provide a robustness check based on the final answer only, or human annotations showing that silent correction is negligible in these models.","section":"§3 and Appendix B.3"},{"comment":"All decomposition, intervention generation, conclusion extraction, and behavior classification are performed by LLM annotators (GPT-4O-MINI and Qwen2.5-Instruct) with no reported human-agreement metrics or inter-annotator consistency checks. Given that the headline findings A1 and A2 are computed from these classifications, and given the rubric asymmetry identified above, the paper should report (a) human agreement on a sample for each classification task, and (b) the marginal frequency of the ambiguous \"recognizes but adopts\" behavior in CONTINUE versus BACKTRACK conditions, so that readers can quantify the impact of the rubric on the results.","section":"§4.1, Appendices A.2 and B.1"}],"minor_comments":[{"comment":"The heading reads \"Mesauring Intra-Draft Faithfulness\"; this is a typo for \"Measuring.\"","section":"§4.2 heading"},{"comment":"The caption text contains \"Webold\" instead of \"We bold.\"","section":"Table 1 and Table 3 captions"},{"comment":"The entry \"f94.32\" contains a stray \"f\" and should be \"94.32.\"","section":"Table 3, MMLU, R1-14B row"},{"comment":"The claim that RLVR-tuned models show the lowest Draft-Answer Consistency rates is supported on GPQA and on MMLU for QwQ, but on MMLU the R1-8B average is lower than OR1's average; the sentence should be qualified to avoid over-generalization.","section":"§4.3.2, A3.2"},{"comment":"The definition of an exploitation block is somewhat ambiguous: it says a block \"starts with a BACKTRACK step and a contiguous sequence of CONTINUE steps that precedes another BACKTRACK step,\" but then notes the first block may contain only CONTINUE steps. Please clarify the intended segmentation rule.","section":"Appendix B.1"}],"recommendation":"major_revision","confidential_remarks":"The draft-to-answer half of the paper, especially the immediate-versus-standard answering comparison in Table 3, is the more robust contribution and could be published even if the intra-draft comparison is substantially repaired. I would ask the authors to re-analyze the intra-draft results with a symmetric classifier and to provide human-agreement estimates before resubmission. If the BACKTRACK advantage disappears under the corrected rubric, the abstract and Section 1 claims must be revised accordingly; the current version overstates the support for that specific conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take on arXiv:2505.13774. What's new here is a systematic counterfactual-intervention framework for faithfulness of LRM thinking drafts, split into intra-draft (do inserted steps affect the draft conclusion?) and draft-to-answer (does the final answer depend on the draft?). It benchmarks six open models on GPQA and MMLU, releases data and code, and reports patterns: models are more faithful to BACKTRACK steps than CONTINUE, the answer stage often overrides the draft conclusion, and RLVR-tuned models (QwQ, OR1) are less draft-to-answer faithful. That last set of findings is genuinely interesting and, if it holds, matters for monitoring and control.\n\nWhat it does well: the intervention design is careful — restating steps, shift-mapping and corrupt-option perturbations, three insertion locations, both self-generated and external drafts. Two-dimensional decomposition of faithfulness is sensible.\n\nNow the soft spots. The biggest one is internal to the classifier, not just missing human validation. In Appendix B.1, the shift-mapping classification prompts define the same behavior differently. For a CONTINUE shift, EXPLICITLY_CORRECTED includes 'reiterates the original mapping' with no requirement to reject the new mapping. For a BACKTRACK shift, EXPLICITLY_CORRECTED requires 'and doesn't adopt the new mapping,' while CONSISTENTLY_FOLLOWED explicitly includes 'recognizes the discrepancy but adopts the new mapping.' So a model that notices the remapping and adopts it gets labeled corrected in the CONTINUE condition and followed in the BACKTRACK condition. Since the δIntra metric scores correction against the original conclusion and following against the shifted one, the same behavior counts as unfaithful in CONTINUE and faithful in BACKTRACK. That asymmetry alone can manufacture the headline CONTINUE-vs-BACKTRACK gap in the shift columns. The other interventions and the corrupt-option prompts don't have quite this problem, but the aggregated A1 claim rests on this comparison.\n\nSecond, no human agreement metrics for the LLM step decomposer, intervention generator, answer extractor, or behavior classifier. The faithfulness definition also counts implicit rejection as unfaithful — Appendix B.3 shows the model reverting to the original mapping without explicitly narrating it, labeled as unfaithful step-following. That is a defensible operationalization if you care about explicit verbalization, but it will systematically lower absolute rates and the paper doesn't validate it.\n\nThird, comparative claims like 'consistent and significant gap' have no significance tests or error bars.\n\nHow much does this matter? The broad conclusion — thinking drafts are selectively faithful and the answer stage can override them — is plausible and directionally supported by multiple metrics. The specific BACKTRACK advantage is the least trustworthy result in the paper. A reanalysis with a unified rubric and human-validated classification could shrink or reverse it. As is, the paper deserves a serious referee, but I would not take the step-type finding at face value.\n\nWho it's for: interpretability and safety folks working on monitoring/control of reasoning models. I'd probably cite the draft-to-answer results with a caveat. Bring to reading group? Maybe — the rubric inconsistency is a good teaching example of how measurement choices shape findings.","headline":"Useful protocol for measuring thinking-draft faithfulness, but the headline BACKTRACK-vs-CONTINUE gap is suspect because the classifier rubrics differ across conditions.","tokens_in":15872,"tokens_out":2826,"would_cite":false,"duration_ms":24340,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large reasoning models are only selectively faithful to the thinking drafts they produce before answering, so those drafts cannot be treated as a reliable record of how the final answer was reached.","keywords":["thinking draft faithfulness","large reasoning models","counterfactual intervention","chain-of-thought","backtracking steps","draft-to-answer consistency","interpretability","monitoring and control"],"falsifier":"Carry out the same counterfactual insertions but have independent human raters judge whether each continuation is faithful reasoning, including cases where the model silently returns to the original mapping. If human raters call those cases faithful and the gap between BACKTRACK and CONTINUE faithfulness rates disappears or reverses, the paper's asymmetry and low faithfulness rates would be artifacts of the classifier rather than properties of the models.","tokens_in":14836,"feed_emoji":"🧠","tokens_out":4390,"duration_ms":40092,"temperature":0.7,"pith_summary":"This paper asks whether the long thinking draft a large reasoning model writes before answering actually drives the final answer. To test this, the authors insert counterfactual reasoning steps into drafts and check whether the model integrates, corrects, or ignores them. They also rewrite the draft's conclusion and compare answers given with and without an explanatory answer stage. Their reported finding is that faithfulness is selective: inserted backtracking steps are more faithfully integrated than ordinary continue steps, and the answer stage often introduces new reasoning that shifts the final answer away from the draft's stated conclusion. If this holds, thinking drafts cannot be trusted as a transparent record for monitoring or controlling models.","feed_headline":"Thinking drafts do not reliably drive final answers","feed_subtitle":"Counterfactual probes show backtracking steps are followed more than ordinary steps, and answer stages add new reasoning.","key_machinery":"The measuring instrument is a counterfactual intervention on the thinking draft. A draft is decomposed into CONTINUE and BACKTRACK steps; the probe inserts a restating step (remapped answer choices or a corrupted option) at the initial, middle, or final block, and an LLM judge labels the model's continuation as explicit correction or step following. Faithfulness is scored by whether the draft conclusion stays the same under correction or changes according to the intervention's mapping under following. For draft-to-answer faithfulness, the draft conclusion is altered directly or plausibly, and the answer is compared under standard answering with a free explanation versus immediate answering with none. These paired conditions isolate whether the draft alone determines the final answer.","core_discovery":"The paper's central claim is that current large reasoning models exhibit only selective faithfulness to their own thinking drafts. On the intra-draft dimension, an inserted counterfactual step—either a mistaken continuation or an explicit backtrack—is often not integrated: the model may silently revert to its original reasoning, and step-following changes the draft conclusion less faithfully than explicit correction. Backtracking insertions produce higher faithfulness rates than continue insertions. On the draft-to-answer dimension, the answer stage performs additional computation beyond summarizing the draft; standard answering changes the answer relative to immediate answering in roughly a third of GPQA cases for most models, and consistency with the draft's altered conclusion is often low, especially for reinforcement-learning-tuned models. The paper concludes that drafts are not a faithful record of the computation that produces the final answer.","pith_inferences":["The measurement may underestimate faithfulness: silently ignoring an inserted error and keeping the correct answer might reflect genuine reasoning that the classifier does not credit, so the reported faithfulness rates are plausibly lower bounds.","A testable extension would prompt or train models to narrate corrections explicitly and check whether faithfulness rates rise; that would show whether the deficit is partly a narration convention rather than a computational failure.","The same counterfactual method could be applied to non-reasoning language models to see whether selective faithfulness is specific to large reasoning models or general to next-token generation.","For control applications, the direct-versus-plausible alternation result suggests controls written as plausible revisions should steer larger models more reliably than bare assertions."],"forward_implications":["Monitoring models by reading their thinking drafts alone is unreliable; the answer-stage reasoning also needs to be inspected.","Inserting thinking content to steer a model will work more reliably for backtrack-style revisions than for ordinary continue steps.","Larger models show higher intra-draft faithfulness, but model scale alone does not guarantee draft-to-answer consistency.","Reinforcement-learning-tuned models are less likely to follow edited draft conclusions, making them harder to control through draft edits.","Suppressing the explanatory answer stage, by forcing an immediate answer, aligns final answers with the draft conclusion more often than free generation does."],"supporting_citations":[{"why":"Supplies the definition of faithfulness as intermediate steps accurately reflecting how the final answer is derived.","marker":"[14]"},{"why":"Earlier method of editing chain-of-thought steps to observe causal influence on final outcomes, which the paper adapts to thinking drafts.","marker":"[17]"},{"why":"Motivates the need for draft faithfulness by showing how monitoring models inspect thinking drafts.","marker":"[5]"},{"why":"Application that controls reasoning by inserting thinking content, whose reliability the paper's findings constrain.","marker":"[26]"},{"why":"Prior evidence that reasoning models do not always say what they think, used as a comparison point for draft-to-answer faithfulness.","marker":"[6]"},{"why":"Related prior finding that chain-of-thought reasoning is not always faithful, extended here to thinking drafts.","marker":"[3]"},{"why":"DeepSeek-R1 provides both a source of benchmarking drafts and the families of evaluated distilled models.","marker":"[11]"},{"why":"Qwen3 provides benchmarking drafts and base models for several evaluated models.","marker":"[22]"},{"why":"Supplies the GPQA Diamond dataset of hard graduate-level multiple-choice questions used in the experiments.","marker":"[20]"},{"why":"Supplies the MMLU dataset used for simpler factoid multiple-choice questions.","marker":"[13]"}],"fun_headline_variants":["Thinking drafts don't reliably steer AI final answers","AI models ignore parts of their reasoning drafts","When AI drafts stray, answers often ignore them","Selective faithfulness in AI thinking drafts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework counts a model as faithful only when an inserted error is either explicitly rejected or adopted with a matching change of conclusion; silently ignoring or reverting the insertion while keeping the right answer is classified as unfaithful.","fun_headline_variants_meta":{"raw":{"variants":["Thinking drafts don't reliably steer AI final answers","AI models ignore parts of their reasoning drafts","When AI drafts stray, answers often ignore them","Selective faithfulness in AI thinking drafts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1375,"prompt_tokens":880,"completion_tokens":495,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":439}},"tokens_in":496,"tokens_out":495,"duration_ms":5357,"temperature":1.0,"reasoning_tokens":439,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:10:35.174425+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Carry out the same counterfactual insertions but have independent human raters judge whether each continuation is faithful reasoning, including cases where the model silently returns to the original mapping. If human raters call those cases faithful and the gap between BACKTRACK and CONTINUE faithfulness rates disappears or reverses, the paper's asymmetry and low faithfulness rates would be artifacts of the classifier rather than properties of the models.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of faithfulness as intermediate steps accurately reflecting how the final answer is derived."},{"cited_title":"Reasoning models don’t always say what they think","cited_arxiv_id":null,"evidence_quote":"Prior evidence that reasoning models do not always say what they think, used as a comparison point for draft-to-answer faithfulness."}],"review_version":1}