{"id":"5abb51d8-3618-4fcf-bbfc-cd279871b2c9","arxiv_id":"2605.28913","paper_version":1,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Cross-model CoT transfer works through distinct mechanisms—explicit answer extraction, partial reasoning scaffolding, or receiver competence—varying by benchmark and whether the receiver must answer from prefixes or may continue generating.","lead":"This paper examines how chain-of-thought reasoning traces from one model can help another model solve the same problem using a controlled provider-receiver setup. The findings on different transfer mechanisms could guide more efficient use of reasoning traces in multi-model AI systems.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Performance differences may reflect generation artifacts or benchmark formats rather than distinct CoT mechanisms","rationale":"The reader's weakest assumption directly identifies the key vulnerability for the central claim of non-single-phenomenon transfer. The abstract-only review correctly flags the missing isolation of semantic content from artifacts as the load-bearing point; no stronger internal inconsistency is visible from the provided description.","tokens_in":1715,"tokens_out":298,"duration_ms":28138,"concrete_test":"Recompute force-answer accuracy curves on AIME after replacing answer tokens in full prefixes with neutral placeholders while preserving prefix length and structure; if the performance jump disappears or shifts to match short-prefix baselines, the extraction mechanism inference is supported; repeat the masking on MMLU-Pro and ZebraLogic to test whether benchmark distinctions survive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that cross-model CoT transfer reflects multiple distinct mechanisms (answer extraction on AIME force-answer, receiver competence on MMLU-Pro, partial structured info on ZebraLogic, scaffolding in free-generation) depends on interpreting prefix-length performance curves as evidence of semantic contributions from the reasoning content. This requires that observed mode and benchmark differences arise from the actual reasoning steps in the prefixes rather than model-specific generation patterns (e.g., answer formatting or token distributions in provider outputs) or benchmark idiosyncrasies that receivers exploit regardless of prefix semantics. The abstract provides no indication of controls that isolate these factors.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces a provider-receiver experimental framework to dissect cross-model chain-of-thought (CoT) transfer. Providers generate reasoning traces on benchmarks including AIME, MMLU-Pro, and ZebraLogic; receivers are given progressively longer prefixes under two modes (force-answer, where the receiver must answer directly, versus free-generation, where it may continue reasoning). The central claim is that successful transfer is not monolithic but reflects distinct mechanisms: explicit answer extraction (AIME force-answer), receiver competence (MMLU-Pro), partial structured information (ZebraLogic), and scaffolding (free-generation across benchmarks). Receiver agreement is proposed as a gold-free signal for early stopping of provider reasoning.","tokens_in":1839,"tokens_out":506,"duration_ms":30352,"significance":"If the mechanistic distinctions survive controls for confounds, the work would meaningfully advance the field by replacing binary views of CoT transfer with a more granular account. The prefix-length curves, mode comparisons, and multi-benchmark design provide a reusable template for probing transfer; the agreement-based stopping rule is a concrete, immediately usable contribution. The absence of parameter fitting or invented formalisms keeps the claims tightly tied to observable behavior.","major_comments":[{"comment":"Experimental setup / methods (abstract and §3–4): The claim that performance trajectories reveal distinct semantic mechanisms (answer extraction vs. scaffolding vs. competence) is load-bearing and rests on the assumption that differences arise from CoT content rather than generation artifacts (answer formatting, token distributions, or benchmark idiosyncrasies). The provided abstract gives no indication of controls that isolate these factors (e.g., shuffled prefixes, format-normalized baselines, or token-distribution matching), leaving the central interpretation vulnerable to the alternative explanation raised in the stress-test note.","section":"Experimental setup / methods (§3–4)"}],"minor_comments":[{"comment":"The abstract introduces 'force-answer' and 'free-generation' without a one-sentence operational definition; adding this would improve accessibility for readers outside the immediate sub-area.","section":"Abstract"},{"comment":"Figure captions or §4 should explicitly state the number of providers/receivers, model sizes, and statistical tests used for the reported accuracy differences; these details are referenced in the reader's soundness assessment as currently underspecified.","section":"Results (§4)"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive review. Below we address the major comment directly and describe the revisions we will undertake.","responses":[{"response":"We agree that the mechanistic interpretation would be strengthened by explicit controls for generation artifacts. The current design already provides partial protection against uniform artifacts because the identical provider generation process produces qualitatively different prefix-length trajectories across benchmarks (explicit-answer leakage on AIME, competence-driven gains on MMLU-Pro, structured-information effects on ZebraLogic). Nevertheless, benchmark-specific idiosyncrasies or formatting effects remain possible confounds. In the revision we will add (i) shuffled-prefix baselines that preserve length and token statistics but destroy semantic order, (ii) format-normalized controls that strip or standardize answer formatting, and (iii) token-distribution matching where feasible. These additions will be reported in an expanded §4 and will directly test whether performance gains depend on CoT content rather than surface features. The stress-test note already flags the artifact concern; the new experiments will quantify its practical impact.","revision_made":"yes","referee_comment":"[Experimental setup / methods (§3–4)] Experimental setup / methods (abstract and §3–4): The claim that performance trajectories reveal distinct semantic mechanisms (answer extraction vs. scaffolding vs. competence) is load-bearing and rests on the assumption that differences arise from CoT content rather than generation artifacts (answer formatting, token distributions, or benchmark idiosyncrasies). The provided abstract gives no indication of controls that isolate these factors (e.g., shuffled prefixes, format-normalized baselines, or token-distribution matching), leaving the central interpretation vulnerable to the alternative explanation raised in the stress-test note."}],"tokens_in":1414,"tokens_out":361,"duration_ms":128322,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core finding is that cross-model CoT transfer is not one process. Prefix-length curves in force-answer mode point to answer leakage on AIME, receiver skill on MMLU-Pro, and structured partial information on ZebraLogic, while free-generation mode shows scaffolding benefits across tasks. They also note that receiver agreement can flag when to cut provider reasoning early.\n\nThe setup is straightforward and useful: one model generates traces, another receives growing prefixes under two generation rules. That controlled comparison plus the benchmark breakdown is the actual addition over prior transfer papers. The agreement signal is a practical side result.\n\nThe main uncertainty is whether the curves track semantic content or just generation artifacts and benchmark formats. The abstract gives no sign of controls like prefix shuffling, paraphrasing, or token-distribution matching, so the mechanistic labels rest on interpretation of the performance gaps. If the full paper adds those checks or statistical tests on the trajectories, the claims tighten; otherwise they stay plausible but not fully isolated.\n\nThis is for people already working on CoT efficiency or interpretability who want a diagnostic lens rather than a new method. The experimental frame is clean enough that a referee could usefully pressure the artifact question and the exact model sizes and prefix definitions. I would send it to review.","headline":"Prefix trajectories show CoT transfer splits into benchmark-specific mechanisms (answer leakage on AIME, competence on MMLU-Pro, partial structure on ZebraLogic), plus an agreement early-stop signal, but artifact controls remain the open question.","tokens_in":2329,"tokens_out":347,"would_cite":false,"duration_ms":20193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Cross-model chain-of-thought transfer works through answer extraction, reasoning scaffolding, or receiver competence depending on the task.","keywords":["chain-of-thought","cross-model transfer","reasoning transfer","large language models","provider-receiver framework","force-answer","free-generation"],"falsifier":"Replacing the actual words in the prefixes with random but length-matched text and finding that the same performance patterns across modes and benchmarks remain unchanged would show the differences are not caused by reasoning content.","tokens_in":2622,"feed_emoji":"🧠","tokens_out":671,"duration_ms":25666,"temperature":0.7,"pith_summary":"The paper tests how reasoning traces from one model help another solve the same problems by feeding progressively longer prefixes to a receiver model. It separates two response modes: one where the receiver must answer immediately from the prefix and one where it can keep reasoning first. Across benchmarks the same full trace succeeds for different reasons, sometimes because the answer is already visible, sometimes because the prefix guides further steps, and sometimes because the receiver already knows enough to use the prefix. This distinction matters because treating all CoT transfer as the same process would hide when shorter traces suffice or when transfer adds little value. The results also show that agreement among multiple receivers can mark the point at which additional provider reasoning stops helping.","feed_headline":"CoT transfer splits into answer extraction, scaffolding, or competence","feed_subtitle":"Prefix tests across benchmarks show each mechanism dominates on different tasks and response modes","key_machinery":"The provider-receiver framework that supplies increasingly long CoT prefixes to a receiver under force-answer versus free-generation conditions.","core_discovery":"In the provider-receiver setup, full traces transfer successfully, yet prefix analysis shows distinct mechanisms: force-answer transfer on AIME is driven mainly by explicit answer availability, on MMLU-Pro by receiver competence, and on ZebraLogic by partial structured-answer information. In free-generation mode partial prefixes improve performance by guiding continued reasoning, and answer agreement among receivers supplies a signal for stopping provider reasoning early. Overall, cross-model CoT transfer is not a single phenomenon.","pith_inferences":["For tasks where answer extraction dominates, very short prefixes that contain only the final answer might achieve most of the gain.","The same framework could be used to test transfer of other explicit artifacts such as code comments or proof steps.","If receiver competence is the main driver on some benchmarks, fine-tuning the receiver on the task might reduce the value of any transferred trace."],"forward_implications":["Full traces enable successful transfer on the tested benchmarks.","Partial prefixes can scaffold continued reasoning when the receiver is allowed to generate further steps.","Receiver agreement provides a usable signal for cutting provider reasoning short without losing transfer benefit.","Different tasks rely on different parts of the trace, so the same prefix length does not help equally everywhere."],"fun_headline_variants":["CoT transfer mechanisms vary across tasks and modes","Prefix tests show CoT transfer is answer or reasoning driven","Cross model CoT transfer depends on partial information","Full traces transfer but mechanisms differ by setup"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Performance differences between force-answer and free-generation modes, and across benchmarks, are driven by the semantic content of the CoT prefixes rather than by model-specific generation artifacts or benchmark quirks.","fun_headline_variants_meta":{"raw":{"variants":["CoT transfer mechanisms vary across tasks and modes","Prefix tests show CoT transfer is answer or reasoning driven","Cross model CoT transfer depends on partial information","Full traces transfer but mechanisms differ by setup"]},"model":"grok-4.3","cost_usd":0.006017,"raw_usage":{"total_tokens":2860,"prompt_tokens":691,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":60174500,"prompt_tokens_details":{"text_tokens":691,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2111,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":691,"tokens_out":58,"duration_ms":22590,"temperature":1.0,"reasoning_tokens":2111,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T13:09:32.604300+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Replacing the actual words in the prefixes with random but length-matched text and finding that the same performance patterns across modes and benchmarks remain unchanged would show the differences are not caused by reasoning content.","supporting_citations":[],"review_version":1}