{"id":"c05f0883-db14-4846-a397-125f23155484","arxiv_id":"2607.05722","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Joint AR–diffusion training yields one tri-mode LM that switches AR, diffusion, and self-speculation, beating open AR/diffusion models on accuracy and tokens-per-forward.","lead":"A single language model can run as autoregressive, diffusion, or self-speculative decoding by training AR and diffusion losses together. That flexibility yields higher tokens-per-forward and real-device throughput than strong open AR and diffusion baselines at 3B–14B scale.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Outperformance vs. open AR/diffusion/MTP baselines is only partially isolated from unmatched continuous-pretraining data, recipe, and starting checkpoint (Ministral3).","rationale":"The reader correctly isolates the weakest link for a systems paper whose strongest claim mixes architectural novelty with large-scale empirical superiority. The controlled ablations (global loss averaging, two-stage, α=0.3, AR loss) are clean and support complementarity; real-device SGLang numbers on GB200/RTX/DGX and the SOL construction are carefully defined. The residual risk is only that full-scale SOTA tables (Tab. 5–8, Fig. 1, Tab. 10) compare against heterogeneous open models, so the precise contribution of joint training versus data/recipe cannot be read off those tables alone. Because the paper already supplies the matched small-scale evidence and the systems results stand on their own measurements, the ACCEPT verdict with medium correctness risk and the noted caveat remains appropriate; no stronger internal inconsistency or metric error rises to the same load-bearing level.","tokens_in":30958,"tokens_out":633,"duration_ms":44302,"concrete_test":"From the identical Ministral3-8B checkpoint, run a pure-AR control for the same Stage-2 budget (300 B tokens, α=0) and identical SFT (45 B tokens); evaluate AR-mode accuracy on the Tab. 5 suite and attach an Eagle3-style MTP head for SPEED-Bench acceptance/throughput. If the joint model’s AR accuracy advantage and linear-SS acceptance length (Tab. 10) shrink by more than ~1–2 points / ~1 token relative to this matched pure-AR control, the architecture-driven portion of the headline gains is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim (8B instruct matches/exceeds Qwen3-8B accuracy at ~6× TPF / ~4× SPEED-Bench throughput, self-speculation beats Eagle3/MTP, SOL headroom) rests on models that receive 1 T pure-AR + 300 B joint tokens from Ministral3 plus proprietary-style SFT (Sec. 5.1–5.2). Baselines (Qwen3, LLaDA, Dream, SDAR, Eagle3) differ in data mixture, total tokens, and often evaluation harness (NeMo-Skills vs. official diffusion pipelines). Controlled 25 B-token ablations (Tab. 1–3) and AR-only vs. joint SFT (Tab. 3) support complementarity under matched data, but the full-scale SOTA comparisons do not; therefore the fraction of the reported accuracy/TPF/throughput gains that is truly caused by the joint objective + tri-mode inference (rather than stronger provenance) remains unquantified. This softens causal attribution for the strongest claim while leaving the systems packaging and real-device measurements intact.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper presents Nemotron-Labs-Diffusion, a family of 3B/8B/14B language models (base, instruct, and VLM) trained with a joint AR–diffusion objective (Eq. 3, α=0.3) under a two-stage recipe and a dual-stream attention pattern that keeps the clean stream strictly causal. The resulting checkpoint supports three inference modes—standard AR, block-wise diffusion denoising (with optional learned sampler), and self-speculation (diffusion draft + AR verify; linear and quadratic variants, with optional LoRA draft alignment)—without architectural forks. Empirically, the 8B instruct model matches or exceeds Qwen3-8B accuracy while reporting ~6× tokens-per-forward under linear self-speculation and ~4× SPEED-Bench throughput vs Qwen3-8B-Eagle3 on GB200/SGLang; progressive ablations (Tab. 1–3), multi-scale tables (Tab. 5–9), acceptance-length comparisons to Eagle3/MTP (Tab. 10), and multi-GPU device measurements (Fig. 1, 9) support complementarity of the two losses and practical efficiency of self-speculation. A speed-of-light (SOL) analysis via recursive dynamic compaction estimates up to 76.5% more real TPF than linear self-speculation under an optimal diffusion sampler.","tokens_in":31355,"tokens_out":1587,"duration_ms":22006,"significance":"If the results hold under the stated training and evaluation conditions, the work is a substantial systems contribution: it packages AR, parallel diffusion, and self-speculation into one drop-in checkpoint that adapts across concurrency regimes, and it shows self-speculation can beat auxiliary-head MTP (Eagle3) in acceptance length and real-device throughput. The controlled 25B-token ablations and AR-with/without-diffusion SFT controls give credible evidence that AR and diffusion losses are complementary rather than zero-sum. The SOL construction and multi-GPU SPEED-Bench measurements are concrete, falsifiable artifacts that clarify headroom beyond current samplers. Model family release and Megatron Bridge pipeline further raise the work’s utility to the community.","major_comments":[{"comment":"Sec. 5.1–5.2 and Tab. 5–8: Full-scale SOTA accuracy claims (e.g., NLD-8B vs Qwen3-8B / LLaDA / Dream / SDAR) rest on continuous pretraining from Ministral3 (1T pure-AR + 300B joint tokens) plus proprietary-style SFT, while baselines differ in data mixture, total tokens, and often evaluation harness (NeMo-Skills vs official diffusion pipelines). The 25B-token ablations (Tab. 1–3) and matched AR-only vs joint SFT (Tab. 3) isolate complementarity under controlled data, but they do not quantify what fraction of the headline accuracy/TPF/throughput gains at full scale is due to the joint objective and tri-mode inference versus stronger provenance. The manuscript should explicitly bound causal attribution for the strongest claim and, where possible, add a matched-data or matched-checkpoint comparison at a scale closer to the released models.","section":"Sec. 5.1–5.2, Tab. 5–8"},{"comment":"Sec. 4.2 and Fig. 7: The 76.5% real-TPF advantage of SOL over linear self-speculation mixes two different correctness targets (SOL matches the diffusion mode’s own serial highest-confidence path; linear SS matches AR verification) and two different cost models (one vs two forwards, plus prefix-only acceptance). The paper notes this, but the abstract and intro still present 76.5% as a clean headroom figure for “diffusion under an optimal sampler.” Please restate the claim so that acceptance-rate proximity to SOL (~10% gap) and real-TPF gap (two-forward + prefix truncation) are separated, and clarify that SOL is not an AR-accuracy ceiling.","section":"Sec. 4.2, Fig. 7, Abstract"},{"comment":"Tab. 5 and Sec. 6.1: Diffusion baselines are evaluated with their official pipelines while NLD and AR baselines use NeMo-Skills; decoding hyperparameters (block size, confidence thresholds, thinking vs non-thinking) are only partially aligned. For load-bearing accuracy comparisons against LLaDA/Dream/SDAR, report a sensitivity check under a single harness or document that residual gaps survive re-evaluation under the authors’ pipeline.","section":"Sec. 6.1, Tab. 5"}],"minor_comments":[{"comment":"Fig. 1(b) and caption: Symbol sizes for diffusion block sizes (8/16/32) are hard to read; add an explicit legend entry for block size and for Linear vs Quad SS.","section":"Fig. 1"},{"comment":"Eq. (2): The 1/t reweighting and its interaction with global vs sequence averaging (Eq. 4–5) is well motivated in text; a short note on whether t is continuous or discretized in practice would help reproducibility.","section":"Sec. 2.1, Eq. (2)"},{"comment":"Tab. 6: Small accuracy differences between AR and self-speculation are attributed to “kernel mismatches between 1-token decoding and multi-token prefilling”; quantify or cite the kernel path so readers can judge whether this is numerical noise or a systematic bias.","section":"Tab. 6, Sec. 6.1"},{"comment":"Appendix A: Sampler feature list (144-d, PCA top-3, entropy, etc.) is useful; state the PCA basis source and whether features are frozen across model scales.","section":"Appendix A"},{"comment":"Related work (Sec. 7): Cite and briefly position against concurrent joint AR–diffusion / set-block decoding lines already mentioned in Sec. 2.2 ([14], [7]) so novelty of tri-mode inference and SOL is sharper.","section":"Sec. 7"},{"comment":"Typos / polish: “Y onggan F u”, “W u”, “T uruvekere”, “Y e Y u” spacing in author list; “AsshowninTab.3” missing spaces (Sec. 2.4); consistent “tokens per forward” vs “TPF” on first use in abstract.","section":"Front matter, Sec. 2.4"}],"recommendation":"minor_revision","confidential_remarks":"The systems packaging, device measurements, and controlled ablations are strong enough for a top venue after minor revision. The main risk is over-claiming causal credit for joint training in full-scale SOTA tables; requiring clearer qualification (and ideally one matched-data check) is proportionate and fixable without new model families. I would not reject on data-provenance grounds alone—that is endemic in open LM comparisons—but the abstract’s 6×/4× framing should not outrun what the controlled evidence isolates."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a real systems paper, not a pure restatement of block diffusion. They ship one checkpoint that actually runs AR, block diffusion, and diffusion-draft + AR-verify self-speculation, with LoRA draft alignment, a small trained commit sampler, multi-scale base/instruct/VLM weights, and measured throughput on H100/GB200/RTX/DGX Spark under SGLang. The 8B linear SS numbers (TPF ~6, acceptance above Eagle3/MTP on SPEED-Bench, ~4× throughput vs Qwen3-8B-Eagle3 at low concurrency) and the SOL analysis (recursive dynamic compaction, ~76.5% more real TPF than linear SS if you could commit non-prefix safe sets) are the parts I would actually use.\n\nWhat is new is the packaging and the measurement, not the joint objective itself. Block diffusion, AR-initialized conversion, and joint AR–diffusion attention already exist in the papers they cite. Credit where due: the progressive ablations (global loss avg, two-stage, α=0.3, AR loss on) and the matched-token AR-only vs joint controls are clean and support complementarity under fixed data. The LoRA o_proj alignment and the SOL construction are concrete engineering contributions. Citations look honest; they do not hide the prior line.\n\nThe soft spot is exactly the stress-test note, and it is real but not fatal. Full-scale wins vs Qwen3/LLaDA/Dream/SDAR rest on continuous pretraining from Ministral3 (1T AR + 300B joint) plus their SFT stack, while baselines differ in data, tokens, and sometimes harness. The 25B ablations isolate the training tricks; the headline SOTA tables do not fully isolate architecture from provenance. That softens causal claims about “why” accuracy is high; it does not erase the device measurements or the self-speculation vs Eagle3 acceptance gap under their own stack. Quadratic SS TPF looks good on paper and weaker on real kernels—they say so. Free parameters (α, block size, LoRA, sampler threshold) are many but standard for this genre.\n\nThis is for people who care about inference systems, speculative decoding, and whether diffusion is more than a research toy. Not for pure theory. I would bring it to reading group, cite the SOL and self-speculation results, and send it to peer review. Accept with the provenance caveat named, not desk-reject.","headline":"Solid systems packaging of joint AR–diffusion into a real tri-mode family with device numbers and a useful SOL ceiling; SOTA accuracy claims are only partly isolated from Ministral3 provenance.","tokens_in":32093,"tokens_out":620,"would_cite":true,"duration_ms":8848,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"One model trained jointly for AR and diffusion can switch among autoregressive, parallel diffusion, and self-speculation decoding, matching strong open baselines while decoding about six tokens per forward.","keywords":["tri-mode language model","joint AR-diffusion training","block diffusion","self-speculation decoding","tokens per forward","speed-of-light analysis","multi-token prediction"],"falsifier":"A controlled experiment that continuous-pretrains and SFT-matches an otherwise identical AR-only baseline on the exact same token budget, data mixture, and evaluation harness as the joint model; if the joint model then loses its accuracy or tokens-per-forward advantage, the complementarity claim fails.","tokens_in":31843,"feed_emoji":"⚡","tokens_out":702,"duration_ms":8038,"temperature":0.7,"pith_summary":"This paper argues that autoregressive and diffusion language modeling need not compete. A single network trained with a weighted joint next-token and block-diffusion objective can run in three inference modes: ordinary left-to-right decoding, block-wise parallel diffusion, and self-speculation in which diffusion drafts and AR verifies. The authors claim the two losses are complementary—diffusion strengthens lookahead planning while AR supplies left-to-right linguistic priors—and that self-speculation already beats multi-token-prediction heads on acceptance length and measured throughput. A speed-of-light construction further claims that an ideal diffusion sampler could still deliver roughly 76 percent more real tokens per forward than today’s best self-speculation path. The resulting 3B/8B/14B base, instruct, and vision-language family is reported to match or exceed open AR and diffusion peers on accuracy while substantially raising tokens-per-forward and system throughput, so one set of weights can serve high-concurrency cloud and low-concurrency personal inference without architecture changes.","feed_headline":"One model does AR, diffusion, and self-speculation","feed_subtitle":"Joint training yields ~6× tokens per forward at matched accuracy and higher real-device throughput.","key_machinery":"The joint objective (AR next-token loss plus α times a block-wise diffusion denoising loss) together with a dual-stream attention pattern that keeps the clean stream strictly causal. That pattern lets both losses be computed in one forward–backward pass and, at inference, lets the same weights act as AR decoder, diffusion denoiser, or diffusion drafter plus AR verifier.","core_discovery":"Joint AR–diffusion training with a carefully chosen diffusion weight (α = 0.3), global loss averaging, and a two-stage AR-then-joint schedule produces a single model that fully preserves AR accuracy, supports native block diffusion, and enables high-acceptance self-speculation without auxiliary prediction heads; at 8B instruct scale this yields roughly 6× tokens per forward versus a comparable AR baseline at matched accuracy, with diffusion’s theoretical upper bound still substantially higher.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Tri-mode LM unifies AR, diffusion, self-speculation","Joint AR-diffusion training yields 6× tokens per forward","Diffusion drafts, AR verifies in high-acceptance self-speculation","Single architecture switches AR, diffusion, self-speculation modes","Nemotron-Labs-Diffusion preserves AR accuracy with block diffusion"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the reported gains over open AR, diffusion, and multi-token-prediction baselines come mainly from the joint objective and tri-mode inference rather than from differences in pretraining data volume, recipe, and evaluation harness.","fun_headline_variants_meta":{"raw":{"variants":["Tri-mode LM unifies AR, diffusion, self-speculation","Joint AR-diffusion training yields 6× tokens per forward","Diffusion drafts, AR verifies in high-acceptance self-speculation","Single architecture switches AR, diffusion, self-speculation modes","Nemotron-Labs-Diffusion preserves AR accuracy with block diffusion"]},"model":"grok-4.5","effort":"low","cost_usd":0.00715,"raw_usage":{"total_tokens":1785,"prompt_tokens":842,"num_sources_used":0,"completion_tokens":89,"cost_in_usd_ticks":71500000,"prompt_tokens_details":{"text_tokens":842,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":854,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":842,"tokens_out":89,"duration_ms":6499,"temperature":1.0,"reasoning_tokens":854,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T03:02:27.210887+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled experiment that continuous-pretrains and SFT-matches an otherwise identical AR-only baseline on the exact same token budget, data mixture, and evaluation harness as the joint model; if the joint model then loses its accuracy or tokens-per-forward advantage, the complementarity claim fails.","supporting_citations":[],"review_version":1}