{"id":"0f18600e-b79b-4346-89f3-234522c85230","arxiv_id":"2607.09185","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Three lightweight LAM fine-tuning objectives (foreground-weighted reconstruction, primitive contrastive learning, zero-transition calibration) debias latent actions and yield stronger, cheaper robot-action following in large ACWMs.","lead":"Reconstruction-only latent action models leak backgrounds and other non-action cues into the action condition of robot world models. CD-LAM adds three short fine-tuning losses that cut action-following error and cut robot-action adaptation cost by more than 12× on 2B and 14B backbones.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"FDCE and training both depend on SAM3 masks, so reported action-following gains may partly be metric-aligned rather than pure causal debiasing of A_t.","rationale":"The reader correctly isolates the weakest assumption: that SAM3 masks plus verb clusters are faithful proxies for A_t versus C_t/V_t. I sharpen the same point by noting the concrete circularity between training loss (Eq. 8–9) and the primary success metric (FDCE), which the paper never breaks. Ablations (Table V) and interventions still give independent support that something useful is happening, so the empirical efficiency claim remains credible; the causal language and the precise magnitude of the FDCE numbers do not. Hence the verdict stays CONDITIONAL rather than moving to ACCEPT or REJECT. No stronger internal inconsistency appears: architecture, conditioning format, and adaptation protocol are held fixed, and multi-scale / multi-tier results are coherent. The single highest-value check is therefore an independent-mask re-evaluation of the main table.","tokens_in":18526,"tokens_out":605,"duration_ms":7808,"concrete_test":"Recompute Table III FDCE (and the zero-action residual) on the same 300 AgiBot rollouts using an independent foreground extractor never seen in training—e.g., Grounded-SAM with robot-arm text prompts, or AgiBot’s own robot-link masks if available—while keeping CoWTracker fixed. If mean FDCE reductions fall below ~15 % (or residual zero-action FDCE no longer halves), the published gains are substantially metric-aligned and the causal-debiasing claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the three Stage-1 objectives remove action-irrelevant confounding from z_t so that ACWMs follow embodiment dynamics more faithfully (lower FDCE, fewer adaptation steps). The load-bearing link is that the supervision used to define “embodiment” is independent of the metric that certifies success. It is not. Embodiment-centric reconstruction (Eq. 8–9) reweights pixels by SAM3 masks M_t; FDCE (Appendix A) also seeds tracks inside SAM3 foreground masks and scores only those tracks. Action-centric contrast further uses the same 12-way caption-verb clusters that appear in the shortcut-leakage diagnostic. Consequently a model that simply concentrates capacity on SAM3-selected regions and verb-cluster neighborhoods can improve FDCE and the Table I diagnostics without necessarily purifying the causal factor A_t from C_t/V_t. The zero-action and target-transfer interventions still show real behavioral change, but the headline percentage reductions (35 % / 30 % FDCE, 12× fewer updates) rest on a metric that shares tooling with the training signal; that is the softest point in the causal story.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that reconstruction-only latent action models (LAMs) encode action-irrelevant confounders (background, non-interacted objects, camera-like factors) into the latent action z_t, which then confounds action-conditioned world models (ACWMs). It proposes CD-LAM, a three-stage fine-tuning pipeline whose Stage-1 LAM objectives—embodiment-centric weighted reconstruction (Eqs. 8–9), action-centric contrastive learning over 12-way caption-verb primitives (Eq. 10), and latent-space calibration via free-bit KL plus zero-transition anchoring (Eqs. 11–12)—produce debiased latents. On 2B and 14B DreamDojo-style ACWMs, the method reports lower FDCE after latent-action and robot-action conditioning, higher PSNR/SSIM, stronger zero-action and target-action interventions, and matching of the DreamDojo 50k-update reference with more than 12× fewer robot-action adaptation steps (final 3k/6k checkpoints). Supporting evidence includes a LAM confounding audit (Table I), multi-stage rollout tables (II–III), data-tier scaling (Table IV), objective ablations (Table V), and qualitative rollouts.","tokens_in":18943,"tokens_out":1228,"duration_ms":20996,"significance":"If the results hold under independent scrutiny, the work is a practically useful contribution to embodied world models: it isolates a concrete failure mode of reconstruction-trained LAMs, supplies diagnostic metrics (zero-transition response, camera-shift response, shortcut leakage, FDCE), and shows that a short, targeted LAM fine-tune can improve controllability and cut robot-action adaptation cost by more than an order of magnitude at both 2B and 14B. The multi-stage evaluation design (latent-only rollouts, robot-action adaptation, zero-action and target-transfer interventions), objective ablations that map each loss to a distinct failure mode, and the promised release of debiased LAMs/ACWMs, protocols, and code are genuine strengths. The efficiency finding—that debiasing the condition is cheaper than unlearning a confounded condition downstream—is of clear engineering value for robot world-model pipelines that rely on unlabeled video pretraining.","major_comments":[{"comment":"The primary action-following metric and a core training signal share the same tooling. Embodiment-centric reconstruction reweights pixels by SAM3 masks M_t (Eqs. 8–9); FDCE seeds tracks inside SAM3 foreground masks and scores only those tracks (Appendix A, Eq. A.4). Action-centric contrast and the shortcut-leakage diagnostic further share the 12-way caption-verb clusters (Appendix B). A model that concentrates capacity on SAM3-selected regions and verb-cluster neighborhoods can therefore improve headline FDCE and Table I diagnostics without necessarily purifying the causal factor A_t from C_t/V_t. Zero-action residual FDCE and full-frame PSNR provide partially independent evidence, but the claimed 35%/30% FDCE reductions and the “causally debiased” framing rest on a metric aligned with the training signal. Please add at least one mask-independent motion metric (e.g., full-frame or random","section":null},{"comment":"Tables II and III (and the efficiency curves in Fig. 8) report point estimates only—no standard errors, bootstrap intervals, or multi-seed variance—despite multi-scale claims and percentage reductions that are central to the abstract. With 300 evaluation clips, seed or clip-level variability is estimable. Without it, it is hard to judge whether the 2B→14B baseline FDCE worsening, the 12× efficiency claim, or the per-action breakdowns in Fig. A.1 are stable. Please report uncertainty for the main FDCE/PSNR numbers and for the step at which CD-LAM crosses the DreamDojo reference.","section":null},{"comment":"The causal analysis in §III (Eq. 5, Fig. 1d) and the title/abstract language (“causally debiased,” “confounding path”) go beyond what the experiments strictly establish. The interventions show improved sensitivity to the supplied action, and Table I shows reduced responses to static pairs and synthetic shifts, but there is no identification argument or interventional test that separates removal of C_t/V_t from reweighting toward correlated embodiment appearance. Table V’s footnote already notes that removing zero-transition calibration can lower FDCE while failing the camera-shift diagnostic—evidence that FDCE alone does not certify causal purity. Soften or operationalize the causal claims (e.g., “reduces measured action-irrelevant responses and improves action following under fixed context”) unless additional identification-style evidence is added.","section":null},{"comment":"Comparisons are limited to DreamDojo with its original reconstruction-trained LAM (§V-A). The related-work section cites Genie, LAPO/LAPA, AdaWorld, Moto, IGOR, and ConLA, but none appear as empirical baselines for the LAM audit or for Stage-2/3 rollouts. At minimum, a reconstruction-only LAM re-finetuned for the same 1k steps without the three CD-LAM terms (or with only L_emb) should be reported as a compute-matched control beyond the partial ablations in Table V, so that gains are not confounded with extra Stage-1/2 fine-tuning budget alone.","section":null}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: reconstruction-only LAMs leak background and context into the action condition, and a short three-objective fine-tune (foreground-weighted recon, coarse-primitive contrast, zero-transition calibration) measurably cleans that up. On DreamDojo-style 2B/14B ACWMs they cut mean FDCE ~30–35% after robot adaptation, raise PSNR, and match the 50k-update baseline with 3k–6k steps. That efficiency claim is the part worth caring about.\n\nWhat is new is not any single loss—weighted MSE, contrastive structure, free-bits + zero anchoring are familiar—but the packaging as a staged LAM debiasing recipe plus an explicit audit suite (zero-transition response, camera-shift, shortcut leakage) and the FDCE motion metric. The multi-scale tables, latent-only vs robot-action stages, zero-action and target-transfer interventions, data-tier scaling (1h already gets most of the gain), and objective ablations all point the same way. The zero-action residual motion drop (especially at 14B) is hard to dismiss as pure metric gaming; the model actually stays still when told to. Code and models are promised, which helps.\n\nSoft spots, in proportion. The stress-test is right that SAM3 masks appear in both L_emb and FDCE, and the 12-way caption verbs appear in both L_ctr and the shortcut diagnostic, so some of the headline percentage is metric-aligned rather than pure causal purification of A_t. The paper’s “causally debiased” language outruns the identification strategy; this is better described as targeted reweighting and calibration. Single baseline family (DreamDojo), no error bars on main tables, free parameters (λ schedules, α_fg/bg, m_zero, etc.) are ordinary for this literature but keep the claim conditional. None of that erases the behavioral interventions or the adaptation-efficiency curve.\n\nThis is for people building or using LAM-conditioned robot world models who care about labeled-data budgets. It is not a foundational theory paper. I would bring it to reading group, cite the efficiency and audit results if I work in the area, and send it to peer review—referees should push on the metric independence and the causal wording, not desk-reject.","headline":"Solid empirical fix for confounded latent actions in LAM-based world models; the efficiency and intervention results are real, even if the causal packaging and SAM3-shared metrics overclaim a bit.","tokens_in":19596,"tokens_out":579,"would_cite":true,"duration_ms":5857,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Debiasing latent actions from unlabeled video makes robot world models follow commands with far less labeled data.","keywords":["latent action models","action-conditioned world models","causal debiasing","embodied AI","robot action following","video world models","adaptation efficiency"],"falsifier":"Hold the world-model architecture fixed, swap only the latent-action encoder for a reconstruction-only baseline, and check whether mean foreground displacement error under identical robot-action sequences still drops by roughly thirty percent and whether zero-action and camera-shift diagnostics remain low; if the gains vanish or the diagnostics stay high, the causal-debiasing claim fails.","tokens_in":19442,"feed_emoji":"🤖","tokens_out":719,"duration_ms":8283,"temperature":0.7,"pith_summary":"Action-conditioned world models can simulate how a robot and its scene would evolve under a sequence of controls, but they usually need large amounts of costly action-labeled robot video. Latent action models try to dodge that cost by inventing compact action codes from unlabeled videos alone. This paper argues that the usual reconstruction objective for those codes is the problem: it lets backgrounds, camera-like shifts, and non-interacted objects leak into the latent action, so the world model is not really being conditioned on embodiment dynamics. CD-LAM is a three-stage fine-tuning recipe that re-trains the latent action space with embodiment-weighted reconstruction, action-primitive contrastive learning, and zero-transition calibration, then carries the cleaned latents into the world model and finally maps real robot commands into the same space. On 2B and 14B backbones the method cuts action-following error, raises visual fidelity, and matches a much longer baseline adaptation budget with more than twelve times fewer robot-action updates. A sympathetic reader cares because the bottleneck for controllable simulators is often labeled data, and the paper claims a short, targeted clean-up of the latent action is enough to unlock that controllability.","feed_headline":"Debiased latent actions cut robot data needs 12\times","feed_subtitle":"Three short fine-tuning losses make world models follow commands with far less labeled robot video.","key_machinery":"CD-LAM: a three-objective LAM fine-tuning loss (embodiment-centric weighted reconstruction, action-centric contrastive learning over coarse verb primitives, and latent-space calibration with free-bit KL plus zero-transition anchoring) applied in a three-stage pipeline that first debiases the latent action, then debiases the world model on those latents, then bridges executable robot actions into the same space.","core_discovery":"Reconstruction-only latent actions entangle embodiment dynamics with action-irrelevant visual factors, confounding the downstream world model; three short fine-tuning objectives that force embodiment focus, action-aware neighborhoods, and calibrated non-collapse produce debiased latents that measurably improve action following, visual fidelity, and robot-action adaptation efficiency on both 2B and 14B action-conditioned world models.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["CD-LAM debiasing cuts robot labels 12× for action world models","Three short losses free latent actions from visual entanglement","Embodiment focus and contrast make latents follow robot commands","Causal debias of LAMs yields tighter action following, 6k steps","Debiased latent actions need far less labeled robot video to control"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The method assumes that automatic embodiment masks and coarse caption-verb clusters are faithful enough proxies for true action factors that the three losses remove confounding rather than merely re-weighting toward another correlated visual cue.","fun_headline_variants_meta":{"raw":{"variants":["CD-LAM debiasing cuts robot labels 12× for action world models","Three short losses free latent actions from visual entanglement","Embodiment focus and contrast make latents follow robot commands","Causal debias of LAMs yields tighter action following, 6k steps","Debiased latent actions need far less labeled robot video to control"]},"model":"grok-4.5","effort":"low","cost_usd":0.005478,"raw_usage":{"total_tokens":1521,"prompt_tokens":816,"num_sources_used":0,"completion_tokens":92,"cost_in_usd_ticks":54780000,"prompt_tokens_details":{"text_tokens":816,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":613,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":816,"tokens_out":92,"duration_ms":6149,"temperature":1.0,"reasoning_tokens":613,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T04:49:37.483516+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold the world-model architecture fixed, swap only the latent-action encoder for a reconstruction-only baseline, and check whether mean foreground displacement error under identical robot-action sequences still drops by roughly thirty percent and whether zero-action and camera-shift diagnostics remain low; if the gains vanish or the diagnostics stay high, the causal-debiasing claim fails.","supporting_citations":[],"review_version":1}