{"id":"96deb809-0dde-4fed-a749-346e0b34bd19","arxiv_id":"2607.01736","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CROF, led by Reward Observability Fraction, offline-selects RSSM checkpoints whose closed-loop MPC and imagination A2C returns stay high after validation loss has collapsed on LunarLander.","lead":"A structural score called CROF, built from reward-observability and control ranks, picks world-model checkpoints that keep working for planning and imagination RL after ordinary validation loss has already overfit. On LunarLander it yields a model-based policy that beats a model-free baseline with far fewer real interactions.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the paper's own Reacher-bounded scope.","rationale":"The central claim is carefully scoped to LunarLander-v3 under non-Markovian/shaped rewards. The paper supplies the exact negative control (Reacher) that bounds generalization, so the reader's weakest_assumption is already the paper's own limitation statement rather than an unacknowledged soft spot. The empirical chain—40-metric ranking, fixed-vs-time-varying ablation, checkpoint-selection table, and data-efficiency comparison—is coherent and reproducible via the released code. Secondary issues (composite weights, single seed, purposive rather than exhaustive A2C sweep) justify the existing CONDITIONAL label and MODERATE confidence but do not require a further downgrade. Therefore the stress-test leaves the reader's verdict unchanged.","tokens_in":14888,"tokens_out":500,"duration_ms":6053,"concrete_test":"Independently recompute jac_rof_combined and CROF-A/B on the released checkpoints using the paper's Jacobian definitions (Eqs. 3–8) and the same good/bad validation splits; confirm that the raw CROF minimum still lands at epoch 280 and that Spearman ρs against MA-7 MPC remains ≤ −0.55. If the ranking or correlation collapses under reimplementation, the offline-selection claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly flags the non-Markovian reward premise, but the paper already treats this as a scoped structural condition rather than a domain-agnostic rule (Section 4.8 and Conclusion). The LunarLander evidence for the strongest claim is internally consistent: ROF/CROF track the MPC U-shape that standard losses miss (Table 1, Figures 2–3,5), the raw CROF pick (WM 280) yields the reported A2C gains at the stated data cost (Tables 3–4), and the Reacher contrast (MLP R^{2}=0.9998, opposite-sign ρs) is presented as a negative control. No hidden inconsistency in the Jacobian construction (Eqs. 5–10) or evaluation protocol undermines the single-environment finding. The remaining risks (hand-chosen CROF weights, single seed, purposive A2C sample) are secondary and already reflected in the CONDITIONAL verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies offline checkpoint selection for latent world models (RSSM) when standard validation losses and multi-step prediction errors continue to improve after closed-loop performance collapses. On LunarLander-v3, whose reward is only partially recoverable from (o_t, a_t) due to potential-based shaping and terminal events gated by simulator-internal flags, the authors evaluate 40 validation-time diagnostics against a CEM-MPC return oracle. The strongest single predictor is the Reward Observability Fraction (ROF; Eq. 7), the fraction of the reward gradient lying in the observable subspace of the linearized dynamics. ROF is combined with controllability/observability ranks and open-loop observation RMSE into Composite Reward Observability Fraction (CROF-A/B; Eqs. 9–10). The CROF-selected checkpoint yields a model-based A2C policy that exceeds a model-free A2C baseline by ~24.5 return points at ~65× fewer real-environment steps, and also supports strong zero-shot CEM-MPC. A Reacher negative control (fully Markovian reward) shows the opposite-sign, near-zero ROF–MPC correlation, scoping the claim to non-Markovian/shaped rewards.","tokens_in":15277,"tokens_out":1289,"duration_ms":10095,"significance":"If the single-environment result holds, the work supplies a concrete, control-theoretic diagnostic for a practically common failure mode of model-based RL: objective mismatch under shaped or partially observed rewards. The contribution is complementary to training-time representation methods (DeepMDP, SOLAR, bisimulation embeddings) because it is post-hoc and does not alter the world-model objective. Strengths that raise the paper’s value include: a large fixed-seed checkpoint sweep (100 WMs), an independent closed-loop oracle (CEM-MPC + real-env A2C), fixed vs. time-varying Jacobian ablation, good/bad state split, explicit Reacher negative control, and released code/data. The ~65× data-efficiency claim against a fairly evaluated model-free baseline is a clear, falsifiable headline result within the stated scope.","major_comments":[{"comment":"The central claim is scoped to non-Markovian/shaped rewards (Section 4.8 and Conclusion), yet the only positive evidence is LunarLander-v3 and the only negative control is Reacher. For a journal contribution that introduces ROF/CROF as structural diagnostics, at least one additional environment with a different non-Markovian structure (e.g., delayed reward, POMDP flag, or alternative PBRS) is needed to show that the Jacobian subspace geometry, not LunarLander-specific reward algebra, is doing the work. Without it the generalization premise remains an untested assumption of the paper’s own framing.","section":null},{"comment":"CROF-A/B (Eqs. 9–10) use hand-chosen regularizer weights (1.0 and 0.5) and a fixed α=0.5 good/bad mix (Eq. 8). Table 2 shows that pure jac_rof_combined already lands near the MPC plateau once smoothed, so the composite’s added value is modest and weight-sensitive. A short leave-one-regularizer-out or weight-sweep ablation against the same MA-7 oracle would establish that the three regularizers are load-bearing rather than post-hoc stabilizers of a single-run minimum.","section":null},{"comment":"A2C results (Tables 3–4) rest on a purposive sample of nine world-model checkpoints rather than a uniform or CROF-ranked sweep, and all runs share a single seed (12345). The headline +24.5 return / 65× data claim therefore depends on the particular CROF raw pick (WM 280) and on the chosen model-free baseline horizon (1000 vs. 600 steps). Reporting variance over a few seeds for both the world-model training and the A2C stage, or at least evaluating the next-best CROF-ranked checkpoints, is required before the efficiency comparison can be treated as robust.","section":null}],"minor_comments":[{"comment":"Table 1 truncates ~18 of the 40 metrics; Appendix Figure 6 shows the full bar chart, but the main text should state explicitly that every unlisted metric falls in the weak band so readers do not have to cross-check the appendix.","section":null},{"comment":"Notation for the composite (]ROF, ekc, eko, eeobs in Eqs. 9–10) is slightly inconsistent with the surrounding text (jac_rof_combined, jac_ctrl_rank, etc.); a single symbol table would help.","section":null},{"comment":"Figure 2’s light min–max band is described as clipped; the caption should quantify how many points were clipped so the envelope is interpretable.","section":null},{"comment":"The effective-rank threshold 10^{-3} (Section 3.4.3) is stated without sensitivity analysis; a one-sentence note that ranks are stable under 10^{-2}–10^{-4} would suffice.","section":null},{"comment":"Related-work discussion of objective mismatch (Lambert et al.) is accurate but could briefly note whether any prior work already used controllability/observability ranks for world-model selection, even if only in linear settings.","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript is careful and self-aware about its single-environment scope; the Reacher control is a genuine strength. The main risk for the journal is that the contribution is currently a thorough case study rather than a validated method. If the authors can add one more non-Markovian domain and a minimal weight/seed ablation, the paper becomes a solid methods contribution; otherwise it may fit better as a workshop/extended abstract. Code release is a plus and should be verified at acceptance."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: on LunarLander, standard validation loss and multi-step RMSE keep improving while CEM-MPC returns collapse after ~epoch 380, and the paper’s Reward Observability Fraction (fraction of the reward gradient in the observable subspace) is the metric that actually tracks that U-shape (Spearman ~-0.71). CROF packages it with three cheap regularizers and offline-selects a checkpoint that yields a model-based A2C beating a fairly run model-free A2C by ~24.5 return at ~65× fewer real steps.\n\nWhat is new is not Jacobian controllability/observability (Kalman, Moore, Sussillo & Barak) or the objective-mismatch diagnosis (Lambert et al.), but the concrete post-hoc ROF/CROF construction as a validation-time checkpoint score for RSSMs under non-Markovian/shaped rewards, plus the careful 100-checkpoint sweep that shows it works where the usual losses fail. The paper does the empirical work cleanly: fixed seed, MA-7 oracle, fixed vs time-varying ablation, good/bad state split, 40-metric table, and a Reacher contrast where reward is fully Markovian (MLP R^{2}≈1) and ROF flips sign and becomes useless. Code and data are released. That is real, reproducible evidence for the LunarLander claim.\n\nSoft spots are real but already mostly owned by the author. Everything is one environment; the composite weights and α=0.5 are hand-chosen; A2C is only run on a purposive sample of nine WMs; single seed. The load-bearing premise is that the reward is not recoverable from (o_t,a_t) alone, so the reward head must sit outside the observation-decoder column space—exactly the regime LunarLander’s PBRS-plus-terminal-flag reward creates and Reacher does not. The paper states this scope in Section 4.8 and the conclusion rather than overselling generality. No circularity: the oracle is closed-loop return, not the metric itself. Math is standard linearization; citations are appropriate.\n\nThis is for people who train world models on shaped or partially observed rewards and need an offline selection rule. It deserves a serious referee. I would bring it to reading group and would cite the diagnostic construction when working on objective mismatch. Send it out.","headline":"Solid single-env empirical paper: ROF/CROF actually track the MPC collapse that validation loss misses on LunarLander, with code and a clean Reacher negative control; scope is the real limit, not a hidden flaw.","tokens_in":15871,"tokens_out":594,"would_cite":true,"duration_ms":5904,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A control-theoretic score predicts which world-model checkpoint will actually plan and control well, without running the real environment.","keywords":["latent world models","RSSM","checkpoint selection","model predictive control","model-based RL","observability","non-Markovian reward","LunarLander"],"falsifier":"Train the same RSSM on another non-Markovian or shaped-reward control task, compute ROF/CROF across checkpoints, and check whether the CROF-selected checkpoint still yields high closed-loop MPC or imagination-trained policy return while ordinary prediction metrics select a collapsed one.","tokens_in":15778,"feed_emoji":"🚀","tokens_out":632,"duration_ms":5828,"temperature":0.7,"pith_summary":"World models are usually chosen by how accurately they predict the next observation or reward. On LunarLander, those ordinary losses keep improving long after the model has become useless for closed-loop control: planning and imagination-based policy training both collapse. The paper shows that a structural diagnostic drawn from classical control theory—the Reward Observability Fraction—tracks the true closed-loop quality. It measures how much of the reward gradient lives inside the latent directions that the observation decoder can see. When the reward is only partly recoverable from the current observation and action (as it is under LunarLander’s shaped reward), a low value of this fraction signals that the reward head has learned to read the right latent subspace and will stay accurate during open-loop imagination. Combining that fraction with three simple regularizers yields a single offline score, CROF, that selects a checkpoint whose model-based policy beats a strong model-free baseline by roughly 25 return points while using about 65 times fewer real interactions.","feed_headline":"One score picks the world model that still plans well","feed_subtitle":"It tracks closed-loop return when ordinary prediction loss keeps improving after collapse","key_machinery":"Reward Observability Fraction (ROF): the fraction of the reward-head gradient that lies in the observable subspace of the latent dynamics (obtained from the SVD of the stacked observability matrix). CROF is ROF plus three normalized regularizers on controllability rank, observability rank, and open-loop observation error.","core_discovery":"On LunarLander-v3 the Reward Observability Fraction (and its composite CROF) is a strong offline predictor of per-checkpoint CEM-MPC return and of the quality of an actor-critic policy trained entirely inside the world model; ordinary validation loss and multi-step prediction error are not.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["CROF offline score selects world models that still plan well","Reward Observability Fraction predicts CEM-MPC closed-loop return","Validation loss keeps falling after world-model planning collapses","One composite score finds the RSSM checkpoint that drives strong MPC","Structural diagnostics beat RMSE for LunarLander checkpoint selection"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The score only works when the reward cannot be fully recovered from the current observation and action alone, so the reward head must learn to read outside the observation-decoder subspace; the paper itself shows the opposite behaviour on a fully Markovian reward task.","fun_headline_variants_meta":{"raw":{"variants":["CROF offline score selects world models that still plan well","Reward Observability Fraction predicts CEM-MPC closed-loop return","Validation loss keeps falling after world-model planning collapses","One composite score finds the RSSM checkpoint that drives strong MPC","Structural diagnostics beat RMSE for LunarLander checkpoint selection"]},"model":"grok-4.5","effort":"low","cost_usd":0.004926,"raw_usage":{"total_tokens":1415,"prompt_tokens":795,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":49260000,"prompt_tokens_details":{"text_tokens":795,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":537,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":795,"tokens_out":83,"duration_ms":4549,"temperature":1.0,"reasoning_tokens":537,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T08:38:05.508164+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train the same RSSM on another non-Markovian or shaped-reward control task, compute ROF/CROF across checkpoints, and check whether the CROF-selected checkpoint still yields high closed-loop MPC or imagination-trained policy return while ordinary prediction metrics select a collapsed one.","supporting_citations":[],"review_version":2}