{"id":"bb42444a-9110-44b5-82e9-d8edba05a817","arxiv_id":"2607.10362","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Planner suboptimality in latent world models is bounded by twice the predicted-versus-true plan-cost gap; data-averaged prediction error neither bounds nor tracks that gap once actions leave the data manifold.","lead":"Prediction error on training data does not govern control quality in latent world models, because planners query off the data manifold. The paper proves suboptimality is bounded by the predicted-versus-true plan-cost gap and shows that gap is dominated by off-manifold divergence, not on-manifold MSE.","discovery_kind":"first_principles","skeptic_critique":{"model":"grok-4.5","headline":"Lemma 1 proves MSE does not bound the deployment gap, but the jump to “does not track / ρ→0” is not forced by the theory and rests on underpowered three-seed evidence.","rationale":"The reader correctly flags the linear-control premise and linear residual growth as the weakest assumption for the pricing theorems (2–4) and for turning g into computable taxes. Those assumptions are not required for Theorem 1 or for the pure “does not bound” half of Lemma 1, which are the heart of the strongest claim. My concern is therefore adjacent rather than identical: the paper’s strongest claim as written includes “neither bounds nor tracks” and the informal ρ→0 statement, and the “tracks” half is the least secure piece of that claim. Empirics are thin exactly as the reader notes (three seeds, Jacobian proxy, acknowledged need for replication, negative model-selection result). I do not see an internal contradiction or a reason to move off CONDITIONAL: the structural core is sound, synthetic pricing checks are clean, and the authors already flag the empirical limits. Verdict stays CONDITIONAL pending stronger multi-seed evidence that MSE truly fails to track success rather than merely failing to bound it.","tokens_in":39079,"tokens_out":746,"duration_ms":55979,"concrete_test":"Replicate the matched-window latent-MPC comparison on PUSHT and TWOROOM with ≥5 seeds per objective. Report Spearman ρ(val MSE, success) and ρ(ν-rollout R², success) with bootstrap CIs, and recompute Table 1 after dropping or winsorizing the weakest seed. If |ρ(MSE, success)| stays near 0 with CI excluding moderate correlation and ν-fidelity remains clearly higher, the tracking claim stands; if ρ rises once the 78% outlier is averaged out, the “essentially uncorrelated” half of the strongest claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim pairs a tight theorem with a looser empirical slogan. Theorem 1 (suboptimality ≤ 2g) and the change-of-measure half of Lemma 1 (δ_μ places no upper bound on δ_ν once ν has off-support mass) are elementary and correct for unrestricted residuals. The paper then writes that success is monotone in δ_ν, so “ρ(MSE, success) → 0” (Eq. 4 / Lemma 1). That correlation claim does not follow from the absence of an upper bound: shared factors (better encoders, lower L, better reachability geometry) can keep MSE and success correlated even when MSE alone cannot certify success. Under the paper’s own residual model ∥ε∥ ≤ δ₀ + L r(z), Theorem 4 still has δ₀ inside the bound; off-manifold terms can dominate, but that is a quantitative, not structural, statement. The only evidence offered for “essentially uncorrelated” is rank correlation 0.024 across a handful of runs (three seeds; PUSHT ONE_STEP success 78–92% with losses within ±8%, one weak 78% seed), while ν-side fidelity reaches only 0.62 and failed as a model-selection rule. Thus the load-bearing soft spot for the stated strongest claim is not the linear-control premise (which prices the gap but is not required for Theorem 1) but the unsupported leap from “does not bound” to “does not track.”","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that for latent world models used inside planners, data-averaged prediction error is the wrong control objective: a planner queries states reached by candidate actions (measure ν), which generally leave the training support (μ), so μ-averaged MSE neither bounds nor tracks control quality. Theorem 1 bounds planner suboptimality by twice the supremum gap between predicted and true plan-cost over the candidate set. Three lemmas locate why MSE fails (off-support decoupling, one-sided holes, realizable nonlinearity). Under a linear-control premise the gap is priced as an on-manifold residual (spectral tax from non-normality of the latent transition) plus an off-manifold rollout tax; the latter is claimed to be binding. Synthetic operators recover spectral/Jordan formulas; latent-MPC experiments on TwoRoom and PushT report near-zero rank correlation of validation MSE with success and higher correlation for a ν-side fidelity score, plus a linear state-readout intervention.","tokens_in":39475,"tokens_out":1382,"duration_ms":16662,"significance":"If the central framing holds, it is a useful corrective for world-model practice and model selection: optimize predicted-versus-true plan-cost fidelity on planner-reachable states rather than held-out MSE. Theorem 1 is elementary but correctly identifies the control objective; the change-of-measure decoupling and one-sided hole analysis are clean. The spectral-tax development (pseudospectra, Jordan staircase, self-adjoint zero-tax boundary recovering LeJEPA-style guarantees) is a substantive operator-theoretic contribution with synthetic verification of closed-form exponents. The theory-guided linear state readout and the explicit falsifiable predictions are strengths. The work is significant for latent MPC / JEPA-style planning even if the empirical correlation claims need tightening.","major_comments":[{"comment":"Lemma 1 / Eq. (4) and the abstract: the rigorous half is that once ν places mass off supp μ, δ_μ places no upper bound on δ_ν (and hence on g). The further claim that ρ(MSE, success) → 0 and that MSE “neither bounds nor tracks” success does not follow from absence of an upper bound alone—shared factors (encoder quality, L, reachability) can keep MSE and success correlated. Under the paper’s residual model, Theorem 4 still includes δ_0 inside the bound; off-manifold dominance is quantitative. Please separate the theorem (no bound) from the empirical slogan (does not track), and rephrase abstract/intro claims accordingly.","section":null},{"comment":"§5.2 and Table 1: the “essentially uncorrelated” claim rests on rank correlation 0.024 across a small set of runs (three seeds; PushT ONE_STEP success 78–92% with losses within ±8%, including one weak 78% seed). The ν-side fidelity correlation is only 0.62 and “did not outperform selection by success.” The manuscript itself flags the need for five-seed replication. Either expand the seed count and report confidence intervals / permutation tests, or demote the empirical decoupling language to “directional / consistent with” rather than confirmatory of Eq. (4).","section":null},{"comment":"§3.1, Theorems 2–4 and Appendix F: pricing and the clean on/off-manifold split assume linear deployed dynamics z_{t+1}=Âz_t+B̂u_t+ε_t and linear residual growth ∥ε∥≤δ_0+L r(z). Deployed predictors in the experiments are autoregressive latent models (MLP-class), not the linear model used for pricing. Clarify which conclusions (Theorem 1, Lemmas 1–3) are premise-free versus which (spectral/rollout taxes, optimal σ_u⋆, readout amplification argument) require the linear-control premise, and state when superlinear residual growth or strong off-support nonlinearity would invalidate the constants in (8).","section":null},{"comment":"§5.2 off-manifold probe: bL is a Jacobian-based amplification proxy, not the residual growth L in (12)/(53), and counterfactual states are not available. The claim that “the off-manifold amplification is the informative diagnostic” and that it “tracks the binding constraint” is therefore only loosely tied to Theorem 4’s Π_off. Either measure a closer proxy (e.g., residual on planner-rolled states vs. on μ) or qualify that bL is a mechanism diagnostic, not an estimator of L or of the bound.","section":null}],"minor_comments":[{"comment":"Notation drift: plan-cost is D̂/D in the main text (Eq. 2) but Ĵ/J and ΔJ in Appendix C; terminal cost is sometimes distance, sometimes squared distance. Unify.","section":null},{"comment":"Figure 1 right panel reports Pearson r=−0.91 for gain vs base success, while the caption text says “Pearson r=0.91” in one place in the manuscript body—fix the sign consistency.","section":null},{"comment":"Table 3 bound column is extremely loose (e.g. suboptimality 11.9 vs bound 564). Note that the bound is order-of-magnitude only, or tighten constants, so readers do not over-read “the bound holds.”","section":null},{"comment":"Related work: exposure bias and scheduled sampling are cited; a brief pointer to distributionally robust / pessimistic MPC and offline RL pessimism would better situate Proposition H6 (ˆJ+λˆr).","section":null},{"comment":"Typos / formatting: “A CONTROLTHEORY OFPREDICTABILITY INLATENTWORLDMODELS” title spacing; “Mezi ´c”; occasional missing spaces in compound words in the preprint header.","section":null}],"recommendation":"major_revision","confidential_remarks":"The mathematical core (Theorem 1, change-of-measure decoupling, spectral-tax frontier) is publishable and above the bar for a theory-leaning ML venue once the abstract’s “does not track” language is aligned with what is proved versus what is observed on three seeds. I would not reject on novelty grounds; the LeJEPA zero-tax positioning is a genuine contribution. Main risk is overclaiming empirical decoupling relative to sample size—standard major-revision territory, not a soundness failure of the theorems."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is simple and right: once a planner queries off the training support, data-averaged prediction error cannot certify control quality. Theorem 1 is elementary argmin regret (suboptimality ≤ 2g), but it correctly names the target—predicted plan-cost fidelity at the committed plan. The three lemmas then make the failure mode precise: change-of-measure decoupling when ν is not absolutely continuous w.r.t. μ, one-sided holes, and the need for realizable nonlinearity. That packaging is new enough to matter for JEPA/Dreamer-style work, and the authors are honest about what they are doing.\n\nWhat they do well is the pricing under the linear-control premise. The on-manifold spectral tax (pseudospectra, Jordan staircase, zero-tax self-adjoint boundary recovering LeJEPA) and the off-manifold rollout tax are carefully derived; synthetic operators recover the claimed exponents closely. The linear state-readout intervention is a clean, theory-guided lever that acts on amplification rather than pretending a μ-loss can kill L. Citations are appropriate (JEPA, Koopman, Trefethen, exposure bias) without padding.\n\nSoft spots, in proportion. The stress-test is right on the slogan: “does not bound” does not force “ρ \to 0.” Shared factors can keep MSE and success correlated even when MSE alone cannot certify success; under their own residual model δ₀ still sits inside Theorem 4. The only evidence for “essentially uncorrelated” is rank correlation 0.024 on a handful of runs (three seeds, one weak 78% seed, Jacobian proxy for L, ν-fidelity only 0.62 and failed as a selection rule). They flag the limits themselves. The linear-control premise and linear residual growth are strong; if residual growth is superlinear or the manifold highly curved, the clean split needs re-derivation. No code or data.\n\nThis is for people who train latent planners and care about what to optimize and monitor. The math is solid enough, the conceptual claim is load-bearing and mostly correct, and the empirics are directional rather than decisive. I would send it to peer review; it deserves a serious referee who will demand multi-seed replication and a cleaner statement of what the theory actually forces versus what the runs suggest. Engage with the theory; treat the correlation slogan as provisional.","headline":"Solid control-theoretic reframing of latent world-model objectives; the bound and decoupling lemmas are real, the empirics are thin, and the “does not track” slogan overreaches the theory.","tokens_in":40175,"tokens_out":637,"would_cite":true,"duration_ms":7625,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A planner's regret is bounded by how badly predicted plan-cost diverges from true plan-cost at the plan it commits to—not by data-averaged prediction error.","keywords":["World Models","Joint-Embedding Prediction","Koopman Operators","Pseudospectra","Model-Predictive Control","Latent Planning","Off-Manifold Error","Spectral Tax"],"falsifier":"Across matched training seeds of a latent model-predictive controller, if single-step validation error on the expert distribution ordered models by control success as tightly as a fidelity score measured on the planner-reachable states, the claimed decoupling would fail; the reported near-zero rank correlation of MSE with success and the positive correlation of the reachable-measure fidelity would have to reverse.","tokens_in":39913,"feed_emoji":"🎯","tokens_out":1117,"duration_ms":9832,"temperature":0.7,"pith_summary":"Latent world models are trained to predict future states and then used inside planners that pick actions by rolling those predictions forward. Standard practice treats lower prediction error on held-out data as the path to better control. This paper argues that this is the wrong objective for a structural reason: the planner queries the model on states that candidate actions reach, which typically leave the data manifold, so an error averaged on the training distribution cannot govern what happens off it. The authors reframe the objective as the gap between predicted and true plan-cost at the plan the planner actually commits to, and prove that planner suboptimality is at most twice that gap. Under a linear-control premise the gap splits into a small on-manifold residual, priced by a spectral tax on non-normality of the latent transition operator, and an off-manifold divergence that no data-averaged error bounds and that experiments identify as the binding term. Synthetic operators recover the pricing formulas, and latent model-predictive control runs show that single-step validation error is essentially uncorrelated with control success while a fidelity score on the planner-reachable measure tracks it.","feed_headline":"Prediction error does not govern latent-world control","feed_subtitle":"Planner regret tracks cost gap on reachable states; data-averaged MSE neither bounds nor predicts it","key_machinery":"The control tax: the supremum gap g = sup |D̂(u) − D(u)| between predicted and true plan-cost on the candidate set. Theorem 1 shows planner regret is at most 2g. Under the linear-control premise this gap separates into an on-manifold residual (priced by a spectral tax from non-normality of the latent transition) and an off-manifold divergence (a rollout tax that compounds residual along action-selected trajectories and is bounded by no data average).","core_discovery":"For a planner that selects the action of lowest predicted plan-cost, suboptimality is bounded by twice the supremum gap between predicted and true plan-cost over the candidate set. The correct control objective is therefore to make predicted cost track true cost at the committed plan. Once candidate actions carry the query off the data support, the data-averaged prediction error neither bounds nor tracks this gap, so lowering it is neither necessary nor sufficient for control.","pith_inferences":["Any latent planner whose candidate actions systematically leave the training support will show the same MSE–success decoupling; the phenomenon is not specific to the two control tasks reported.","Online monitors built from off-manifold amplification proxies could serve as early-warning signals for control collapse even when validation loss looks healthy.","Combining a linear readout (metric tax) with counterfactual or pessimistic hole-filling (extrapolation tax) is the natural next objective; relative landing on the two-tax surface is a direct test of the theory's two-region structure.","The same on/off-manifold split may reappear in any learned simulator used for trajectory optimization, not only joint-embedding predictors."],"forward_implications":["Model selection and training for latent planners should target predicted-versus-true plan-cost agreement on the reachable measure, not held-out prediction loss alone.","Once the planner leaves the data manifold, lowering on-manifold MSE does not reduce the binding extrapolation tax; interventions must fill holes (pessimism or counterfactual data) or shrink nonlinear growth off support.","The linear state readout is a minimal on-manifold intervention that can shrink goal-relevant amplification without changing the off-manifold residual itself.","Safe control is geometrically characterized: suboptimality is zero when the manifold is invariant and actions stay tangent, independent of residual size on the manifold.","The spectral tax vanishes exactly on the self-adjoint reversible boundary, recovering prior ideal-case guarantees as the zero-cost corner of a larger cost surface."],"fun_headline_variants":["Data MSE neither bounds nor tracks latent-model planner suboptimality","Plan-cost gap on reachable states bounds control, not prediction error","Off-manifold divergence, not on-data error, drives latent control failure","Latent MPC success tracks planner-reachable fidelity, not validation MSE","Predicted-true cost gap at committed plan, not MSE, governs control"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"The pricing formulas rest on treating the deployed latent dynamics as a learned linear model whose residual grows at most linearly with distance from the data manifold; if residual growth is superlinear, the manifold is highly curved, or the predictor is strongly nonlinear off support, the clean split and the stated constants need not hold.","fun_headline_variants_meta":{"raw":{"variants":["Data MSE neither bounds nor tracks latent-model planner suboptimality","Plan-cost gap on reachable states bounds control, not prediction error","Off-manifold divergence, not on-data error, drives latent control failure","Latent MPC success tracks planner-reachable fidelity, not validation MSE","Predicted-true cost gap at committed plan, not MSE, governs control"]},"model":"grok-4.5","effort":"low","cost_usd":0.00389,"raw_usage":{"total_tokens":1311,"prompt_tokens":850,"num_sources_used":0,"completion_tokens":96,"cost_in_usd_ticks":38900000,"prompt_tokens_details":{"text_tokens":850,"audio_tokens":0,"image_tokens":0,"cached_tokens":384},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":365,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":850,"tokens_out":96,"duration_ms":5400,"temperature":1.0,"reasoning_tokens":365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T12:18:39.427482+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Across matched training seeds of a latent model-predictive controller, if single-step validation error on the expert distribution ordered models by control success as tightly as a fidelity score measured on the planner-reachable states, the claimed decoupling would fail; the reported near-zero rank correlation of MSE with success and the positive correlation of the reachable-measure fidelity would have to reverse.","supporting_citations":[],"review_version":1}