Pith. sign in

REVIEW 3 major objections 4 minor 45 references

Metric Non-Collapse in Learned World Models for Control: Approximation Theory, Finite-Sample Geometric Guarantees, and Deterministic Planning Transfer

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Under structural assumptions and enough samples, a local–global metric hinge makes every near-minimizer of a regularized latent world model pointwise injective, uniformly predictive, and transferable to deterministic control.

desk verdict A serious conditional theorem with an uninstantiated sufficient condition; worth refereeing despite the practical gap. read the letter →

arxiv 2608.07265 v1 pith:B53DDJJR submitted 2026-08-07 math.OC

classification math.OC MSC 93C1093B1741A1568T05
keywords metricnon-collapseworldmodelsjoint-embeddingpredictivearchitecturesco-Lipschitzencodercontrolledsemiconjugacyfinite-sampleguaranteesmodelcontrolB-splineapproximation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to prove a training-to-control chain: if a latent world model is fitted with a specific one-sided metric regularizer, then under structural assumptions and enough samples every approximate optimiser is guaranteed to keep distinct observable states apart, to predict future latent states uniformly well, and to transfer to finite-horizon planning with an explicit suboptimality bound. The authors show that pure forward prediction is insufficient, because constant encoders solve it exactly, and that covariance or spectral-spread penalties are insufficient, because folded non-injective encoders can still have large latent variance. Their regularizer is an encoder-only local–global metric hinge that penalises infinitesimal directional collapse and overlap of states that are at least a set distance apart, using state-metric supervision from simulator, proprioceptive, or state-estimation settings. The main theorem applies to every $\varepsilon_{\mathrm{train}}$-approximate empirical minimizer above an explicit regularization threshold once the finite-sample statistical deviation lies below a metric-margin threshold. If the claim is right, a learned representation can carry a verifiable certificate of geometric faithfulness and a worst-case deterministic planning guarantee at the same time.

What carries the argument

The load-bearing mechanism is the encoder-only local–global metric hinge $N_{\mathrm{met}}(\Phi)=N_{\mathrm{loc}}(\Phi)+N_{\mathrm{glob}}(\Phi)$, with $N_{\mathrm{loc}}=\int (\kappa-\|D\psi(s)[v]\|_Z)_+^2\,d\omega$ and $N_{\mathrm{glob}}=\int (\alpha-\|\psi(s)-\psi(s')\|_Z)_+^2\,d\nu_\rho$, where $\psi=\Phi\circ H$. The local term penalizes infinitesimal directional collapse below slope $\kappa$; the separated-pair term penalizes identified states that are at least $\rho$ apart below separation $\alpha$. The theorem then uses: an exact $C^{1,1}$ realization $(\Phi_0,F_0)$ built from a bi-Lipschitz observable embedding via extension theorems and local inverse charts; a margin-clearing competitor with zero metric penalty and prediction loss $\beta(W)$; uniform statistical deviation $\varepsilon_{\mathrm{stat}}$ from covering numbers; Lipschitz–Ahlfors $L^2$-to-$L^\infty$ interpolation to convert averaged defect bounds into pointwise margin certificates; and a deterministic simulation inequality $e_{t+1}\le\eta+L_F e_t$ that propagates semiconjugacy to trajectory and cost errors.

What would settle it

On a deterministic toy system satisfying the assumptions (the saturated pendulum of Experiment C is one), fix margins inside the admissible regime and $\lambda$ above $\lambda_{\min}$, and train many seeds on frozen samples. The theorem says violations — a pair $s,s'$ with $\|\Phi(H(s))-\Phi(H(s'))\| < c_*|s-s'|$, or a state-action pair with residual above $\eta$ — occur with probability at most $\delta$. If a large run of seeds produces violations with empirical frequency clearly above $\delta$ while the training tolerance is certified small, the central claim is refuted; the paper's own dense-grid diagnostics make this test directly computable.

Watch

Extended reading notes

Core claim

The central claim is the finite-sample implication objective-to-geometry-to-dynamics-to-control: approximate optimisation of an empirical prediction loss plus $\lambda$ times a local–global metric hinge forces a pointwise co-Lipschitz encoder, a uniform approximate controlled-semiconjugacy estimate, and then a deterministic worst-case finite-horizon optimizer-transfer bound, all for the same learned model and with approximation, statistical, training, cost-head, and planner errors kept separate. The proof compares any approximate empirical minimizer against a margin-clearing competitor constructed from a smooth exact realization; because the competitor clears the local margin $\kappa<\kappa_0$ and global margin $\alpha<\alpha_0(\rho)$, its metric penalty is zero and its prediction loss $\beta(W)\to 0$ as capacity grows. Choosing $\lambda\ge(\beta(W)+\varepsilon_{\mathrm{stat}}+\varepsilon_{\mathrm{train}})/(P^*-\varepsilon_{\mathrm{stat}})$ forces the minimizer's population metric penalty below the threshold $P^*$, and a Lipschitz–Ahlfors $L^2$-to-$L^\infty$ interpolation turns the averaged hinge defects into pointwise margin certificates, giving $c_*=\min\{\kappa/4,\ \alpha/(2\operatorname{diam} S)\}$ and residual modulus $\eta=\Theta_M(\beta(W)+2\varepsilon_{\mathrm{stat}}+\varepsilon_{\mathrm{train}})$. The paper also proves that the exponent $1/(d_m+2)$ in that interpolation is sharp under Lipschitz regularity alone and verifies the required finite-capacity $C^{1,1}$ approximation property for norm-constrained tensor-product B-spline classes.

Load-bearing premise

The load-bearing premise is that the true system admits a smooth, distance-preserving-in-both-directions embedding of its observable state into the latent box, with the chosen margins lying below unknown constants $\kappa_0$ and $\alpha_0(\rho)$ of that embedding; the paper provides no way to compute those constants, and without such an embedded reference point the comparison argument — and with it the guarantee — collapses.

Editorial extensions

If this is right

  • At sufficient capacity and sample size, every approximate minimizer of the regularized objective is quantitatively injective: distinct observable states map to latent codes separated by at least $c_*$ times their state distance, so representations cannot fold or identify states with different future consequences.
  • The same learned model satisfies a uniform approximate semiconjugacy bound: encoding-then-evolving and evolving-then-encoding agree to within $\eta$ uniformly over state-action pairs, not just on average.
  • For Lipschitz stage and terminal costs, any $\xi_{\mathrm{plan}}$-optimal action sequence for the induced latent problem achieves true cost within $2\Delta_{T,L}(\eta)+\xi_{\mathrm{plan}}$ of optimal, with the geometric constant $c_*$ controlling the induced latent-cost Lipschitz constants.
  • The admissible regularization strengths form the entire half-line $[\lambda_{\min},\infty)$; when approximation, statistical, and training errors all vanish, every fixed $\lambda>0$ eventually becomes admissible.
  • Learned latent cost heads are covered by an explicit additive term $2(T\varepsilon_\ell+\varepsilon_g)$, so error from imperfect cost modelling remains separated from transition error in the planner bound.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Our inference: the theorem yields a practical two-sided audit — track the sampled lower chord ratio together with an upper chord ratio, since the compact-range and uniform-budget hypotheses imply that one-sided resolution bought by scaling up the latent code is an artifact; the paper's Experiment A3 exhibits exactly this inflation failure.
  • Our inference: the sharpness of the $L^2$-to-$L^\infty$ exponent suggests that any substantial dimension improvement must come from extra structure — controlling the residual on the planner's reachable set, higher-order regularity, or trajectory-aware objectives — because Lipschitz regularity alone is shown to be exhausted by the paper's bound.
  • Our inference: because the sufficient condition requires $\varepsilon_{\mathrm{stat}} < P^*$, and the paper's own finite-capacity proxy reports $\varepsilon_{\mathrm{stat}}/P^* > 10^5$ at the tested sample sizes, practical adoption will need either much larger sample regimes or tighter statistical estimates than the uniform covering-number bound.
  • Our inference: the analysis stops at pixel-only observations; the metric hinge explicitly requires observable-state distances and tangent directions, so transferring the guarantee to a purely observational setting demands a separate method for estimating those geometric quantities.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper develops a three-part mathematical theory for learned latent world models in deterministic control. It constructs smooth exact latent realizations under a regular observable-factor assumption, introduces an encoder-only local–global metric hinge that penalizes infinitesimal directional collapse and pairwise overlap at a separation scale, and proves a finite-sample representation theorem (Theorem 7.15). Under Assumptions A1--A4, the realizability hypotheses of Lemma 7.3, and the parameter regime of Lemma 7.8, the theorem states that if the statistical deviation satisfies ε_stat,W(n,δ) < P* and the regularization strength is at least an explicit λ_min, then with probability at least 1−δ every ε_train-approximate empirical minimizer is pointwise co-Lipschitz, satisfies a uniform approximate controlled-semiconjugacy bound, and yields deterministic finite-horizon planning-transfer guarantees. The proof chain runs through a margin-clearing competitor, an L²-to-L∞ interpolation lemma under lower Ahlfors coverage, a small-arc argument, and a rollout recursion. The paper also proves sharpness of the exponent 1/(d_m+2), provides a spline construction verifying the finite-capacity approximation hypothesis, and reports experiments on analytic obstructions, a spline proxy, and a pendulum planning benchmark.

Significance. The conditional theorem is the paper's central contribution and appears mathematically sound. I checked the main chain of Section 7: the margin-clearing comparison, the L²-to-L∞ interpolation via Lemma 6.1, the small-arc argument producing the lower co-Lipschitz constant, and the recursive rollout bound leading to optimizer transfer are internally consistent. If the sufficient condition ε_stat,W(n,δ) < P* were instantiated in a practical regime, the result would be a substantial contribution: it connects an approximately optimized finite-sample metric objective to pointwise geometry, uniform semiconjugacy, and worst-case planning certificates. The paper is also unusually honest about its limitations: it reports the folded-basin failures, states that the spline proxy does not verify Assumption A4, reports that no reliable ranking among regularized objectives was found, and explicitly records in Table 10 that the theorem's sufficient condition is not met in the experiments. The exact-rational certificate and archived reproducible code are additional strengths.

major comments (3)
  1. [10.2, Table 10; Theorem 7.15, Eq. (7.1)] The sufficient condition ε_stat,W(n,δ) < P* that gates the representation and planning guarantees is not met in any reported configuration: the plug-in diagnostic shows ε_stat/P* > 10^5 for every capacity and sample size tested, and λ_min^thm is therefore undefined. This means the paper's headline end-to-end finite-sample guarantee is uninstantiated at all tested sample sizes. The authors are transparent about this, but the abstract and introduction present the finite-sample implication as the principal contribution without noting that the sufficient condition is far outside the experimental regime. Since the theorem is conditional, this is not a mathematical error, but the paper should either exhibit a quantitative regime in which the condition holds (using Corollary 7.11) or explicitly qualify the claim of a finite-sample guarantee as conditional on an unrealized sufficient condition.
  2. [7.4, Proposition 7.7 and Lemma 7.8] The admissible parameter region is defined through the unknown exact realization ψ0: the margins must satisfy κ < κ0 and α < α0(ρ), where κ0 and α0(ρ) are infima over ψ0. The theorem provides no computable certificate for these constants, so a user cannot verify the main hypothesis without already knowing the system. Experiment B chooses margins for the known realization ψ0(s) = s, so the numerical study does not certify the hypothesis. I view this as an input condition rather than a flaw in the proof, but it is load-bearing for application; please add a discussion of how κ0 and α0(ρ) could be estimated or certified from data, or state prominently that the guarantee is conditional on unverified constants.
  3. [7.3, Lemma 7.10 and Corollary 7.11] The uniform covering-number bound for ε_stat is the bottleneck that makes the theorem's sufficient condition unreachable at the tested sample sizes, while the paper's own diagnostic shows that observed finite-sample deviations are orders of magnitude smaller. Since every downstream conclusion is gated by ε_stat < P*, the conservativeness is not a cosmetic issue. The paper should investigate tighter concentration arguments for the hinge classes, such as localized Rademacher complexities or chaining, or provide a quantitative sample-size requirement at which the current bound becomes informative.
minor comments (4)
  1. [Theorem 1.1] The informal theorem states that the conclusion holds for sufficiently large capacity W and sample size n, but omits the threshold condition ε_stat,W(n,δ) < P*. Since the formal conclusion is vacuous without that condition, the informal statement should include it to avoid overstating the result.
  2. [Table 10] The λ_min^thm column is undefined for every row; the table should state explicitly in the body of the table, not only in the caption, that λ_min is defined only when ε_stat < P*.
  3. [Section 10.3, Eq. (10.5)] The quantity B_eval_T is called an evaluation-set theorem proxy, but the text sometimes reads as if it were a bound. Consider consistently calling it an evaluation-set diagnostic or proxy and explicitly stating that it is not a certified bound, since it omits ξ_plan and uses finite-set suprema.
  4. [Section 9.2, Table 2] The variable r in the spike example is used before its role is made explicit; please define r as the support radius of the spike and state the normalization of Lebesgue measure on [0,1]^{d_m} before the table.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the theorem is a conditional implication with explicit inputs, derived constants, and no fitted quantity renamed as a prediction.

full rationale

The paper's derivation chain is self-contained as a conditional proof. The metric hinge margins κ, α, ρ are user-chosen inputs, not fitted parameters, and the co-Lipschitz constant c* = min(κ/4, α/(2 diam(S))) is derived from them through the L2-to-L∞ interpolation lemma under lower Ahlfors coverage, rather than being imposed by definition. The margin-clearing competitor of Proposition 7.7 is an existence input: it requires an assumed C^{1,1} bi-Lipschitz embedding ψ0 of the observable factor with positive constants κ0 and α0(ρ). That is a sufficient condition stated before the theorem, not a conclusion smuggled in, and the theorem's guarantees are conditional on it. The statistical deviation ε_stat,W(n,δ) is bounded by covering-number arguments, and the admissible regularization threshold λ_min = (β(W)+ε_stat+ε_train)/(P*−ε_stat) is a computed sufficient condition; the predicted semiconjugacy modulus η_W,n,δ(ε_train) is an explicit function of approximation, statistical, and optimization errors, so no fitted input is renamed as a prediction. The paper is also explicitly honest about the limits of its numerical study: it states that the implemented spline proxy is not the uniform-capacity C^{1,1} class of Proposition 7.6, that the plug-in threshold diagnostic shows ε_stat/P* exceeds 10^5 at all tested sample sizes, and that no neural or spline run is presented as verifying the theorem. These are limitations that affect instantiation, not circularity. There is no load-bearing self-citation: the cited classical results (McShane extension, B-spline quasi-interpolation, Hermann–Krener observability, Kearns–Singh simulation bound) are external and independently established, and the paper does not invoke an author-specific uniqueness theorem to force its construction. No equation in the paper reduces the conclusion to its own inputs by construction, and no known empirical pattern is merely renamed. Score 0 is therefore appropriate.

Assumptions & free parameters 4 free parameters · 7 assumptions · 0 invented entities

The central theorem rests on the sampling and regularity assumptions A1-A4 plus an explicit realizability and margin-selection condition. These are domain assumptions rather than standard math; they are stated clearly, and the spline construction in Proposition 7.6 shows A4 is satisfiable. The most restrictive are A2/A3 (coverage of all state-action and sampling spaces) and the realizability of a smooth bi-Lipschitz embedding with known margins.

free parameters (4)
  • κ (local directional margin) = 0.4 in Experiment B; in general must satisfy 0 < κ < κ_0
    Hand-chosen margin for the directional hinge. Its admissible upper bound κ_0 depends on the unknown exact realization, so choosing a valid κ requires knowledge not provided by the theorem.
  • α (separated-pair margin) = 0.3 in Experiment B; in general must satisfy 0 < α < α_0(ρ)
    Hand-chosen margin for the separated-pair hinge, bounded above by the unknown minimum separation of the exact realization on P_ρ.
  • ρ (separation scale) = 0.35 in Experiment B; in general must satisfy 0 < ρ < min(diam(S), κ/(4 B_loc))
    Hand-chosen scale separating local and global regimes; also enters α_0(ρ) and P_ρ.
  • λ (regularization strength) = swept in experiments (e.g., 0.03 to 30)
    Hyperparameter of the training objective; theorem requires λ ≥ λ_min(β, ε_stat, ε_train, P*), where all quantities are assumed known.
assumptions (7)
  • domain assumption A1: Regular observable Markov factor with H_o bi-Lipschitz onto its image.
    Requires a finite-dimensional observable factor q separating observationally distinct states, with dynamics, observation, and costs factoring through q and H_o bi-Lipschitz onto its image. Invoked in Assumption A1, Section 3. Without it, the Euclidean co-Lipschitz estimates have no target space.
  • domain assumption A2: Lower Ahlfors coverage of the training measure μ_o on S_o × A.
    μ_o is lower Ahlfors (d_o+d_a)-regular on S_o × A. Used in Lemma 6.1 to convert the L2 residual bound into a sup-norm semiconjugacy bound. Requires the training measure to cover all of state-action space.
  • domain assumption A3: Metric sampling geometry lower Ahlfors for ω on U and ν_ρ on P_ρ.
    ω on U and ν_ρ on P_ρ are lower Ahlfors regular with exponents 2d_s-1 and 2d_s. Used to uniformize the local and separated-pair hinge defects from L2 to L∞.
  • domain assumption A4: Finite-capacity C1 approximation with uniform C1,1 budgets.
    Assumes admissible classes A_Φ(W), A_F(W) with W-independent budgets and O(δ(W)) C1 approximation of the exact realization. Proposition 7.6 verifies it for norm-constrained tensor-product cubic B-spline classes.
  • ad hoc to paper Realizability of an exact C1,1 bi-Lipschitz embedding ψ_0 inside the latent box with a range-enforcing map Π_Z.
    Lemma 7.3 and Lemma 7.8 require a smooth ψ_0 with c_{ψ_0}>0 embedded in the interior of K_Z, plus a range-enforcing map Π_Z identity on a neighborhood. This is an additional structural assumption on the true system, not implied by A1-A4.
  • domain assumption Lipschitz continuity of stage and terminal costs as in (7.3).
    Inequality (7.3) requires stage and terminal costs Lipschitz in state and action; needed for Lemma 4.1 induced latent costs and the optimizer-transfer bound.
  • standard math Standard mathematical facts: McShane extension, Hoeffding's inequality, B-spline quasi-interpolation.
    McShane extension (Lemma 4.1, 7.1), Hoeffding inequality plus union bound (Lemma 7.10), and cubic B-spline quasi-interpolation estimates (Lemma 7.4, citing Schumaker) are used as background.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Metric Non-Collapse in Learned World Models for Control: Approximation Theory, Finite-Sample Geometric Guarantees, and Deterministic Planning Transfer." pith.science (2026). https://pith.science/paper/B53DDJJR

@misc{pith2026260807265,
  author       = {Pith},
  title        = {Pith review of: Metric Non-Collapse in Learned World Models for Control: Approximation Theory, Finite-Sample Geometric Guarantees, and Deterministic Planning Transfer},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/B53DDJJR}},
  note         = {Machine review of arXiv:2608.07265}
}
abstract

We develop a three-part mathematical theory for metrically faithful learned world models for nonlinear deterministic control systems. First, on the approximation-theoretic side, we construct smooth exact latent realizations and verify the required finite-capacity $C^{1,1}$ approximation property for norm-constrained tensor-product B-spline classes with capacity-independent regularity budgets. Second, on the finite-sample geometric side, we introduce an encoder-only local--global metric hinge whose directional and separated-pair terms prevent infinitesimal collapse and global folding. Although the prediction loss is observation-based, evaluation of this penalty uses state-metric supervision through observable-state distances and tangent directions. Under a regular observable-factor assumption, lower-Ahlfors coverage, and uniform $C^{1,1}$ budgets, provided the explicit finite-sample deviation lies below the metric-margin threshold, every approximate empirical minimizer above a computable one-sided regularization threshold is pointwise co-Lipschitz and satisfies a uniform approximate controlled-semiconjugacy estimate, with approximation, statistical, and training-optimization errors kept separate. The $L^2$-to-$L^\infty$ exponent used in this step is sharp under Lipschitz regularity. Third, for deterministic planning transfer, metric non-collapse induces compatible Lipschitz latent costs, while semiconjugacy gives uniform trajectory and finite-horizon cost bounds and an optimizer-transfer guarantee; learned latent cost heads enter through explicit compatibility errors. Numerical experiments use bounded latent ranges, archived fixed metric samples, a finite-capacity coefficient-box spline proxy motivated by the theorem, and latent model-predictive control on a controlled pendulum.

Figures

Figures reproduced from arXiv: 2608.07265 by the authors.

Figure 1
Figure 1. Experiment A3. Learned encoder profiles s 7→ ψ(s) under a common com￾pact latent range, for the seven objectives and 12 paired seeds in the folded-basin initial￾ization regime. Dashed curves are noninjective on the evaluation grid. The right panel plots the sampled lower against the sampled upper chord ratio: points far to the right are metrically faithful, points high up are metrically inflated, and only the lower-… view at source ↗
Figure 2
Figure 2. Experiment B. Dense-grid local and separated-pair metric defects and the uniform residual against the regularization strength, for several sample sizes, for the norm-constrained spline class. Bands are min–max over seeds. Markers on the right panel distinguish the fixed-sample training used for the theorem-facing runs from the two resampled variants of the ablation. do not treat the experiment as a numerical verific… view at source ↗
Figure 3
Figure 3. Experiment C. Worst-case paired closed-loop cost difference against the geometry/residual ratio ηbeval/bcmin (top) and against the evaluation-set theorem proxy Beval T of (10.5) (bottom), colored by method and faceted by horizon. metric measures the window rather than the representation, and we treat it as supplementary. The evaluation-set proxy Beval T of (10.5) is very loose — a median factor of 81.95 above the ob… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Experiment C. Representative true closed-loop trajectory and control signal under exact-model MPC and under latent MPC with the covariance, pure-prediction and proposed-hinge representations. method bcmin bcmax κbgeom Pure prediction 1.75×10−4 [1.36×10−4 , 2.53×10−4 ] …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

45 extracted references · 32 canonical work pages

  1. [1]

    R. S. Sutton. Dyna, an integrated architecture for learning, planning, and reacting.ACM SIGART Bulletin, 2(4):160–163, 1991

  2. [2]

    Ha and J

    D. Ha and J. Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018

  3. [3]

    Hafner, T

    D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. InProceedings of the 36th International Conference on Machine Learning, volume 97 ofProceedings of Machine Learning Research, pages 2555–2565. PMLR, 09–15 Jun 2019

  4. [4]

    Hafner, T

    D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi. Dream to control: Learning behaviors by latent imagination. InInternational Conference on Learning Representations (ICLR), 2020

  5. [5]

    Hafner, J

    D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse control tasks through world models.Nature, 640(8057):647–653, 2025

  6. [6]

    Y. LeCun. A path towards autonomous machine intelligence.Open Review, 2022

  7. [7]

    Assran, Q

    M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, N. Ballas, and Y. LeCun. Self-supervised learning from images with a joint-embedding predictive architecture. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15619– 15629, 2023. 72 ALAIN BENSOUSSAN, MINH-NHAT PHUNG, AND MINH-BINH TRAN

  8. [8]

    Bardes, Q

    A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y. LeCun, M. Assran, and N. Ballas. Re- visiting feature prediction for learning visual representations from video.Transactions on Machine Learning Research, 2024

Show all 45 references
  1. [9]

    Assran, A

    M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Komeili, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, S. Arnaud, A. Gejji, A. Martin, F. R. Hogan, D. Dugas, P. Bojanowski, V. Khalidov, P. Labatut, F. Massa, M. Szafraniec, K. Krishnakumar, Y. Li, X. Ma, S. Chan...

  2. [10]

    Grill, F

    J.-B. Grill, F. Strub, F. Altch´ e, C. Tallec, P. H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. Valko. Bootstrap your own latent: A new approach to self-supervised learning. InProceedings of the 34th Int...

  3. [11]

    Chen and K

    X. Chen and K. He. Exploring simple siamese representation learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15750–15758, 2021

  4. [12]

    Zbontar, L

    J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny. Barlow twins: Self-supervised learning via redundancy reduction. InProceedings of the 38th International Conference on Machine Learning (ICML), volume 139 ofProceedings of Machine Learning Research, pages 12310–12320. PMLR, 2021

  5. [13]

    Bardes, J

    A. Bardes, J. Ponce, and Y. LeCun. Variance-invariance-covariance regularization for self- supervised learning. InInternational Conference on Learning Representations (ICLR), 2022

  6. [14]

    Givan, T

    R. Givan, T. Dean, and M. Greig. Equivalence notions and model minimization in Markov decision processes.Artificial Intelligence, 147(1–2):163–223, 2003

  7. [15]

    Ravindran and A

    B. Ravindran and A. G. Barto. SMDP homomorphisms: An algebraic approach to abstraction in semi-Markov decision processes. InProceedings of the 18th International Joint Conference on Artificial Intelligence (IJCAI 2003), pages 1011–1016, 2003

  8. [16]

    L. Li, T. J. Walsh, and M. L. Littman. Towards a unified theory of state abstraction for MDPs. InProceedings of the 9th International Symposium on Artificial Intelligence and Mathematics (AI&M), 2006

  9. [17]

    Ferns, P

    N. Ferns, P. Panangaden, and D. Precup. Metrics for finite Markov decision processes. InProceed- ings of the 20th Conference on Uncertainty in Artificial Intelligence (UAI 2004), pages 162–169, 2004

  10. [18]

    Ferns, P

    N. Ferns, P. Panangaden, and D. Precup. Bisimulation metrics for continuous Markov decision processes.SIAM Journal on Computing, 40(6):1662–1714, 2011

  11. [19]

    Gelada, S

    C. Gelada, S. Kumar, J. Buckman, O. Nachum, and M. G. Bellemare. DeepMDP: Learning continuous latent space models for representation learning. InProceedings of the 36th International Conference on Machine Learning (ICML 2019), pages 2170–2179, 2019

  12. [20]

    Zhang, R

    A. Zhang, R. McAllister, R. Calandra, Y. Gal, and S. Levine. Learning invariant representations for reinforcement learning without reconstruction. InProceedings of the 9th International Conference on Learning Representations (ICLR 2021), 2021

  13. [21]

    Kemertas and T

    M. Kemertas and T. Aumentado-Armstrong. Towards robust bisimulation metric learning. InProceedings of the 35th International Conference on Neural Information Processing Systems (NeurIPS), 2021. METRIC NON-COLLAPSE IN WORLD MODELS FOR CONTROL 73

  14. [22]

    Q. Zhan, Z. Zhou, Z. Wang, Q. Long, and L. Shen. Bi-lipschitz autoencoder with injectivity guarantee. InInternational Conference on Learning Representations (ICLR), 2026

  15. [23]

    van Amersfoort, L

    J. van Amersfoort, L. Smith, A. Jesson, O. Key, and Y. Gal. On feature collapse and deep kernel learning for single forward pass uncertainty.arXiv preprint arXiv:2102.11409, 2021

  16. [24]

    J. Cui, Q. Zhang, H. Wen, and Y. Wang. A generalization theory for JEPA-based world models. arXiv preprint arXiv:2606.27014, 2026

  17. [25]

    Klindt, Y

    D. Klindt, Y. LeCun, and R. Balestriero. When does LeJEPA learn a world model?arXiv preprint arXiv:2605.26379, 2026

  18. [26]

    Zhang, Y

    X. Zhang, Y. Guan, B. Zhang, Y.-Q. Zhang, and S. E. Li. On the identifiability of controlled world models.arXiv preprint arXiv:2607.22430, 2026

  19. [27]

    Destrade, O

    M. Destrade, O. Bounou, Q. Le Lidec, J. Ponce, and Y. LeCun. Value-guided action planning with JEPA world models.arXiv preprint arXiv:2601.00844, 2026

  20. [28]

    W. Li, G. Li, K. Maeda, T. Ogawa, and M. Haseyama. Predictive but not plannable: RC-aux for latent world models.arXiv preprint arXiv:2605.07278, 2026

  21. [29]

    Bai and J

    J. Bai and J. Xiong. Temporal-distance JEPA: Plan-aware representation learning for latent world model predictive control.arXiv preprint arXiv:2607.25337, 2026

  22. [30]

    Z. Liu, Y. Zhuang, Y. Li, W. Gou, J. Liu, M. Zhou, and M. Yang. ProWorld: Progress-aware hyperbolic world models for long-horizon visual goal reaching.arXiv preprint arXiv:2608.01926, 2026

  23. [31]

    Z. Gan, Z. Zeng, J. Cheng, Y. Song, Y. Tang, and X. Wang. ActSWM: Action-sensitive world models for long-horizon planning in open-world games.arXiv preprint arXiv:2607.26712, 2026

  24. [32]

    Bhand and A

    K. Bhand and A. Joshi. Multi-horizon consistency as geometry: When latent dynamics contract, and when they do not.arXiv preprint arXiv:2607.21645, 2026

  25. [33]

    H. Wang. Scale buys interpolation, structure buys a horizon: Certified predictability for equivariant world models.arXiv preprint arXiv:2606.13092, 2026

  26. [34]

    H. You, Y. Zhang, M. Ran, Z. Yang, Z. Zhang, W. Xue, J. Song, X. Tian, and Y. Guo. A control theory of predictability in latent world models.arXiv preprint arXiv:2607.10362, 2026

  27. [35]

    Smolyanskiy

    N. Smolyanskiy. Predicting closed-loop performance of latent world models: Offline checkpoint selection for MPC and model-based RL under non-markovian rewards in LunarLander.arXiv preprint arXiv:2607.01736, 2026

  28. [36]

    J. Yu, S. Chen, M. Liu, N. Horiuchi, V. Braverman, Z. Xu, D. Haramati, and R. Balestriero. Why and how auxiliary tasks improve jepa representations.arXiv preprint arXiv:2509.12249, 2025

  29. [37]

    Bagatella, M

    M. Bagatella, M. Pirotta, A. Touati, A. Lazaric, and A. Tirinzoni. TD-JEPA: Latent-predictive representations for zero-shot reinforcement learning.arXiv preprint arXiv:2510.00739, 2025

  30. [38]

    Y. Huang. VJEPA: Variational joint embedding predictive architectures as probabilistic world models.arXiv preprint arXiv:2601.14354, 2026

  31. [39]

    L. Maes, Q. Le Lidec, D. Scieur, Y. LeCun, and R. Balestriero. LeWorldModel: Stable end-to-end joint-embedding predictive architecture from pixels.arXiv preprint arXiv:2603.19312, 2026

  32. [40]

    Hermann and A

    R. Hermann and A. J. Krener. Nonlinear controllability and observability.IEEE Transactions on Automatic Control, 22(5):728–740, 1977

  33. [41]

    Burago, Y

    D. Burago, Y. Burago, and S. Ivanov.A Course in Metric Geometry, volume 33 ofGraduate Studies in Mathematics. American Mathematical Society, Providence, Rhode Island, 2001

  34. [42]

    E. J. McShane. Extension of range of functions.Bulletin of the American Mathematical Society, 40(12):837–842, 1934. 74 ALAIN BENSOUSSAN, MINH-NHAT PHUNG, AND MINH-BINH TRAN

  35. [43]

    L. L. Schumaker.Spline Functions: Basic Theory. Cambridge Mathematical Library. Cambridge University Press, Cambridge, 3 edition, 2007

  36. [44]

    Asadi, D

    K. Asadi, D. Misra, and M. L. Littman. Lipschitz continuity in model-based reinforcement learning. InProceedings of the 35th International Conference on Machine Learning (ICML 2018), volume 80 ofProceedings of Machine Learning Research, pages 264–273. PMLR, 2018

  37. [45]

    Kearns and S

    M. Kearns and S. Singh. Near-optimal reinforcement learning in polynomial time.Machine Learn- ing, 49(2–3):209–232, 2002. (Alain Bensoussan)Naveen Jindal School of Management, University of Texas at Dallas, Richardson, TX, 75080 USA Email address:alain.bensoussan@utdallas.edu ...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.