Pith. sign in

REVIEW 16 cited by

Temporal Difference Learning for Model Predictive Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.04955 v2 pith:NZD34S44 submitted 2022-03-09 cs.LG cs.RO

classification cs.LGcs.RO
keywords modelcontrollearnedlearningdifferenceefficiencymethodsmodel-free
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data-driven model predictive control has two key advantages over model-free methods: a potential for improved sample efficiency through model learning, and better performance as computational budget for planning increases. However, it is both costly to plan over long horizons and challenging to obtain an accurate model of the environment. In this work, we combine the strengths of model-free and model-based methods. We use a learned task-oriented latent dynamics model for local trajectory optimization over a short horizon, and use a learned terminal value function to estimate long-term return, both of which are learned jointly by temporal difference learning. Our method, TD-MPC, achieves superior sample efficiency and asymptotic performance over prior work on both state and image-based continuous control tasks from DMControl and Meta-World. Code and video results are available at https://nicklashansen.github.io/td-mpc.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies

    cs.AI 2026-07 conditional novelty 7.0 of 10

    A counterfactual audit separates same-state headroom from recoverable state-allocation gain, returning NO-GO or ABSTAIN for learned command adapters on frozen Go2 and H1 locomotion policies at 1% thresholds.

  2. Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning

    cs.LG 2026-03 conditional novelty 7.0 of 10

    A symplectic LQR layer inserted as an adapter into pretrained LLMs yields large gains on MATH-500, AMC and AIME by solving a latent optimal-control problem at inference time.

  3. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  4. World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Pairing VLM-generated action proposals with rollouts from a pose-image-conditioned video world model yields high success rates in novel simulated manipulation tasks without end-to-end policy retraining.

  5. Toward Goal-Agnostic Joint-Embedding Predictive Control of Partial Differential Equations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A goal-agnostic latent-dynamics controller for 2D Navier-Stokes improves tracking by planning against a learned kinetic-energy probe rather than raw latent-space distance.

  6. DriftWorld: Fast World Modeling through Drifting

    cs.RO 2026-07 conditional novelty 6.0 of 10

    An action-conditioned world model trained with drifting generates robot rollout videos in one forward pass, matching diffusion quality while running 2.8-478x faster per-table, and raises GPC-RANK Push-T IoU from 0.635...

  7. Next-Latent Prediction Transformers Learn Compact World Models

    cs.LG 2025-11 unverdicted novelty 6.0 of 10

    NextLat augments next-token prediction with latent next-state prediction, theoretically converging latents to belief states and showing empirical gains in world modeling, reasoning, planning, and faster inference via ...

  8. Arnold: a generalist muscle transformer policy

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A single transformer policy with a compositional sensorimotor vocabulary achieves expert or super-expert performance on 14 musculoskeletal control tasks spanning four embodiments.

  9. Latent Policy Barrier: Learning Robust Visuomotor Policies by Staying In-Distribution

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Latent Policy Barrier improves behavior-cloned visuomotor policies by using a latent dynamics model trained on expert and rollout data to guide actions back toward in-distribution expert states.

  10. WoMAP: World Models For Embodied Open-Vocabulary Object Localization

    cs.RO 2025-06 conditional novelty 6.0 of 10

    WoMAP generates training data from Gaussian Splatting scenes, distills detector confidence into a latent world model, and uses that model to refine vision-language action proposals for open-vocabulary object localization.

  11. CoRL-MPPI: Enhancing MPPI With Learnable Behaviours For Efficient And Provably-Safe Multi-Robot Collision Avoidance

    cs.RO 2025-11 conditional novelty 5.0 of 10

    CoRL-MPPI injects a learned cooperative policy into MPPI's sampling distribution to speed up and make safer multi-robot navigation in dense simulations.

  12. Sample-Efficient Reinforcement Learning Controller for Deep Brain Stimulation in Parkinson's Disease

    cs.LG 2025-07 reject novelty 5.0 of 10

    A DDPG-based adaptive DBS controller with a predictive reward model and Gumbel-Softmax exploration suppresses beta power faster than standard DDPG in simulation and survives FP16 quantization.

  13. M3PO: Massively Multi-Task Model-Based Policy Optimization

    cs.LG 2025-06 reject novelty 4.0 of 10

    A hybrid model-based RL algorithm combining TDMPC2-style world models, PPO clipped updates, and POME-style exploration bonuses reports strong benchmark results but omits reproducible evidence.

  14. Investigating Lagrangian Neural Networks for Infinite Horizon Planning in Quadrupedal Locomotion

    cs.RO 2025-06 conditional novelty 4.0 of 10

    In Isaac Gym simulation, Lagrangian-structured dynamics models learn faster and predict motion more accurately than a plain MLP, but the headline 10x and 2-10x numbers are not fully supported.

  15. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

  16. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools