Pith. sign in

REVIEW 5 cited by

Humanoid Locomotion as Next Token Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.19469 v1 pith:OKA42BU4 submitted 2024-02-29 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords datamodelnextpredictiontokentrajectorieshumanoidcontrol
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We cast real-world humanoid control as a next token prediction problem, akin to predicting the next word in language. Our model is a causal transformer trained via autoregressive prediction of sensorimotor trajectories. To account for the multi-modal nature of the data, we perform prediction in a modality-aligned way, and for each input token predict the next token from the same modality. This general formulation enables us to leverage data with missing modalities, like video trajectories without actions. We train our model on a collection of simulated trajectories coming from prior neural network policies, model-based controllers, motion capture data, and YouTube videos of humans. We show that our model enables a full-sized humanoid to walk in San Francisco zero-shot. Our model can transfer to the real world even when trained on only 27 hours of walking data, and can generalize to commands not seen during training like walking backward. These findings suggest a promising path toward learning challenging real-world control tasks by generative modeling of sensorimotor trajectories.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Deep Sensorimotor Control by Imitating Predictive Models of Human Motion

    cs.RO 2025-08 conditional novelty 7.0 of 10

    A predictive model of human hand motion, trained on human interaction data, can reward a robot policy for tracking predicted future keypoints and enable learning of dexterous manipulation from sparse rewards.

  2. Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Auras, a perception-generation disaggregation framework with a public context buffer and asynchronous pipeline executor, raises embodied-agent throughput by 2.54x on average without losing accuracy (102.7%).

  3. SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SkillBlender pretrains reusable goal-conditioned skills and blends them with softmax per-joint weights to solve simulated humanoid loco-manipulation tasks with one or two reward terms.

  4. Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.

  5. Poly-Autoregressive Prediction for Modeling Interactions

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A single transformer training recipe, poly-autoregressive prediction, improves multi-agent ego forecasting over autoregressive baselines on three distinct tasks.

Pith tools