Pith. sign in

REVIEW 17 cited by

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.14795 v2 pith:SSKMQM3S submitted 2025-02-20 cs.RO cs.CV

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

classification cs.RO cs.CV
keywords motioncontroldatahumanoid-vlahumanoiduniversalegocentricenabling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a novel framework that integrates language understanding, egocentric scene perception, and motion control, enabling universal humanoid control. Humanoid-VLA begins with language-motion pre-alignment using non-egocentric human motion datasets paired with textual descriptions, allowing the model to learn universal motion patterns and action semantics. We then incorporate egocentric visual context through a parameter efficient video-conditioned fine-tuning, enabling context-aware motion generation. Furthermore, we introduce a self-supervised data augmentation strategy that automatically generates pseudoannotations directly derived from motion data. This process converts raw motion sequences into informative question-answer pairs, facilitating the effective use of large-scale unlabeled video data. Built upon whole-body control architectures, extensive experiments show that Humanoid-VLA achieves object interaction and environment exploration tasks with enhanced contextual awareness, demonstrating a more human-like capacity for adaptive and intelligent engagement.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

    cs.RO 2026-06 unverdicted novelty 7.0

    HumanoidArena is a new benchmark of 7 leg-critical HOI/HSI tasks that evaluates egocentric hierarchical whole-body policies in humanoids and finds performance is strongly conditioned on the low-level GMT used.

  2. Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers

    cs.CV 2026-05 unverdicted novelty 7.0

    A two-stage framework augments HOI data with dynamic priors and blends pre-trained dynamic motion and static interaction agents via a composer network to enable long-term dynamic human-object interactions with higher ...

  3. VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

    cs.RO 2026-06 unverdicted novelty 6.0

    Generates 48,000 synthetic VLK trajectories in 3D-reconstructed scenes to train a policy for egocentric perception-based humanoid navigation and object transport, shown on physical Unitree G1 robot.

  4. What Matters in Orchestrating Robot Policies: A Systematic Study of Hierarchical VLA Agents

    cs.RO 2026-06 unverdicted novelty 6.0

    A systematic study of hierarchical VLA agents identifies design principles that improve robot manipulation performance over flat and naive hierarchical baselines in simulation and real-world experiments.

  5. Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot

    cs.RO 2026-04 unverdicted novelty 6.0

    The Weightlessness Mechanism lets humanoid robots imitate non-self-stabilizing motions by dynamically relaxing specific joints to exploit passive environmental contacts, generalizing from single demonstrations to vari...

  6. Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot

    cs.RO 2026-04 unverdicted novelty 6.0

    A weightlessness mechanism enables humanoid robots to dynamically relax joints for stable, contact-rich motions across diverse environments without task-specific tuning.

  7. HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

    cs.RO 2026-04 unverdicted novelty 6.0

    HEX is a new framework with humanoid-aligned state representation, mixture-of-experts proprioceptive predictor, history tokens, and residual-gated fusion that achieves state-of-the-art success and generalization on re...

  8. HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

    cs.RO 2026-04 unverdicted novelty 6.0

    HEX introduces a state-centric framework with humanoid-aligned representations and mixture-of-experts proprioceptive prediction for coordinated whole-body control on bipedal humanoids.

  9. VLANeXt: Recipes for Building Strong VLA Models

    cs.CV 2026-02 conditional novelty 6.0

    VLANeXt distills 12 design insights from a unified VLA study into a model that outperforms prior methods on LIBERO benchmarks while releasing code for further exploration.

  10. EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration

    cs.RO 2026-02 conditional novelty 6.0

    Co-training a vision-language-action humanoid policy on aligned egocentric human demonstrations plus limited robot data improves real-world loco-manipulation success by 20% in-domain and 51% in environments the robot ...

  11. Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning

    cs.RO 2026-02 unverdicted novelty 6.0

    R&B-EnCoRe uses self-supervised importance-weighted variational inference to distill action-predictive reasoning datasets that improve VLA performance on manipulation, navigation, and driving tasks without external verifiers.

  12. Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary

    cs.RO 2025-11 unverdicted novelty 6.0

    Humanoid-LLA converts unconstrained natural language commands into stable whole-body motions for humanoid robots using a unified motion vocabulary and two-stage supervised-plus-reinforcement fine-tuning.

  13. AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation

    cs.CV 2025-07 unverdicted novelty 6.0

    AnyPos automates task-agnostic action collection and inverse-dynamics modeling with arm/end-effector decoupling plus a direction-aware decoder, delivering 51% higher test accuracy and 30-40% better success rates on bi...

  14. OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    OASIS generates scalable simulation data for humanoid loco-manipulation via 3D generative asset reconstruction and domain randomization, yielding a policy with higher zero-shot real-world success than real-robot teleo...

  15. Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum

    cs.RO 2026-05 unverdicted novelty 5.0

    A multi-agent large-model framework (Active Spatial Brain + Generalizable Action Cerebellum) enables spatial-aware humanoid whole-body manipulation without task-specific real-robot data.

  16. Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum

    cs.RO 2026-05 unverdicted novelty 5.0

    A multi-agent LLM framework for humanoid loco-manipulation that separates active spatial perception and task planning from generalizable action generation without task-specific real-robot data.

  17. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.