REVIEW 5 cited by
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of actionable affordances. However, transferring such knowledge remains challenging due to: 1) a lack of large-scale datasets with precise affordance annotations, and 2) insufficient exploration of affordances in diverse manipulation contexts. To address these gaps, we introduce HOVA-500K, a large-scale, affordance-annotated dataset comprising 500,000 images across 1,726 object categories and 675 actions. We also release a standardized benchmarking suite for multi-modal affordance reasoning. Built upon HOVA-500K, we present GLOVER++, a global-to-local affordance training framework that effectively transfers actionable affordance knowledge from human demonstrations to downstream open-vocabulary reasoning tasks. GLOVER++ achieves state-of-the-art results on the HOVA-500K benchmark and demonstrates strong generalization across diverse downstream robotic manipulation tasks. By explicitly modeling actionable affordances, GLOVER++ facilitates robust transfer across scenes, modalities, and tasks. We hope that HOVA-500K and the GLOVER++ framework will serve as valuable resources for bridging the gap between human demonstrations and robotic manipulation capabilities.
Forward citations
Cited by 5 Pith papers
-
RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation
From a single egocentric RGB-D frame and a language instruction, RoboReact generates a human manipulation video, distills it into object-relative keyframe skills, refines them through a vision-language-model trial loo...
-
FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry
FrameVGGT replaces token-level KV retention with frame-level segments and prototypes to bound memory while preserving geometric coherence in streaming VGGT.
-
VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
From 204K egocentric human videos, the authors automatically extract visual, grasp, and trajectory affordances and train one vision-language model, VLAff, that predicts all three for robot manipulation.
-
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.
-
Omni-Perception: Omnidirectional Collision Avoidance for Legged Locomotion in Dynamic Environments
Omni-Perception is an end-to-end RL policy for legged robots that processes raw LiDAR point clouds with PD-RiskNet to achieve omnidirectional collision avoidance, validated in simulation and on a Unitree G1.
Discussion (0). Continue with ORCID to comment.