Pith. sign in

REVIEW 33 cited by

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.10721 v1 pith:NAQCE2ZP submitted 2024-06-15 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords robopointvlmslanguagerobotaffordancedatadownstreammodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

From rearranging objects on a table to putting groceries into shelves, robots must plan precise action points to perform tasks accurately and reliably. In spite of the recent adoption of vision language models (VLMs) to control robot behavior, VLMs struggle to precisely articulate robot actions using language. We introduce an automatic synthetic data generation pipeline that instruction-tunes VLMs to robotic domains and needs. Using the pipeline, we train RoboPoint, a VLM that predicts image keypoint affordances given language instructions. Compared to alternative approaches, our method requires no real-world data collection or human demonstration, making it much more scalable to diverse environments and viewpoints. In addition, RoboPoint is a general model that enables several downstream applications such as robot navigation, manipulation, and augmented reality (AR) assistance. Our experiments demonstrate that RoboPoint outperforms state-of-the-art VLMs (GPT-4o) and visual prompting techniques (PIVOT) by 21.8% in the accuracy of predicting spatial affordance and by 30.5% in the success rate of downstream tasks. Project website: https://robo-point.github.io.

Discussion (0). Sign in to comment.

Forward citations

Cited by 33 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

    cs.CV 2026-01 conditional novelty 7.0 of 10

    Using a simple action-token adapter, nine VLMs are compared as robot policy backbones, showing general VLM ability transfers poorly to control and the vision encoder is the key bottleneck.

  2. GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

    cs.CV 2025-10 conditional novelty 7.0 of 10

    A frozen, large-scale pretrained diffusion policy converts sparse 3D waypoints from a VLM into dense robot actions, enabling zero-shot reuse of the action expert on new tasks and environments.

  3. Weakly-Supervised Learning of Dense Functional Correspondences

    cs.CV 2025-09 conditional novelty 7.0 of 10

    A weakly-supervised pipeline that distills VLM functional part knowledge and multi-view spatial structure into a model for dense cross-category functional correspondence, outperforming baselines on new synthetic and r...

  4. Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    BadSem shows that semantic mismatches between images and text can serve as stealthy backdoor triggers for VLMs, achieving near-perfect attack success with low poisoning rates.

  5. PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

    cs.RO 2025-05 conditional novelty 7.0 of 10

    PartInstruct is a new large-scale simulated benchmark with part-level language instructions and training demonstrations; current robot policies achieve at most 31.72% average success on it.

  6. BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Adding stage-wise temporal and spatial memory to a heatmap-prediction 3D VLA policy yields strong results on memory-dependent manipulation benchmarks while keeping data efficiency.

  7. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  8. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  9. RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A unified transformer model generates language-and-image planning sequences for embodied tasks, and shows real-robot manipulation without large-scale action pretraining.

  10. Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Projecting 3D gripper keypoints onto camera pixels and classifying those pixels yields millimeter-precise, multi-modal closed-loop manipulation faster than diffusion policies.

  11. EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos

    cs.RO 2026-06 accept novelty 6.0 of 10

    A full-stack system curates 9.6K hours of egocentric video into language-aligned action priors that, after robot post-training and DAgger, enable free-form steerable dexterous manipulation at ~75% success across 40+ tasks.

  12. RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    Training a 4B vision-language model to emit identity-tracked, visually grounded reasoning anchors improves embodied spatial, multi-view, and pointing task performance over 7B baselines.

  13. Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    cs.CV 2026-03 reject novelty 6.0 of 10

    SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...

  14. Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 6K-question embodied-reasoning benchmark plus a flow-matching action tokenizer let one 3B vision-language model reason and manipulate better than continuous- or discrete-action VLA baselines.

  15. Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

    cs.RO 2025-12 conditional novelty 6.0 of 10

    A 3D-aware VLM, RoboTracer, generates metric-grounded spatial traces for robot manipulation using scale supervision and metric-sensitive reinforcement rewards.

  16. SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A two-phase interactive RL framework (DIRL) lets a 3B VLM learn to coordinate multiple vision and robot tools, reaching top benchmark scores and 86% real-robot pick-and-place success.

  17. SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Dense scene-graph-grounded rewards let a 7B multimodal LLM trained on 7K synthetic questions beat SFT and sparse-RL baselines and outscore GPT-4o on average across 12 spatial/real-world benchmarks.

  18. FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A VLM trained on auto-generated failure trajectories with executable correction actions helps VLA models detect and fix manipulation errors, lifting success rates by up to 22.6 percentage points.

  19. O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A one-shot training regime with DINOv2-enriched point clouds and joint cross-attention predicts 3D object-to-object affordance maps that guide optimization-based robotic manipulation.

  20. Robix: A Unified Model for Robot Interaction, Reasoning and Planning

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A three-stage-trained VLM unifies robot planning and dialogue, and beats commercial VLMs on the authors' interactive-task benchmarks.

  21. AimBot: A Simple Auxiliary Visual Cue to Enhance Spatial Awareness of Visuomotor Policies

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Overlaying end-effector-derived shooting lines and reticles on RGB images consistently raises success rates of visuomotor policies, especially on long-horizon manipulation tasks.

  22. PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    PASG automatically extracts object keypoints and axes and couples them through a fine-tuned vision-language model to task semantics, claiming manipulation performance comparable to manual annotations.

  23. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.

  24. RoboBrain 2.0 Technical Report

    cs.RO 2025-07 conditional novelty 6.0 of 10

    RoboBrain 2.0, a 7B/32B embodied vision-language model built on Qwen2.5-VL, reports state-of-the-art or near-top scores on several spatial and temporal reasoning benchmarks for robotics.

  25. Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Gondola generates multi-view segmentation-mask-grounded next-step plans for robotic manipulation and reports improved generalization on the GemBench benchmark over a prior LLM-based planner.

  26. GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    GenManip is a benchmark and simulation platform with LLM-generated scene graphs for testing how robot policies generalize to new instructions, layouts, and objects.

  27. UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    UAD distills affordance knowledge from vision-language models and DINOv2 features into a lightweight task-conditioned model that predicts pixel-level manipulation regions and improves few-shot imitation learning gener...

  28. VideoMolmo: Spatio-Temporal Grounding Meets Pointing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video language model that conditions each frame on earlier frames via a temporal attention module, predicts text-requested object points, and uses SAM2-based bidirectional mask fusion to outperform prior models on v...

  29. Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    Embodied-R1.5 is an 8B EFM achieving SOTA on 16 of 24 embodied VLM benchmarks, fine-tunable to outperform leading VLAs, with claimed zero-shot real-robot generalization.

  30. Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A three-layer hierarchical system using a VLM planner and VLM skill monitor with imitation-learned skills and an RL tracking policy achieved 73% success on a real humanoid pick-and-place task.

  31. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A vision-language model fine-tuned on 572K synthetic simulation examples improves open-world mobile manipulation action decisions and object grounding over GPT-4o, with 21.9% full-task success in simulation and 90% ac...

  32. On the Dual-Use Dilemma in Physical Reasoning and Force

    cs.RO 2025-05 conditional novelty 5.0 of 10

    Adding Asimov-style safety prompts to vision-language models lowers both harmful and helpful force generation for contact-rich robotic tasks.

  33. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06

Pith tools