Pith. sign in

REVIEW 5 cited by

A3VLM: Actionable Articulation-Aware Vision Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07549 v2 pith:FP7IYUKK submitted 2024-06-11 cs.RO

classification cs.RO
keywords a3vlmlanguageroboticsvisionvlmsactionactionableactions
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Vision Language Models (VLMs) have received significant attention in recent years in the robotics community. VLMs are shown to be able to perform complex visual reasoning and scene understanding tasks, which makes them regarded as a potential universal solution for general robotics problems such as manipulation and navigation. However, previous VLMs for robotics such as RT-1, RT-2, and ManipLLM have focused on directly learning robot-centric actions. Such approaches require collecting a significant amount of robot interaction data, which is extremely costly in the real world. Thus, we propose A3VLM, an object-centric, actionable, articulation-aware vision language model. A3VLM focuses on the articulation structure and action affordances of objects. Its representation is robot-agnostic and can be translated into robot actions using simple action primitives. Extensive experiments in both simulation benchmarks and real-world settings demonstrate the effectiveness and stability of A3VLM. We release our code and other materials at https://github.com/changhaonan/A3VLM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ShowUI: One Vision-Language-Action Model for GUI Visual Agent

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ShowUI is a lightweight 2B vision-language-action model that uses UI-guided token pruning and a curated 256K dataset to reach 75.1% zero-shot screenshot grounding accuracy.

  2. CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

    cs.RO 2025-05 conditional novelty 5.0 of 10

    CrayonRobo trains a vision-language-action model to read colored 2D prompt overlays (contact point, end-effector axes, movement direction) and output SE(3) contact poses, enabling step-by-step and long-horizon robotic...

  3. Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A small dataset of sensor images plus positive and negative answer examples markedly improves VLM performance on thermal, depth, and X-ray understanding without retraining the model architecture.

  4. CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

    cs.RO 2024-12 conditional novelty 5.0 of 10

    Adding a visual and textual chain-of-affordance reasoning step to a vision-language-action model improves robot manipulation success rates and generalization in the paper's evaluations.

  5. GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping

    cs.RO 2024-11 conditional novelty 5.0 of 10

    GLOVER predicts open-vocabulary graspable regions on objects from one RGB image and estimates grasp poses from those regions, reporting large speedups and higher success rates than prior systems.

Pith tools