Pith. sign in

REVIEW 3 cited by

Language-Driven 6-DoF Grasp Detection Using Negative Prompt Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13842 v2 pith:P7DU3WAI submitted 2024-07-18 cs.RO cs.CV

classification cs.ROcs.CV
keywords graspdetectionlanguagelanguage-drivennegativepromptapproachcluttered
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

6-DoF grasp detection has been a fundamental and challenging problem in robotic vision. While previous works have focused on ensuring grasp stability, they often do not consider human intention conveyed through natural language, hindering effective collaboration between robots and users in complex 3D environments. In this paper, we present a new approach for language-driven 6-DoF grasp detection in cluttered point clouds. We first introduce Grasp-Anything-6D, a large-scale dataset for the language-driven 6-DoF grasp detection task with 1M point cloud scenes and more than 200M language-associated 3D grasp poses. We further introduce a novel diffusion model that incorporates a new negative prompt guidance learning strategy. The proposed negative prompt strategy directs the detection process toward the desired object while steering away from unwanted ones given the language input. Our method enables an end-to-end framework where humans can command the robot to grasp desired objects in a cluttered scene using natural language. Intensive experimental results show the effectiveness of our method in both benchmarking experiments and real-world scenarios, surpassing other baselines. In addition, we demonstrate the practicality of our approach in real-world robotic applications. Our project is available at https://airvlab.github.io/grasp-anything.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Goal State Generation for Robotic Manipulation Based on Linguistically Guided Hybrid Gaussian Diffusion

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A language-conditioned hybrid Gaussian diffusion network generates mug-hanging poses in simulation, then uses a gravity-based overlap removal step to produce collision-free target states.

  2. Grasp Diffusion Network: Learning Grasp Generators from Partial Point Clouds with Diffusion Models in SO(3)xR3

    cs.RO 2024-12 conditional novelty 4.0 of 10

    A conditional diffusion model over SE(3) poses generates successful grasps from partial point clouds, using a collision-avoidance guidance step and fast DDIM sampling.

  3. Data Pyramid for Embodied Manipulation: A Survey

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools