Pith. sign in

REVIEW 7 cited by

VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.13860 v1 pith:RTF5O4LC submitted 2024-10-17 cs.CV cs.RO

classification cs.CVcs.RO
keywords vlm-groundergroundingmethodszero-shotvisualdatasetsnr3dobject
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D visual grounding is crucial for robots, requiring integration of natural language and 3D scene understanding. Traditional methods depending on supervised learning with 3D point clouds are limited by scarce datasets. Recently zero-shot methods leveraging LLMs have been proposed to address the data issue. While effective, these methods only use object-centric information, limiting their ability to handle complex queries. In this work, we present VLM-Grounder, a novel framework using vision-language models (VLMs) for zero-shot 3D visual grounding based solely on 2D images. VLM-Grounder dynamically stitches image sequences, employs a grounding and feedback scheme to find the target object, and uses a multi-view ensemble projection to accurately estimate 3D bounding boxes. Experiments on ScanRefer and Nr3D datasets show VLM-Grounder outperforms previous zero-shot methods, achieving 51.6% Acc@0.25 on ScanRefer and 48.0% Acc on Nr3D, without relying on 3D geometry or object priors. Codes are available at https://github.com/OpenRobotLab/VLM-Grounder .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. G$^2$TAM: Geometry Grounded Track Anything Model

    cs.CV 2026-07 accept novelty 6.5 of 10

    Spatially aligned geometric features serve as implicit memory so one model reconstructs scenes and produces promptable, cross-view consistent instance masks from unordered RGB only.

  2. Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A tri-sensor camera-LiDAR-radar 3D visual grounding benchmark and a language-routed fusion model, TSFormer, that improves grounding accuracy over prior baselines.

  3. TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A training-free pipeline that disambiguates text queries and infers viewpoints improves zero-shot 3D visual grounding, reaching 64.06% Acc@0.5 on ScanRefer.

  4. PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    Pretraining a vision-language model to output discrete 3D pose tokens on large non-robotic data, before training a robot action head, improves downstream manipulation success and data efficiency.

  5. T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A zero-training framework that adaptively selects spatial representation extractors per object and per task stage improves real-world robot manipulation success and efficiency over fixed-representation baselines.

  6. SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.

  7. Zero-Shot 3D Visual Grounding from Vision-Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    SeeGround localizes objects in 3D scenes from natural language without 3D-specific training, using query-aligned rendered views and spatially enriched text fed to a 2D vision-language model.

Pith tools