REVIEW 7 cited by
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
3D visual grounding is crucial for robots, requiring integration of natural language and 3D scene understanding. Traditional methods depending on supervised learning with 3D point clouds are limited by scarce datasets. Recently zero-shot methods leveraging LLMs have been proposed to address the data issue. While effective, these methods only use object-centric information, limiting their ability to handle complex queries. In this work, we present VLM-Grounder, a novel framework using vision-language models (VLMs) for zero-shot 3D visual grounding based solely on 2D images. VLM-Grounder dynamically stitches image sequences, employs a grounding and feedback scheme to find the target object, and uses a multi-view ensemble projection to accurately estimate 3D bounding boxes. Experiments on ScanRefer and Nr3D datasets show VLM-Grounder outperforms previous zero-shot methods, achieving 51.6% Acc@0.25 on ScanRefer and 48.0% Acc on Nr3D, without relying on 3D geometry or object priors. Codes are available at https://github.com/OpenRobotLab/VLM-Grounder .
Forward citations
Cited by 7 Pith papers
-
G$^2$TAM: Geometry Grounded Track Anything Model
Spatially aligned geometric features serve as implicit memory so one model reconstructs scenes and produces promptable, cross-view consistent instance masks from unordered RGB only.
-
Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching
A tri-sensor camera-LiDAR-radar 3D visual grounding benchmark and a language-routed fusion model, TSFormer, that improves grounding accuracy over prior baselines.
-
TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding
A training-free pipeline that disambiguates text queries and infers viewpoints improves zero-shot 3D visual grounding, reaching 64.06% Acc@0.5 on ScanRefer.
-
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
Pretraining a vision-language model to output discrete 3D pose tokens on large non-robotic data, before training a robot action head, improves downstream manipulation success and data efficiency.
-
T-Rex: Task-Adaptive Spatial Representation Extraction for Robotic Manipulation with Vision-Language Models
A zero-training framework that adaptively selects spatial representation extractors per object and per task stage improves real-world robot manipulation success and efficiency over fixed-representation baselines.
-
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
Proposal-guided multi-view projection with iterative VLM selection achieves 55.6% and 53.2% Acc@0.25 on ScanRefer and Nr3D, a new zero-shot 3D visual grounding state of the art.
-
Zero-Shot 3D Visual Grounding from Vision-Language Models
SeeGround localizes objects in 3D scenes from natural language without 3D-specific training, using query-aligned rendered views and spatially enriched text fed to a 2D vision-language model.
Discussion (0). Continue with ORCID to comment.