REVIEW 12 cited by
GraspSplats: Efficient Manipulation with 3D Feature Splatting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The ability for robots to perform efficient and zero-shot grasping of object parts is crucial for practical applications and is becoming prevalent with recent advances in Vision-Language Models (VLMs). To bridge the 2D-to-3D gap for representations to support such a capability, existing methods rely on neural fields (NeRFs) via differentiable rendering or point-based projection methods. However, we demonstrate that NeRFs are inappropriate for scene changes due to their implicitness and point-based methods are inaccurate for part localization without rendering-based optimization. To amend these issues, we propose GraspSplats. Using depth supervision and a novel reference feature computation method, GraspSplats generates high-quality scene representations in under 60 seconds. We further validate the advantages of Gaussian-based representation by showing that the explicit and optimized geometry in GraspSplats is sufficient to natively support (1) real-time grasp sampling and (2) dynamic and articulated object manipulation with point trackers. With extensive experiments on a Franka robot, we demonstrate that GraspSplats significantly outperforms existing methods under diverse task settings. In particular, GraspSplats outperforms NeRF-based methods like F3RM and LERF-TOGO, and 2D detection methods.
Forward citations
Cited by 12 Pith papers
-
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
Feature-space Gaussian Splat Feature Adapter (GS-Adapter) grounds camera-controlled video diffusion in 3D Gaussians, improving geometric consistency and controllability over SEVA and CameraCtrl without retraining geom...
-
ReMoSPLAT: Reactive Mobile Manipulation Control on a Gaussian Splat
ReMoSPLAT achieves reactive mobile-manipulation collision avoidance by querying distances from a Gaussian Splat reconstruction, matching a ground-truth-SDF controller in simulation.
-
Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
A robot policy trained on one real demonstration plus AI-generated 3D views succeeds from novel initial poses, including opposite-side starts, across six real manipulation tasks.
-
InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
InstaScene combines Gaussian-based instance decomposition with generative completion to produce complete, scene-aligned 3D object models from cluttered scenes.
-
GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
GeoProg3D combines a georeferenced hierarchical 3D language field, geographic vision APIs, and LLM-generated programs to answer natural-language queries about city-scale 3D scenes, and includes a new 952-query benchma...
-
Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation
RoboSplat edits 3D Gaussian scene reconstructions to synthesize diverse robot demonstrations from one expert trajectory, and behavior-cloned policies trained on this data generalize robustly across six disturbance typ...
-
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
SIREN uses semantic features in Gaussian Splatting maps to register and fuse maps from multiple robots without camera poses, images, or an initial transform.
-
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
Multi-GraspLLM uses a single multimodal LLM, trained on a new 140k-grasp, 1.1M-dialogue dataset, to generate semantic grasp poses for five different robotic hands.
-
Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance
The viewpoint-agnostic grasp pipeline using VLM and partial observation handling achieves 90% success (9/10 trials) in cluttered tabletop scenarios on a real quadruped robot, outperforming a view-dependent baseline at...
-
RoboPearls: Editable Video Simulation for Robot Manipulation
RoboPearls is a 3D Gaussian Splatting based framework that edits demonstration videos into varied photorealistic simulations, and training on them improves robot manipulation success rates on RLBench and COLOSSEUM.
-
Re$^3$Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
A reconstruction and neural-rendering pipeline converts real tabletop scenes into photorealistic robot simulations, and policies trained only on simulated data transfer zero-shot to the real robot with an average succ...
-
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
A review of navigation and manipulation simulators, datasets, and methods, framed around the sim-to-real gap.
Discussion (0). Continue with ORCID to comment.