REVIEW 7 cited by
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image features. This work bridges this 2D-to-3D gap for robotic manipulation by leveraging distilled feature fields to combine accurate 3D geometry with rich semantics from 2D foundation models. We present a few-shot learning method for 6-DOF grasping and placing that harnesses these strong spatial and semantic priors to achieve in-the-wild generalization to unseen objects. Using features distilled from a vision-language model, CLIP, we present a way to designate novel objects for manipulation via free-text natural language, and demonstrate its ability to generalize to unseen expressions and novel categories of objects.
Forward citations
Cited by 7 Pith papers
-
MeshFM: 2D Features Are All You Need for 3D Shape Understanding
A feedforward network trained only on 2D foundation-model features predicts rotation-robust, general-purpose 3D mesh features that work zero-shot for segmentation, correspondence, and deformation.
-
One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation
Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.
-
DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF
A NeRF-based method jointly performs unsupervised semantic clustering and CLIP-guided relevancy to discover and segment query-relevant sub-concepts in 3D scenes, with a new Replica benchmark.
-
SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps
SIREN uses semantic features in Gaussian Splatting maps to register and fuse maps from multiple robots without camera poses, images, or an initial transform.
-
Beyond Point-Attached Semantics: Object-Centric Semantic Fields for Generalizable Manipulation
An object-conditioned continuous semantic field queried at explicit 3D locations yields more stable part cues and higher manipulation success than point-attached 2D/3D features.
-
Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels
A supervised 3D U-Net predicts per-voxel material fields from CLIP feature grids, enabling fast MPM-based animation, but the reported evidence depends on pseudo-labels and a VLM judge from the same model family as the...
-
Gaussian Process-Based Active Exploration Strategies in Vision and Touch
A robot arm uses Gaussian Process Distance Fields to fuse RGBD vision and tactile contacts, actively choosing next views and touch points to reduce shape uncertainty, while material classification remains near chance.
Discussion (0). Continue with ORCID to comment.