Pith. sign in

REVIEW 7 cited by

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.07931 v2 pith:ZJQTJ6DG submitted 2023-07-27 cs.CV cs.AIcs.CLcs.LGcs.RO

classification cs.CVcs.AIcs.CLcs.LGcs.RO
keywords distilledmanipulationobjectsfeaturefeaturesfew-shotfieldsgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image features. This work bridges this 2D-to-3D gap for robotic manipulation by leveraging distilled feature fields to combine accurate 3D geometry with rich semantics from 2D foundation models. We present a few-shot learning method for 6-DOF grasping and placing that harnesses these strong spatial and semantic priors to achieve in-the-wild generalization to unseen objects. Using features distilled from a vision-language model, CLIP, we present a way to designate novel objects for manipulation via free-text natural language, and demonstrate its ability to generalize to unseen expressions and novel categories of objects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MeshFM: 2D Features Are All You Need for 3D Shape Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A feedforward network trained only on 2D foundation-model features predicts rotation-robust, general-purpose 3D mesh features that work zero-shot for segmentation, correspondence, and deformation.

  2. One View, Many Worlds: Single-Image to 3D Object Meets Generative Domain Randomization for One-Shot 6D Pose Estimation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Given one RGB-D photo of an unseen object, an AI-generated 3D mesh, aligned jointly in metric scale and pose, yields state-of-the-art one-shot 6D pose estimation on YCBInEOAT, TOYL, and LM-O.

  3. DiSCO-3D : Discovering and segmenting Sub-Concepts from Open-vocabulary queries in NeRF

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A NeRF-based method jointly performs unsupervised semantic clustering and CLIP-guided relevancy to discover and segment query-relevant sub-concepts in 3D scenes, with a new Replica benchmark.

  4. SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps

    cs.RO 2025-02 conditional novelty 6.0 of 10

    SIREN uses semantic features in Gaussian Splatting maps to register and fuse maps from multiple robots without camera poses, images, or an initial transform.

  5. Beyond Point-Attached Semantics: Object-Centric Semantic Fields for Generalizable Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    An object-conditioned continuous semantic field queried at explicit 3D locations yields more stable part cues and higher manipulation success than point-attached 2D/3D features.

  6. Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels

    cs.CV 2025-08 reject novelty 5.0 of 10

    A supervised 3D U-Net predicts per-voxel material fields from CLIP feature grids, enabling fast MPM-based animation, but the reported evidence depends on pseudo-labels and a VLM judge from the same model family as the...

  7. Gaussian Process-Based Active Exploration Strategies in Vision and Touch

    cs.RO 2025-07 conditional novelty 4.0 of 10

    A robot arm uses Gaussian Process Distance Fields to fuse RGBD vision and tactile contacts, actively choosing next views and touch points to reduce shape uncertainty, while material classification remains near chance.

Pith tools