Pith. sign in

REVIEW 7 cited by

Learning Generalizable Feature Fields for Mobile Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07563 v2 pith:JKOV33SJ submitted 2024-03-12 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords manipulationfeaturegeffgeneralizablenavigationcapturingfieldsmobile
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while understanding fine-grained semantics, whereas the former involves capturing the complexity inherent at an expansive physical scale. In this work, we present GeFF (Generalizable Feature Fields), a scene-level generalizable neural feature field that acts as a unified representation for both navigation and manipulation that performs in real-time. To do so, we treat generative novel view synthesis as a pre-training task, and then align the resulting rich scene priors with natural language via CLIP feature distillation. We demonstrate the effectiveness of this approach by deploying GeFF on a quadrupedal robot equipped with a manipulator. We quantitatively evaluate GeFF's ability for open-vocabulary object-/part-level manipulation and show that GeFF outperforms point-based baselines in runtime and storage-accuracy trade-offs, with qualitative examples of semantics-aware navigation and articulated object manipulation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QuASH: Using Natural-Language Heuristics to Query Visual-Language Robotic Maps

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Querying VLM robot maps with an SVM trained on LLM-generated synonym/antonym embeddings outperforms cosine-threshold and single-antonym baselines on images and OpenSeg maps, but not consistently on LSeg maps.

  2. Learning Multi-Stage Pick-and-Place with a Legged Mobile Manipulator

    cs.RO 2025-09 accept novelty 6.0 of 10

    A simulation-trained teacher-student policy with progressive policy expansion achieves 78.3% real-world success on a long-horizon mobile pick-and-place task.

  3. SIREN: Semantic, Initialization-Free Registration of Multi-Robot Gaussian Splatting Maps

    cs.RO 2025-02 conditional novelty 6.0 of 10

    SIREN uses semantic features in Gaussian Splatting maps to register and fuse maps from multiple robots without camera poses, images, or an initial transform.

  4. Representative Volume Element: Existence and Extent in Cracked Heterogeneous Medium

    cs.CE 2025-08 unverdicted novelty 5.0 of 10

    Modified periodic boundary conditions that add strain periodicity to displacement periodicity are claimed to reduce mesh and size sensitivity in cracked-composite RVE simulations, tested on 1,200 samples.

  5. OWMM-Agent: Open World Mobile Manipulation With Multi-modal Agentic Data Synthesis

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A vision-language model fine-tuned on 572K synthetic simulation examples improves open-world mobile manipulation action decisions and object grounding over GPT-4o, with 21.9% full-task success in simulation and 90% ac...

  6. MoTo: A Zero-shot Plug-in Interaction-aware Navigation for General Mobile Manipulation

    cs.RO 2025-09 conditional novelty 4.0 of 10

    MoTo turns existing fixed-base manipulation models into mobile manipulators by using VLM-picked contact keypoints and trajectory optimization to find docking points, with no training of MoTo itself.

  7. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06

Pith tools