Pith. sign in

REVIEW 5 cited by

Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.20179 v2 pith:BVQLIHY5 submitted 2024-07-29 cs.RO cs.AIcs.CVcs.LG

classification cs.ROcs.AIcs.CVcs.LG
keywords learningrobotmodelstheiavisualvisiondiversefoundation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-based robot policy learning, which maps visual inputs to actions, necessitates a holistic understanding of diverse visual tasks beyond single-task needs like classification or segmentation. Inspired by this, we introduce Theia, a vision foundation model for robot learning that distills multiple off-the-shelf vision foundation models trained on varied vision tasks. Theia's rich visual representations encode diverse visual knowledge, enhancing downstream robot learning. Extensive experiments demonstrate that Theia outperforms its teacher models and prior robot learning models using less training data and smaller model sizes. Additionally, we quantify the quality of pre-trained visual representations and hypothesize that higher entropy in feature norm distributions leads to improved robot learning performance. Code, models, and demo are available at https://theia.theaiinstitute.com.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  2. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  3. Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Distilling three pattern-specific alignments from a large human-centric teacher yields a 5M-parameter student that approaches teacher-level generalization on many downstream tasks.

  4. ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A co-training framework that maps retargeted human hand trajectories to robot demonstrations with dynamic time warping and MixUp interpolation improves robot manipulation success rates and smoothness across four embodiments.

  5. Is an object-centric representation beneficial for robotic manipulation ?

    cs.AI 2025-06 reject novelty 4.0 of 10

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...

Pith tools