Pith. sign in

Paper Citation Record · LEDGER

Theia: Distilling Diverse Vision Foundation Models for Robot Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2407.20179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.20179 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:11:50.078534Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.580056Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5cec6f46-9634-477d-81a8-74278cc13a85 · inbound

Is an object-centric representation beneficial for robotic manipulation ? cites this paper.

Is an object-centric representation beneficial for robotic manipulation ? Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:50.078534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:50.078534Z digest=sha256:ff471f6f9e0bdffeb180798030e4a79aa91167663d30d2b81719835fa08fcf6c

Observation a9cb98dc-1036-4994-8d70-8f919ae0bd3c · inbound

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge cites this paper.

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:42:41.660719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:42:41.363422Z digest=sha256:834d1fbba1bab5a204c529952555d26f6c4abeb338c7620c02d84d9a972a50d8

Observation 26caf29b-1180-4d8a-aa97-7c020eed7bff · inbound

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models cites this paper.

Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T22:22:58.808101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:22:58.808101Z digest=sha256:05a1a8bbebe243cd5ff6e139b89d505d8c737e569afbd6105ef1dc8fac34fc4f

Observation 3c575d66-de10-419b-9a9a-87d97cf82a59 · inbound

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation cites this paper.

ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:19.552910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:19.552910Z digest=sha256:ba1cea3511170a5a5b66cffaea4ae5c83dd1fc37f87143baaf9ca8330691756c

Observation 890b4ed2-661c-42f5-88be-4dcaf068aabd · inbound

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models cites this paper.

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:24:04.765662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:23:11.261860Z digest=sha256:ac276067a24724616fc6637bfa02507912c85641a0383561f7f5f667884f011e

Observation 319d9c38-1d2e-40b0-b065-2741b47c2424 · inbound

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models cites this paper.

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:23:23.731777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T20:21:47.430865Z digest=sha256:988dc21e276e7c0d695cc82083da2bdc1d28627d291fef7135aef348bbd9578a

Observation 72cff02c-9387-4abf-9791-edfbdb9ef77d · inbound

Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation cites this paper.

Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T07:04:44.923245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:04:44.923245Z digest=sha256:008c8baf50e4cd10bcb4359393c75533089895f901392d55c123f82ca7c88492

Observation 4923df3b-62c9-4fc4-8bfb-0c93c0e8c268 · inbound

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors cites this paper.

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:39:55.177952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:37:02.168898Z digest=sha256:a97eabb3a17d61dedc9f899c7760150601878caac97ed5a85341f16fd1b21000

Observation 181c552a-0578-4513-8992-5c66a83521f3 · inbound

Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation cites this paper.

Efficient Image Annotation via Semi-Supervised Object Segmentation with Label Propagation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:07.173665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:30:21.475036Z digest=sha256:b1d4cc90ad6c7bf1c4ab5d9a5921c2aee776a71c7162b75b12e19a50b49b64fd

Observation 5129e6c7-3db3-46b5-a89c-b4c1535e013f · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.613444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:58fba47b31075800ae7fb754a05afbefe69815cfbcf0ff214eca060cd8993ab2

Observation e36865a4-f893-4107-be7b-acd44ab738e4 · inbound

LACE: Latent Visual Representation for Cross-Embodiment Learning cites this paper.

LACE: Latent Visual Representation for Cross-Embodiment Learning Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:42:48.316408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T21:38:33.937153Z digest=sha256:3dc5109878f0d6acc38015e918b3d071f59b0163b73305f1624fe29894ecd111

Observation c1c81271-8e74-4de3-a5e8-286af5926ef0 · inbound

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization cites this paper.

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.241421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:47:33.591670Z digest=sha256:0aa44a127959c9f7b93dca1442019e2a9d2beb66ceab30b5b31abbad6fa1f630

Observation f9a1dc47-02e1-47ac-a0b8-0e32d2a9388e · inbound

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation cites this paper.

Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:29.281950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:28:41.952330Z digest=sha256:9c988abf0206234676be1828754930577f89fafa68ddd3e21fd0143df5946abc

Observation 0b1cabc2-4a61-4189-a742-e4820622a1cd · inbound

Action-Effect Memory Pretraining for Robot Manipulation cites this paper.

Action-Effect Memory Pretraining for Robot Manipulation Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:48:04.840744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:22:17.435427Z digest=sha256:00e60c7f1c9602ffe5e5b77d195f11b7b1d7b5ce33df1b70da317122d2a84494

Observation ba4c843e-e302-4175-beb8-53ae8518e5cb · inbound

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning cites this paper.

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.582106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:25:13.494219Z digest=sha256:e8074c98b0d442df178ee6afad4d52b8abf70edfecfd4109517926f4a7f0eeea

Observation f41f797f-ca3e-441a-9b29-09198ee0bab1 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:35.210298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:35.210298Z digest=sha256:e3c3c86b320c09e0580e34e7cf5e03d19a8e0819baf5aac81887192edc881097