Pith. sign in

Paper Citation Record · LEDGER

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2607.09825.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09825 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:16:18.444295Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d21c3f39-ff4c-46c0-9377-1f9ed95c6293 · outbound

This paper cites an unresolved cited work.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:341ecaa409f0f16b68f7e9ef77258c1f82dfcf65b84d07881f35323dafffb70b

Observation 1a9e1690-711b-4d77-bd24-dc28788264d8 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Emerging properties in self-supervised vision transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:0f653c4fdecc908c2737bfe5e8f85a342d3157edf2eb81dfe0c4d8094a4f497d

Observation 7c6c0035-98c9-452b-8e11-ab9342987ca1 · outbound

This paper cites Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:a003145afaf1491ad3552720def8e0e61f51dabb23b6c3e9595d332af0010cbc

Observation 5769fcdf-36d2-4dd9-b8f2-5a501c0880a8 · outbound

This paper cites Dif- fusion policy: Visuomotor policy learning via action diffusion.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Dif- fusion policy: Visuomotor policy learning via action diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:33f4947aabc72c96c866142ce40c0a0d0af488053c895e30575dd252019a7b2b

Observation e6c20047-e6d0-4a8c-accd-46a412e1df54 · outbound

This paper cites Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Georg Heigold, and Thomas Kipf.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff, Georg Heigold, and Thomas Kipf

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:a15dfea925dc53ef589480368fa28d252911eedd114e1dc21497862c9104bc57

Observation cbe21540-d6f7-44d4-9164-b0532f9a1f18 · outbound

This paper cites SPOT: Self-training with patch-order permutation for object-centric learning with autoregressive transformers.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning SPOT: Self-training with patch-order permutation for object-centric learning with autoregressive transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:34e2a4db3bfc416d21a754a487dc700b282da658d4d9fc218c1ac6903078a194

Observation 6b0de0dc-0086-47d3-b7dc-de6be97c2b15 · outbound

This paper cites Elsayed, Aravindh Ma- hendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Gr- eff.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Elsayed, Aravindh Ma- hendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, and Klaus Gr- eff

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:dea8dbd60dcaabe9d554d61f4f17af0d7808a2dab27fb736e3395ea94ef870dd

Observation 4da1e232-2bc9-421a-bcae-8f28171c86d6 · outbound

This paper cites Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Jin Kim, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:7a6b25c4f884bd695b022d2876ac527d35dd569ce9abc0d185d09f5739a44d2a

Observation df1fab8f-408d-4515-9c0f-28ea657caf90 · outbound

This paper cites Object-centric learning with slot attention.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Object-centric learning with slot attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:6094df2e0ffdaa2e477486c416c6fb2848f56c3e00db169e394ea4907f737a55

Observation 61571298-dec9-4341-b172-d642a6e0da31 · outbound

This paper cites R3M: A Universal Visual Representation for Robot Manipulation.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning R3M: A Universal Visual Representation for Robot Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:804fd92792ab975991bed56f84f4c565e45ec3da669fd6a459e32b09884795df

Observation 1fcd72bf-0c41-490c-a563-164046556eb9 · outbound

This paper cites UniGaze: Towards universal gaze estimation via large- scale pre-training.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning UniGaze: Towards universal gaze estimation via large- scale pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:065958aeff6ab48d0c8fb91612ac4eff6f0b2baa21c21fbdcb313e4ac8b753fe

Observation d605f260-4107-4637-bc09-db7c4425f4b9 · outbound

This paper cites Real-world robot learning with masked visual pre-training.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Real-world robot learning with masked visual pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:044b5adc43d90e54f2e7c7cf1f02976dc4cea16c8b19f185deede5fda1955189

Observation 5c483134-84b9-4cc7-b440-9996ad2feb10 · outbound

This paper cites Bridging the gap to real-world object-centric learning.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Bridging the gap to real-world object-centric learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:7130ac413d679e4133bc85824a72ef9848913abd5e5b5e184cfdd83909690b06

Observation e6c6b9df-c8a5-45c5-8dbe-e6cfd0759729 · outbound

This paper cites Vision trans- formers need more than registers.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Vision trans- formers need more than registers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:39fef4c01b3914eac98c1cfd5e57ee6ccf84f4e1a94b8323f99c4664b176caf4

Observation 8cc58953-8bec-40d8-b3e8-9ad5779c8318 · outbound

This paper cites ManiSkill3: GPU parallelized robotics simulation and rendering for generalizable embodied AI.Robotics: Science and Systems, 2025.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning ManiSkill3: GPU parallelized robotics simulation and rendering for generalizable embodied AI.Robotics: Science and Systems, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:239fc13a216616f3fc3803c5aabda7edf5c28af0f0b285de46e9919b24fdbc18

Observation 914fc1cb-d5ec-48ba-bc94-af4a3d423731 · outbound

This paper cites SlotDiffusion: Object-centric generative modeling with diffusion models.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning SlotDiffusion: Object-centric generative modeling with diffusion models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:7b7286438aba6fab45c04ad5a2cfc11a654b1898f110abf83160a07d8724ae19

Observation 4ec39184-f587-406e-ad99-e5921c56af02 · outbound

This paper cites Utonia: Toward one encoder for all point clouds.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Utonia: Toward one encoder for all point clouds

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:23eb0fb07df890dafe8a3a5a4e16ce3b2650addb26363ca34d4afee1c4e9c9e9

Observation 2bc5bf99-b5d1-4e6f-8beb-6a4449b77384 · outbound

This paper cites Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn.

More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T15:16:18.444295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:16:18.444295Z digest=sha256:dbb1e7308a184ffcf2b7f3f56317f5ceeab378a2f505ec89ffcee7c2a130ca73

Pith citing papers

No inbound Pith citation observations are available.