Pith. sign in

REVIEW 12 cited by

Bridging the Gap to Real-World Object-Centric Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.14860 v2 pith:ZEHNRAT5 submitted 2022-09-29 cs.CV cs.LG

classification cs.CVcs.LG
keywords object-centriclearningunsuperviseddatadinosaurmodelsreal-worldsimulated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Humans naturally decompose their environment into entities at the appropriate level of abstraction to act in the world. Allowing machine learning algorithms to derive this decomposition in an unsupervised way has become an important line of research. However, current methods are restricted to simulated data or require additional information in the form of motion or depth in order to successfully discover objects. In this work, we overcome this limitation by showing that reconstructing features from models trained in a self-supervised manner is a sufficient training signal for object-centric representations to arise in a fully unsupervised way. Our approach, DINOSAUR, significantly out-performs existing image-based object-centric learning models on simulated data and is the first unsupervised object-centric model that scales to real-world datasets such as COCO and PASCAL VOC. DINOSAUR is conceptually simple and shows competitive performance compared to more involved pipelines from the computer vision literature.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CObL: Toward Zero-Shot Ordinal Layering without User Prompting

    cs.CV 2025-08 conditional novelty 8.0 of 10

    CObL uses multiple linked Stable Diffusion models to decompose an image into occlusion-ordered object layers, guided at inference so the layers reproduce the input.

  2. Object-level Self-Distillation for Vision Pretraining

    cs.CV 2025-06 conditional novelty 7.0 of 10

    ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.

  3. CoLa: Chinese Character Decomposition with Compositional Latent Components

    cs.CV 2025-06 conditional novelty 7.0 of 10

    CoLa learns compositional latent components of Chinese characters via slot attention and matches them to printed templates, achieving strong zero-shot Chinese character recognition without human-defined decomposition.

  4. TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    TSA learns per-slot, per-frame activation scores that gate state updates and decoder attention, preserving object identity through occlusion in unsupervised video object-centric learning.

  5. ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A cycle-consistent GAN that translates between images and object lists matches state-of-the-art detection on synthetic scenes and detects low-contrast cells where slot-attention models fail.

  6. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  7. Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SlotSAR fuses wavelet scattering features with a SAR foundation model's semantic features to make slot attention separate targets from clutter in SAR images, improving segmentation metrics on ATRNet-STAR.

  8. Dyn-O: Building Structured World Models with Object-Centric Representations

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Dyn-O learns object-centric world models directly from pixels in complex Procgen games, using SAM2-guided slot attention and Mamba state-space dynamics, and reports better rollout prediction than DreamerV3.

  9. Identifiable Object Representations under Spatial Ambiguities

    cs.LG 2025-06 reject novelty 6.0 of 10

    VISA learns view-invariant object representations by aggregating probabilistic slots across multiple unlabeled viewpoints, with an identifiability analysis up to affine and permutation equivalence.

  10. Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A ViT layer-decomposition based self-disentanglement and re-composition method improves cross-domain few-shot segmentation, beating prior state-of-the-art by 1.92 (1-shot) and 1.88 (5-shot) average mIoU.

  11. Compositional Scene Understanding through Inverse Generative Modeling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.

  12. Is an object-centric representation beneficial for robotic manipulation ?

    cs.AI 2025-06 reject novelty 4.0 of 10

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...

Pith tools