REVIEW 12 cited by
Bridging the Gap to Real-World Object-Centric Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Humans naturally decompose their environment into entities at the appropriate level of abstraction to act in the world. Allowing machine learning algorithms to derive this decomposition in an unsupervised way has become an important line of research. However, current methods are restricted to simulated data or require additional information in the form of motion or depth in order to successfully discover objects. In this work, we overcome this limitation by showing that reconstructing features from models trained in a self-supervised manner is a sufficient training signal for object-centric representations to arise in a fully unsupervised way. Our approach, DINOSAUR, significantly out-performs existing image-based object-centric learning models on simulated data and is the first unsupervised object-centric model that scales to real-world datasets such as COCO and PASCAL VOC. DINOSAUR is conceptually simple and shows competitive performance compared to more involved pipelines from the computer vision literature.
Forward citations
Cited by 12 Pith papers
-
CObL: Toward Zero-Shot Ordinal Layering without User Prompting
CObL uses multiple linked Stable Diffusion models to decompose an image into occlusion-ordered object layers, guided at inference so the layers reproduce the input.
-
Object-level Self-Distillation for Vision Pretraining
ODIS replaces image-level self-distillation with object-level distillation using segmentation-guided cropping and masked attention, improving image- and patch-level benchmarks over iBOT.
-
CoLa: Chinese Character Decomposition with Compositional Latent Components
CoLa learns compositional latent components of Chinese characters via slot attention and matches them to printed templates, achieving strong zero-shot Chinese character recognition without human-defined decomposition.
-
TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation
TSA learns per-slot, per-frame activation scores that gate state updates and decoder attention, preserving object identity through occlusion in unsupervised video object-centric learning.
-
ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks
A cycle-consistent GAN that translates between images and object lists matches state-of-the-art detection on synthetic scenes and detects low-contrast cells where slot-attention models fail.
-
Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation
Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.
-
Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
SlotSAR fuses wavelet scattering features with a SAR foundation model's semantic features to make slot attention separate targets from clutter in SAR images, improving segmentation metrics on ATRNet-STAR.
-
Dyn-O: Building Structured World Models with Object-Centric Representations
Dyn-O learns object-centric world models directly from pixels in complex Procgen games, using SAM2-guided slot attention and Mamba state-space dynamics, and reports better rollout prediction than DreamerV3.
-
Identifiable Object Representations under Spatial Ambiguities
VISA learns view-invariant object representations by aggregating probabilistic slots across multiple unlabeled viewpoints, with an identifiability analysis up to affine and permutation equivalence.
-
Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot Segmentation
A ViT layer-decomposition based self-disentanglement and re-composition method improves cross-domain few-shot segmentation, beating prior state-of-the-art by 1.92 (1-shot) and 1.88 (5-shot) average mIoU.
-
Compositional Scene Understanding through Inverse Generative Modeling
Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.
-
Is an object-centric representation beneficial for robotic manipulation ?
Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...
Discussion (0). Continue with ORCID to comment.