REVIEW 3 cited by
Unsupervised Object Learning via Common Fate
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learning generative object models from unlabelled videos is a long standing problem and required for causal scene modeling. We decompose this problem into three easier subtasks, and provide candidate solutions for each of them. Inspired by the Common Fate Principle of Gestalt Psychology, we first extract (noisy) masks of moving objects via unsupervised motion segmentation. Second, generative models are trained on the masks of the background and the moving objects, respectively. Third, background and foreground models are combined in a conditional "dead leaves" scene model to sample novel scene configurations where occlusions and depth layering arise naturally. To evaluate the individual stages, we introduce the Fishbowl dataset positioned between complex real-world scenes and common object-centric benchmarks of simplistic objects. We show that our approach allows learning generative models that generalize beyond the occlusions present in the input videos, and represent scenes in a modular fashion that allows sampling plausible scenes outside the training distribution by permitting, for instance, object numbers or densities not observed in the training set.
Forward citations
Cited by 3 Pith papers
-
Successes and Limitations of Object-centric Models at Compositional Generalisation
Object-centric models handle novel combinations of object properties when all local features are present, and can even extrapolate to unseen shapes on the Pentomino dataset.
-
Temporally Consistent Object-Centric Learning by Contrasting Slots
Slot-slot temporal contrast improves temporal consistency and object discovery in unsupervised object-centric video models, reaching state-of-the-art FG-ARI on MOVi-E and YouTube-VIS.
-
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
A2VIS integrates amodal, full-shape masks into video instance segmentation via global prototypes and a spatiotemporal-prior mask head, improving occlusion-robust tracking on synthetic benchmarks.
Discussion (0). Continue with ORCID to comment.