{"id":"ccdb5f8f-deb3-45d3-8d08-be650ee5d43f","arxiv_id":"2412.18743","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Object-centric models handle novel combinations of object properties when all local features are present, and can even extrapolate to unseen shapes on the Pentomino dataset.","lead":"This paper tests whether object-centric neural networks can combine familiar object properties in new ways, such as seeing a known shape at an unseen rotation. It finds that these models succeed when the training data contains all the local image details, and this points to local feature availability as a key ingredient.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The source-of-skills claim rests on an unmeasured assumption that Pentomino rotations introduce no novel local pixel features; aliasing could falsify it.","rationale":"The reader's weakest_assumption identifies the same load-bearing issue: the Pentomino dataset is meant to cleanly provide 'no novel local features,' but this is never measured. My stress-test refines that concern by making it concrete: rasterization aliasing can create pixel-level features not seen in training, and the cross-dataset comparison (dSprites failure vs. Pentomino success) is confounded unless local feature availability is quantified. This is genuinely load-bearing because the abstract and the discussion attribute the model's success to this factor; if the assumption fails, the source-of-skills conclusion is not supported, even though the 3DShapes recombination result may stand. The paper has real strengths—the 3DShapes result is a large gap over WAE, the WAE control on Pentomino helps, and the extrapolation figures are visually striking—so I am not recommending rejection. The concern is an unverified empirical premise, not an internal inconsistency, and it can be settled by the proposed patch-density test. Since the reader's verdict was already CONDITIONAL and this concern is essentially the same one, no adjustment is needed.","tokens_in":8715,"tokens_out":8578,"duration_ms":90914,"concrete_test":"Train a patch-based density model (e.g., a small VAE or k-NN over normalized 7x7 patches) on all training images of the Pentomino dataset, then compare average log-likelihood or nearest-neighbor distance on two held-out sets: (1) test images of held-in shapes at held-out rotations, and (2) test images of held-out shapes at held-out rotations. If the held-out shape-rotation patches have substantially lower density than the held-in shape patches, then novel rotations do introduce novel local pixel features, and the paper's clean isolation of global configuration from local features is falsified. The same measurement should be run on dSprites to confirm that the contrast between the two datasets is driven by local feature novelty rather than by other dataset differences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central attribution—that SlotAttention/FgSeg succeeds when novel combinations do not require novel local features—depends on the Section 2.1 claim that Pentomino low-level features (straight lines, convex and concave right angles) remain unchanged under rotation. This is stated as 'we can be more confident,' not measured. In rasterized images, rotating a polyomino through the 40 discrete angles (9-degree spacing) produces anti-aliased edge fragments, corner patterns, and subpixel phase shifts whose local pixel neighborhoods are not guaranteed to already appear in training, even if the geometric primitives are the same. More importantly, the theory is tested by contrasting a failure on dSprites with a success on Pentomino, but the datasets differ in shape count, rendering style, and feature complexity, so the success does not isolate 'local feature availability' unless that equivalence is empirically established. If the Pentomino test patches are actually novel, the FgSeg success could be explained by dataset ease or interpolation, and the dSprites diagnosis would be unsupported. This concern does not target the existence result on 3DShapes, but it does target the paper's stated explanation of the source of the skills.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether object-centric autoencoders (SlotAttention and a one-slot variant FgSeg) support compositional generalisation over object properties, extending prior work on scene composition. On 3DShapes, SlotAttention reconstructs excluded combinations of pill shape and colour; on dSprites it partially fails for novel heart rotations. To localise the failure, the authors introduce a Pentomino dataset of twelve block shapes and show that FgSeg generalises to novel shape–rotation combinations and even to completely held-out shapes, whereas a Wasserstein autoencoder fails. They attribute success to figure-ground grouping plus the availability of all local pixel features during training, and report probe experiments suggesting that the latent representations are not readily abstract.","tokens_in":8947,"tokens_out":4940,"duration_ms":43815,"significance":"The main contribution is a set of positive existence results plus a new diagnostic dataset. If the results hold, they significantly extend the evidence that object-centric grouping helps combinatorial generalisation beyond scene composition, and they give a concrete hypothesis about the role of local feature availability. The paper is honest about limitations (single seed, uninformative latent visualisations) and includes a baseline control. However, the central explanatory claim—that local-feature availability is the source of the success—is not yet tested, and the quantitative support for the extrapolation claim is statistically thin. The strengths are the clear qualitative demonstrations and the introduction of a controlled dataset; the weak points are the unmeasured rotation-aliasing assumption and the absence of ablations.","major_comments":[{"comment":"The claim that rotating pentominoes in rasterised image space does not create novel local pixel features is asserted rather than measured. Anti-aliasing and resampling at arbitrary angles can produce subpixel patterns not present in the training set; this is precisely the quantity on which the paper's explanation rests. Please either quantify the novelty of local patches between training and test rotations or run a control in which rotations are generated by the same rendering style but for curved shapes, to support the attribution.","section":"§2.1, Fig. 2"},{"comment":"All quantitative claims rest on a single seed per experiment, and Tables 1–3 report no variance. The statement in Section 2.3 that there is 'no significant drop' from one to three novel shapes is not statistically supported (Table 3: 1.65 vs 1.29 train; 2.47 vs 2.63 test). Please report multiple seeds and error bars, or explicitly mark the extrapolation result as a single-run qualitative demonstration.","section":"Appendix A"},{"comment":"The claim that SlotAttention's position embeddings and per-object decoder are the source of its compositional skills is stated as if established ('it is also easy to pinpoint'), but no ablation manipulates these components. FgSeg is itself a simplification; the contribution of each component remains untested. An ablation removing position embeddings or replacing the per-object decoder with a monolithic one would be needed to support the causal attribution.","section":"§2.1, §2.2"},{"comment":"The Discussion says the probe results 'suggest that the models learn abstract representations,' while Appendix D concludes that because linear and non-linear probes fail on held-out combinations, 'the model's representations are not abstract after all.' This contradiction should be resolved; the current text overstates the probe evidence.","section":"§3 vs Appendix D"},{"comment":"The manuscript does not specify model architectures, hyperparameters, training budgets, or optimisation details for SlotAttention, FgSeg, or the WAE baseline. The Pentomino experiments are said to use 'the same architecture and training configuration' as dSprites, but that configuration is not given anywhere. Without these details the quantitative comparisons cannot be independently assessed.","section":"Main text and appendices"}],"minor_comments":[{"comment":"The Scale row lists six values (1.5, 1.8, 2.1, 2.4, 2.7, 3.0) while the text says five values, and the stated total of 960,000 images implies five values; please correct the table.","section":"Table 4"},{"comment":"Equation (1) does not define a normalised sampling distribution: summing the stated probabilities over all training images yields the number of shape classes, not one. Please clarify the reweighting procedure.","section":"Appendix B, Eq. (1)"},{"comment":"There are several typographical errors, including 'Sciecne' in the affiliation, 'this area this area', 'generalsation', 'preeliminary', 'traning', 'angels', and 'dDsprites' in Appendix A.","section":"Throughout"},{"comment":"The text refers to 'Figure 5, panel b' when discussing low-level features, but the relevant figure in the main text is Figure 2; the reference should be updated.","section":"§2.1"},{"comment":"The phrase 'An even these models cannot achieve significant prediction accuracy' is garbled; also, the appendix reports a two-layer MLP, so the sentence 'we need at least a two-layer MLP' should be checked for consistency.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-format manuscript with promising but preliminary evidence. I recommend major revision rather than rejection because the existence results are potentially valuable and the main gap—the local-feature assumption—is empirically addressable within the scope of the paper. The single-seed policy and missing experimental details should be fixed before archival publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper shows SlotAttention can handle a recombination-to-range task on 3DShapes that disentangled VAEs fail, and a single-slot variant (FgSeg) generalizes to novel shape-rotation combos on a new Pentomino set where a WAE doesn't. The central explanation—that success depends on novel combinations not requiring novel local features—is plausible but rests on an unmeasured assumption about the Pentomino rendering.\n\nWhat's actually new: previous tests of object-centric models were mostly about scene composition; this paper goes after property-level recombination, which is a different and harder problem. The Pentomino dataset is a useful idea: all shapes share straight lines and right angles, so you can vary global configuration while keeping low-level primitives constant. The extrapolation result (reconstructing shapes never seen in training, up to 3 of 12) is a real finding, and the gap vs WAE is big: reconstruction error 2.15 vs 10.55 on the shape-rotation generalization split.\n\nWhere it's soft: the load-bearing assumption in Section 2.1 is that rotating a pentomino through the 40 discrete angles doesn't create novel local pixel features. That's plausible for ideal geometry, but in rasterized images anti-aliasing can introduce edge fragments and corner patterns that were absent in training. They never measure this. Also, the dSprites-failure vs Pentomino-success comparison changes multiple things at once—shape count, rendering style, number of low-level features—so it doesn't isolate local feature availability as cleanly as the paper claims. The Appendix admits one seed per experiment, and there are no error bars, no ablations isolating the segmentation component, and no code or dataset release. These are real gaps, but they don't undermine the 3DShapes existence result, which shows a large, consistent gap against the WAE baseline.\n\nThe paper is honest about its own limitations, which is good. For a workshop version this is fine; for a full venue it would need repeated seeds, error bars, and preferably a direct test of the local-feature story (e.g., compare patch statistics between training and test rotations, or add an ablation that deliberately creates novel local features).\n\nVerdict: I'd send this to peer review. The core result is worth refereeing, and the dataset is a useful contribution. The explanation needs more support, but that's a revision problem, not a desk-reject problem.","headline":"A clean existence proof that object-centric models can recombine properties on 3DShapes, with a plausible but under-tested attribution to local feature availability.","tokens_in":9453,"tokens_out":2803,"would_cite":true,"duration_ms":24632,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Object-centric models can recombine familiar object properties into novel combinations that defeat standard disentangled models, and can even reconstruct entirely novel shapes when all local image features were seen during training.","keywords":["object-centric models","Slot Attention","compositional generalization","disentanglement","recombination-to-range","figure-ground segmentation","pentomino dataset","extrapolation"],"falsifier":"Measure the set of local image patches actually present in the pentomino training set and in the test set for excluded rotations: if any excluded rotation contains an oriented edge or corner patch that never appears in training, the paper's claim that 'low-level features are not novel' is false. A complementary test: train FgSeg on a pentomino variant whose rotations introduce curved or diagonal features and check whether generalisation to excluded shape-rotation combinations collapses.","tokens_in":8511,"feed_emoji":"🧩","tokens_out":8241,"duration_ms":61988,"temperature":0.7,"pith_summary":"This paper asks whether object-centric generative models, which first segment an image into objects, can generalise to novel combinations of an object's own properties, such as a familiar shape in a colour or rotation that was never seen with it. The authors show that a Slot Attention model succeeds on recombination-to-range tasks for shape and colour in 3DShapes, and that a stripped-down one-slot variant (FgSeg) succeeds on shape-rotation combinations in a new pentomino dataset where standard disentangled autoencoders and a Wasserstein auto-encoder fail badly. They argue the earlier failures of disentangled models arose not from a missing compositional mechanism but from training data missing the local pixel features needed to reconstruct a rotated shape. When those local features are shared across shapes, as in the pentominoes, the model even extrapolates to completely novel shapes. The paper also documents a limitation: the learned representations are not abstract, since a linear probe cannot decode shape and nonlinear probes fail on held-out combinations.","feed_headline":"Object-centric models recombine familiar features into novel shapes","feed_subtitle":"A one-slot attention model generalizes to unseen shape-rotation and shape-color pairs, and even generates new shapes.","key_machinery":"The load-bearing mechanism is the object-centric bottleneck: Slot Attention and its simplified one-slot version FgSeg, which forces the latent to reconstruct only the foreground figure through a Sigmoid rather than a Softmax slot competition. Position embeddings let the model track how patches move under rotation, and the per-object slot decoder reinforces grouping of patches that should be manipulated together. The pentomino dataset is the control: every shape is built from the same three low-level features (straight lines, convex right angles, concave right angles), so excluded shape-rotation combinations never contain unseen local patches. The Wasserstein auto-encoder serves as the baseline that isolates the contribution of this grouping bias.","core_discovery":"The central claim is that object-centric models with figure-ground segmentation can perform compositional generalisation over object properties, a setting where standard disentangled generative models fail. On 3DShapes, Slot Attention reconstructs pills in colours that were never paired with the pill shape during training. On dSprites, the same model partially handles hearts in novel rotations, and the authors attribute the residual failure to novel local features introduced by rotation in image space. To test this diagnosis, they introduce a pentomino dataset whose twelve shapes are all composed of straight lines and right angles, so a novel rotation only rearranges already-seen features. FgSeg, a single-slot attention model, reconstructs held-out shape-rotation combinations and, when up to three of twelve shapes are withheld, reconstructs entirely novel shapes from the shared local features. A Wasserstein auto-encoder control fails on the same conditions, supporting the conclusion that the success is due to the object-centric bottleneck rather than dataset ease.","pith_inferences":["If the diagnosis is correct, a dataset's local-feature coverage, not the model's inductive bias alone, determines when compositional generalisation is possible; one can test this by computing the novelty of oriented edge patches in any visual dataset before predicting generalisation success.","The pentomino result suggests object-centric models could serve as generative part-based synthesizers, but the probe results imply the generated output does not come with an explicit, reusable shape label, so generative competence and conceptual knowledge remain decoupled.","A natural extension is to apply the recombination-to-range test to natural-image object crops while controlling for local feature novelty, predicting that Slot Attention succeeds when held-out combinations reuse seen local patches and fails otherwise."],"forward_implications":["Object-centric inductive biases, rather than stronger symbolic or invariant priors, are sufficient for recombination-to-range generalisation over object properties.","Failures of disentangled VAEs at this task are best explained by missing figure-ground segmentation, not by insufficient disentanglement.","Ensuring that all local pixel features appear during training can turn a previously unsolvable generalisation condition into a solvable one.","Extrapolation to unseen shapes is achievable when a model has learned to recombine a small set of shared local features.","The representations learned by these models are not abstract concepts: linear probes fail, and nonlinear probes do not transfer to held-out combinations."],"supporting_citations":[{"why":"Introduces the Slot Attention architecture that is the object-centric model under test.","marker":"Locatello et al., 2020"},{"why":"Defined the recombination-to-range conditions and catalogued failures of disentangled models that this work aims to beat.","marker":"Montero et al., 2022"},{"why":"Provides the Wasserstein auto-encoder baseline that controls for dataset difficulty.","marker":"Tolstikhin et al., 2017"},{"why":"Earlier evidence that object-centric models compose novel scenes, which this paper extends to object properties.","marker":"Singh et al., 2021"},{"why":"Supplies the 3DShapes dataset used for the shape-colour recombination experiment.","marker":"Burgess and Kim, 2018"},{"why":"Supplies the dSprites dataset used for the shape-rotation recombination experiment.","marker":"Matthey et al., 2017"}],"fun_headline_variants":["Object-centric models recombine features beyond training pairs","Single-slot attention generalizes to unseen shape and color combos","Object-centric bottleneck beats disentangled models at recombining","Object-centric models generalize to new shape-rotation pairs","Object-centric models turn familiar parts into unseen wholes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that rotating a pentomino in rasterised image space never creates novel local pixel features, a claim the paper states informally and never measures; if aliasing or rasterisation produces unseen corners or edge patterns, the pentomino success could be explained by dataset ease rather than by the model's perceptual grouping.","fun_headline_variants_meta":{"raw":{"variants":["Object-centric models recombine features beyond training pairs","Single-slot attention generalizes to unseen shape and color combos","Object-centric bottleneck beats disentangled models at recombining","Object-centric models generalize to new shape-rotation pairs","Object-centric models turn familiar parts into unseen wholes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001041,"raw_usage":{"total_tokens":4350,"prompt_tokens":886,"completion_tokens":3464,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":3385}},"tokens_in":502,"tokens_out":3464,"duration_ms":25237,"temperature":1.0,"reasoning_tokens":3385,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:31:12.363648+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the set of local image patches actually present in the pentomino training set and in the test set for excluded rotations: if any excluded rotation contains an oriented edge or corner patch that never appears in training, the paper's claim that 'low-level features are not novel' is false. A complementary test: train FgSeg on a pentomino variant whose rotations introduce curved or diagonal features and check whether generalisation to excluded shape-rotation combinations collapses.","supporting_citations":[{"cited_title":"dSprites : Disentanglement testing Sprites dataset","cited_arxiv_id":null,"evidence_quote":"Supplies the dSprites dataset used for the shape-rotation recombination experiment."}],"review_version":1}