REVIEW 12 cited by
On the Binding Problem in Artificial Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contemporary neural networks still fall short of human-level generalization, which extends far beyond our direct experiences. In this paper, we argue that the underlying cause for this shortcoming is their inability to dynamically and flexibly bind information that is distributed throughout the network. This binding problem affects their capacity to acquire a compositional understanding of the world in terms of symbol-like entities (like objects), which is crucial for generalizing in predictable and systematic ways. To address this issue, we propose a unifying framework that revolves around forming meaningful entities from unstructured sensory inputs (segregation), maintaining this separation of information at a representational level (representation), and using these entities to construct new inferences, predictions, and behaviors (composition). Our analysis draws inspiration from a wealth of research in neuroscience and cognitive psychology, and surveys relevant mechanisms from the machine learning literature, to help identify a combination of inductive biases that allow symbolic information processing to emerge naturally in neural networks. We believe that a compositional approach to AI, in terms of grounded symbol-like representations, is of fundamental importance for realizing human-level generalization, and we hope that this paper may contribute towards that goal as a reference and inspiration.
Forward citations
Cited by 12 Pith papers
-
ActionParty: Multi-Subject Action Binding in Generative Video Games
ActionParty binds discrete actions to individual subjects in a single generated video by jointly modeling subject state tokens and video latents, controlling up to seven players across 46 Melting Pot games.
-
Human-like Object Grouping in Self-supervised Vision Transformers
DINO self-supervised transformers best predict human same/different object RTs; object-centric patch affinity and Gram-matrix distillation explain and transfer the alignment.
-
EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations
EditCLEVR uses before/after scene pairs with one known object edit to show that object-centric models fail to decode exactly the intended single-object attribute change under a CoGenT color-shape shift, even with grou...
-
HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition
One hierarchical slot-attention model with 10% labels jointly yields holistic, semantic, and panoptic scene decompositions that outperform three separate flat baselines by large ARI margins.
-
Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study
In information-matched tiny transformers, zero-shot compositional binding fails for every route, while few-shot efficiency is governed by input-pathway sharing and code readability.
-
ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks
A cycle-consistent GAN that translates between images and object lists matches state-of-the-art detection on synthetic scenes and detects low-contrast cells where slot-attention models fail.
-
Compositional Video Synthesis by Temporal Object-Centric Learning
An object-centric video model learns temporally consistent slots that condition a frozen diffusion decoder, enabling reconstruction and compositional editing of real-world videos.
-
Identifiable Object Representations under Spatial Ambiguities
VISA learns view-invariant object representations by aggregating probabilistic slots across multiple unlabeled viewpoints, with an identifiability analysis up to affine and permutation equivalence.
-
Bound by semanticity: universal laws governing the generalization-identification tradeoff
Finite-resolution similarity functions force a universal tradeoff between identification and generalization, with a predicted 1/n collapse of multi-input capacity.
-
Compositional Scene Understanding through Inverse Generative Modeling
Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.
-
Dynamics Reveals Structure: Challenging the Linear Propagation Assumption
Under the linear propagation assumption, first-order feature geometry cannot support both negation and composition: the only feature map satisfying both is the zero map.
-
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
A survey showing that common systematic generalization benchmarks measure behavioural systematicity, not the representational systematicity that Fodor and Pylyshyn's challenge requires, and mapping them onto Hadley's ...
Discussion (0). Continue with ORCID to comment.