Pith. sign in

REVIEW 12 cited by

On the Binding Problem in Artificial Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.05208 v1 pith:2VBBHERS submitted 2020-12-09 cs.NE cs.AIcs.LG

classification cs.NEcs.AIcs.LG
keywords entitiesinformationnetworksneuralbindingcompositionalgeneralizationhuman-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contemporary neural networks still fall short of human-level generalization, which extends far beyond our direct experiences. In this paper, we argue that the underlying cause for this shortcoming is their inability to dynamically and flexibly bind information that is distributed throughout the network. This binding problem affects their capacity to acquire a compositional understanding of the world in terms of symbol-like entities (like objects), which is crucial for generalizing in predictable and systematic ways. To address this issue, we propose a unifying framework that revolves around forming meaningful entities from unstructured sensory inputs (segregation), maintaining this separation of information at a representational level (representation), and using these entities to construct new inferences, predictions, and behaviors (composition). Our analysis draws inspiration from a wealth of research in neuroscience and cognitive psychology, and surveys relevant mechanisms from the machine learning literature, to help identify a combination of inductive biases that allow symbolic information processing to emerge naturally in neural networks. We believe that a compositional approach to AI, in terms of grounded symbol-like representations, is of fundamental importance for realizing human-level generalization, and we hope that this paper may contribute towards that goal as a reference and inspiration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ActionParty: Multi-Subject Action Binding in Generative Video Games

    cs.CV 2026-04 conditional novelty 7.0 of 10

    ActionParty binds discrete actions to individual subjects in a single generated video by jointly modeling subject state tokens and video latents, controlling up to seven players across 46 Melting Pot games.

  2. Human-like Object Grouping in Self-supervised Vision Transformers

    cs.CV 2026-03 conditional novelty 6.5 of 10

    DINO self-supervised transformers best predict human same/different object RTs; object-centric patch affinity and Gram-matrix distillation explain and transfer the alignment.

  3. EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    EditCLEVR uses before/after scene pairs with one known object edit to show that object-centric models fail to decode exactly the intended single-object attribute change under a CoGenT color-shape shift, even with grou...

  4. HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition

    cs.CV 2026-07 conditional novelty 6.0 of 10

    One hierarchical slot-attention model with 10% labels jointly yields holistic, semantic, and panoptic scene decompositions that outperform three separate flat baselines by large ARI margins.

  5. Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study

    cs.LG 2026-07 accept novelty 6.0 of 10

    In information-matched tiny transformers, zero-shot compositional binding fails for every route, while few-shot efficiency is governed by input-pathway sharing and code readability.

  6. ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A cycle-consistent GAN that translates between images and object lists matches state-of-the-art detection on synthetic scenes and detects low-contrast cells where slot-attention models fail.

  7. Compositional Video Synthesis by Temporal Object-Centric Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An object-centric video model learns temporally consistent slots that condition a frozen diffusion decoder, enabling reconstruction and compositional editing of real-world videos.

  8. Identifiable Object Representations under Spatial Ambiguities

    cs.LG 2025-06 reject novelty 6.0 of 10

    VISA learns view-invariant object representations by aggregating probabilistic slots across multiple unlabeled viewpoints, with an identifiability analysis up to affine and permutation equivalence.

  9. Bound by semanticity: universal laws governing the generalization-identification tradeoff

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Finite-resolution similarity functions force a universal tradeoff between identification and generalization, with a predicted 1/n collapse of multi-input capacity.

  10. Compositional Scene Understanding through Inverse Generative Modeling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Composing per-concept diffusion models and inverting them with denoising loss enables multi-object scene understanding that generalizes beyond the training distribution.

  11. Dynamics Reveals Structure: Challenging the Linear Propagation Assumption

    cs.LG 2026-01 conditional novelty 5.0 of 10

    Under the linear propagation assumption, first-order feature geometry cannot support both negation and composition: the only feature map satisfying both is the zero map.

  12. Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey

    cs.LG 2025-06 conditional novelty 5.0 of 10

    A survey showing that common systematic generalization benchmarks measure behavioural systematicity, not the representational systematicity that Fodor and Pylyshyn's challenge requires, and mapping them onto Hadley's ...

Pith tools