REVIEW 2 major objections 4 minor 98 references
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Systematic behaviour on a benchmark does not show that a model has systematic internal representations, and most current benchmarks do not require the strong form of systematicity that comes closest to human generalization.
desk verdict A useful survey that cleanly separates behavioural from representational systematicity, but its classification of visual benchmarks under a syntax-free Hadley taxonomy is underdetermined. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is Hadley's (1994) three-level taxonomy of systematicity—weak, quasi, and strong—which classifies what a training set has already shown a learner. Weak systematicity permits novel combinations of familiar words only in familiar syntactic positions; quasi-systematicity adds recursion over structurally familiar clauses; strong systematicity requires handling words in syntactic positions never seen in training, the level Hadley ties to human generalization. The authors use this scale to operationalize Fodor and Pylyshyn's representational systematicity, and they pair it with the competence/performance distinction: behavioural success is evidence, not proof, and mechanistic interpretability supplies the causal evidence needed to move from behaviour to representations.
What would settle it
Enumerate the train and test splits of a benchmark the paper classifies as testing only weak systematicity, for example SCAN Split 1, and count the test items in which some word appears in a syntactic position never attested in any training sentence; if that fraction is substantial rather than near zero, the weak-systematicity classification is false. The same position-coverage calculation, published for each benchmark, would settle which of the survey's classifications are correct.
Extended reading notes
Core claim
The central claim is that Fodor and Pylyshyn defined compositionality through systematicity of representations, not behaviour, and that the machine-learning literature invoking them has largely substituted behavioural systematicity for representational systematicity. On the authors' reading, a model can pass systematicity benchmarks through memorization or task-specific heuristics while lacking the structured internal representations Fodor and Pylyshyn argued are necessary. The survey adopts Hadley's weak/quasi/strong taxonomy as the operationalization of that distinction and uses it to classify benchmarks: SCAN's Split 1 and PCFG SET's systematicity split test weak systematicity; productivity-oriented splits require quasi-systematicity; COGS and ReCOGS-style structural generalization approach strong systematicity; most visual benchmarks do not reach it. The conclusion is that claims to have answered the Fodor–Pylyshyn challenge should be backed by mechanistic interpretability of the model's representations, not by benchmark scores alone.
Load-bearing premise
Everything downstream rests on the assumption that Hadley's 1994 three-level scale, built for connectionist language learning, transfers to modern end-to-end models and especially to visual benchmarks, where the paper concedes there is no agreed syntax of images to anchor the levels.
Editorial extensions
If this is right
- Benchmark reports should state which level of Hadley's taxonomy the train/test split actually demands, rather than treating any compositional split as evidence of Fodor–Pylyshyn systematicity.
- A claim that a model addresses the Fodor–Pylyshyn challenge should be accompanied by mechanistic evidence that the identified representations are causally responsible for the systematic behaviour.
- New language benchmarks should be built to target strong systematicity by controlling which words appear in which syntactic positions during training.
- For large pre-trained models, claims of strong behavioural systematicity are hard to support because the training distribution is not controlled.
- Visual systematicity benchmarks cannot be fully scored on Hadley's scale until a theory of the syntax of images is available.
Reading between the lines
- A natural extension the authors do not develop is a quantitative distance-to-strong measure: for any train/test split, compute the fraction of test items whose words appear only in syntactic positions absent from all training sentences; benchmarks could then be ordered by how much beyond weak systematicity they demand.
- If the survey is right, cross-benchmark model rankings should not be read as rankings of systematic generalization; the same model may exploit different shortcuts on different benchmarks, which is a direct explanation for the disagreement across benchmarks the paper cites.
- The paper's position implies a testable criterion for representational systematicity: causal interventions on internal representations, for example swapping role or binding vectors, should change the model's output exactly as the compositional operation predicts; where they do not, the behavioural success is probably shortcut-driven.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper argues that much of the machine-learning literature on systematicity measures behavioural systematicity, whereas Fodor and Pylyshyn's original challenge concerned systematicity of representations. The authors adopt Hadley's (1994) weak/quasi/strong taxonomy as an operational framework, apply it to a selection of language and vision benchmarks, and conclude that many current benchmarks fall short of strong systematicity, the level closest to human systematic generalization. They then review evidence for and against representational systematicity in end-to-end Transformer models, drawing on mechanistic interpretability, and close with recommendations for evaluating systematicity claims.
Significance. If the central distinction is accepted, the paper provides an important corrective to the field's frequent invocation of Fodor and Pylyshyn. Its strengths are the clear separation of behavioural from representational systematicity, the concrete examples illustrating the dangers of conflating the two (e.g., Bastings et al.'s SCAN shortcuts and ReCOGS's impact on COGS results), and the constructive emphasis on mechanistic interpretability as a complement to behavioural benchmarks. The paper is an opinionated survey, so its soundness rests on the accuracy and internal consistency of its interpretive claims; the main argument is well supported by the historical discussion and the case studies in language, though the visual-benchmark analysis is less secure.
major comments (2)
- [§4.2] The classification of visual benchmarks under Hadley's taxonomy is underdetermined because the paper itself acknowledges that no strong theory of image syntax exists. The text states, 'Lacking a strong theory of the syntax of images, it is unclear how Hadley's levels of systematicity apply to visual concepts,' yet it immediately assigns 'weak systematicity' to disentanglement datasets on the grounds that they 'lack a hierarchical structure' and do not place concepts in novel contexts. Hierarchical structure, syntactic position, and novel context are not defined for image contents. As a result, the weak/quasi/strong labels for visual benchmarks are not well grounded, and the Section 6 conclusion that 'many current benchmarks fall short of testing strong systematicity' is not supported for the visual half of the survey. I recommend either providing a working visual syntax (e.g., scene graphs or a grammar over generative factors) or explicitly limiting the Hadley-based classification to linguistic benchmarks and presenting the visual discussion as exploratory.
- [§2.3 and §4.1] The classification of productivity splits as requiring at least quasi-systematicity rests on an unproven conjecture. The paper asserts that weak systematicity yields no productivity because a sentence of unseen length 'will necessarily' contain at least one word in a novel syntactic position. This is not established for arbitrary phrase-structure grammars; under recursive grammars, arguments may occupy the same positions at all depths, and whether a word is in a novel position depends on the grammar's definition of positions. The classification of the PCFG SET Productivity split and SCAN's Split 2 depends on this step. Please either prove the claim under explicit grammar assumptions or soften the classification to 'may require at least quasi-systematicity'.
minor comments (4)
- [§3.1] The informal phrase 'Mike Young and colleagues' should be replaced with a formal citation, since the text later cites Young and Wasserman (1997, 2001) by name.
- [References] Several reference entries contain formatting errors in author names (e.g., 'DeV os', 'V on Kügelgen'); these should be corrected before publication.
- [§4.2] The sentence 'ARC may constitute a test for strong compositionality but cannot be judged on an objective basis' is ambiguous; please separate the hedged claim from the justification that the construction process is not systematically described.
- [§4.1] For SCAN Split 1, the classification as weak systematicity relies on the statistical claim that even 2% of the training data exhibits all possible commands in all possible positions; please state explicitly that this is an empirical judgment about the dataset rather than a formal property of the split.
Circularity Check
No significant circularity: the survey's central distinction and benchmark classifications are imported from external sources (Fodor & Pylyshyn; Hadley), not derived from its own conclusion.
full rationale
This is an opinionated survey rather than a derivation, and its central claims are backed by external, non-self-authored sources. The behavioural-vs-representational distinction is attributed to Fodor and Pylyshyn (1988), and the weak/quasi/strong taxonomy is Hadley's (1994); the paper explicitly adopts these frameworks rather than defining them in terms of its own target conclusion. The benchmark classifications in Sections 4.1 and 4.2 are applications of Hadley's externally defined levels to specific datasets, and the observation that many benchmarks 'fall short of testing strong systematicity' follows from inspecting those datasets under Hadley's definitions; it is a classification, not a forced equivalence. The paper itself flags the main threat to this application in Section 4.2: 'lacking a strong theory of the syntax of images, it is unclear how Hadley's levels of systematicity apply to visual concepts,' which is an acknowledged limitation on the validity of the visual extension, not a circular step. Self-citations (Lewis et al. 2024; Doumas et al. 2022) appear only as examples or as out-of-scope footnotes and are not load-bearing; they do not supply the framework or the uniqueness of the taxonomy. No fitted parameters are renamed as predictions, and no uniqueness theorem from the authors' prior work is invoked. The productivity conjectures in Section 2.3 are consequences of the definitions of weak/quasi/strong systematicity and unseen-length sentences, not circular inputs. The least secure part of the paper is the transfer of Hadley's linguistic taxonomy to visual benchmarks, but that is an underdetermination/correctness concern, not circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Fodor and Pylyshyn's systematicity is fundamentally a property of representations, not behaviour.
- domain assumption Hadley's three levels of systematicity provide a valid operationalization of Fodor and Pylyshyn's representational systematicity.
- domain assumption A valid behavioural operationalisation requires principled control over dataset syntax and semantics.
- domain assumption Behavioural success under a valid operationalisation is evidence of representational systematicity only if Fodor and Pylyshyn's position is accepted.
Cite this review
Pith. "Pith review of Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey." pith.science (2026). https://pith.science/paper/IMUHCEJ4
@misc{pith2026250604461,
author = {Pith},
title = {Pith review of: Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/IMUHCEJ4}},
note = {Machine review of arXiv:2506.04461}
}
read the original abstract
A core aspect of compositionality, systematicity is a desirable property in ML models as it enables strong generalization to novel contexts. This has led to numerous studies proposing benchmarks to assess systematic generalization, as well as models and training regimes designed to enhance it. Many of these efforts are framed as addressing the challenge posed by Fodor and Pylyshyn. However, while they argue for systematicity of representations, existing benchmarks and models primarily focus on the systematicity of behaviour. We emphasize the crucial nature of this distinction. Furthermore, building on Hadley's (1994) taxonomy of systematic generalization, we analyze the extent to which behavioural systematicity is tested by key benchmarks in the literature across language and vision. Finally, we highlight ways of assessing systematicity of representations in ML models as practiced in the field of mechanistic interpretability.
Reference graph
Works this paper leans on
-
[1]
Nura Aljaafari, Danilo S. Carvalho, and André Freitas. 2025. https://doi.org/10.48550/arXiv.2410.12924 Interpreting token compositionality in LLMs : A robustness analysis . arXiv preprint. ArXiv:2410.12924 [cs] version: 2
-
[2]
Jacob Andreas. 2019. https://doi.org/10.48550/arXiv.1902.07181 Measuring Compositionality in Representation Learning . arXiv preprint. ArXiv:1902.07181 [cs]
-
[3]
Rim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal, and Adriana Romero-Soriano. 2024. Oc-clip: Object-centric binding in contrastive language-image pretraining. In NeurIPS 2024 Workshop on Compositional Learning: Perspectives, Methods, and Paths Forward
2024
-
[4]
Rim Assouel, Pau Rodriguez, Perouz Taslakian, David Vazquez, and Yoshua Bengio. 2022. Object-centric compositional imagination for visual abstract reasoning. In ICLR2022 Workshop on the Elements of Reasoning: Objects, Structure and Causality
2022
-
[5]
Samy Badreddine, Artur d'Avila Garcez, Luciano Serafini, and Michael Spranger. 2022. https://doi.org/10.1016/j.artint.2021.103649 Logic Tensor Networks . Artificial Intelligence, 303:103649. ArXiv:2012.13635 [cs]
arXiv 2022
-
[6]
Renée Baillargeon and Julie DeVos. 1991. https://doi.org/10.1111/j.1467-8624.1991.tb01602.x Object Permanence in Young Infants : Further Evidence . Child Development, 62(6):1227--1246. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1467-8624.1991.tb01602.x
arXiv 1991
-
[7]
Jasmijn Bastings, Marco Baroni, Jason Weston, Kyunghyun Cho, and Douwe Kiela. 2018. https://doi.org/10.18653/v1/W18-5407 Jump to better conclusions: SCAN both left and right . In Proceedings of the 2018 EMNLP Workshop BlackboxNLP : Analyzing and Interpreting Neural Networks for NLP , pages 47--55, Brussels, Belgium. Association for Computational Linguistics
-
[8]
Yonatan Belinkov. 2022. https://doi.org/10.1162/coli_a_00422 Probing Classifiers : Promises , Shortcomings , and Advances . Computational Linguistics, 48(1):207--219
Show all 98 references
-
[9]
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen,...
2023
-
[10]
Elisabeth Camp. 2007. https://doi.org/10.1111/j.1520-8583.2007.00124.x Thinking with Maps . Philosophical Perspectives, 21(1):145--182. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1520-8583.2007.00124.x
2007
-
[11]
Patrick Cavanagh. 2021. https://doi.org/10.1177/0301006621991491 The Language of Vision * . Perception, 50(3):195--215. Publisher: SAGE Publications Ltd STM
2021 doi
-
[12]
Fran c ois Chollet. 2019. On the measure of intelligence. arXiv preprint arXiv:1911.01547
2019 arXiv
-
[13]
Noam Chomsky. 1965. https://www.jstor.org/stable/j.ctt17kk81z Aspects of the Theory of Syntax , 50 edition. The MIT Press
1965
- [14]
-
[15]
Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.49 The Devil is in the Detail : Simple Tricks Improve Systematic Generalization of Transformers . In Proceedings of the 2021 Conference on Empirical Methods in Natural Langu...
2021 doi
-
[16]
Robert Cummins. 1996. Systematicity. The Journal of Philosophy, 93(12):591--614
1996
- [17]
- [18]
-
[19]
Aniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Charles Blundell, Philippe Beaudoin, Nicolas Heess, Michael C Mozer, and Yoshua Bengio. 2021. Neural production systems. Advances in Neural Information Processing Systems, 34:25673--25687
2021
-
[20]
Anuj Diwan, Layne Berry, Eunsol Choi, David Harwath, and Kyle Mahowald. 2022. Why is winoground hard? investigating failures in visuolinguistic compositionality. arXiv preprint arXiv:2211.00768
2022 arXiv
-
[21]
Leonidas A. A. Doumas, Guillermo Puebla, Andrea E. Martin, and John E. Hummel. 2022. https://doi.org/10.1037/rev0000346 A theory of relation learning and cross-domain generalization . Psychological Review, 129:999--1041. Place: US Publisher: American Psychological Association
2022 doi
-
[22]
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy models of sup...
2022
-
[23]
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dari...
2021
-
[24]
Joël Fagot and Carole Parron. 2010. https://doi.org/10.1037/a0017169 Relational matching in baboons ( Papio papio) with reduced grouping requirements . Journal of Experimental Psychology: Animal Behavior Processes, 36(2):184--193. Place: US Publisher: American Psychological As...
2010 doi
-
[25]
Wasserman, and Michael E
Joël Fagot, Edward A. Wasserman, and Michael E. Young. 2001. https://doi.org/10.1037/0097-7403.27.4.316 Discriminating the relation between relations: The role of entropy in abstract conceptualization by baboons ( Papio papio) and humans ( Homo sapiens) . Journal of Experiment...
2001 doi
-
[26]
Jerome Feldman. 2013. https://doi.org/10.1007/s11571-012-9219-8 The neural binding problem(s) . Cognitive Neurodynamics, 7(1):1--11
2013 doi
-
[27]
Jiahai Feng, Stuart Russell, and Jacob Steinhardt. 2024. https://openreview.net/forum?id=0yvZm2AjUr Monitoring Latent World States in Language Models with Propositional Probes
2024
- [28]
-
[29]
McLaughlin
Jerry Fodor and Brian P. McLaughlin. 1990. https://doi.org/10.1016/0010-0277(90)90014-B Connectionism and the problem of systematicity: Why Smolensky 's solution doesn't work . Cognition, 35(2):183--204
1990 doi
-
[30]
Fodor and Zenon W
Jerry A. Fodor and Zenon W. Pylyshyn. 1988. https://doi.org/10.1016/0010-0277(88)90031-5 Connectionism and cognitive architecture: A critical analysis . Cognition, 28(1):3--71
1988 doi
-
[31]
Gottlob Frege. 1892. Über Sinn und Bedeutung , 1. auflage edition. Zeitschrift für Philosophie und philosophische Kritik , Neue Folge . Pfeffer, Leipzig
-
[32]
Pullum, and Ivan A
Gerald Gazdar, Evan Klein, Geoffrey K. Pullum, and Ivan A. Sag. 1985. https://www.hup.harvard.edu/books/9780674344563 Generalized Phrase Structure Grammar . Harvard University Press
1985
-
[33]
Wichmann
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020. https://doi.org/10.1038/s42256-020-00257-z Shortcut learning in deep neural networks . Nature Machine Intelligence, 2(11):665--673. Publisher:...
2020 doi
-
[34]
Adele Goldberg. 1995. Constructions: A construction grammar approach to argument structure . University of Chicago Press, Chicago ; London. Series Title: Cognitive theory of language and culture
1995
- [35]
-
[36]
Robert F. Hadley. 1994. https://doi.org/10.1111/j.1468-0017.1994.tb00225.x Systematicity in Connectionist Language Learning . Mind & Language, 9(3):247--272. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1468-0017.1994.tb00225.x
1994
-
[37]
Per-Kristian Halvorsen and William A. Ladusaw. 1979. https://doi.org/10.1007/BF00126510 Montague's ‘universal grammar’: An introduction for the linguist . Linguistics and Philosophy, 3(2):185--223
1979 doi
-
[38]
Jaakko Hintikka. 1979. https://doi.org/10.1007/978-1-4020-4108-2_1 Language- Games . In Esa Saarinen, editor, Game- Theoretical Semantics : Essays on Semantics by Hintikka , Carlson , Peacocke , Rantala , and Saarinen , pages 1--26. Springer Netherlands, Dordrecht
1979 doi
-
[39]
Wilhelm von Humboldt. 1836. On Language : On the Diversity of Human Language Construction and Its Influence on the Mental Development of the Human Species . Google-Books-ID: \_UODbGlD4WUC
- [40]
-
[41]
Theo M. V. Janssen and Barbara H. Partee. 1997. https://doi.org/10.1016/B978-044481714-3/50011-4 Chapter 7 - Compositionality . In Johan van Benthem and Alice ter Meulen, editors, Handbook of Logic and Language , pages 417--473. North-Holland, Amsterdam
1997 doi
-
[42]
Lawrence Zitnick, and Ross Girshick
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. https://doi.org/10.1109/CVPR.2017.215 CLEVR : A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning . In 2017 IEEE Conference on Compu...
2017 doi
-
[43]
Aravind K. Joshi. 2005. https://doi.org/10.1093/oxfordhb/9780199276349.013.0026 Tree- Adjoining Grammars . In Ruslan Mitkov, editor, The Oxford Handbook of Computational Linguistics , page 0. Oxford University Press
2005
-
[44]
jylin04 , JackS, Adam Karvonen, and Can . 2024. https://www.alignmentforum.org/posts/gcpNuEZnxAPayaKBY/othellogpt-learned-a-bag-of-heuristics-1 OthelloGPT learned a bag of heuristics
2024
-
[46]
Najoung Kim and Tal Linzen. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.731 COGS : A Compositional Generalization Challenge Based on Semantic Interpretation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing ( EMNLP ) , pages 9087...
2020 doi
-
[47]
Yeongbin Kim, Gautam Singh, Junyeong Park, Caglar Gulcehre, and Sungjin Ahn. 2023. Imagine the unseen world: a benchmark for systematic generalization in visual world models. Advances in Neural Information Processing Systems, 36:27880--27896
2023
-
[48]
Seijin Kobayashi, Simon Schug, Yassir Akram, Florian Redhardt, Johannes von Oswald, Razvan Pascanu, Guillaume Lajoie, and João Sacramento. 2024. https://doi.org/10.48550/arXiv.2407.12275 When can transformers compositionally generalize in-context? arXiv preprint. ArXiv:2407.12275 [cs]
-
[49]
Brenden Lake and Marco Baroni. 2018. https://proceedings.mlr.press/v80/lake18a.html Generalization without Systematicity : On the Compositional Skills of Sequence -to- Sequence Recurrent Networks . In Proceedings of the 35th International Conference on Machine Learning , pages...
2018
-
[50]
Kevin J. Lande. 2021. https://doi.org/10.1111/nous.12324 Mental structures . Noûs, 55(3):649--677. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1111/nous.12324
2021 doi
-
[51]
Kevin J. Lande. 2024. https://doi.org/10.1002/wcs.1691 Compositionality in perception: A framework . WIREs Cognitive Science, 15(6):e1691. \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/wcs.1691
2024 doi
-
[52]
Martha Lewis, Nihal Nayak, Peilin Yu, Jack Merullo, Qinan Yu, Stephen Bach, and Ellie Pavlick. 2024. https://aclanthology.org/2024.findings-eacl.101/ Does CLIP Bind Concepts ? Probing Compositionality in Large Image Models . In Findings of the Association for Computational Lin...
2024
-
[53]
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, and Najoung Kim. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.194 SLOG : A Structural Generalization Benchmark for Semantic Parsing . In Proceedings of the 2023 Conference on Empirical Methods in Natur...
2023 doi
-
[54]
Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024. https://doi.org/10.48550/arXiv.2210.13382 Emergent World Representations : Exploring a Sequence Model Trained on a Synthetic Task . arXiv preprint. ArXiv:2210.13382 [cs]
- [55]
-
[56]
Weiduo Liao, Ying Wei, Mingchen Jiang, Qingfu Zhang, and Hisao Ishibuchi. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/hash/6a42b45af2b72e6e5b5e3a6fe695809f-Abstract-Datasets_and_Benchmarks.html Does Continual Learning Meet Compositionality ? New Benchmarks and ...
2023
-
[57]
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. 2020. Object-centric learning with slot attention. Advances in neural information processing systems, 33:11525--11538
2020
-
[58]
R Duncan Luce. 1996. The ongoing dialog between empirical science and measurement theory. journal of mathematical psychology, 40(1):78--98
1996
-
[59]
Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi, Irena Gao, and Ranjay Krishna. 2023. https://doi.org/10.1109/CVPR52729.2023.01050 CREPE : Can Vision - Language Foundation Models Reason Compositionally ? In 2023 IEEE / CVF Conference on Computer Vision and Pattern Recogni...
2023
-
[60]
Tenenbaum, and Jiajun Wu
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum, and Jiajun Wu. 2019. https://doi.org/10.48550/arXiv.1904.12584 The Neuro - Symbolic Concept Learner : Interpreting Scenes , Words , and Sentences From Natural Supervision . arXiv preprint. ArXiv:1904.12584 [cs]
-
[61]
Thomas McCoy, Tal Linzen, Ewan Dunbar, and Paul Smolensky
R. Thomas McCoy, Tal Linzen, Ewan Dunbar, and Paul Smolensky. 2020. https://aclanthology.org/2020.scil-1.34 Tensor Product Decomposition Networks : Uncovering Representations of Structure Learned by Neural Networks . In Proceedings of the Society for Computation in Linguistics...
2020
-
[62]
Kate McCurdy, Paul Soulos, Paul Smolensky, Roland Fernandez, and Jianfeng Gao. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.524 Toward Compositional Behavior in Neural Models : A Survey of Current Views . In Proceedings of the 2024 Conference on Empirical Methods in Natur...
2024 doi
-
[63]
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2024. https://doi.org/10.18653/v1/2024.naacl-long.281 Language Models Implement Simple Word2Vec -style Vector Arithmetic . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computationa...
2024 doi
-
[64]
Arseny Moskvichev, Victor Vikram Odouard, and Melanie Mitchell. 2023. The conceptarc benchmark: Evaluating understanding and generalization in the arc domain. arXiv preprint arXiv:2305.07141
2023 arXiv
- [65]
-
[66]
nostalgebraist . 2020. https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens interpreting GPT : the logit lens
2020
-
[67]
Victor Vikram Odouard and Melanie Mitchell. 2022. Evaluating understanding on conceptual abstraction benchmarks. arXiv preprint arXiv:2206.14187
2022 arXiv
-
[68]
Dick, and Hidenori Tanaka
Maya Okawa, Ekdeep Singh Lubana, Robert P. Dick, and Hidenori Tanaka. 2023. https://openreview.net/forum?id=ZXH8KUgFx3#all Compositional Abilities Emerge Multiplicatively : Exploring Diffusion Models on a Synthetic Task
2023
-
[69]
Chris Olah. 2023. https://transformer-circuits.pub/2023/superposition-composition/index.html Distributed Representations : Composition & Superposition
2023
- [70]
-
[71]
Kiho Park, Yo Joong Choe, and Victor Veitch. 2024. The linear representation hypothesis and the geometry of large language models. In Proceedings of the 41st International Conference on Machine Learning , volume 235 of ICML '24 , pages 39643--39666, Vienna, Austria. JMLR.org
2024
-
[72]
Barbara H. Partee. 2004. https://doi.org/10.1002/9780470751305.ch7 Compositionality . In Compositionality in Formal Semantics , pages 153--181. John Wiley & Sons, Ltd. Section: 7 \_eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/9780470751305.ch7
2004 doi
-
[73]
Ellie Pavlick. 2023. https://doi.org/10.1098/rsta.2022.0041 Symbols and grounding in large language models . Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences, 381(2251):20220041. Publisher: Royal Society
2023
-
[74]
Jean Piaget. 2013. https://doi.org/10.4324/9781315009650 The Construction Of Reality In The Child . Routledge, London
2013 doi
-
[75]
David Premack and Guy Woodruff. 1978. https://doi.org/10.1017/S0140525X00076512 Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515--526
1978 doi
-
[76]
Jake Quilty-Dunn, Nicolas Porot, and Eric Mandelbaum. 2023. https://doi.org/10.1017/S0140525X22002849 The best game in town: The reemergence of the language-of-thought hypothesis across the cognitive sciences . Behavioral and Brain Sciences, 46:e261
2023 doi
- [77]
-
[78]
Williams, and Lotem Elber-Dorozko
Jacob Russin, Sam Whitman McGrath, Danielle J. Williams, and Lotem Elber-Dorozko. 2024. https://doi.org/10.48550/arXiv.2405.15164 From Frege to chatGPT : Compositionality in language, cognition, and deep neural networks . arXiv preprint. ArXiv:2405.15164 [cs]
-
[79]
Lukas Schott, Julius Von Kügelgen, Frederik Träuble, Peter Vincent Gehler, Chris Russell, Matthias Bethge, Bernhard Schölkopf, Francesco Locatello, and Wieland Brendel. 2021. https://openreview.net/forum?id=9RUHPlladgh Visual Representation Learning Does Not Generalize Strongl...
2021
-
[80]
Prithviraj Sen, Breno WSR de Carvalho, Ryan Riegel, and Alexander Gray. 2022. Neuro-symbolic inductive logic programming with logical neural networks. In Proceedings of the AAAI conference on artificial intelligence, volume 36, pages 8212--8219
2022
-
[81]
Michaud, Stephen Casper, Max Tegmark, William Saunders, David Bau, Eric Todd, Atticus Geiger, Mor Geva, Jesse Hoogland, Daniel Murfet, and Tom McGrath
Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, Stella Biderman, Adria Garriga-Alonso, Arthur Conmy, Neel Nanda, Jessica Rumbelow, Martin Wattenberg, Nandi Schoots, ...
-
[82]
Paul Soulos, Henry Conklin, Mattia Opper, Paul Smolensky, Jianfeng Gao, and Roland Fernandez. 2024. Compositional generalization across distributional shifts with sparse tree operations. arXiv preprint arXiv:2412.14076
2024 arXiv
-
[83]
Mark Steedman. 2019. https://doi.org/10.1515/9783110540253-014 Combinatory Categorial Grammar . In Current Approaches to Syntax . Publication Title: Current Approaches to Syntax
2019 doi
-
[84]
Stanley Smith Stevens. 1946. On the theory of scales of measurement. Science, 103(2684):677--680
1946
-
[85]
Kaiser Sun, Adina Williams, and Dieuwke Hupkes. 2023. https://doi.org/10.18653/v1/2023.conll-1.19 The Validity of Evaluation Results : Assessing Concurrence Across Compositionality Benchmarks . In Proceedings of the 27th Conference on Computational Natural Language Learning ( ...
2023 doi
-
[86]
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 a . https://doi.org/10.18653/v1/P19-1452 BERT Rediscovers the Classical NLP Pipeline . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4593--4601, Florence, Italy. Association ...
2019 doi
-
[87]
Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 b . https://doi.org/10.48550/arXiv.1905.06316 What do you learn from context? Probing for sentence structure ...
-
[88]
Jonathan Thomm, Giacomo Camposampiero, Aleksandar Terzic, Michael Hersche, Bernhard Sch \"o lkopf, and Abbas Rahimi. 2024. Limits of transformer language models on learning to compose algorithms. In The Thirty-eighth Annual Conference on Neural Information Processing Systems
2024
-
[89]
Roger K. R. Thompson, David L. Oden, and Sarah T. Boysen. 1997. https://doi.org/10.1037/0097-7403.23.1.31 Language-naive chimpanzees ( Pan troglodytes) judge relations between relations in a conceptual matching-to-sample task . Journal of Experimental Psychology: Animal Behavi...
1997 doi
-
[90]
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. 2022. https://doi.org/10.1109/CVPR52688.2022.00517 Winoground: Probing Vision and Language Models for Visio - Linguistic Compositionality . In 2022 IEEE / CVF Conference on...
2022
- [91]
-
[92]
Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan
Keyon Vafa, Justin Y. Chen, Ashesh Rambachan, Jon Kleinberg, and Sendhil Mullainathan. 2024. https://doi.org/10.48550/arXiv.2406.03689 Evaluating the World Model Implicit in a Generative Model . arXiv preprint. ArXiv:2406.03689 [cs]
- [93]
-
[94]
Dag Westerståhl. 1998. https://www.jstor.org/stable/25001726 On Mathematical Proofs of the Vacuity of Compositionality . Linguistics and Philosophy, 21(6):635--643. Publisher: Springer
1998
-
[95]
Manning, and Christopher Potts
Zhengxuan Wu, Christopher D. Manning, and Christopher Potts. 2023. https://doi.org/10.1162/tacl_a_00623 ReCOGS : How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation . Transactions of the Association for Computational Linguistics, 11:1719--1733
2023 doi
-
[96]
Young and Edward A
Michael E. Young and Edward A. Wasserman. 1997. https://doi.org/10.1037/0097-7403.23.2.157 Entropy detection by pigeons: Response to mixed visual displays after same–different discrimination training . Journal of Experimental Psychology: Animal Behavior Processes, 23(2):157--1...
1997 doi
-
[97]
Young and Edward A
Michael E. Young and Edward A. Wasserman. 2001. https://doi.org/10.1037/0278-7393.27.1.278 Entropy and variability discrimination . Journal of Experimental Psychology: Learning, Memory, and Cognition, 27(1):278--293. Place: US Publisher: American Psychological Association
2001 doi
-
[98]
Aimen Zerroug, Mohit Vaishnav, Julien Colin, Sebastian Musslick, and Thomas Serre. 2022. A benchmark for compositional visual reasoning. Advances in neural information processing systems, 35:29776--29788
2022
-
[99]
Chi Zhang, Feng Gao, Baoxiong Jia, Yixin Zhu, and Song-Chun Zhu. 2019. Raven: A dataset for relational and analogical visual reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5317--5327
2019
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.