REVIEW 3 major objections 4 minor 1 cited by
Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read DALL-E 3 fails most logic prompts, say human judges
desk verdict Solid human evaluation of DALL-E 3's logic failures, but the abstract overclaims and an uncalibrated threshold plus unreliable auto-counts need fixing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the logical probe: a minimal prompt that pairs an everyday object with one relation, one negation, or one integer, rendered 18 times per prompt. The measuring instrument is human agreement—participants select all, some, or none of the 18 images as matching the prompt. Around this, the paper adds three analytic devices: a fixed prefix that stops the model's internal language model from rewriting a prompt, so the raw model is tested; an N-gram frequency correlation that tests whether relational success tracks training-data statistics; and an approximate-numeracy analysis using object-detection counts to estimate scalar variability and ratio dependence in generated quantities.
What would settle it
Run the same 18-image selection task with deliberately non-matching prompts as catch trials and compare selection rates: if people endorse non-matching images at or above the rates observed for relations (45%), the claim that relations fail falls to a measurement artifact. A second check would ask humans to count objects in images from the auto-count follow-up and compare their counts to the detection-model estimates.
Extended reading notes
Core claim
The paper's central claim is that DALL-E 3 does not reliably deploy basic logical operators in image generation, and that no probe family—relations, negations, or numbers—consistently clears the 50% human-agreement mark. Negation fails most completely: when the prompt is left unmodified, it almost always renders the very object it forbids. Number generation is exact for small counts and approximate beyond three, with agreement collapsing from 75% at "one" to 9% at "six"; a follow-up analysis with object detectors finds ratio-dependent, scalar-variable error patterns like an approximate number system rather than exact counting. The paper also claims that a grounded pipeline using structured layout graphs does not fix these problems and is judged worse than DALL-E 3 on the same relation prompts, because its intermediate representations miss physics such as occlusion-in-depth.
Load-bearing premise
The headline "none greater than 50%" treats 50% agreement as the success threshold, but the task offered no baseline or chance condition showing how often people endorse images when a prompt has no valid match.
Editorial extensions
If this is right
- Text-to-image "prompt following" should not be read as compositional understanding; simple operator probes reveal systematic gaps that survive scale.
- Plain negation is not just imperfect but inverted: an unmodified "not X" prompt tends to produce X, which means systems need an explicit rewriting step that replaces the forbidden object.
- Exact count generation beyond three or four objects is outside current capability, and the error pattern is ratio-dependent rather than exact-integer-like.
- Structured intermediate representations such as scene graphs can hurt rather than help when they omit physical constraints like occlusion, so grounding alone is not sufficient.
- Prompt frequency predicts relational success, implying that reported gains in relations may reflect dataset statistics rather than relational reasoning.
Reading between the lines
- Because no baseline or chance condition was run, the headline figures are best read as relative rankings of probes, not calibrated success rates; adding catch trials with no matching image would tell where the true floor sits.
- The same probe battery could be run on other image generators to map whether the relation-frequency and count-collapse patterns are universal or specific to DALL-E 3.
- The approximate-numeracy finding suggests generative models could serve as a platform for studying the emergence of number representations, but the detection-model-based counts need validation against human counting on the same images.
- If relation success is driven by N-gram frequency, benchmarks should balance prompt frequencies before comparing models, otherwise reported compositional gains may be frequency effects in disguise.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports four behavioral experiments (total N=178) plus auxiliary analyses in which human participants judged whether images generated by DALL·E 3 (and, in Experiment 4, an LLM-grounded diffusion pipeline, LMD+) matched prompts built on physical relations, negations, and exact numbers. The headline claim is that no condition reliably produces human agreement scores above 50%, with relations at about 45% agreement, unmodified negation at about 12.3%, and number agreement declining from roughly 75% for 'one' to about 9% for 'six'. A follow-up auto-count analysis using twelve object detectors reports scalar variability and ratio dependence in DALL·E 3's generated counts, and an n-gram frequency analysis connects prompt frequency to perceived match. The paper interprets these results as evidence of persistent limitations in compositional and logical control in current text-to-image models.
Significance. The question of whether modern text-to-image models can handle logical operators is timely and important, and this paper contributes a simple, transparent probe set with open data and code, attention-checked human judgments, and a direct comparison between DALL·E 3 and a grounded-diffusion pipeline. The negation failure and the sharp decline for larger numbers are well supported by the human data. However, the headline 'none greater than 50%' claim is not calibrated against any chance-level baseline, and it is internally contradicted by the 75% agreement for 'one' in Experiment 3. The auto-count analysis is auxiliary rather than load-bearing for the central human-judgment claim, but its detector-based counts are not validated against human counts and should not be treated as strong evidence on their own. The paper's strengths are its clear experimental design, replication-oriented presentation, and open materials; its main weaknesses are the ungrounded absolute threshold and an overbroad abstract.
major comments (3)
- [Abstract and Experiment 3 (Results)] The abstract's statement that 'none reliably produce human agreement scores greater than 50%' is contradicted by the paper's own Experiment 3 result that the average agreement for 'one' entity is approximately 75%. If 'none' is meant to range over all prompt-level averages, the claim is false; if it is meant only for negation and numbers beyond three, the abstract must be rewritten to say so. This sentence is the central takeaway of the paper, so the mismatch between the abstract and the reported data should be fixed before publication.
- [Experimental Design and Experiment 1] The agreement measure lacks a chance-level or baseline condition. Participants were allowed to select all, some, or none of the 18 images per grid, so the absolute value of the reported agreement percentage depends on participants' overall selection rate under uncertainty. Without a neutral-prompt condition or a grid of images that are not matched to the prompt, the values 45% for relations and 12.3% for negation cannot be interpreted as being above or below chance, and the 50% threshold used throughout the abstract is not calibrated. The authors should add a baseline condition, such as prompts paired with random or clearly mismatched image grids, and report the distribution of selection counts over trials.
- [A.4, Table A.1, and Figure 7] The auto-count analysis treats the counts produced by twelve object detectors as reliable estimates of the number of objects in each generated image, but the reported inter-rater agreement among these detectors is near zero (mean Cohen's kappa between -0.001 and 0.038). With this level of agreement, the estimated scalar-variability slope of 1.7 and the ratio-dependence breakpoint of approximately 3.33 are not trustworthy measures of DALL·E 3's output counts. At minimum, a subset of images should be counted by human annotators so that detector counts can be validated, and the analysis should report agreement between detector counts and human counts before drawing conclusions about 'approximate numeracy'.
minor comments (4)
- [Appendix A.1] The section on 'Experiments 5-9' says '119 participants were recruited for Experiment 4'; this should refer to Experiments 5-9, since Experiment 4 is the LMD+ study with 30 participants. The total N=178 in the abstract does not include these 119 participants, so the numbering and the participant totals should be clarified.
- [Experiment 4 (Results)] The LMD+ comparison reports an average agreement of 29.4% without confidence intervals, whereas the earlier experiments report means with 95% intervals; add the intervals for this condition so the reader can assess the precision of the comparison.
- [Figure 7 caption] The caption says 'A shows a density plot' and 'B provides further detail', but the panel labels are not explicitly marked in the described figure; please label the panels A and B in the figure itself and describe what the axes of each panel represent.
- [Experiment 2 (Methods)] The categorization of modified negation prompts into Replacement, Addition, and No Change is based on the authors' own judgment, with approximate percentages but no inter-rater reliability or coding scheme; report a second coder and agreement, or explicitly describe the categorization as informal.
Circularity Check
No significant circularity: the human-agreement results are direct measurements, the auxiliary analyses are post-hoc descriptions, and self-citations are contextual rather than load-bearing.
full rationale
This paper is an empirical evaluation rather than a derivation. The central results—relations at 45% agreement, plain negation at 12.3%, and number agreement falling from ~75% for 'one' to ~9% for 'six'—are direct human ratings of generated images against prompted logical-operator sentences. Those percentages are the evidence, not the output of a fitted equation, and they do not reduce to any parameter of the paper. The auxiliary analyses (the Google N-gram frequency correlation of rho = 0.43 and the scalar-variability/ratio-dependence fits) are descriptive, post-hoc interpretations of the same empirical results; they are not used as inputs to generate the headline claim, so no fitted input is relabeled as a prediction. The paper's self-citations, chiefly Conwell & Ullman (2022), are used for prompt-design continuity and as a comparison point to DALL·E 2; those citations are contextual and do not carry the argument. Even where prior work is cited to justify excluding 'agentic' relations, that is a design rationale, not a proof step. The abstract's blanket 'none greater than 50%' is internally inconsistent with the ~75% agreement for 'one,' and the lack of a chance-level baseline makes the 50% threshold hard to interpret; these are methodological and framing concerns about the strength of the claim, not circularity. No equation or definition is shown to be equivalent to its own input, and no self-citation is invoked as an unverified substitute for the new human-judgment data.
Assumptions & free parameters
free parameters (2)
- scalar variability slope (log variance ~ log mean) =
1.7 [1.48, 1.91]
- ratio dependence breakpoint =
3.33 [3.01, 3.72]
assumptions (3)
- domain assumption Human agreement proportion is a valid, interpretable measure of image-prompt match without a baseline condition.
- domain assumption Google Books N-gram frequency approximates the training-data frequency of the visual/textual concepts in DALL-E 3.
- domain assumption The bank of 12 object detection models provides valid counts of objects in generated images for the approximate numeracy analysis.
Cite this review
Pith. "Pith review of Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models." pith.science (2026). https://pith.science/paper/727RCSC4
@misc{pith2026241117066,
author = {Pith},
title = {Pith review of: Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/727RCSC4}},
note = {Machine review of arXiv:2411.17066}
}
read the original abstract
Despite remarkable progress in multi-modal AI research, there is a salient domain in which modern AI continues to lag considerably behind even human children: the reliable deployment of logical operators. Here, we examine three forms of logical operators: relations, negations, and discrete numbers. We asked human respondents (N=178 in total) to evaluate images generated by a state-of-the-art image-generating AI (DALL-E 3) prompted with these `logical probes', and find that none reliably produce human agreement scores greater than 50\%. The negation probes and numbers (beyond 3) fail most frequently. In a 4th experiment, we assess a `grounded diffusion' pipeline that leverages targeted prompt engineering and structured intermediate representations for greater compositional control, but find its performance is judged even worse than that of DALL-E 3 across prompts. To provide further clarity on potential sources of success and failure in these text-to-image systems, we supplement our 4 core experiments with multiple auxiliary analyses and schematic diagrams, directly quantifying, for example, the relationship between the N-gram frequency of relational prompts and the average match to generated images; the success rates for 3 different prompt modification strategies in the rendering of negation prompts; and the scalar variability / ratio dependence (`approximate numeracy') of prompts involving integers. We conclude by discussing the limitations inherent to `grounded' multimodal learning systems whose grounding relies heavily on vector-based semantics (e.g. DALL-E 3), or under-specified syntactical constraints (e.g. `grounded diffusion'), and propose minimal modifications (inspired by development, based in imagery) that could help to bridge the lingering compositional gap between scale and structure. All data and code is available at https://github.com/ColinConwell/T2I-Probology
Figures
Figures from the paper (6 more)
Forward citations
Cited by 1 Pith paper
-
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
A new evaluation framework, DIMCIM, measures default-mode diversity and prompted generalization in text-to-image models, finding a scale trade-off and a 0.85 correlation between default diversity and training data diversity.
Reference graph
Works this paper leans on
-
[1]
Analogs of linguistic structure in deep representations
Jacob Andreas and Dan Klein. Analogs of linguistic structure in deep representations. arXiv preprint arXiv:1707.08139, 2017
arXiv 2017
-
[2]
Improving image generation with better captions
James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023
2023
-
[3]
Language development: Form and function in emerging grammars
Lois Masket Bloom. Language development: Form and function in emerging grammars. ERIC, 1968
1968
-
[4]
brms: Bayesian Regression Models using Stan, 2024
Paul-Christian B¨urkner. brms: Bayesian Regression Models using Stan, 2024. URL https: //github.com/paul-buerkner/brms. R package version 2.21.0
2024
-
[5]
Bootstrapping & the origin of concepts
Susan Carey. Bootstrapping & the origin of concepts. Daedalus, 133(1):59–68, 2004
2004
-
[6]
Ontogenetic origins of human integer representations
Susan Carey and David Barner. Ontogenetic origins of human integer representations. Trends in cognitive sciences, 23(10):823–835, 2019
2019
-
[7]
Topological structure in visual perception
Lin Chen. Topological structure in visual perception. Science, 218(4573):699–700, 1982
1982
-
[8]
A unified account of numerosity perception.Nature human behaviour, 4(12):1265–1272, 2020
Samuel J Cheyette and Steven T Piantadosi. A unified account of numerosity perception.Nature human behaviour, 4(12):1265–1272, 2020
2020
Show all 94 references
-
[9]
Aspects of the Theory of Syntax
Noam Chomsky. Aspects of the Theory of Syntax. Number 11. MIT press, 2014
2014
-
[10]
Alternative representations of time, number, and rate
Russell M Church and Hilary A Broadbent. Alternative representations of time, number, and rate. Cognition, 37(1-2):55–81, 1990
1990
-
[11]
Testing relational understanding in text-guided image generation
Colin Conwell and Tomer Ullman. Testing relational understanding in text-guided image generation. arXiv preprint arXiv:2208.00005, 2022. 14
2022 arXiv
-
[12]
Vqgan-clip: Open domain image generation and editing with natural language guidance
Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. Vqgan-clip: Open domain image generation and editing with natural language guidance. arXiv preprint arXiv:2204.08583, 2022
2022 arXiv
-
[13]
de Saint-Exup ´ery and I
A. de Saint-Exup ´ery and I. Testot-Ferry. The Little Prince . A Harvest/HBJ book. Wordsworth, 1995. ISBN 9781853261589. URL https://books.google.com/books? id=CQYg20lTHtMC
1995
-
[14]
The number sense: How the mind creates mathematics
Stanislas Dehaene. The number sense: How the mind creates mathematics. OUP USA, 2011
2011
-
[15]
Development of elementary numerical abilities: A neuronal model
Stanislas Dehaene and Jean-Pierre Changeux. Development of elementary numerical abilities: A neuronal model. Journal of cognitive neuroscience, 5(4):390–407, 1993
1993
-
[16]
Three parietal circuits for number processing
Stanislas Dehaene, Manuela Piazza, Philippe Pinel, and Laurent Cohen. Three parietal circuits for number processing. In The handbook of mathematical cognition, pages 433–453. Psychology Press, 2005
2005
-
[17]
Describing scenes hardly seen
Christian Dobel, Heidi Gumnior, Jens B¨olte, and Pienie Zwitserlood. Describing scenes hardly seen. Acta psychologica, 125(2):129–143, 2007
2007
-
[18]
Relate: Physically plausible multi-object scene synthesis using structured latent spaces
S´ebastien Ehrhardt, Oliver Groth, Aron Monszpart, Martin Engelcke, Ingmar Posner, Niloy Mitra, and Andrea Vedaldi. Relate: Physically plausible multi-object scene synthesis using structured latent spaces. Advances in Neural Information Processing Systems, 33:11202–11213, 2020
2020
-
[19]
Cultural constraints on grammar and cognition in pirah ˜a: Another look at the design features of human language
DanielL Everett. Cultural constraints on grammar and cognition in pirah ˜a: Another look at the design features of human language. Current anthropology, 46(4):621–646, 2005
2005
-
[20]
Perspective (in) consistency of paint by text
Hany Farid. Perspective (in) consistency of paint by text. arXiv preprint arXiv:2206.14617, 2022
2022 arXiv
-
[21]
no” and “not
Roman Feiman, Shilpa Mody, Sophia Sanborn, and Susan Carey. What do you mean, no? toddlers’ comprehension of logical “no” and “not”. Language Learning and Development, 13 (4):430–450, 2017
2017
-
[22]
Seeing physics in the blink of an eye
Chaz Firestone and Brian Scholl. Seeing physics in the blink of an eye. Journal of Vision, 17 (10):203–203, 2017
2017
-
[23]
Bridging the data gap between children and large language models
Michael C Frank. Bridging the data gap between children and large language models. Trends in Cognitive Sciences, 2023
2023
-
[24]
The psychophysics of chasing: A case study in the perception of animacy
Tao Gao, George E Newman, and Brian J Scholl. The psychophysics of chasing: A case study in the perception of animacy. Cognitive psychology, 59(2):154–179, 2009
2009
-
[25]
Mapping the early language environment using all-day recordings and automated analysis
Jill Gilkerson, Jeffrey A Richards, Steven F Warren, Judith K Montgomery, Charles R Green- wood, D Kimbrough Oller, John HL Hansen, and Terrance D Paul. Mapping the early language environment using all-day recordings and automated analysis. American journal of speech- language...
2017
-
[26]
Rapid apprehension of the coherence of action scenes
Reinhild Glanemann, Pienie Zwitserlood, Jens B¨olte, and Christian Dobel. Rapid apprehension of the coherence of action scenes. Psychonomic bulletin & review, 23(5):1566–1575, 2016
2016
-
[27]
It’s not just what we don’t know: The mapping problem in the acquisition of negation
Victor Gomes, Rebecca Doherty, Daniel Smits, Susan Goldin-Meadow, John C Trueswell, and Roman Feiman. It’s not just what we don’t know: The mapping problem in the acquisition of negation. Cognitive Psychology, 145:101592, 2023
2023
-
[28]
Seeing what’s possible: Disconnected visual parts are confused for their potential wholes
Chenxiao Guan and Chaz Firestone. Seeing what’s possible: Disconnected visual parts are confused for their potential wholes. Journal of experimental psychology: general, 149(3):590, 2020
2020
-
[29]
The perception of relations
Alon Hafri and Chaz Firestone. The perception of relations. Trends in Cognitive Sciences, 25 (6):475–492, 2021
2021
-
[30]
A phone in a basket looks like a knife in a cup: The perception of abstract relations
Alon Hafri, Michael F Bonner, Barbara Landau, and Chaz Firestone. A phone in a basket looks like a knife in a cup: The perception of abstract relations. PsyArXiv, 2020. 15
2020
-
[31]
Number sense across the lifespan as revealed by a massive internet-based sample
Justin Halberda, Ryan Ly, Jeremy B Wilmer, Daniel Q Naiman, and Laura Germine. Number sense across the lifespan as revealed by a massive internet-based sample. Proceedings of the National Academy of Sciences, 109(28):11116–11120, 2012
2012
-
[32]
Social evaluation by preverbal infants
J Kiley Hamlin, Karen Wynn, and Paul Bloom. Social evaluation by preverbal infants. Nature, 450(7169):557–559, 2007
2007
-
[33]
Conceptual precursors to language
Susan J Hespos and Elizabeth S Spelke. Conceptual precursors to language. Nature, 430(6998): 453–456, 2004
2004
-
[34]
Structural ambiguity and lexical relations
Donald Hindle and Mats Rooth. Structural ambiguity and lexical relations. Computational linguistics, 19(1):103–120, 1993
1993
-
[35]
A natural history of negation
Laurence R Horn. A natural history of negation. University of Chicago Press, 1989
1989
-
[36]
T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation. Advances in Neural Information Processing Systems, 36:78723–78747, 2023
2023
-
[37]
On the origins of denial negation
Peter Hummer, Heinz Wimmer, and Gertraud Antes. On the origins of denial negation. Journal of child language, 20(3):607–618, 1993
1993
-
[38]
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognitio...
2017
-
[39]
Understanding negation: Issues in the processing of negation
Barbara Kaup and Carolin Dudschig. Understanding negation: Issues in the processing of negation. Oxford University Press, 2020
2020
-
[40]
The experiential view of language comprehen- sion: How is negation represented
Barbara Kaup, Rolf A Zwaan, and Jana L¨udtke. The experiential view of language comprehen- sion: How is negation represented. Higher level language processes in the brain: Inference and comprehension processes, pages 255–288, 2007
2007
-
[41]
Perception of partly occluded objects in infancy
Philip J Kellman and Elizabeth S Spelke. Perception of partly occluded objects in infancy. Cognitive psychology, 15(4):483–524, 1983
1983
-
[42]
Visual number sense in untrained deep neural networks
Gwangsu Kim, Jaeson Jang, Seungdae Baek, Min Song, and Se-Bum Paik. Visual number sense in untrained deep neural networks. Science advances, 7(1):eabd6127, 2021
2021
-
[43]
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253, 2017
2017
-
[44]
Holistic evaluation of text-to-image models
Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, et al. Holistic evaluation of text-to-image models. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[45]
Gligen: Open-set grounded text-to-image generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. Gligen: Open-set grounded text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22511–22521, 2023
2023
-
[46]
Llm-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models
Long Lian, Boyi Li, Adam Yala, and Trevor Darrell. Llm-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models. arXiv preprint arXiv:2305.13655, 2023
2023 arXiv
-
[47]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...
2014
-
[48]
Compositional visual generation with composable diffusion models
Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. arXiv preprint arXiv:2206.01714, 2022. 16
2022 arXiv
-
[49]
Training priors predict text-to-image model performance
Charles Lovering and Ellie Pavlick. Training priors predict text-to-image model performance. arXiv preprint arXiv:2306.01755, 2023
2023 arXiv
-
[50]
Topological relations between objects are categorically coded
Andrew Lovett and Steven L Franconeri. Topological relations between objects are categorically coded. Psychological science, 28(10):1408–1418, 2017
2017
-
[51]
sjPlot: Data Visualization for Statistics in Social Science , 2024
Daniel L ¨udecke. sjPlot: Data Visualization for Statistics in Social Science , 2024. URL https://strengejacke.github.io/sjPlot/. R package version 2.8.16
2024
-
[52]
Dissociating language and thought in large language models
Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. Dissociating language and thought in large language models. Trends in Cognitive Sciences, 2024
2024
-
[53]
A very preliminary analysis of dall-e 2
Gary Marcus, Ernest Davis, and Scott Aaronson. A very preliminary analysis of dall-e 2. arXiv preprint arXiv:2204.13807, 2022
2022 arXiv
-
[54]
A mode control model of counting and timing processes
Warren H Meck and Russell M Church. A mode control model of counting and timing processes. Journal of experimental psychology: animal behavior processes, 9(3):320, 1983
1983
-
[55]
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal. The illusion of state in state-space models. arXiv preprint arXiv:2404.08819, 2024
2024 arXiv
-
[56]
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229, 2024
-
[57]
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021
2021 arXiv
-
[58]
Compo- sitional text-to-image generation with dense blob representations
Weili Nie, Sifei Liu, Morteza Mardani, Chao Liu, Benjamin Eckart, and Arash Vahdat. Compo- sitional text-to-image generation with dense blob representations. In International Conference on Machine Learning (ICML), 2024
2024
-
[59]
An introduction to the approximate number system
Darko Odic and Ariel Starr. An introduction to the approximate number system. Child Development Perspectives, 12(4):223–229, 2018
2018
-
[60]
Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task
Maya Okawa, Ekdeep S Lubana, Robert Dick, and Hidenori Tanaka. Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[61]
Numerical cognition in bees and other insects
Mario Pahl, Aung Si, and Shaowu Zhang. Numerical cognition in bees and other insects. Frontiers in psychology, 4:33732, 2013
2013
-
[62]
The human imagination: the cognitive neuroscience of visual mental imagery
Joel Pearson. The human imagination: the cognitive neuroscience of visual mental imagery. Nature Reviews Neuroscience, 20(10):624–634, 2019
2019
-
[63]
Beyond the turk: Alternative platforms for crowdsourcing behavioral research
Eyal Peer, Laura Brandimarte, Sonam Samat, and Alessandro Acquisti. Beyond the turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology, 70:153–163, 2017
2017
-
[64]
Siemenn, Saisamrit Surbehera, Zad Chin, Keith Tyser, Gregory Hunter, Arvind Raghavan, Yann Hicke, Bryan A
Vitali Petsiuk, Alexander E. Siemenn, Saisamrit Surbehera, Zad Chin, Keith Tyser, Gregory Hunter, Arvind Raghavan, Yann Hicke, Bryan A. Plummer, Ori Kerret, Tonio Buonassisi, Kate Saenko, Armando Solar-Lezama, and Iddo Drori. Human evaluation of text-to-image models on a multi...
2022 arXiv
-
[65]
Exact number concepts are limited to the verbal count range
Benjamin Pitt, Edward Gibson, and Steven T Piantadosi. Exact number concepts are limited to the verbal count range. Psychological Science, 33(3):371–381, 2022
2022
-
[66]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pa...
2021
-
[67]
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022
2022 arXiv
-
[68]
Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46, 2024
Sunayana Rane, Alexander Ku, Jason Baldridge, Ian Tenney, Tom Griffiths, and Been Kim. Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46, 2024
2024
-
[69]
Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment
Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, and Gal Chechik. Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[70]
Impact of pretraining term frequencies on few-shot reasoning
Yasaman Razeghi, Robert L Logan IV , Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022
2022 arXiv
-
[71]
High- resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High- resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022
2022
-
[72]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22500...
2023
-
[73]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv...
2022 arXiv
-
[74]
Verification of text ideas during reading
Murray Singer. Verification of text ideas during reading. Journal of Memory and Language, 54 (4):574–591, 2006
2006
-
[75]
Objectstitch: Object compositing with diffusion model
Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian Price, Jianming Zhang, Soo Ye Kim, and Daniel Aliaga. Objectstitch: Object compositing with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18310–18319, 2023
2023
-
[76]
Initial knowledge: Six suggestions
Elizabeth Spelke. Initial knowledge: Six suggestions. Cognition, 50(1-3):431–445, 1994
1994
-
[77]
What babies know: Core knowledge and composition volume 1, volume 1
Elizabeth S Spelke. What babies know: Core knowledge and composition volume 1, volume 1. Oxford University Press, 2022
2022
-
[78]
Core knowledge
Elizabeth S Spelke and Katherine D Kinzler. Core knowledge. Developmental science, 10(1): 89–96, 2007
2007
-
[79]
Event completion: Event based inferences distort memory in a matter of seconds
Brent Strickland and Frank Keil. Event completion: Event based inferences distort memory in a matter of seconds. Cognition, 121(3):409–415, 2011
2011
-
[80]
What formal lan- guages can transformers express? a survey
Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin. What formal lan- guages can transformers express? a survey. Transactions of the Association for Computational Linguistics, 12:543–561, 2024
2024
-
[81]
Lexicalization patterns: Semantic structure in lexical forms.Language typology and syntactic description, 3(99):36–149, 1985
Leonard Talmy. Lexicalization patterns: Semantic structure in lexical forms.Language typology and syntactic description, 3(99):36–149, 1985
1985
-
[82]
Winoground: Probing vision and language models for visio-linguistic compositionality
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...
2022
-
[83]
Computing machinery and intelligence
Alan M Turing. Computing machinery and intelligence. Mind, 1950
1950
-
[84]
No” zero-shot” without exponential data: Pretraining con- cept frequency determines multimodal model performance
Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma, Philip HS Torr, Adel Bibi, Samuel Albanie, and Matthias Bethge. No” zero-shot” without exponential data: Pretraining con- cept frequency determines multimodal model performance. arXiv preprint arXiv:2404.04125, 2024. 18
2024 arXiv
-
[85]
Help or hinder: Bayesian models of social goal inference
Tomer Ullman, Chris Baker, Owen Macindoe, Owain Evans, Noah Goodman, and Joshua Tenenbaum. Help or hinder: Bayesian models of social goal inference. Advances in neural information processing systems, 22, 2009
2009
-
[86]
The automaticity of perceiving animacy: Goal-directed motion in simple shapes influences visuomotor behavior even when task-irrelevant
Benjamin van Buren, Stefan Uddenberg, and Brian J Scholl. The automaticity of perceiving animacy: Goal-directed motion in simple shapes influences visuomotor behavior even when task-irrelevant. Psychonomic bulletin & review, 23(3):797–802, 2016
2016
-
[87]
Response to affirmative and negative binary statements
Peter C Wason. Response to affirmative and negative binary statements. British Journal of Psychology, 52(2):133–142, 1961
1961
-
[88]
Paradoxical effects of thought suppression
Daniel M Wegner, David J Schneider, Samuel R Carter, and Teri L White. Paradoxical effects of thought suppression. Journal of personality and social psychology, 53(1):5, 1987
1987
-
[89]
Conceptmix: A compositional image generation benchmark with controllable difficulty
Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, and Sanjeev Arora. Conceptmix: A compositional image generation benchmark with controllable difficulty. arXiv preprint arXiv:2408.14339, 2024
2024 arXiv
-
[90]
Boxdiff: Text-to-image synthesis with training-free box-constrained diffu- sion
Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou. Boxdiff: Text-to-image synthesis with training-free box-constrained diffu- sion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7452–7461, 2023
2023
-
[91]
Diffusion models: A comprehensive survey of methods and applications
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys, 56(4):1–39, 2023
2023
-
[92]
Perceiving fully occluded objects via physical simulation
Ilker Yildirim, Max H Siegel, and Joshua B Tenenbaum. Perceiving fully occluded objects via physical simulation. In Proceedings of the 38th annual conference of the cognitive science society, 2016
2016
-
[93]
When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2022
Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2022
2022
-
[94]
I NEED to test how the tool works with extremely simple prompts. DO NOT add any detail, just use it AS-IS:
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. A Appendix (Supplementary Information) A.1 Details on Behavioral Experi...
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.