Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read DALL-E 3 fails most logic prompts, say human judges

desk verdict Solid human evaluation of DALL-E 3's logic failures, but the abstract overclaims and an uncalibrated threshold plus unreliable auto-counts need fixing. read the letter →

arxiv 2411.17066 v1 pith:727RCSC4 submitted 2024-11-26 cs.CV cs.CLcs.SC

classification cs.CVcs.CLcs.SC
keywords text-to-imagegenerationlogicaloperatorsDALL-E3compositionalhumanevaluationnegationnumbercognitionrelations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a state-of-the-art text-to-image system can do what young children do without effort: turn simple logical descriptions into pictures. Using human judges who pick matching images from 18-image grids, it finds that prompts built on physical relations, plain negation, and exact integers mostly do not produce images people agree match. Relations average 45% agreement, plain negation 12.3%, and counts fall from about 75% for "one" to about 9% for "six"; a grounded diffusion pipeline with scene-graph layouts scores even lower on the same relation prompts. The point of the exercise is not just to grade one model, but to show that logical composition remains a distinct failure mode in systems whose gains come from scale and vector-based grounding.

What carries the argument

The carrying object is the logical probe: a minimal prompt that pairs an everyday object with one relation, one negation, or one integer, rendered 18 times per prompt. The measuring instrument is human agreement—participants select all, some, or none of the 18 images as matching the prompt. Around this, the paper adds three analytic devices: a fixed prefix that stops the model's internal language model from rewriting a prompt, so the raw model is tested; an N-gram frequency correlation that tests whether relational success tracks training-data statistics; and an approximate-numeracy analysis using object-detection counts to estimate scalar variability and ratio dependence in generated quantities.

What would settle it

Run the same 18-image selection task with deliberately non-matching prompts as catch trials and compare selection rates: if people endorse non-matching images at or above the rates observed for relations (45%), the claim that relations fail falls to a measurement artifact. A second check would ask humans to count objects in images from the auto-count follow-up and compare their counts to the detection-model estimates.

Watch

Extended reading notes

Core claim

The paper's central claim is that DALL-E 3 does not reliably deploy basic logical operators in image generation, and that no probe family—relations, negations, or numbers—consistently clears the 50% human-agreement mark. Negation fails most completely: when the prompt is left unmodified, it almost always renders the very object it forbids. Number generation is exact for small counts and approximate beyond three, with agreement collapsing from 75% at "one" to 9% at "six"; a follow-up analysis with object detectors finds ratio-dependent, scalar-variable error patterns like an approximate number system rather than exact counting. The paper also claims that a grounded pipeline using structured layout graphs does not fix these problems and is judged worse than DALL-E 3 on the same relation prompts, because its intermediate representations miss physics such as occlusion-in-depth.

Load-bearing premise

The headline "none greater than 50%" treats 50% agreement as the success threshold, but the task offered no baseline or chance condition showing how often people endorse images when a prompt has no valid match.

Editorial extensions

If this is right

  • Text-to-image "prompt following" should not be read as compositional understanding; simple operator probes reveal systematic gaps that survive scale.
  • Plain negation is not just imperfect but inverted: an unmodified "not X" prompt tends to produce X, which means systems need an explicit rewriting step that replaces the forbidden object.
  • Exact count generation beyond three or four objects is outside current capability, and the error pattern is ratio-dependent rather than exact-integer-like.
  • Structured intermediate representations such as scene graphs can hurt rather than help when they omit physical constraints like occlusion, so grounding alone is not sufficient.
  • Prompt frequency predicts relational success, implying that reported gains in relations may reflect dataset statistics rather than relational reasoning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because no baseline or chance condition was run, the headline figures are best read as relative rankings of probes, not calibrated success rates; adding catch trials with no matching image would tell where the true floor sits.
  • The same probe battery could be run on other image generators to map whether the relation-frequency and count-collapse patterns are universal or specific to DALL-E 3.
  • The approximate-numeracy finding suggests generative models could serve as a platform for studying the emergence of number representations, but the detection-model-based counts need validation against human counting on the same images.
  • If relation success is driven by N-gram frequency, benchmarks should balance prompt frequencies before comparing models, otherwise reported compositional gains may be frequency effects in disguise.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports four behavioral experiments (total N=178) plus auxiliary analyses in which human participants judged whether images generated by DALL·E 3 (and, in Experiment 4, an LLM-grounded diffusion pipeline, LMD+) matched prompts built on physical relations, negations, and exact numbers. The headline claim is that no condition reliably produces human agreement scores above 50%, with relations at about 45% agreement, unmodified negation at about 12.3%, and number agreement declining from roughly 75% for 'one' to about 9% for 'six'. A follow-up auto-count analysis using twelve object detectors reports scalar variability and ratio dependence in DALL·E 3's generated counts, and an n-gram frequency analysis connects prompt frequency to perceived match. The paper interprets these results as evidence of persistent limitations in compositional and logical control in current text-to-image models.

Significance. The question of whether modern text-to-image models can handle logical operators is timely and important, and this paper contributes a simple, transparent probe set with open data and code, attention-checked human judgments, and a direct comparison between DALL·E 3 and a grounded-diffusion pipeline. The negation failure and the sharp decline for larger numbers are well supported by the human data. However, the headline 'none greater than 50%' claim is not calibrated against any chance-level baseline, and it is internally contradicted by the 75% agreement for 'one' in Experiment 3. The auto-count analysis is auxiliary rather than load-bearing for the central human-judgment claim, but its detector-based counts are not validated against human counts and should not be treated as strong evidence on their own. The paper's strengths are its clear experimental design, replication-oriented presentation, and open materials; its main weaknesses are the ungrounded absolute threshold and an overbroad abstract.

major comments (3)
  1. [Abstract and Experiment 3 (Results)] The abstract's statement that 'none reliably produce human agreement scores greater than 50%' is contradicted by the paper's own Experiment 3 result that the average agreement for 'one' entity is approximately 75%. If 'none' is meant to range over all prompt-level averages, the claim is false; if it is meant only for negation and numbers beyond three, the abstract must be rewritten to say so. This sentence is the central takeaway of the paper, so the mismatch between the abstract and the reported data should be fixed before publication.
  2. [Experimental Design and Experiment 1] The agreement measure lacks a chance-level or baseline condition. Participants were allowed to select all, some, or none of the 18 images per grid, so the absolute value of the reported agreement percentage depends on participants' overall selection rate under uncertainty. Without a neutral-prompt condition or a grid of images that are not matched to the prompt, the values 45% for relations and 12.3% for negation cannot be interpreted as being above or below chance, and the 50% threshold used throughout the abstract is not calibrated. The authors should add a baseline condition, such as prompts paired with random or clearly mismatched image grids, and report the distribution of selection counts over trials.
  3. [A.4, Table A.1, and Figure 7] The auto-count analysis treats the counts produced by twelve object detectors as reliable estimates of the number of objects in each generated image, but the reported inter-rater agreement among these detectors is near zero (mean Cohen's kappa between -0.001 and 0.038). With this level of agreement, the estimated scalar-variability slope of 1.7 and the ratio-dependence breakpoint of approximately 3.33 are not trustworthy measures of DALL·E 3's output counts. At minimum, a subset of images should be counted by human annotators so that detector counts can be validated, and the analysis should report agreement between detector counts and human counts before drawing conclusions about 'approximate numeracy'.
minor comments (4)
  1. [Appendix A.1] The section on 'Experiments 5-9' says '119 participants were recruited for Experiment 4'; this should refer to Experiments 5-9, since Experiment 4 is the LMD+ study with 30 participants. The total N=178 in the abstract does not include these 119 participants, so the numbering and the participant totals should be clarified.
  2. [Experiment 4 (Results)] The LMD+ comparison reports an average agreement of 29.4% without confidence intervals, whereas the earlier experiments report means with 95% intervals; add the intervals for this condition so the reader can assess the precision of the comparison.
  3. [Figure 7 caption] The caption says 'A shows a density plot' and 'B provides further detail', but the panel labels are not explicitly marked in the described figure; please label the panels A and B in the figure itself and describe what the axes of each panel represent.
  4. [Experiment 2 (Methods)] The categorization of modified negation prompts into Replacement, Addition, and No Change is based on the authors' own judgment, with approximate percentages but no inter-rater reliability or coding scheme; report a second coder and agreement, or explicitly describe the categorization as informal.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the human-agreement results are direct measurements, the auxiliary analyses are post-hoc descriptions, and self-citations are contextual rather than load-bearing.

full rationale

This paper is an empirical evaluation rather than a derivation. The central results—relations at 45% agreement, plain negation at 12.3%, and number agreement falling from ~75% for 'one' to ~9% for 'six'—are direct human ratings of generated images against prompted logical-operator sentences. Those percentages are the evidence, not the output of a fitted equation, and they do not reduce to any parameter of the paper. The auxiliary analyses (the Google N-gram frequency correlation of rho = 0.43 and the scalar-variability/ratio-dependence fits) are descriptive, post-hoc interpretations of the same empirical results; they are not used as inputs to generate the headline claim, so no fitted input is relabeled as a prediction. The paper's self-citations, chiefly Conwell & Ullman (2022), are used for prompt-design continuity and as a comparison point to DALL·E 2; those citations are contextual and do not carry the argument. Even where prior work is cited to justify excluding 'agentic' relations, that is a design rationale, not a proof step. The abstract's blanket 'none greater than 50%' is internally inconsistent with the ~75% agreement for 'one,' and the lack of a chance-level baseline makes the 50% threshold hard to interpret; these are methodological and framing concerns about the strength of the claim, not circularity. No equation or definition is shown to be equivalent to its own input, and no self-citation is invoked as an unverified substitute for the new human-judgment data.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The paper is an empirical evaluation, not a derivation; the only fitted quantities are descriptive parameters in auxiliary analyses.

free parameters (2)
  • scalar variability slope (log variance ~ log mean) = 1.7 [1.48, 1.91]
    Fitted via Bayesian hierarchical mixed-effects model on object-detector counts (Appendix A.4, Table A.2); used to claim approximate numeracy.
  • ratio dependence breakpoint = 3.33 [3.01, 3.72]
    Fitted via single-breakpoint segmented regression on pairwise distribution overlaps of detector counts (Appendix A.4).
assumptions (3)
  • domain assumption Human agreement proportion is a valid, interpretable measure of image-prompt match without a baseline condition.
    The paper's conclusions treat agreement above or below 50% as success/failure, but no control condition calibrates the selection rate.
  • domain assumption Google Books N-gram frequency approximates the training-data frequency of the visual/textual concepts in DALL-E 3.
    Used in Experiment 1 analysis (Section 'Analysis' and Figure A.2) to explain the relation between prompt frequency and perceived match.
  • domain assumption The bank of 12 object detection models provides valid counts of objects in generated images for the approximate numeracy analysis.
    Appendix A.4 uses these counts to estimate scalar variability and ratio dependence; Table A.1 shows near-zero inter-rater agreement (kappa 0.001-0.038), casting doubt on this assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models." pith.science (2026). https://pith.science/paper/727RCSC4

@misc{pith2026241117066,
  author       = {Pith},
  title        = {Pith review of: Relations, Negations, and Numbers: Looking for Logic in Generative Text-to-Image Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/727RCSC4}},
  note         = {Machine review of arXiv:2411.17066}
}
read the original abstract

Despite remarkable progress in multi-modal AI research, there is a salient domain in which modern AI continues to lag considerably behind even human children: the reliable deployment of logical operators. Here, we examine three forms of logical operators: relations, negations, and discrete numbers. We asked human respondents (N=178 in total) to evaluate images generated by a state-of-the-art image-generating AI (DALL-E 3) prompted with these `logical probes', and find that none reliably produce human agreement scores greater than 50\%. The negation probes and numbers (beyond 3) fail most frequently. In a 4th experiment, we assess a `grounded diffusion' pipeline that leverages targeted prompt engineering and structured intermediate representations for greater compositional control, but find its performance is judged even worse than that of DALL-E 3 across prompts. To provide further clarity on potential sources of success and failure in these text-to-image systems, we supplement our 4 core experiments with multiple auxiliary analyses and schematic diagrams, directly quantifying, for example, the relationship between the N-gram frequency of relational prompts and the average match to generated images; the success rates for 3 different prompt modification strategies in the rendering of negation prompts; and the scalar variability / ratio dependence (`approximate numeracy') of prompts involving integers. We conclude by discussing the limitations inherent to `grounded' multimodal learning systems whose grounding relies heavily on vector-based semantics (e.g. DALL-E 3), or under-specified syntactical constraints (e.g. `grounded diffusion'), and propose minimal modifications (inspired by development, based in imagery) that could help to bridge the lingering compositional gap between scale and structure. All data and code is available at https://github.com/ColinConwell/T2I-Probology

Figures

Figures reproduced from arXiv: 2411.17066 by the authors.

Figure 1
Figure 1. Example layout of a typical trial in all behavioral experiments: Participants were [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Results for Experiment 1 (Relations): participant agreement that images matched a [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Examples of ‘Modified’ and ‘Unmodified’ stimuli used in the Negation Experiment. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Results for Experiment 2 (Negation): Participant agreement that images matched a [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Illustrative examples of number progressions. Each row shows an entity class, and each [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Results for Experiment 3 (Numbers): participant agreement that images matched a [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Results from a follow-up analysis on DALL [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Examples of the intermediate ‘scene layout’ generated as part of the GLIGEN-augmented [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Results from an experiment with the use of a ‘grounded‘ diffusion pipeline that first [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new evaluation framework, DIMCIM, measures default-mode diversity and prompted generalization in text-to-image models, finding a scale trade-off and a 0.85 correlation between default diversity and training data diversity.

Reference graph

Works this paper leans on

94 extracted references · 52 canonical work pages · cited by 1 Pith paper

  1. [1]

    Analogs of linguistic structure in deep representations

    Jacob Andreas and Dan Klein. Analogs of linguistic structure in deep representations. arXiv preprint arXiv:1707.08139, 2017

  2. [2]

    Improving image generation with better captions

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023

  3. [3]

    Language development: Form and function in emerging grammars

    Lois Masket Bloom. Language development: Form and function in emerging grammars. ERIC, 1968

  4. [4]

    brms: Bayesian Regression Models using Stan, 2024

    Paul-Christian B¨urkner. brms: Bayesian Regression Models using Stan, 2024. URL https: //github.com/paul-buerkner/brms. R package version 2.21.0

  5. [5]

    Bootstrapping & the origin of concepts

    Susan Carey. Bootstrapping & the origin of concepts. Daedalus, 133(1):59–68, 2004

  6. [6]

    Ontogenetic origins of human integer representations

    Susan Carey and David Barner. Ontogenetic origins of human integer representations. Trends in cognitive sciences, 23(10):823–835, 2019

  7. [7]

    Topological structure in visual perception

    Lin Chen. Topological structure in visual perception. Science, 218(4573):699–700, 1982

  8. [8]

    A unified account of numerosity perception.Nature human behaviour, 4(12):1265–1272, 2020

    Samuel J Cheyette and Steven T Piantadosi. A unified account of numerosity perception.Nature human behaviour, 4(12):1265–1272, 2020

Show all 94 references
  1. [9]

    Aspects of the Theory of Syntax

    Noam Chomsky. Aspects of the Theory of Syntax. Number 11. MIT press, 2014

  2. [10]

    Alternative representations of time, number, and rate

    Russell M Church and Hilary A Broadbent. Alternative representations of time, number, and rate. Cognition, 37(1-2):55–81, 1990

  3. [11]

    Testing relational understanding in text-guided image generation

    Colin Conwell and Tomer Ullman. Testing relational understanding in text-guided image generation. arXiv preprint arXiv:2208.00005, 2022. 14

  4. [12]

    Vqgan-clip: Open domain image generation and editing with natural language guidance

    Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. Vqgan-clip: Open domain image generation and editing with natural language guidance. arXiv preprint arXiv:2204.08583, 2022

  5. [13]

    de Saint-Exup ´ery and I

    A. de Saint-Exup ´ery and I. Testot-Ferry. The Little Prince . A Harvest/HBJ book. Wordsworth, 1995. ISBN 9781853261589. URL https://books.google.com/books? id=CQYg20lTHtMC

  6. [14]

    The number sense: How the mind creates mathematics

    Stanislas Dehaene. The number sense: How the mind creates mathematics. OUP USA, 2011

  7. [15]

    Development of elementary numerical abilities: A neuronal model

    Stanislas Dehaene and Jean-Pierre Changeux. Development of elementary numerical abilities: A neuronal model. Journal of cognitive neuroscience, 5(4):390–407, 1993

  8. [16]

    Three parietal circuits for number processing

    Stanislas Dehaene, Manuela Piazza, Philippe Pinel, and Laurent Cohen. Three parietal circuits for number processing. In The handbook of mathematical cognition, pages 433–453. Psychology Press, 2005

  9. [17]

    Describing scenes hardly seen

    Christian Dobel, Heidi Gumnior, Jens B¨olte, and Pienie Zwitserlood. Describing scenes hardly seen. Acta psychologica, 125(2):129–143, 2007

  10. [18]

    Relate: Physically plausible multi-object scene synthesis using structured latent spaces

    S´ebastien Ehrhardt, Oliver Groth, Aron Monszpart, Martin Engelcke, Ingmar Posner, Niloy Mitra, and Andrea Vedaldi. Relate: Physically plausible multi-object scene synthesis using structured latent spaces. Advances in Neural Information Processing Systems, 33:11202–11213, 2020

  11. [19]

    Cultural constraints on grammar and cognition in pirah ˜a: Another look at the design features of human language

    DanielL Everett. Cultural constraints on grammar and cognition in pirah ˜a: Another look at the design features of human language. Current anthropology, 46(4):621–646, 2005

  12. [20]

    Perspective (in) consistency of paint by text

    Hany Farid. Perspective (in) consistency of paint by text. arXiv preprint arXiv:2206.14617, 2022

  13. [21]

    no” and “not

    Roman Feiman, Shilpa Mody, Sophia Sanborn, and Susan Carey. What do you mean, no? toddlers’ comprehension of logical “no” and “not”. Language Learning and Development, 13 (4):430–450, 2017

  14. [22]

    Seeing physics in the blink of an eye

    Chaz Firestone and Brian Scholl. Seeing physics in the blink of an eye. Journal of Vision, 17 (10):203–203, 2017

  15. [23]

    Bridging the data gap between children and large language models

    Michael C Frank. Bridging the data gap between children and large language models. Trends in Cognitive Sciences, 2023

  16. [24]

    The psychophysics of chasing: A case study in the perception of animacy

    Tao Gao, George E Newman, and Brian J Scholl. The psychophysics of chasing: A case study in the perception of animacy. Cognitive psychology, 59(2):154–179, 2009

  17. [25]

    Mapping the early language environment using all-day recordings and automated analysis

    Jill Gilkerson, Jeffrey A Richards, Steven F Warren, Judith K Montgomery, Charles R Green- wood, D Kimbrough Oller, John HL Hansen, and Terrance D Paul. Mapping the early language environment using all-day recordings and automated analysis. American journal of speech- language...

  18. [26]

    Rapid apprehension of the coherence of action scenes

    Reinhild Glanemann, Pienie Zwitserlood, Jens B¨olte, and Christian Dobel. Rapid apprehension of the coherence of action scenes. Psychonomic bulletin & review, 23(5):1566–1575, 2016

  19. [27]

    It’s not just what we don’t know: The mapping problem in the acquisition of negation

    Victor Gomes, Rebecca Doherty, Daniel Smits, Susan Goldin-Meadow, John C Trueswell, and Roman Feiman. It’s not just what we don’t know: The mapping problem in the acquisition of negation. Cognitive Psychology, 145:101592, 2023

  20. [28]

    Seeing what’s possible: Disconnected visual parts are confused for their potential wholes

    Chenxiao Guan and Chaz Firestone. Seeing what’s possible: Disconnected visual parts are confused for their potential wholes. Journal of experimental psychology: general, 149(3):590, 2020

  21. [29]

    The perception of relations

    Alon Hafri and Chaz Firestone. The perception of relations. Trends in Cognitive Sciences, 25 (6):475–492, 2021

  22. [30]

    A phone in a basket looks like a knife in a cup: The perception of abstract relations

    Alon Hafri, Michael F Bonner, Barbara Landau, and Chaz Firestone. A phone in a basket looks like a knife in a cup: The perception of abstract relations. PsyArXiv, 2020. 15

  23. [31]

    Number sense across the lifespan as revealed by a massive internet-based sample

    Justin Halberda, Ryan Ly, Jeremy B Wilmer, Daniel Q Naiman, and Laura Germine. Number sense across the lifespan as revealed by a massive internet-based sample. Proceedings of the National Academy of Sciences, 109(28):11116–11120, 2012

  24. [32]

    Social evaluation by preverbal infants

    J Kiley Hamlin, Karen Wynn, and Paul Bloom. Social evaluation by preverbal infants. Nature, 450(7169):557–559, 2007

  25. [33]

    Conceptual precursors to language

    Susan J Hespos and Elizabeth S Spelke. Conceptual precursors to language. Nature, 430(6998): 453–456, 2004

  26. [34]

    Structural ambiguity and lexical relations

    Donald Hindle and Mats Rooth. Structural ambiguity and lexical relations. Computational linguistics, 19(1):103–120, 1993

  27. [35]

    A natural history of negation

    Laurence R Horn. A natural history of negation. University of Chicago Press, 1989

  28. [36]

    T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation

    Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation. Advances in Neural Information Processing Systems, 36:78723–78747, 2023

  29. [37]

    On the origins of denial negation

    Peter Hummer, Heinz Wimmer, and Gertraud Antes. On the origins of denial negation. Journal of child language, 20(3):607–618, 1993

  30. [38]

    Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

    Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceedings of the IEEE conference on computer vision and pattern recognitio...

  31. [39]

    Understanding negation: Issues in the processing of negation

    Barbara Kaup and Carolin Dudschig. Understanding negation: Issues in the processing of negation. Oxford University Press, 2020

  32. [40]

    The experiential view of language comprehen- sion: How is negation represented

    Barbara Kaup, Rolf A Zwaan, and Jana L¨udtke. The experiential view of language comprehen- sion: How is negation represented. Higher level language processes in the brain: Inference and comprehension processes, pages 255–288, 2007

  33. [41]

    Perception of partly occluded objects in infancy

    Philip J Kellman and Elizabeth S Spelke. Perception of partly occluded objects in infancy. Cognitive psychology, 15(4):483–524, 1983

  34. [42]

    Visual number sense in untrained deep neural networks

    Gwangsu Kim, Jaeson Jang, Seungdae Baek, Min Song, and Se-Bum Paik. Visual number sense in untrained deep neural networks. Science advances, 7(1):eabd6127, 2021

  35. [43]

    Building machines that learn and think like people

    Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253, 2017

  36. [44]

    Holistic evaluation of text-to-image models

    Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, et al. Holistic evaluation of text-to-image models. Advances in Neural Information Processing Systems, 36, 2024

  37. [45]

    Gligen: Open-set grounded text-to-image generation

    Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. Gligen: Open-set grounded text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22511–22521, 2023

  38. [46]

    Llm-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

    Long Lian, Boyi Li, Adam Yala, and Trevor Darrell. Llm-grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models. arXiv preprint arXiv:2305.13655, 2023

  39. [47]

    Microsoft coco: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceeding...

  40. [48]

    Compositional visual generation with composable diffusion models

    Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. arXiv preprint arXiv:2206.01714, 2022. 16

  41. [49]

    Training priors predict text-to-image model performance

    Charles Lovering and Ellie Pavlick. Training priors predict text-to-image model performance. arXiv preprint arXiv:2306.01755, 2023

  42. [50]

    Topological relations between objects are categorically coded

    Andrew Lovett and Steven L Franconeri. Topological relations between objects are categorically coded. Psychological science, 28(10):1408–1418, 2017

  43. [51]

    sjPlot: Data Visualization for Statistics in Social Science , 2024

    Daniel L ¨udecke. sjPlot: Data Visualization for Statistics in Social Science , 2024. URL https://strengejacke.github.io/sjPlot/. R package version 2.8.16

  44. [52]

    Dissociating language and thought in large language models

    Kyle Mahowald, Anna A Ivanova, Idan A Blank, Nancy Kanwisher, Joshua B Tenenbaum, and Evelina Fedorenko. Dissociating language and thought in large language models. Trends in Cognitive Sciences, 2024

  45. [53]

    A very preliminary analysis of dall-e 2

    Gary Marcus, Ernest Davis, and Scott Aaronson. A very preliminary analysis of dall-e 2. arXiv preprint arXiv:2204.13807, 2022

  46. [54]

    A mode control model of counting and timing processes

    Warren H Meck and Russell M Church. A mode control model of counting and timing processes. Journal of experimental psychology: animal behavior processes, 9(3):320, 1983

  47. [55]

    The illusion of state in state-space models

    William Merrill, Jackson Petty, and Ashish Sabharwal. The illusion of state in state-space models. arXiv preprint arXiv:2404.08819, 2024

  48. [56]

    Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models

    Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel, Samy Bengio, and Mehrdad Farajtabar. Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. arXiv preprint arXiv:2410.05229, 2024

  49. [57]

    Glide: Towards photorealistic image generation and editing with text-guided diffusion models

    Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021

  50. [58]

    Compo- sitional text-to-image generation with dense blob representations

    Weili Nie, Sifei Liu, Morteza Mardani, Chao Liu, Benjamin Eckart, and Arash Vahdat. Compo- sitional text-to-image generation with dense blob representations. In International Conference on Machine Learning (ICML), 2024

  51. [59]

    An introduction to the approximate number system

    Darko Odic and Ariel Starr. An introduction to the approximate number system. Child Development Perspectives, 12(4):223–229, 2018

  52. [60]

    Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task

    Maya Okawa, Ekdeep S Lubana, Robert Dick, and Hidenori Tanaka. Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task. Advances in Neural Information Processing Systems, 36, 2024

  53. [61]

    Numerical cognition in bees and other insects

    Mario Pahl, Aung Si, and Shaowu Zhang. Numerical cognition in bees and other insects. Frontiers in psychology, 4:33732, 2013

  54. [62]

    The human imagination: the cognitive neuroscience of visual mental imagery

    Joel Pearson. The human imagination: the cognitive neuroscience of visual mental imagery. Nature Reviews Neuroscience, 20(10):624–634, 2019

  55. [63]

    Beyond the turk: Alternative platforms for crowdsourcing behavioral research

    Eyal Peer, Laura Brandimarte, Sonam Samat, and Alessandro Acquisti. Beyond the turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology, 70:153–163, 2017

  56. [64]

    Siemenn, Saisamrit Surbehera, Zad Chin, Keith Tyser, Gregory Hunter, Arvind Raghavan, Yann Hicke, Bryan A

    Vitali Petsiuk, Alexander E. Siemenn, Saisamrit Surbehera, Zad Chin, Keith Tyser, Gregory Hunter, Arvind Raghavan, Yann Hicke, Bryan A. Plummer, Ori Kerret, Tonio Buonassisi, Kate Saenko, Armando Solar-Lezama, and Iddo Drori. Human evaluation of text-to-image models on a multi...

  57. [65]

    Exact number concepts are limited to the verbal count range

    Benjamin Pitt, Edward Gibson, and Steven T Piantadosi. Exact number concepts are limited to the verbal count range. Psychological Science, 33(3):371–381, 2022

  58. [66]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pa...

  59. [67]

    Hierarchical text-conditional image generation with clip latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022

  60. [68]

    Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46, 2024

    Sunayana Rane, Alexander Ku, Jason Baldridge, Ian Tenney, Tom Griffiths, and Been Kim. Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46, 2024

  61. [69]

    Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment

    Royi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel, Yoav Goldberg, and Gal Chechik. Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment. Advances in Neural Information Processing Systems, 36, 2024

  62. [70]

    Impact of pretraining term frequencies on few-shot reasoning

    Yasaman Razeghi, Robert L Logan IV , Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022

  63. [71]

    High- resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High- resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022

  64. [72]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22500...

  65. [73]

    Photorealistic text-to-image diffusion models with deep language understanding

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv...

  66. [74]

    Verification of text ideas during reading

    Murray Singer. Verification of text ideas during reading. Journal of Memory and Language, 54 (4):574–591, 2006

  67. [75]

    Objectstitch: Object compositing with diffusion model

    Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian Price, Jianming Zhang, Soo Ye Kim, and Daniel Aliaga. Objectstitch: Object compositing with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18310–18319, 2023

  68. [76]

    Initial knowledge: Six suggestions

    Elizabeth Spelke. Initial knowledge: Six suggestions. Cognition, 50(1-3):431–445, 1994

  69. [77]

    What babies know: Core knowledge and composition volume 1, volume 1

    Elizabeth S Spelke. What babies know: Core knowledge and composition volume 1, volume 1. Oxford University Press, 2022

  70. [78]

    Core knowledge

    Elizabeth S Spelke and Katherine D Kinzler. Core knowledge. Developmental science, 10(1): 89–96, 2007

  71. [79]

    Event completion: Event based inferences distort memory in a matter of seconds

    Brent Strickland and Frank Keil. Event completion: Event based inferences distort memory in a matter of seconds. Cognition, 121(3):409–415, 2011

  72. [80]

    What formal lan- guages can transformers express? a survey

    Lena Strobl, William Merrill, Gail Weiss, David Chiang, and Dana Angluin. What formal lan- guages can transformers express? a survey. Transactions of the Association for Computational Linguistics, 12:543–561, 2024

  73. [81]

    Lexicalization patterns: Semantic structure in lexical forms.Language typology and syntactic description, 3(99):36–149, 1985

    Leonard Talmy. Lexicalization patterns: Semantic structure in lexical forms.Language typology and syntactic description, 3(99):36–149, 1985

  74. [82]

    Winoground: Probing vision and language models for visio-linguistic compositionality

    Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing vision and language models for visio-linguistic compositionality. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, ...

  75. [83]

    Computing machinery and intelligence

    Alan M Turing. Computing machinery and intelligence. Mind, 1950

  76. [84]

    No” zero-shot” without exponential data: Pretraining con- cept frequency determines multimodal model performance

    Vishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma, Philip HS Torr, Adel Bibi, Samuel Albanie, and Matthias Bethge. No” zero-shot” without exponential data: Pretraining con- cept frequency determines multimodal model performance. arXiv preprint arXiv:2404.04125, 2024. 18

  77. [85]

    Help or hinder: Bayesian models of social goal inference

    Tomer Ullman, Chris Baker, Owen Macindoe, Owain Evans, Noah Goodman, and Joshua Tenenbaum. Help or hinder: Bayesian models of social goal inference. Advances in neural information processing systems, 22, 2009

  78. [86]

    The automaticity of perceiving animacy: Goal-directed motion in simple shapes influences visuomotor behavior even when task-irrelevant

    Benjamin van Buren, Stefan Uddenberg, and Brian J Scholl. The automaticity of perceiving animacy: Goal-directed motion in simple shapes influences visuomotor behavior even when task-irrelevant. Psychonomic bulletin & review, 23(3):797–802, 2016

  79. [87]

    Response to affirmative and negative binary statements

    Peter C Wason. Response to affirmative and negative binary statements. British Journal of Psychology, 52(2):133–142, 1961

  80. [88]

    Paradoxical effects of thought suppression

    Daniel M Wegner, David J Schneider, Samuel R Carter, and Teri L White. Paradoxical effects of thought suppression. Journal of personality and social psychology, 53(1):5, 1987

  81. [89]

    Conceptmix: A compositional image generation benchmark with controllable difficulty

    Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, and Sanjeev Arora. Conceptmix: A compositional image generation benchmark with controllable difficulty. arXiv preprint arXiv:2408.14339, 2024

  82. [90]

    Boxdiff: Text-to-image synthesis with training-free box-constrained diffu- sion

    Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou. Boxdiff: Text-to-image synthesis with training-free box-constrained diffu- sion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7452–7461, 2023

  83. [91]

    Diffusion models: A comprehensive survey of methods and applications

    Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys, 56(4):1–39, 2023

  84. [92]

    Perceiving fully occluded objects via physical simulation

    Ilker Yildirim, Max H Siegel, and Joshua B Tenenbaum. Perceiving fully occluded objects via physical simulation. In Proceedings of the 38th annual conference of the cognitive science society, 2016

  85. [93]

    When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2022

    Mert Yuksekgonul, Federico Bianchi, Pratyusha Kalluri, Dan Jurafsky, and James Zou. When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2022

  86. [94]

    I NEED to test how the tool works with extremely simple prompts. DO NOT add any detail, just use it AS-IS:

    Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. A Appendix (Supplementary Information) A.1 Details on Behavioral Experi...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.