{"id":"c497c9a1-dcbe-4bed-b148-b8b2a257bc06","arxiv_id":"2411.17066","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"DALL-E 3 fails to reliably produce images matching prompts about physical relations, negations, and exact numbers beyond three, and a grounded diffusion pipeline does worse on relations.","lead":"Researchers asked people to judge images generated by DALL-E 3 from simple prompts about relationships, negations, and counts, and found the model often fails, especially at negation and numbers above three. The study offers a systematic snapshot of where today's text-to-image models still fall short of basic human logic.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 'none >50%' claim is internally contradicted by the 75% agreement for 'one' and is not calibrated against a chance baseline, so the headline scope should be narrowed.","rationale":"The reader's weakest assumption was exactly this: the lack of a baseline/chance condition undermines the interpretation of the absolute agreement rates. I agree with that assessment and have identified it as the most load-bearing concern because it directly targets the headline claim. In addition, I have flagged an internal inconsistency in the abstract that the reader also noted ('the abstract overstates the findings (the 'none >50%' claim conflicts with the 75% agreement for 'one')'). The concern is not that the empirical results are fabricated or that the direction of the effects is wrong; the negation and high-number results are likely qualitatively correct. The concern is that the paper's central quantitative claim, as stated in the abstract, is both (a) internally contradicted by one of the paper's own data points and (b) not calibrated against any chance baseline. The fix is straightforward: narrow the claim to the specific conditions that actually failed and add a control condition or otherwise report the selection-rate distribution so that 45% and 12.3% can be interpreted. This warrants a CONDITIONAL rather than a REJECT or ACCEPT because the underlying experimental design is sound and the main negative findings are likely robust; only the headline and its interpretation need to be corrected. I do not think the detector-based auto-count analysis is the single most load-bearing issue here, even though the near-zero kappa values are concerning, because that analysis is presented as a supplementary 'follow-up' and the main conclusions do not rest on it. The baseline issue, by contrast, sits at the center of the paper's core claim.","tokens_in":19983,"tokens_out":2067,"duration_ms":16160,"concrete_test":"Recompute the headline statistics after excluding the 'one' condition from the numbers experiment, and re-report the agreement rates relative to a chance baseline derived from the actual distribution of participant selection counts. Concretely: (1) Re-read Experiment 3 and state the agreement for 'one' separately, and revise the abstract to say 'none of the negation or large-number prompts exceed 50%' or similar. (2) Re-run a small control condition (e.g., 30 participants, 10 trials) in which each 18-image grid is paired with a semantically neutral or mismatched prompt (e.g., a relation prompt matched against images generated for a different relation, or a prompt like 'a picture of something' against random images), and measure the mean number of images selected per grid.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim as stated in the abstract — 'none reliably produce human agreement scores greater than 50%' — is not supported by the paper's own data. In Experiment 3, the average agreement for 'one' is approximately 75%, which is well above 50%. The authors acknowledge this in the Results section ('The average agreement for 'one' entity is approximately 75%'), so the abstract's blanket claim is internally inconsistent unless it is restricted to the specific negation and large-number conditions. This matters because the headline is the claim that a reader takes away, and it is the basis for the paper's rhetorical conclusion that DALL-E 3 'does not reliably produce' images matching logical operators. The claim is also not calibrated: participants could select all, some, or none of 18 images per grid, with no baseline or chance-level condition. Without knowing the distribution of selection counts under a 'neutral' prompt (e.g., a prompt that is neither clearly satisfiable nor clearly unsatisfiable, or a prompt matched to a random image set), the 45% agreement for relations and the 12.3% for negation cannot be interpreted as 'failure' relative to chance. For example, if participants on average select 4-5 images per grid when uncertain, then 45% agreement could be at or above the rate expected from guessing, whereas if they select only 1-2 images on average, 45% indicates a strong signal. The 'none >50%' formulation is therefore both internally inconsistent with the numbers data and ungrounded without a baseline. This is a measurement/interpretation issue that directly affects the central claim, not a peripheral concern.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports four behavioral experiments (total N=178) plus auxiliary analyses in which human participants judged whether images generated by DALL·E 3 (and, in Experiment 4, an LLM-grounded diffusion pipeline, LMD+) matched prompts built on physical relations, negations, and exact numbers. The headline claim is that no condition reliably produces human agreement scores above 50%, with relations at about 45% agreement, unmodified negation at about 12.3%, and number agreement declining from roughly 75% for 'one' to about 9% for 'six'. A follow-up auto-count analysis using twelve object detectors reports scalar variability and ratio dependence in DALL·E 3's generated counts, and an n-gram frequency analysis connects prompt frequency to perceived match. The paper interprets these results as evidence of persistent limitations in compositional and logical control in current text-to-image models.","tokens_in":20257,"tokens_out":5539,"duration_ms":54272,"significance":"The question of whether modern text-to-image models can handle logical operators is timely and important, and this paper contributes a simple, transparent probe set with open data and code, attention-checked human judgments, and a direct comparison between DALL·E 3 and a grounded-diffusion pipeline. The negation failure and the sharp decline for larger numbers are well supported by the human data. However, the headline 'none greater than 50%' claim is not calibrated against any chance-level baseline, and it is internally contradicted by the 75% agreement for 'one' in Experiment 3. The auto-count analysis is auxiliary rather than load-bearing for the central human-judgment claim, but its detector-based counts are not validated against human counts and should not be treated as strong evidence on their own. The paper's strengths are its clear experimental design, replication-oriented presentation, and open materials; its main weaknesses are the ungrounded absolute threshold and an overbroad abstract.","major_comments":[{"comment":"The abstract's statement that 'none reliably produce human agreement scores greater than 50%' is contradicted by the paper's own Experiment 3 result that the average agreement for 'one' entity is approximately 75%. If 'none' is meant to range over all prompt-level averages, the claim is false; if it is meant only for negation and numbers beyond three, the abstract must be rewritten to say so. This sentence is the central takeaway of the paper, so the mismatch between the abstract and the reported data should be fixed before publication.","section":"Abstract and Experiment 3 (Results)"},{"comment":"The agreement measure lacks a chance-level or baseline condition. Participants were allowed to select all, some, or none of the 18 images per grid, so the absolute value of the reported agreement percentage depends on participants' overall selection rate under uncertainty. Without a neutral-prompt condition or a grid of images that are not matched to the prompt, the values 45% for relations and 12.3% for negation cannot be interpreted as being above or below chance, and the 50% threshold used throughout the abstract is not calibrated. The authors should add a baseline condition, such as prompts paired with random or clearly mismatched image grids, and report the distribution of selection counts over trials.","section":"Experimental Design and Experiment 1"},{"comment":"The auto-count analysis treats the counts produced by twelve object detectors as reliable estimates of the number of objects in each generated image, but the reported inter-rater agreement among these detectors is near zero (mean Cohen's kappa between -0.001 and 0.038). With this level of agreement, the estimated scalar-variability slope of 1.7 and the ratio-dependence breakpoint of approximately 3.33 are not trustworthy measures of DALL·E 3's output counts. At minimum, a subset of images should be counted by human annotators so that detector counts can be validated, and the analysis should report agreement between detector counts and human counts before drawing conclusions about 'approximate numeracy'.","section":"A.4, Table A.1, and Figure 7"}],"minor_comments":[{"comment":"The section on 'Experiments 5-9' says '119 participants were recruited for Experiment 4'; this should refer to Experiments 5-9, since Experiment 4 is the LMD+ study with 30 participants. The total N=178 in the abstract does not include these 119 participants, so the numbering and the participant totals should be clarified.","section":"Appendix A.1"},{"comment":"The LMD+ comparison reports an average agreement of 29.4% without confidence intervals, whereas the earlier experiments report means with 95% intervals; add the intervals for this condition so the reader can assess the precision of the comparison.","section":"Experiment 4 (Results)"},{"comment":"The caption says 'A shows a density plot' and 'B provides further detail', but the panel labels are not explicitly marked in the described figure; please label the panels A and B in the figure itself and describe what the axes of each panel represent.","section":"Figure 7 caption"},{"comment":"The categorization of modified negation prompts into Replacement, Addition, and No Change is based on the authors' own judgment, with approximate percentages but no inter-rater reliability or coding scheme; report a second coder and agreement, or explicitly describe the categorization as informal.","section":"Experiment 2 (Methods)"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely question, and the human data on negation and large numbers are valuable. I would not recommend rejection, because the core empirical pattern is likely robust even without the 50% framing, and the main issues are fixable: qualify the abstract, add a baseline condition for the agreement metric, and validate the detector counts in the auto-count analysis. The auto-count analysis is auxiliary, so it can be revised without changing the central experiments, but as written it should not be used to support strong claims about approximate numeracy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline result this paper should be remembered for is simpler than its abstract: on a careful human evaluation, DALL-E 3 gets negation badly wrong (12% agreement on plain 'not X' prompts) and number agreement falls off a cliff after three (75% for one down to 9% for six). Those findings look solid and are worth having on record. The paper also does a real service in shipping code and data, running attention checks, reporting confidence intervals, and comparing against a grounded-diffusion pipeline (which does worse on relations — a useful data point).\n\nThe main problem is the abstract's claim that 'none reliably produce human agreement scores greater than 50%.' That is false on the paper's own numbers: average agreement for 'one' is about 75%. The body correctly narrows the claim to negation and numbers beyond three, but the banner claim is what readers will take away, and it is internally contradicted.\n\nThere is also no baseline condition. Participants could select all, some, or none of 18 images per grid; without a control prompt we have no idea what selection rate would arise from uncertainty or response bias. The absolute phrase 'greater than 50%' is therefore uncalibrated. The relative comparisons (negation vs. relations, the number drop-off) do not need a baseline to be meaningful, so this is fixable by reframing rather than invalidating.\n\nThe auto-count 'approximate numeracy' analysis has a more serious issue: the 12 detection models show mean Cohen's kappa near zero (Table A.1), meaning they barely agree with each other on object counts. Fitting scalar variability slopes to those outputs is a house built on sand. I would either validate the counts with human annotation or cut that section entirely.\n\nThe n-gram frequency correlation is exploratory but transparently labeled; that is fine. Self-citation of Conwell & Ullman 2022 is appropriate, not circular.\n\nIn short: this deserves a serious referee, but not in current form. I would send it out with instructions that the abstract must match the data, the 50% threshold needs calibration or qualification, and the auto-count analysis needs validation or removal. A clean revision would make this a useful reference point for text-to-image evaluation work.","headline":"Solid human evaluation of DALL-E 3's logic failures, but the abstract overclaims and an uncalibrated threshold plus unreliable auto-counts need fixing.","tokens_in":20845,"tokens_out":2991,"would_cite":false,"duration_ms":27577,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DALL-E 3 fails most logic prompts, say human judges","keywords":["text-to-image generation","logical operators","DALL-E 3","compositional generation","human evaluation","negation","number cognition","relations"],"falsifier":"Run the same 18-image selection task with deliberately non-matching prompts as catch trials and compare selection rates: if people endorse non-matching images at or above the rates observed for relations (45%), the claim that relations fail falls to a measurement artifact. A second check would ask humans to count objects in images from the auto-count follow-up and compare their counts to the detection-model estimates.","tokens_in":19765,"feed_emoji":"🧩","tokens_out":5989,"duration_ms":53843,"temperature":0.7,"pith_summary":"This paper asks whether a state-of-the-art text-to-image system can do what young children do without effort: turn simple logical descriptions into pictures. Using human judges who pick matching images from 18-image grids, it finds that prompts built on physical relations, plain negation, and exact integers mostly do not produce images people agree match. Relations average 45% agreement, plain negation 12.3%, and counts fall from about 75% for \"one\" to about 9% for \"six\"; a grounded diffusion pipeline with scene-graph layouts scores even lower on the same relation prompts. The point of the exercise is not just to grade one model, but to show that logical composition remains a distinct failure mode in systems whose gains come from scale and vector-based grounding.","feed_headline":"Human judges say DALL-E 3 fails most logic prompts","feed_subtitle":"Relations score 45%, plain negation 12%, and counts collapse from 75% (one) to 9% (six).","key_machinery":"The carrying object is the logical probe: a minimal prompt that pairs an everyday object with one relation, one negation, or one integer, rendered 18 times per prompt. The measuring instrument is human agreement—participants select all, some, or none of the 18 images as matching the prompt. Around this, the paper adds three analytic devices: a fixed prefix that stops the model's internal language model from rewriting a prompt, so the raw model is tested; an N-gram frequency correlation that tests whether relational success tracks training-data statistics; and an approximate-numeracy analysis using object-detection counts to estimate scalar variability and ratio dependence in generated quantities.","core_discovery":"The paper's central claim is that DALL-E 3 does not reliably deploy basic logical operators in image generation, and that no probe family—relations, negations, or numbers—consistently clears the 50% human-agreement mark. Negation fails most completely: when the prompt is left unmodified, it almost always renders the very object it forbids. Number generation is exact for small counts and approximate beyond three, with agreement collapsing from 75% at \"one\" to 9% at \"six\"; a follow-up analysis with object detectors finds ratio-dependent, scalar-variable error patterns like an approximate number system rather than exact counting. The paper also claims that a grounded pipeline using structured layout graphs does not fix these problems and is judged worse than DALL-E 3 on the same relation prompts, because its intermediate representations miss physics such as occlusion-in-depth.","pith_inferences":["Because no baseline or chance condition was run, the headline figures are best read as relative rankings of probes, not calibrated success rates; adding catch trials with no matching image would tell where the true floor sits.","The same probe battery could be run on other image generators to map whether the relation-frequency and count-collapse patterns are universal or specific to DALL-E 3.","The approximate-numeracy finding suggests generative models could serve as a platform for studying the emergence of number representations, but the detection-model-based counts need validation against human counting on the same images.","If relation success is driven by N-gram frequency, benchmarks should balance prompt frequencies before comparing models, otherwise reported compositional gains may be frequency effects in disguise."],"forward_implications":["Text-to-image \"prompt following\" should not be read as compositional understanding; simple operator probes reveal systematic gaps that survive scale.","Plain negation is not just imperfect but inverted: an unmodified \"not X\" prompt tends to produce X, which means systems need an explicit rewriting step that replaces the forbidden object.","Exact count generation beyond three or four objects is outside current capability, and the error pattern is ratio-dependent rather than exact-integer-like.","Structured intermediate representations such as scene graphs can hurt rather than help when they omit physical constraints like occlusion, so grounding alone is not sufficient.","Prompt frequency predicts relational success, implying that reported gains in relations may reflect dataset statistics rather than relational reasoning."],"supporting_citations":[{"why":"Supplies the prior relation-probe method and DALL-E 2 baseline that this work extends.","marker":"[11]"},{"why":"Describes the DALL-E 3 model whose prompt-following behavior is under test.","marker":"[2]"},{"why":"Introduces the LLM-grounded diffusion pipeline used as the comparison model in Experiment 4.","marker":"[46]"},{"why":"Provides the training-prior account that motivates the N-gram frequency analysis of relational prompts.","marker":"[49]"},{"why":"Shows pretraining term frequencies affect reasoning, the basis for comparing prompt frequency to image match.","marker":"[70]"},{"why":"Concurrent work on whether generative multimodal models count, which this paper's number results align with.","marker":"[68]"},{"why":"Defines scalar variability and ratio dependence, the approximate-number-system signatures used in the count follow-up.","marker":"[59]"},{"why":"Supplies the object categories and detection training data for the auto-count analysis.","marker":"[47]"}],"fun_headline_variants":["DALL-E 3 flunks logic: 12% negation, 9% for six objects","AI image gen fails logic: relations 45%, counts drop to 9%","Grounded diffusion can't beat DALL-E 3 on logic prompts","No logic probe over 50%: DALL-E 3 stumbles on counts, negation","DALL-E 3 logic: numbers beyond 3 collapse to 9% agreement"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline \"none greater than 50%\" treats 50% agreement as the success threshold, but the task offered no baseline or chance condition showing how often people endorse images when a prompt has no valid match.","fun_headline_variants_meta":{"raw":{"variants":["DALL-E 3 flunks logic: 12% negation, 9% for six objects","AI image gen fails logic: relations 45%, counts drop to 9%","Grounded diffusion can't beat DALL-E 3 on logic prompts","No logic probe over 50%: DALL-E 3 stumbles on counts, negation","DALL-E 3 logic: numbers beyond 3 collapse to 9% agreement"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":3044,"prompt_tokens":1064,"completion_tokens":1980,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":1868}},"tokens_in":680,"tokens_out":1980,"duration_ms":13735,"temperature":1.0,"reasoning_tokens":1868,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:33:17.593628+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 18-image selection task with deliberately non-matching prompts as catch trials and compare selection rates: if people endorse non-matching images at or above the rates observed for relations (45%), the claim that relations fail falls to a measurement artifact. A second check would ask humans to count objects in images from the auto-count follow-up and compare their counts to the detection-model estimates.","supporting_citations":[{"cited_title":"Training Priors Predict Text-To-Image Model Performance","cited_arxiv_id":"2306.01755","evidence_quote":"Provides the training-prior account that motivates the N-gram frequency analysis of relational prompts."},{"cited_title":"Can generative multimodal models count to ten? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 46, 2024","cited_arxiv_id":null,"evidence_quote":"Concurrent work on whether generative multimodal models count, which this paper's number results align with."},{"cited_title":"An introduction to the approximate number system","cited_arxiv_id":null,"evidence_quote":"Defines scalar variability and ratio dependence, the approximate-number-system signatures used in the count follow-up."}],"review_version":1}