REVIEW 3 major objections 6 minor 56 references
Simile Understanding in Text-to-Image Models: An Evaluation Framework
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Text-to-image models systematically draw the literal object named in a simile instead of transferring its attributes.
desk verdict A useful new benchmark for measuring when T2I models literally render the vehicle of a simile, but the core YOLO metric conflates category presence with literalization and likely inflates the reported rates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the YOLO-Det rate: the fraction of generated images in which a pretrained object detector finds the object category named as the metaphorical vehicle. The framework's controlled construction, with vehicles restricted to 80 YOLO-detectable categories and 14 simile templates, makes each failure countable and avoids reliance on subjective aesthetic judgments. Diffusion Lens visualizations then expose the layer at which the vehicle first becomes visible in the text encoder, allowing the paper to separate failures that persist into the final image from those that appear in intermediate layers but are suppressed.
What would settle it
Generate the same G5 prompts with the five models, take all images where YOLO flags the vehicle, and have annotators label each one as an accidental literal object, a deliberate visual metaphor, or an appropriate scene element; if a substantial share of flagged images are judged deliberate or appropriate, the YOLO-Det rate overstates literalization bias and the reported model ordering would not survive.
Extended reading notes
Core claim
The central discovery is that literalization bias is real, widespread, and measurable: given simile prompts, current text-to-image models frequently depict the metaphorical vehicle as a concrete object. Using a dataset of 1,576 maximum-agreement simile sentences covering all 80 YOLO-detectable categories and all 14 templates, the paper reports YOLO detection rates of the vehicle from 0.298 for Dreamlike to 0.614 for Qwen-Image, with detections appearing in 78 of 80 categories and every template. The YOLO-based detection rate correlates strongly with human ratings of vehicle presence (r = 0.826, rho = 0.798), while CLIPScore and PickScore correlate weakly with both vehicle presence and overall prompt-image relevance. Diffusion Lens analysis shows the literal vehicle tends to emerge early and persist in CLIP-based branches, while appearing later but still persisting in Qwen-Image, and the paper shows that both random regeneration and layer-based regeneration can reduce detection rates without modifying the prompt, though the reductions vary by model and text encoder branch.
Load-bearing premise
The measurement treats any YOLO detection of the vehicle category as a literalization failure, which presupposes that no correct rendering of these similes may contain an object of that category; a deliberately visual metaphor or a scene where the vehicle object is part of the intended image would be counted as a failure.
Editorial extensions
If this is right
- Simile-based prompting cannot be assumed to transfer attributes; a prompt like 'as hard as stone' tends to put a stone in the image instead of conveying hardness.
- General-purpose alignment metrics such as CLIPScore and PickScore will not surface this failure, so object-level grounding checks are needed to benchmark figurative prompts.
- Literalization is not confined to a few templates or vehicles: it appears in all 14 templates and 78 of 80 YOLO categories, indicating a general tendency rather than an artifact of a small prompt set.
- Both random seed variation and layer-based regeneration can lower detection rates without changing the prompt, giving a cheap partial mitigation, but their effectiveness is model- and text-encoder-dependent.
- Lower detection of the vehicle does not by itself prove simile comprehension; the intended attribute transfer must be evaluated separately.
Reading between the lines
- A natural extension is to pair vehicle detection with attribute-transfer probes, for example asking whether the target actually looks harder or differently shaped; if regeneration removes the vehicle without producing the attribute, then literalization suppression is not comprehension.
- The 78-of-80 coverage suggests testing the same framework on other figurative constructions, such as metonymy and hyperbole, and on multilingual similes, where the YOLO vocabulary may not be culturally portable.
- The layer-based results imply that intervention at shallow text-encoder layers might reduce literalization more directly than seed sampling in CLIP-based models, though the paper only tests selecting existing layers rather than modifying representations.
- If literalization bias is as general as reported, image-generation benchmarks that reward only overall text-image similarity will systematically rank literalizing models too high for figurative prompts; a literalization rate could be added as a standard reporting axis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes an evaluation framework for measuring 'literalization bias' in text-to-image models, defined as the tendency to depict the metaphorical vehicle of a simile as a literal object. It constructs simile prompts whose vehicles are drawn from YOLO's 80 object categories, generates images with five t2i models (Dreamlike, PixArt, FLUX, SD3.5, Qwen-Image), measures YOLO detection of the vehicle category in generated images, validates this metric against human ratings of vehicle presence, and uses Diffusion Lens to track vehicle appearances across text encoder layers. The paper reports detection rates of 0.298–0.614 across models and proposes random and layer-based regeneration to reduce the bias.
Significance. The strength of the paper is the construction of a controlled, multi-model evaluation pipeline and the high correlation (r=0.826, ρ=0.798) between YOLO detection and human judgments of vehicle presence, which shows the metric reliably measures one thing: whether the vehicle appears. The paper also honestly states in Section 4.4 that absence of detection does not prove correct interpretation. However, the central claim that models 'misinterpret' or 'confuse' the vehicle with the object is not established, because the paper never tests whether detected vehicles are literal objects or the result of apt attribute transfer (e.g., clouds shaped like horses). The regeneration experiments are also circular because they select against the same YOLO metric. If the definitional gap is closed, the framework could be a useful diagnostic, but as it stands the headline claim is overstated.
major comments (3)
- [§2.2.1, §2.4.1, §4.4] The paper defines literalization bias as the literal appearance of the metaphorical vehicle and operationalizes it as YOLO detection of the vehicle category (Section 2.4.1), but it never establishes that a YOLO detection cannot arise from a correct visual simile. For example, a correct rendering of 'The clouds drifted exactly like horses' could show horse-shaped clouds, and YOLO would detect 'horse'; a branch shaped like a baseball bat would likewise trigger detection of 'baseball bat'. The human validation (Section 2.4.2) only asks whether the vehicle is clearly present (Q2) or overall relevance (Q3); it does not ask whether a detected vehicle is a literal object or a transferred attribute. Section 4.4 acknowledges that absence of detection does not prove correct interpretation, but it never addresses the converse. The reported rates (0.298–0.614) therefore conflate literalization with any depiction of the vehicle category, and the abstract's claim that models 'confuse it with the object' is not supported. Please re-annotate a sample of detected images with a judgment of literalness (is the vehicle depicted as an independent object, or is its form transferred to the target?) and report the proportion of detections that are true failures.
- [§3.4.1, §3.4.2] Random regeneration and layer-based regeneration both select the first image in which YOLO does not detect the vehicle; consequently, the reported reductions in YOLO-Det are largely guaranteed by construction and are not independent evidence that literalization bias has decreased. For any prompt whose initial generation contains the vehicle, sampling repeatedly and stopping at the first non-detection will reduce the per-prompt detection rate unless the detection probability is 1. The paper provides no human evaluation or independent metric showing that the regenerated images have better simile understanding; the CLIPScore and PickScore results are acknowledged to be insensitive. The claim that these methods 'reduce literalization bias' should be re-framed as a procedure for reducing YOLO-Det specifically, or validated with human judgments of whether the regenerated image better reflects the simile's meaning.
- [Table 1, Figure 4] The model-level YOLO-Det values in Table 1 are reported without confidence intervals or statistical tests. With approximately 1,576 prompts per model, the differences among Dreamlike (0.298), SD3.5 (0.325), and FLUX (0.349) may or may not be reliable; adding binomial confidence intervals or a paired test would substantiate the claim of 'substantial' variation. Similarly, Figure 4 presents mean first-emergence layers and coverage ratios without error bars, making it hard to evaluate whether the cross-model differences in these metrics are meaningful.
minor comments (6)
- [References [43] and [44]] References [43] and [44] are the same paper (Su et al., Neural Processing Letters) and should be merged.
- [Reproducibility] The paper does not include a data/code availability statement; to support reproducibility, please release the G5 simile dataset and generation/inference scripts.
- [Table 2] In Table 2, Q1 (Context) is defined in Section 2.4.2 but results for Q1 are not reported; clarify whether they appear in the appendices or were excluded.
- [Figure 4] Figure 4 would benefit from error bars or confidence intervals; in addition, the sample sizes behind the counts (e.g., 550 Persistent cases) should be stated in the caption or text.
- [§3.1] Section 3.1 reports a non-significant ANOVA (p=0.11) and then a significant focused comparison; the focused test should be labeled as exploratory and the possibility of multiple comparisons should be addressed.
- [§3.2.2] In Section 3.2.2, the statement that 'higher Q2 scores tend to coincide with lower Q3 scores' is supported by the model-level pattern, but the within-model relationship is not shown; consider reporting per-model correlations.
Circularity Check
The main literalization-bias measurement is validated against external human ratings, but the regeneration experiments select images by the same YOLO-Det criterion they then report as improved, making that improvement partly by construction.
-
fitted input called prediction
[Section 2.5.2 (Regeneration Methods), Section 3.4.1 (Random Regeneration), and Section 3.4.2 (Layer-Based Regeneration)]
"In random regeneration, we repeatedly generate images from the same simile prompt while varying only the random seed. ... After each attempt, YOLO determines whether the metaphorical vehicle is present. The first image in which the metaphorical vehicle is not detected is retained as the regenerated output. ... Random regeneration reduces YOLO-Det for all five models."
The selection rule for the regenerated image is exactly the absence of a YOLO detection of the vehicle, and the reported outcome metric is the YOLO-Det of the retained image. Therefore the observed reduction in YOLO-Det is largely a logical consequence of the selection rule rather than an independent empirical effect of regeneration: whenever a non-detected image exists within the allowed attempts, the method is guaranteed to find and retain it. The same structure applies to layer-based regeneration, where shallower layers are tried until YOLO no longer detects the vehicle. The paper frames the resulting delta as evidence that the methods 'reduce literalization bias,' but the metric used to measure the reduction is identical to the oracle used to select the output.
full rationale
The central finding of the paper, that t2i models frequently depict the metaphorical vehicle literally, is not circular: YOLO-Det is computed on final generated images from five externally defined models, and its image-level correlation with human Q2 ratings (r = 0.826, rho = 0.798) and model-level agreement (r = 0.99, rho = 1.00) provide external anchors. Dataset construction, template design, and LLM-as-a-Judge filtering are independent of the measured outcome, and the Diffusion Lens analysis is descriptive rather than self-validating. The only substantial circularity is in the regeneration experiments, where the selection criterion and the evaluation metric are the same YOLO detector; reporting that the selected outputs have lower YOLO-Det is in part a tautology. Because this issue affects a secondary mitigation claim rather than the core measurement, the overall score is 4 rather than a higher value. A separate construct-validity concern, that a detected vehicle could sometimes be an apt transferred form rather than a literalization, is a measurement validity issue and is not counted as derivation circularity under the stated rules.
Assumptions & free parameters
free parameters (3)
- YOLO detection confidence threshold =
0.25 (detector default)
- Maximum LLM-judge agreement threshold for G5 =
12 (all four binary criteria across three judges)
- Regeneration attempt cap =
5
assumptions (4)
- domain assumption A correct rendering of these simile prompts should not contain the metaphorical vehicle as a literal object.
- domain assumption YOLO detections of the vehicle category correspond to human judgments of literal vehicle presence.
- domain assumption Diffusion Lens visualizations faithfully reflect the content contributed by each text encoder layer.
- domain assumption LLM-as-a-Judge filtering produces a valid and sufficiently diverse set of simile sentences.
Cite this review
Pith. "Pith review of Simile Understanding in Text-to-Image Models: An Evaluation Framework." pith.science (2026). https://pith.science/paper/CENNDUFQ
@misc{pith2026260804750,
author = {Pith},
title = {Pith review of: Simile Understanding in Text-to-Image Models: An Evaluation Framework},
year = {2026},
howpublished = {\url{https://pith.science/paper/CENNDUFQ}},
note = {Machine review of arXiv:2608.04750}
}
read the original abstract
Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models frequently misinterpret the metaphorical vehicle and confuse it with the object. These systematic failures reveal a gap between figurative language and object-level visual grounding in t2i models. To investigate this issue, we propose a scalable evaluation framework for simile understanding. Our framework includes (1) a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates, (2) automatic grounding metrics based on YOLO (You Only Look Once) detection, and (3) text encoder layer analysis using Diffusion Lens to track how metaphorical vehicles emerge during generation. Experiments across architecturally diverse t2i models reveal consistent literalization failure patterns. We further discuss potential mitigation strategies for improving simile grounding in t2i models.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Hewett, Mojan Javaheripi, Piero Kauff- mann, James R
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauff- mann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Y...
arXiv 2024
-
[2]
MetaCLUE: Towards Comprehensive Visual Metaphors Research
Arjun R. Akula, Brendan Driscoll, Pradyumna Narayana, Soravit Changpinyo, Zhiwei Jia, Suyash Damle, Garima Pruthi, Sugato Basu, Leonidas Guibas, William T. Freeman, Yuanzhen Li, and Varun Jampani. 2023. MetaCLUE: To- wards Comprehensive Visual Metaphors Research. arXiv:2212.09898 [cs.CV] https://arxiv.org/abs/2212.09898
work page Pith review arXiv 2023
-
[3]
Dreamlike Art. 2023. Dreamlike Photoreal 2.0. Available at https://huggingface. co/dreamlike-art/dreamlike-photoreal-2.0
work page 2023
-
[4]
Julia Birke and Anoop Sarkar. 2006. A Clustering Approach for Nearly Unsu- pervised Recognition of Nonliteral Language. In11th Conference of the European Chapter of the Association for Computational Linguistics. Association for Compu- tational Linguistics, Trento, Italy, 329–336. https://aclanthology.org/E06-1042/
work page 2006
-
[5]
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Ben- gio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, and Joseph Turian. 2020. Experience Grounds Lan- guage. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Y...
-
[6]
Bowman and George Dahl
Samuel R. Bowman and George Dahl. 2021. What Will it Take to Fix Benchmark- ing in Natural Language Understanding?. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Kristina Toutanova, Anna Rumshisky, Luke Zettle- moyer, Dilek Hakkani-Tur, Iz Beltagy, Steven B...
2021
-
[9]
Tuhin Chakrabarty, Xurui Zhang, Smaranda Muresan, and Nanyun Peng. 2021. MERMAID: Metaphor Generation with Symbolism and Discriminative Decod- ing. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hak...
2021
-
[10]
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhong- dao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. 2023. PixArt-𝛼: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis. arXiv:2310.00426 [cs.CV] https://arxiv.org/abs/2310.00426
arXiv 2023
Show all 56 references
-
[11]
Weijie Chen, Yongzhu Chang, Rongsheng Zhang, Jiashu Pu, Guandan Chen, Le Zhang, Yadong Xi, Yijiang Chen, and Chang Su. 2022. Probing Simile Knowledge from Pre-trained Language Models. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vo...
2022 doi
-
[12]
DiStefano, John D
Paul V. DiStefano, John D. Patterson, and Roger E. Beaty. 2024. Automatic Scoring of Metaphor Creativity with Large Language Models.Creativity Research Journal 37, 4 (2024), 555–569. doi:10.1080/10400419.2024.2326343
2024
-
[13]
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, et al . 2024. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis. arXiv:2403.03206 [cs.CV] http...
2024 arXiv
-
[14]
Ge Gao, Eunsol Choi, Yejin Choi, and Luke Zettlemoyer. 2018. Neural Metaphor Detection in Context. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Association f...
2018 doi
-
[15]
Dedre Gentner. 1983. Structure-mapping: A theoretical framework for analogy. Cognitive Science7, 2 (1983), 155–170. doi:10.1016/S0364-0213(83)80009-3
1983 doi
-
[16]
Dhruba Ghosh, Hanna Hajishirzi, and Ludwig Schmidt. 2023. GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment. arXiv:2310.11513 [cs.CV] https://arxiv.org/abs/2310.11513
2023 arXiv
-
[17]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, A...
2024 arXiv
-
[18]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, Saizhuo Wang, Kun Zhang, Yuanzhuo Wang, Wen Gao, Lionel Ni, and Jian Guo. 2025. A Survey on LLM-as- a-Judge. arXiv:2411.15594 [cs.CL] https://arxiv.org/a...
2025 arXiv
-
[19]
Hamilton, Jure Leskovec, and Dan Jurafsky
William L. Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Katrin Erk and Noah A. Smith (E...
2016 doi
-
[20]
Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie, and Yanghua Xiao. 2022. Can Pre- trained Language Models Interpret Similes as Smart as Human?. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Presla...
2022 doi
-
[21]
Kaiyi Huang, Chengqi Duan, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. 2025. T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation. arXiv:2307.06350 [cs.CV] https: //arxiv.org/abs/2307.06350
2025 arXiv
-
[22]
Nicholas Ichien, Dušan Stamenković, and Keith J. Holyoak. 2024. Large Lan- guage Model Displays Emergent Ability to Interpret Novel Literary Metaphors. arXiv:2308.01497 [cs.CL] https://arxiv.org/abs/2308.01497
2024 arXiv
-
[23]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...
2023 arXiv
-
[24]
Glenn Jocher, Jing Qiu, and Jingyu Peng. 2024. Ultralytics YOLO11. https: //github.com/ultralytics/ultralytics
2024
-
[25]
Ricardo Kleinlein, Cristina Luna-Jiménez, and Fernando Fernández-Martínez
-
[26]
Koushik, Fatemeh Nazarieh, Katherine Birch, Shenbin Qian, and Diptesh Kanojia
Girish A. Koushik, Fatemeh Nazarieh, Katherine Birch, Shenbin Qian, and Diptesh Kanojia. 2025. The Mind’s Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation. arXiv:2508.18569 [cs.CL] https://arxiv.org/abs/ 2508.18569
2025 arXiv
-
[28]
Black Forest Labs. 2024. FLUX. https://github.com/black-forest-labs/flux
2024
-
[29]
1980.Metaphors we live by
George Lakoff and Mark Johnson. 1980.Metaphors we live by. University of Chicago Press, Chicago
1980
-
[30]
Lizhen Liu, Xiao Hu, Wei Song, Ruiji Fu, Ting Liu, and Guoping Hu. 2018. Neural Multitask Learning for Simile Recognition. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsuj...
2018 doi
-
[31]
Isabel Negro, Ester Šorm, and Gerard Steen. 2018. General image understanding in visual metaphor identification.ODISEA(04 2018), 113–131. doi:10.25115/ odisea.v0i18.1900
2018
-
[32]
Shintaro Ozaki, Tomoyuki Jinno, Kazuki Hayashi, Yusuke Sakai, Jingun Kwon, Hidetaka Kamigaito, Katsuhiko Hayashi, Manabu Okumura, and Taro Watan- abe. 2026. TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation. arXiv:2504.1826...
2026 arXiv
-
[33]
Barbara Phillips and Edward Mcquarrie. 2004. Beyond Visual Metaphor: A New Typology of Visual Rhetoric in Advertising.Marketing Theory4 (06 2004), 113–
2004
-
[34]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
-
[35]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research21, 140 (2020), 1–67. MM ’26...
2020
-
[36]
rcland12. 2023. A list of all 80 YOLO classes and its index in JSON format. GitHub Gist. Version 2, created July 14, 2023. https://gist.github.com/rcland12/ dc48e1963268ff98c8b2c4543e7a9be8
2023
-
[37]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752 [cs.CV] https://arxiv.org/abs/2112.10752
2022 arXiv
-
[38]
Arkadiy Saakyan, Shreyas Kulkarni, Tuhin Chakrabarty, and Smaranda Muresan
-
[39]
Elisa Sanchez-Bayona and Rodrigo Agerri. 2025. Metaphor and Large Lan- guage Models: When Surface Features Matter More than Deep Understanding. InFindings of the Association for Computational Linguistics: ACL 2025, Wanxi- ang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad...
2025 doi
-
[40]
Ekaterina Shutova. 2010. Models of Metaphor in NLP. InProceedings of the 48th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Uppsala, Sweden, 688–697. https://aclanthology.org/ P10-1071/
2010
-
[41]
Kevin Stowe, Nils Beck, and Iryna Gurevych. 2021. Exploring Metaphoric Para- phrase Generation. InProceedings of the 25th Conference on Computational Natu- ral Language Learning, Arianna Bisazza and Omri Abend (Eds.). Association for Computational Linguistics, Online, 323–336....
2021 doi
-
[42]
Kevin Stowe, Tuhin Chakrabarty, Nanyun Peng, Smaranda Muresan, and Iryna Gurevych. 2021. Metaphor Generation with Conceptual Mappings. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natur...
2021 doi
-
[44]
Chang Su, Xingyue Wang, Shupin Liu, and Yijiang Chen. 2024. Efficient Vi- sual Metaphor Image Generation Based on Metaphor Understanding.Neural Processing Letters56 (April 2024), 150. doi:10.1007/s11063-024-11609-w
2024 doi
-
[45]
Kaiyue Sun, Rongyao Fang, Chengqi Duan, Xian Liu, and Xihui Liu. 2025. T2I- ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation. arXiv:2508.17472 [cs.CV] https://arxiv.org/abs/2508.17472
2025 arXiv
-
[46]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024 arXiv
-
[47]
Michael Toker, Hadas Orgad, Mor Ventura, Dana Arad, and Yonatan Belinkov
-
[48]
Xiaoyu Tong, Rochelle Choenni, Martha Lewis, and Ekaterina Shutova. 2024. Metaphor Understanding Challenge Dataset for LLMs. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguis...
2024
-
[49]
Xiaoyu Tong, Ekaterina Shutova, and Martha Lewis. 2021. Recent advances in neural metaphor processing: A linguistic, cognitive and social perspec- tive. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human L...
2021
-
[50]
Xiaoyue Wang, Linfeng Song, Xin Liu, Chulun Zhou, Hualin Zeng, and Jinsong Su. 2022. Getting the Most out of Simile Recognition. InFindings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Com...
2022 doi
-
[51]
Zengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong, Man Zhang, and Xiangxiang Chu. 2026. Everything in Its Place: Benchmarking Spatial Intelligence of Text-to- Image Models. arXiv:2601.20354 [cs.CV] https://arxiv.org/abs/2601.20354
2026
-
[52]
Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al . 2025. Qwen-Image Technical Report. arXiv:2508.02324 [cs.CV] https://arxiv.org/abs/2508.02324
2025 arXiv
-
[53]
Yanzhi Xu, Yueying Hua, Shichen Li, and Zhongqing Wang. 2024. Exploring Chain-of-Thought for Multi-modal Metaphor Detection. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek ...
2024 doi
-
[54]
Ron Yosef, Yonatan Bitton, and Dafna Shahaf. 2023. IRFL: Image Recognition of Figurative Language. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 1...
2023 doi
-
[55]
{template_phrase}
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems36 (2023), 46595–46623. A Supp...
2023
-
[136]
doi:10.1177/1470593104044089
-
[2021]
arXiv:2103.00020 [cs.CV] https://arxiv.org/abs/2103.00020
Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.00020 [cs.CV] https://arxiv.org/abs/2103.00020
-
[2022]
arXiv:2210.10578 [cs.CL] https://arxiv.org/abs/2210.10578
Language Does More Than Describe: On The Lack Of Figurative Speech in Text-To-Image Models. arXiv:2210.10578 [cs.CL] https://arxiv.org/abs/2210.10578
-
[2024]
InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguis...
-
[3536]
doi:10.18653/v1/2024.acl-long.193
2024 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.