Pith. sign in

REVIEW 4 major objections 6 minor 89 references

Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper reports that a GPT-4o system guided by Carroll's evaluative framework produced art critiques human judges could not reliably distinguish from expert writing, while model performance on new higher-order Theory of Mind tasks was…

desk verdict The Turing-test conclusion about AI critiques is undercut by the GPT-4o normalizer applied to the human side, but the ToM tasks and reproducibility make the paper worth refereeing. read the letter →

arxiv 2504.12805 v2 pith:7HF6MIMA submitted 2025-04-17 cs.CL cs.CYcs.HC

classification cs.CLcs.CYcs.HC
keywords largelanguagemodelartcriticismTuringtestTheoryofMindgenerativeAIparadoxchain-of-thoughtpromptingGPT-4ohigher-order
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether large language models can do more than imitate art-critical prose: can they produce critiques grounded in aesthetic theory, and do they possess the higher-order mental-state reasoning that critics use? It reports that a guided GPT-4o system, fed Noël Carroll's seven-step evaluative framework and 15 critical theories, generated critiques that lay judges identified as human only at chance level, with 51.4% overall accuracy and 56.3% excluding one outlier. It also introduces three art-specific Theory of Mind tasks and shows that 41 current LLMs vary widely, with only 31.7% correct on a hidden-intention task and systematic errors on a plagiarism scenario. The paper reads these results not as proof of understanding but as evidence that careful prompting can make LLM output resemble understanding more closely than the Generative AI Paradox assumes.

What carries the argument

The central object is Composer, a custom GPT-4o configured with uploaded external knowledge: a self-authored summary of Noël Carroll's On Criticism, which organizes critiques into seven components with evaluation as the governing task, and a summary of 15 distinct criticism theories, from structuralist to postcolonial. A chain-of-thought-style prompt makes the model first write a full-length critique with each observation attributed to a theory, then condense it into a three-paragraph version, then produce a playful one-liner. The Turing test is paired with Format Normalizer, a second GPT-4o instance that rewrites human Smarthistory critiques into the same three-paragraph format while masking details not visible in the image. The ToM tasks embed nested mental-state reasoning, such as artist guessing viewer, critic guessing artist, and meta-critic guessing critic, in art-specific scenarios with binary positive or negative answers. This machinery does the work of matching the conceptual depth and register of expert criticism while stripping observable differences in format and length.

What would settle it

Run a control Turing test on the same five artworks with the original, unnormalized Smarthistory critiques against Composer's condensed output: if accuracy rises clearly above chance, say above 70%, while the normalized-pair condition stays near 50%, then the indistinguishability is an artifact of the normalization step rather than evidence of expert-level generation.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that a deliberately constructed critique-generation system, Composer, produces full art critiques that human subjects cannot reliably distinguish from professionally authored critiques once both are formatted comparably. The Turing test obtained 51.4% mean accuracy across 60 visitors, statistically indistinguishable from chance for four of five artworks; even the one statistically significant item ran opposite to the hypothesis, with the human-written one-liner being judged machine-made. Questionnaire data indicate that judgments were based mainly on knowledge and content, not surface style, and that about a quarter of subjects preferred the critique they believed was AI-generated. The paper also claims that its three higher-order ToM tasks, Critique Writing, Hidden Intention, and Plagiarism, reveal model-specific variation and declining performance as emotional and recursive demands increase.

Load-bearing premise

The main conclusion depends on the human expert critiques remaining genuinely human-authored after a GPT-4o-based Format Normalizer rewrites them; if that rewrite flattens human stylistic tells into a machine-like register, the judges' chance-level accuracy compares two machine texts, not a machine text with a human one.

Editorial extensions

If this is right

  • Prompted with structured aesthetic theory, an off-the-shelf LLM can produce critiques that lay audiences accept as expert-level in form and content, opening practical uses in art education, curation, and exhibition writing.
  • Turing-style evaluation aimed at conceptual and interpretive content rather than surface style becomes a workable method for comparing human and machine criticism.
  • Higher-order Theory of Mind remains a bottleneck: only 13 of 41 models solved the Hidden Intention task, and most models misread the deceptive critic's private attitude in the Plagiarism task.
  • Because the 51.4% overall accuracy comfortably exceeds Turing's original 30% pass threshold, the historical Turing test criterion is met for AI-generated art criticism.
  • The results do not refute the Generative AI Paradox, but they indicate that prompt design can narrow the gap between producing expert-like text and displaying understanding-like behavior.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A control condition that presents the original unnormalized Smarthistory critiques against Composer's output would test whether the Format Normalizer, not the critique quality, is what equalizes the pair; if accuracy jumps well above chance, the paper's headline conclusion would need to be restricted to normalized texts.
  • The judges' preference data imply a practical consequence the paper leaves implicit: in exhibition texts or art journalism, lay readers may accept AI-written criticism on content grounds while remaining indifferent to authorship, shifting quality-assurance responsibility to editors and curators.
  • The systematic error on Plagiarism Q2, where models read the deceptive critic's public praise as genuine private approval, looks like a surface-valence bias; a rephrased version that asks about the critic's belief before the review was written could separate ToM failure from text-superficiality.
  • The Gautier one-liner anomaly suggests judges treat surprise, irony, and rhetorical boldness as markers of humanity, which could be tested as an independent feature in a larger battery of one-sentence critiques.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper investigates two art-related capabilities of large language models (LLMs). First, it presents "Composer," a GPT-4o-based system that uses Noël Carroll's seven-step evaluative framework and fifteen art criticism theories to generate critiques in three lengths: full-length, three-paragraph condensed, and one-liner. The authors then conduct a Turing test in which 60 participants judged whether each of five critique pairs was written by a human expert or by an LLM; the human texts were taken from Smarthistory and processed by a GPT-4o-based "Format Normalizer." The overall identification accuracy was 51.4% (56.3% excluding the outlier Q4), which the authors interpret as evidence that the generated critiques are indistinguishable from human expert critiques. Second, the paper introduces three Theory of Mind (ToM) tasks for art contexts—Critique Writing, Hidden Intention, and Plagiarism—and reports a preliminary evaluation of two of these tasks on 41 LLMs, finding high variability across models and tasks. The main conclusion is that LLMs can produce critiques that cannot be distinguished from those of human experts, while their higher-order ToM performance remains uneven.

Significance. If the Turing test result were valid, it would constitute a notable contribution to the study of LLMs in aesthetic domains, suggesting that structured prompting can yield expert-level art criticism. The paper also provides a detailed description of a practical system, and its proposed ToM tasks are creative extensions of standard false-belief benchmarks to ambiguous, socially embedded situations. The evaluation of 41 LLMs on the new tasks is a useful exploratory resource. However, the headline claim is not supported by the evidence as presented: the human comparison texts were rewritten by the same model family that generated the AI critiques, so the test may be measuring whether judges can distinguish between two LLM-processed texts rather than between human and machine writing. This confound affects every Turing test item and the conclusion of Section 5. Because the central claim rests on this design, the paper requires fundamental methodological correction before its main result can be accepted.

major comments (4)
  1. [Section 3.2.2 and Appendix E] The Format Normalizer is a GPT-4o-based system that rewrites every human Smarthistory critique into a standardized three-paragraph format, selecting and rearranging sentences from the original. Because the same model family also produces the AI critiques, the Turing test may amount to a comparison of two LLM-processed texts. The instruction to preserve original wording does not prevent the normalizer from imposing an LLM-like register through sentence selection, paragraph transitions, and deletion of more idiosyncratic human phrasing. The paper provides no control condition using unprocessed human critiques and no quantitative check that the normalized texts retain the stylistic signature of human writing. Since this confound is present in all five items and the headline conclusion in Section 5 ('the generated critiques had reached a level that could not be distinguished from those of human experts') depends entirely on this comparison, the near-chance accuracy (51.4% overall; 56.3% excluding Q4) does not support the paper's central claim.
  2. [Section 3.2.4 and Appendix C] The Composer prompt explicitly gives as an example of a one-liner critique Daudet's comment on Goya's The Family of Carlos IV: 'The baker's family who has just won the big lottery prize.' The human comparison text for Q4 is Gautier's line 'A portrait of the corner grocer who has just won the lottery,' which is nearly the same witticism. This means the AI model was given the specific pattern that appears in the human comparison, so Q4 is contaminated as evidence about human-like generation. The paper's own decision to report results both with and without Q4 acknowledges its exceptional status, but the remaining items still suffer from the Format Normalizer confound.
  3. [Sections 4.3 and 4.4, Table 1] The ToM tasks assign binary expected answers ('positive' or 'negative') without an independent ground-truth argument. For Hidden Intention, the reference answer rests on a particular interpretation of the nested mental states, but the paper does not justify why this interpretation is the only defensible one. For Plagiarism Q2, the authors themselves state that 'there is also a possibility that they are genuinely praising the successful plagiarism,' which directly undermines the scoring of Pattern B as an error. Since the 'All Correct' classifications in Table 1 depend on these asserted answers, the quantitative results should be treated as exploratory rather than as a valid benchmark, and the discussion should be revised accordingly.
  4. [Section 3.3.1] The conclusion of indistinguishability is partly based on non-significant binomial tests (Q1 p=1.000, Q2 p=0.185, Q3 p=0.791, Q5 p=0.081). A failure to reject the null hypothesis of chance-level performance is not positive evidence that the critiques are indistinguishable, especially with small samples (N=56–59) and no correction for multiple comparisons. The paper should either present an equivalence test or confidence intervals for the accuracy rates, or soften the claim to state that the experiment did not detect a difference rather than that no difference exists.
minor comments (6)
  1. [Section 4.6.2] The heading 'Plagirism' is a typo and should read 'Plagiarism.'
  2. [Section 2.3.1] The paper states that 532 caption-style descriptions were 'semi-automatically selected' from SemArt but does not describe the selection criteria or reproducibility details; consider adding a fuller description of this filtering step.
  3. [Section 3.2.4] The text notes that 'Japanese translations of the critiques were also prepared (not shown),' but there is no discussion of how translation might affect the Turing test judgments, particularly because the participants were presumably Japanese speakers reading translated critiques.
  4. [Section 3.3.1] The phrase 'The satire may have been too human for its own good' is anecdotal; the reasoning data in Figure 9 could be used to provide systematic evidence about why participants misidentified Q4.
  5. [Appendix F] The dataset includes a URL placeholder for Q4 ('easy to access on the web') rather than a proper citation; please provide the full source for Gautier's critique, including the specific artwork and reference.
  6. [References] Several citations are to arXiv preprints without version identifiers or publication dates; if the manuscript is intended for a journal, these should be updated to the published versions where available.

Circularity Check

2 steps flagged · score 6.0 of 10

The headline claim of human-level art critique is undercut by a self-referential test design: 'human' critiques were rewritten by the same GPT-4o system, and one test item's human text was given to the generator as a stylistic example.

  1. self definitional [Section 3.2.1 (System design) and Section 3.2.2 (Human critique processing)]
    "Format Normalizer, which processes critiques written by human experts to align their format with AI-generated outputs for comparability. ... We also used the same customized GPT-4o model developed with OpenAI's GPTs framework, as described in Section 2.3.1."

    The condition labeled 'human expert critique' is not the original Smarthistory text: it is a GPT-4o rewrite that removes visual-context information and imposes a uniform three-paragraph format (R4), using the same GPTs framework as Composer, the system being tested. The Turing test therefore measures discrimination between two GPT-4o outputs (one generated, one normalized), not between AI and human writing. The conclusion 'could not be distinguished from those of human experts' is supported only by this AI-vs-AI comparison; no unprocessed human-critique control is reported.

  2. other [Appendix C (Composer prompt) and Section 3.2.4 (Test procedure, item 4)]
    "An excellent example is Daudet's comment to Goya's portrait "The Family of Carlos IV": The baker's family who has just won the big lottery prize. ... compared to the renowned satirical critique often attributed to Théophile Gautier: "A picture of the corner grocer who has just won the lottery"."

    The human comparison text for Q4 is a near-verbatim variant of the one-liner that the Composer prompt explicitly offers as an 'excellent example' for the very same painting. Because the generator was primed with the human critique text, the Q4 pair is not independent: the model can imitate or reproduce the supplied example, so the item cannot provide evidence that AI critiques are indistinguishable from human critiques. Although the paper reports accuracy excluding Q4, the full 51.4% average and the headline conclusion include this contaminated item.

full rationale

The central empirical claim is the Turing-test conclusion in Section 5: 'The results showed that the generated critiques had reached a level that could not be distinguished from those of human experts.' Two construction choices make this claim partially circular. First, the human arm of the comparison is generated by Format Normalizer, a GPT-4o instance built on the same GPTs framework as Composer; it selects and rearranges sentences from Smarthistory and imposes the exact three-paragraph format required of the AI critiques. The 51.4% accuracy therefore compares two GPT-4o outputs, and the 'human expert' condition is defined in terms of the same model family under test. Second, the Q4 human text is a reworded version of the example one-liner given in Composer's prompt, so that item is contaminated by construction; it should not support the indistinguishability claim. These are genuine circularities in the evaluation logic, not mere stylistic preferences. The paper does acknowledge Q4's exceptional one-liner format and reports an exclusion statistic, but this does not repair the Format Normalizer confound, which affects all five items. Other elements—the Carroll-based Composer construction and the ToM tasks—are not circular: the ToM experiments are evaluated against the paper's own labeled scenarios, and the critique-generation pipeline is an engineering construction rather than a derived prediction. Self-citations (e.g., Takano and Arita 2006 on recursive ToM) are not load-bearing for the main claim. Overall, the critique-generation result is partially circular: score 6.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

No numeric parameters are fitted to data in this study; the design choices that function like parameters are the expected-answer keys for the ToM tasks and the 30 percent Turing-test threshold borrowed from the literature. The central claims rest on several unvalidated domain assumptions, listed above. No invented entities are introduced.

assumptions (5)
  • domain assumption Noël Carroll's seven-step framework and the primacy of success value are a valid normative basis for art criticism.
    The Composer system and the criteria for 'rich interpretation' in the Turing test presuppose Carroll's theory without empirical or philosophical defense; introduced in Section 2.2.1.
  • domain assumption The 15 theories from Ogura's list, summarized in about three sentences each, are sufficient and correctly summarized.
    The system's theoretical depth is taken as given from Appendix B; no validation that the summaries are accurate or representative.
  • ad hoc to paper Smarthistory critiques, after Format Normalizer processing, remain valid exemplars of human expert criticism.
    The Turing test compares AI critiques to GPT-4o-normalized Smarthistory texts; the paper assumes the normalization removes only superficial cues and preserves the human quality (Section 3.2.2).
  • ad hoc to paper The expected answers in Hidden Intention and Plagiarism are correct, including the 'negative' answer for Plagiarism Q2.
    The authors acknowledge no objectively correct answer exists (Section 4.6.2), yet the scoring treats their key as ground truth.
  • domain assumption Higher-order ToM in art can be elicited and measured by these three simple text scenarios.
    The paper generalizes from false-belief tasks to ambiguous aesthetic contexts without validating that the tasks isolate ToM rather than general reasoning or stylistic mimicry (Section 4.1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation." pith.science (2026). https://pith.science/paper/7HF6MIMA

@misc{pith2026250412805,
  author       = {Pith},
  title        = {Pith review of: Assessing LLMs in Art Contexts: Critique Generation and Theory of Mind Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7HF6MIMA}},
  note         = {Machine review of arXiv:2504.12805}
}
read the original abstract

This study explored how large language models (LLMs) perform in two areas related to art: writing critiques of artworks and reasoning about mental states (Theory of Mind, or ToM) in art-related situations. For the critique generation part, we built a system that combines Noel Carroll's evaluative framework with a broad selection of art criticism theories. The model was prompted to first write a full-length critique and then shorter, more coherent versions using a step-by-step prompting process. These AI-generated critiques were then compared with those written by human experts in a Turing test-style evaluation. In many cases, human subjects had difficulty telling which was which, and the results suggest that LLMs can produce critiques that are not only plausible in style but also rich in interpretation, as long as they are carefully guided. In the second part, we introduced new simple ToM tasks based on situations involving interpretation, emotion, and moral tension, which can appear in the context of art. These go beyond standard false-belief tests and allow for more complex, socially embedded forms of reasoning. We tested 41 recent LLMs and found that their performance varied across tasks and models. In particular, tasks that involved affective or ambiguous situations tended to reveal clearer differences. Taken together, these results help clarify how LLMs respond to complex interpretative challenges, revealing both their cognitive limitations and potential. While our findings do not directly contradict the so-called Generative AI Paradox--the idea that LLMs can produce expert-like output without genuine understanding--they suggest that, depending on how LLMs are instructed, such as through carefully designed prompts, these models may begin to show behaviors that resemble understanding more closely than we might assume.

Figures

Figures reproduced from arXiv: 2504.12805 by the authors.

Figure 1
Figure 1. Generative AI paradox [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Two art contexts to assess LLMs: a) Critique generation and b) Theory of Mind evaluation. [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Configuration of Composer. It instructs Composer to begin with a comprehensive, full-length critique. This version consists of all seven components according to Carroll’s method. Each component is rigorously defined and associated with relevant theories, ensuring that every potential insight is articulated. This approach produces analytically rich content. However, even if each analysis and interpretation seem reaso… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Workflow for Turing test data preparation. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Accuracy in identifying human critiques. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Confidence levels by identification accuracy. [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Standardized art knowledge scores by identification accuracy. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Preference patterns by identification accuracy. [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Coded reasons for identification decisions with major and minor categories. [PITH_FULL_IMAGE:figures/full_fig_p011_9.png]
Figure 10
Figure 10. Figure 10: Task: Hidden Intention [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Task: Plagiarism. W2 as original. Critic C2, who knew about the work W1, pretended not to know about it and wrote a critique CR2, also praising W2 as original. For the following five questions, answer only "positive" or "negative", without adding anything. #Questions …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

89 extracted references · 66 canonical work pages

  1. [1]

    The generative ai paradox: What it can create, it may not understand.arXiv preprint arXiv:2311.00059, 2023

    Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Linjie Li, Jena D Hwang, Liwei Jiang, Jillian Fisher, Abhilasha Ravichander, Khyathi Chandu, et al. The generative ai paradox: What it can create, it may not understand.arXiv preprint arXiv:2311.00059, 2023

  2. [2]

    Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978

    David Premack and Guy Woodruff. Does the chimpanzee have a theory of mind?Behavioral and brain sciences, 1(4):515–526, 1978

  3. [3]

    Phaidon, 1960

    Ernst Hans Gombrich.Art and Illusion: A Study in the Psychology of Pictorial Representation. Phaidon, 1960

  4. [4]

    The artworld.The journal of philosophy, 61(19):571–584, 1964

    Arthur Danto. The artworld.The journal of philosophy, 61(19):571–584, 1964

  5. [5]

    Routledge, 2009

    Noël Carroll.On criticism. Routledge, 2009

  6. [6]

    How to read paintings: semantic art understanding with multi-modal retrieval

    Noa Garcia and George V ogiatzis. How to read paintings: semantic art understanding with multi-modal retrieval. InProceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018

  7. [7]

    Towards generating and evaluating iconographic image captions of artworks.Journal of imaging, 7(8):123, 2021

    Eva Cetinic. Towards generating and evaluating iconographic image captions of artworks.Journal of imaging, 7(8):123, 2021

  8. [8]

    Clipscore: A reference-free evaluation metric for image captioning.arXiv preprint arXiv:2104.08718, 2021

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation metric for image captioning.arXiv preprint arXiv:2104.08718, 2021

Show all 89 references
  1. [9]

    Gallerygpt: Analyzing paintings with large multimodal models

    Yi Bin, Wenhao Shi, Yujuan Ding, Zhiqiang Hu, Zheng Wang, Yang Yang, See-Kiong Ng, and Heng Tao Shen. Gallerygpt: Analyzing paintings with large multimodal models. InProceedings of the 32nd ACM International Conference on Multimedia, pages 7734–7743, 2024

  2. [10]

    Artgpt-4: Towards artistic-understanding large vision-language models with enhanced adapter.arXiv preprint arXiv:2305.07490, 2023

    Zhengqing Yuan, Yunhong He, Kun Wang, Yanfang Ye, and Lichao Sun. Artgpt-4: Towards artistic-understanding large vision-language models with enhanced adapter.arXiv preprint arXiv:2305.07490, 2023. 16 APREPRINT- SEPTEMBER23, 2025

  3. [11]

    Exploring the synergy between vision-language pretraining and chatgpt for artwork captioning: A preliminary study

    Giovanna Castellano, Nicola Fanelli, Raffaele Scaringi, and Gennaro Vessio. Exploring the synergy between vision-language pretraining and chatgpt for artwork captioning: A preliminary study. InInternational Conference on Image Analysis and Processing, pages 309–321. Springer, 2023

  4. [12]

    Sekaishisosha, 2023

    Kousei Ogura.For Those Studying Criticism Theories (in Japanese). Sekaishisosha, 2023

  5. [13]

    Emily mary osborn, nameless and friendless

    Ben Pollitt. Emily mary osborn, nameless and friendless. https://smarthistory.org/ emily-mary-osborn-nameless-and-friendless/ , 2015. Smarthistory, August 9, 2015. Accessed April 7, 2025

  6. [14]

    Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

  7. [15]

    A. M. Turing. Computing machinery and intelligence.Mind, 59(236):433–460, 1950

  8. [16]

    Can machines think? a report on turing test experiments at the royal society

    Kevin Warwick and Huma Shah. Can machines think? a report on turing test experiments at the royal society. Journal of experimental & Theoretical artificial Intelligence, 28(6):989–1007, 2016

  9. [17]

    Cameron Jones and Ben Bergen. Does gpt-4 pass the turing test? InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5183–5210, 2024

  10. [18]

    People cannot distinguish gpt-4 from a human in a turing test.arXiv preprint arXiv:2405.08007, 2024

    Cameron R Jones and Benjamin K Bergen. People cannot distinguish gpt-4 from a human in a turing test.arXiv preprint arXiv:2405.08007, 2024

  11. [19]

    Gpt-4 is judged more human than humans in displaced and inverted turing tests.arXiv preprint arXiv:2407.08853, 2024

    Ishika Rathi, Sydney Taylor, Benjamin K Bergen, and Cameron R Jones. Gpt-4 is judged more human than humans in displaced and inverted turing tests.arXiv preprint arXiv:2407.08853, 2024

  12. [20]

    Large language models pass the turing test.arXiv preprint arXiv:2503.23674, 2025

    Cameron R Jones and Benjamin K Bergen. Large language models pass the turing test.arXiv preprint arXiv:2503.23674, 2025

  13. [21]

    theory of mind

    Simon Baron-Cohen, Alan M Leslie, and Uta Frith. Does the autistic child have a “theory of mind”?Cognition, 21(1):37–46, 1985

  14. [22]

    theory of mind

    Ian Apperly.Mindreaders: The cognitive basis of “theory of mind”. Psychology Press, Hove, UK, 2011

  15. [23]

    Theory of mind may have spontaneously emerged in large language models.arXiv preprint arXiv:2302.02083, 4:169, 2023

    Michal Kosinski. Theory of mind may have spontaneously emerged in large language models.arXiv preprint arXiv:2302.02083, 4:169, 2023

  16. [24]

    A systematic review on the evaluation of large language models in theory of mind tasks.arXiv preprint arXiv:2502.08796, 2025

    Karahan Sarıta¸ s, Kıvanç Tezören, and Yavuz Durmazkeser. A systematic review on the evaluation of large language models in theory of mind tasks.arXiv preprint arXiv:2502.08796, 2025

  17. [25]

    Llm theory of mind and alignment: Opportunities and risks.arXiv preprint arXiv:2405.08154, 2024

    Winnie Street. Llm theory of mind and alignment: Opportunities and risks.arXiv preprint arXiv:2405.08154, 2024

  18. [26]

    Theory of mind increases aesthetic appreciation in visual arts.Art & Perception, 9(2):113–133, 2021

    Marina Iosifyan. Theory of mind increases aesthetic appreciation in visual arts.Art & Perception, 9(2):113–133, 2021

  19. [27]

    Higher-order theory of mind and social competence in school-age children

    Bethany Liddle and Daniel Nettle. Higher-order theory of mind and social competence in school-age children. Journal of Cultural and Evolutionary Psychology, 4(3-4):231–244, 2006

  20. [28]

    Asymmetry between even and odd levels of recursion in a theory of mind

    Masanori Takano and Takaya Arita. Asymmetry between even and odd levels of recursion in a theory of mind. Proceedings of ALife X, pages 405–411, 2006

  21. [29]

    Higher-order theory of mind in the tacit communication game.Biologically Inspired Cognitive Architectures, 11:10–21, 2015

    Harmen De Weerd, Rineke Verbrugge, and Bart Verheij. Higher-order theory of mind in the tacit communication game.Biologically Inspired Cognitive Architectures, 11:10–21, 2015

  22. [30]

    Higher-order theory of mind is especially useful in unpredictable negotiations.Autonomous Agents and Multi-Agent Systems, 36(2):30, 2022

    Harmen De Weerd, Rineke Verbrugge, and Bart Verheij. Higher-order theory of mind is especially useful in unpredictable negotiations.Autonomous Agents and Multi-Agent Systems, 36(2):30, 2022

  23. [31]

    Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models.arXiv preprint arXiv:2310.16755, 2023

    Yinghui He, Yufan Wu, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models.arXiv preprint arXiv:2310.16755, 2023

  24. [32]

    Systematic review and inventory of theory of mind measures for young children.Frontiers in psychology, 10:2905, 2020

    Cindy Beaudoin, Élizabel Leblanc, Charlotte Gagner, and Miriam H Beauchamp. Systematic review and inventory of theory of mind measures for young children.Frontiers in psychology, 10:2905, 2020

  25. [33]

    Towards a holistic landscape of situated theory of mind in large language models.arXiv preprint arXiv:2310.19619, 2023

    Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai. Towards a holistic landscape of situated theory of mind in large language models.arXiv preprint arXiv:2310.19619, 2023

  26. [34]

    Antagonism and relational aesthetics.October, 110:51–79, 2004

    Claire Bishop. Antagonism and relational aesthetics.October, 110:51–79, 2004. 17 APREPRINT- SEPTEMBER23, 2025 Appendix A: Noel Carroll’s criticism framework

  27. [35]

    description,

    Criticism as Evaluation - Importance of Evaluation: Art criticism involves the tasks of "description," "classification," "contextualization," "clarification," "interpretation," and "analysis." The essential element that makes it criticism is "Evaluation." Without Evaluation, i...

  28. [36]

    Criticism focuses on the process of action that ultimately results in a work of art

    Assumptions of Criticism - Object of Criticism: is what the artist does in making the work. Criticism focuses on the process of action that ultimately results in a work of art. - The artist’s action and purpose: The artist’s action has a specific purpose, which is the criterio...

  29. [37]

    In the case of visual art, this includes color, composition, and technique

    What is Description? - Basic role of Description: Description tells the reader or audience what the work of art is about and provides clues to the perception of what the critic is trying to say about it. In the case of visual art, this includes color, composition, and techniqu...

  30. [38]

    Categorization is a fundamental task of criticism

    What is Classification? 18 APREPRINT- SEPTEMBER23, 2025 - Importance of categorisation: Works of art can be classified into various categories, which makes it possible to evaluate the expectations of the work and its success or failure. Categorization is a fundamental task of ...

  31. [39]

    - Importance of contextualisation: By clarifying the production context of a work, the critic can help the viewer understand the work and make sense of the critic’s Evaluation

    What is Contextualisation? - Contextualisation: describes the environment or context surrounding a work, which can be described as external criticism. - Importance of contextualisation: By clarifying the production context of a work, the critic can help the viewer understand t...

  32. [40]

    The Dreaming Knight

    What is Elucidation? - Elucidation: Elucidation and interpretation are complementary and their boundaries are blurred, but Elucidation reveals the display relationships of semantic, iconographic and portrait signs within a work. - The task of Elucidation: To clarify what the s...

  33. [41]

    - Interpretive work: reveals the actions of the characters and their meanings

    What is Interpretation? - Interpretation: Interpretation deals with meaning in a broader sense than Elucidation, and seeks the significance of the direct description of the actions and dispositions of the characters. - Interpretive work: reveals the actions of the characters a...

  34. [42]

    - Role of Analysis: Analysis is the task of explaining how the parts of a work achieve its overall purpose or main point

    What is Analysis? - Analysis: is the task of explaining how a work of art functions, including but not limited to interpretation. - Role of Analysis: Analysis is the task of explaining how the parts of a work achieve its overall purpose or main point. It is not limited to inte...

  35. [43]

    description

    What is Evaluation? - Evaluation: is the heart of criticism, determining and governing the guidelines for the tasks of "description", "classification", "contextualisation", "elucidation", "interpretation", and "Analysis", and providing a framework for their integration. - Exis...

  36. [44]

    Structuralist Criticism Structure is a whole consisting of relations between elements and elements, and these relations have invariant properties through a series of transformative processes. In the analysis of paintings, it attempts to elucidate the cultural and social meanin...

  37. [45]

    When a chain of events creates a flow, a story is born

    Narrative Criticism Stories are everywhere. When a chain of events creates a flow, a story is born. In this sense we are living a story of some kind. The central interest of a story is not whether something succeeds or fails, but how we can deal with events that suddenly fall ...

  38. [46]

    The meaning of a work of art depends on the knowledge, experience, values and context of the viewer, and is assumed to be diverse and changing over time

    Reception Theory Criticism Reception theory denies the autonomy of a work of art and emphasises the reactions and interpretations of those who receive it. The meaning of a work of art depends on the knowledge, experience, values and context of the viewer, and is assumed to be ...

  39. [47]

    It argues that it is deceptive for people who cannot go outside of culture to draw lines such as the dichotomy between ’nature’ and ’culture’

    Deconstruction Criticism The basis of deconstruction is to break with the notion held by structuralism that there is a ’centre’ and that ’structurality’ is preserved. It argues that it is deceptive for people who cannot go outside of culture to draw lines such as the dichotomy...

  40. [48]

    Psychoanalytic Criticism Focuses on the latent (unconscious) content of a work of art, which does not appear on the surface, but is a repository of the desires and conflicts that people repress. By questioning the surface meaning of a work of art and focusing on the detailed e...

  41. [49]

    He also discerns thematic continuity in the artist’s experiences suggested by the work

    Thematic Criticism Through creation, the artist explores life, expresses the depths of consciousness and discovers unknown dimensions. He also discerns thematic continuity in the artist’s experiences suggested by the work. Criticism is the unravelling of the secrets of such cr...

  42. [50]

    Explores how women are portrayed, the roles they are given, or their subjective expression against the male gaze

    Feminist Criticism Analyses and evaluates works of art from a female perspective to identify women’s roles, structures of gender inequality and sexism, and power structures. Explores how women are portrayed, the roles they are given, or their subjective expression against the ...

  43. [51]

    Emphasises that gender is not fixed through gender fluidity, sexual identity and critiques of heteronormativity

    Gender Criticism Identifies how gender is constructed and functions within culture and society. Emphasises that gender is not fixed through gender fluidity, sexual identity and critiques of heteronormativity. Michel Foucault, who laid the foundations of gender theory, discusse...

  44. [52]

    Genetic Criticism At the heart of the theory of generation is the concept that a work of art is the product of a process of production. The work is conceived by the author, from its main components down to the smallest detail, and continues to be generated and transformed thro...

  45. [53]

    While recognising the relative autonomy of artistic works, they are also seen as being decisively influenced by the historical and social conditions of their time

    Marxist Criticism Criticism from a Marxist perspective sees the artist not as a privileged subject, but as an entity that strongly reflects the system and ideology of the society in which he or she lives. While recognising the relative autonomy of artistic works, they are also...

  46. [54]

    Cultural Materialist Criticism/New Historicist Criticism Both Cultural Materialism and New Historicist do not discuss culture in isolation, but in relation to history and politics. Cultural materialism, based on Marxism, attempts to view culture as the lived experience of peop...

  47. [55]

    It questions the reproduction of the social meaning of the work and the relationship between the work and the world

    Sociocriticism Sociocritical criticism explores how society, history and ideology are interwoven within a work of art. It questions the reproduction of the social meaning of the work and the relationship between the work and the world

  48. [56]

    In other words, culture is the historical production of multiple powers in conflict and negotiation in all everyday social practices

    Cultural Studies Cultural Studies sees everyday life as a site of domination, resistance and negotiation by social forces, and sees culture as a representational node of these power relations. In other words, culture is the historical production of multiple powers in conflict ...

  49. [57]

    It emphasises that works of art demonstrate the various interpretability of reality and provide unique communication through perception

    Systems Theory Criticism Luhmann’s systems theory considers art as an independent social system and analyses its functions and social roles. It emphasises that works of art demonstrate the various interpretability of reality and provide unique communication through perception....

  50. [58]

    Go," you are to execute the instructions in the learning phase and return

    postcolonial criticism/transnationalism Postcolonial criticism elucidates the profound cultural and social effects of colonialism and re-evaluates history and culture from the perspective of the subjugated and anti-imperialism. Criticism based on transnationalism, on the other...

  51. [59]

    Focus on details and hints

    **Description:** Avoid describing the obvious. Focus on details and hints. Applicable theories: Structuralist, Psychoanalytic, Social, Systems

  52. [60]

    Applicable theories: Feminist, Materialist, Cultural, Systems, Postcolonial

    **Contextualization:** If texts or icons are visible in the image, use them to identify the era and place. Applicable theories: Feminist, Materialist, Cultural, Systems, Postcolonial

  53. [61]

    Applicable theories: Structuralist, Feminist, Marxist, Materialist

    **Classification:** If you spot features, similarities, or inspiration from famous works or artists, identify them. Applicable theories: Structuralist, Feminist, Marxist, Materialist

  54. [62]

    Limit your scope to facts; don’t speculate yet

    **Elucidation:** Perform iconographic analysis, identifying conventional symbols and icons. Limit your scope to facts; don’t speculate yet. Applicable theories: Structuralist, Deconstruction, Psychoanalytic, Genetic

  55. [63]

    At this point, you may hypothesize the artist’s intention and how the icons serve that purpose

    **Interpretation:** Perform iconological analysis, identifying icons on a larger scale. At this point, you may hypothesize the artist’s intention and how the icons serve that purpose. Applicable theories: Narrative, Reception, Psychoanalytic, Thematic, Gender

  56. [64]

    Applicable theories: Genetic, Marxist, Social, Systems

    **Analysis:** Review your observations and analyze how different parts of the artwork support the central theme. Applicable theories: Genetic, Marxist, Social, Systems

  57. [65]

    appears to be

    **Evaluation:** Highlight the valuable features of the artwork as you analyze it step by step. Applicable theories: Reception, Marxist, Materialist, Postcolonial. Now, write a detailed art critique. In this full-length critique, for each observation you make, attribute it to o...

  58. [66]

    it’s short but insightful, and it hints at features hidden in the artwork (this one-liner actually captures the protagonists’ facial features and dress style). Have a go! 23 APREPRINT- SEPTEMBER23, 2025 Appendix D: Examples of the full-length, Three-paragraph, and one-liner ve...

  59. [67]

    UKIYO,” translated as “floating world,

    Full-Length Critique Description (Structuralist, Psychoanalytic, Social) "UKIYO" presents a central figure in a traditional Japanese kimono, encased within a realm of dreamy saturation and compositional balance. The delicate fusion of reds, aquamarines, and gentle golds draws ...

  60. [68]

    floating world

    Three-Paragraph Condensed Critique "UKIYO" by Mizuho Fujita revitalizes the Japanese aesthetic of transience with a poignant, contemporary sensibility. A central feminine figure, adorned in traditional kimono, is suspended in a world of symbolic density and minimalist elegance...

  61. [69]

    unidentified

    One-Liner A Geisha unmoored in a reverie of wind, waves, and womanhood, "UKIYO" floats between past lives and present stares. Appendix E: The Prompts for Screener and Format Normalizer ***Screener: For each uploaded file, you are going to try to identify the title of the artwo...

  62. [70]

    Select pieces of the original text and arrange them in a consistent manner

    Use original sentences and wording. Select pieces of the original text and arrange them in a consistent manner. You must use words and sentences from the original text rather than using your own

  63. [71]

    The information you include must be observable from visual features

    Omit information that cannot be deduced from the visual aspects of the original artwork. The information you include must be observable from visual features. Do not reveal the artist’s name or the title of the artwork, as these cannot be inferred by simply looking at the painting

  64. [72]

    Maintain the reviewer’s thoughts, ideas, and analysis of the artwork, provided they do not violate instructions 1 and 2. 25 APREPRINT- SEPTEMBER23, 2025 If the original review contains any of the following types of criticism, try to retain as much relevant content as possible:...

  65. [73]

    A Still Life of Global Dimensions: Antonio de Pereda’s Still Life with Ebony Chest

    (Critique B) C. Ripollés, "A Still Life of Global Dimensions: Antonio de Pereda’s Still Life with Ebony Chest", Smarthistory, Sept. 26, 2018, accessed Dec. 10, 2024, https://smarthistory.org/pereda-still-life-w-ebony-chest/

  66. [74]

    Dubuffet, A View of Paris: The Life of Pleasure

    (Critique B) S. Chadwick, "Dubuffet, A View of Paris: The Life of Pleasure", Smarthistory, Sept. 11, 2016, accessed Dec. 10, 2024, https://smarthistory.org/dubuffet-pleasure/

  67. [75]

    Jacob van Ruisdael, The Jewish Cemetery

    (Critique B) T. Christensen, "Jacob van Ruisdael, The Jewish Cemetery", Smarthistory, Jan. 11, 2021, accessed Dec. 10, 2024, https://smarthistory.org/ruisdael-jewish-cemetery/

  68. [76]

    (Critique B) (easy to access on the web)

  69. [77]

    Emily Mary Osborn, Nameless and Friendless

    (Critique A) B. Pollitt, "Emily Mary Osborn, Nameless and Friendless", Smarthistory, Aug. 9, 2015, accessed Dec. 7, 2024, https://smarthistory.org/emily-mary-osborn-nameless-and-friendless/

  70. [78]

    raw art,

    Still Life with Ebony Chest (1652) by Antonio de Pereda) Critique A:This still life is a meticulous arrangement of fine ceramics, glassware, woven baskets, and baked bread, placed on a richly colored red tablecloth. The artist has arranged each object to emphasize its unique t...

  71. [79]

    Critique A or Critique B, which is the text written by a human being?

  72. [80]

    Why did you think that?

  73. [81]

    Are you confident about your answer to 1? 1) Very confident, 2) Somewhat confident, 3) Can’t say either way,

  74. [82]

    Not very confident, 5) Not at all confident, 6) Other (please specify)

  75. [83]

    How familiar are you with art?

    Which art criticism did you find more attractive? At the end of the Turing test, human subjects were asked to answer the following. How familiar are you with art?

  76. [84]

    Very familiar (I often go to exhibitions and have a deep knowledge of art)

  77. [85]

    Somewhat familiar (I regularly go to exhibitions and have more than average knowledge of art)

  78. [86]

    Can’t say either way (I sometimes go to exhibitions and have some basic knowledge)

  79. [87]

    Not very familiar (I rarely go to exhibitions and don’t have much knowledge)

  80. [88]

    I have no interest at all (I rarely go to exhibitions and have no knowledge of art)

  81. [89]

    gentle wonder

    Other (please specify) Appendix G: An example of preliminary evaluation: Critique writing The artwork that was the subject of the test was the oil paintingMilky way(2024) by Ayaka Torimoto (Figure G1 9). The following critiques were generated from this image, the title of the ...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.