{"id":"f3b33504-9909-4614-a3e4-01c9770f0395","arxiv_id":"2508.00239","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Audiences rated AI-generated visuals in a dance performance as more artistic when they were not told beforehand that AI was involved.","lead":"This study tested how telling audiences that generative AI helped create a live dance performance changes their ratings of its artistic value. Viewers who learned about the AI after the show rated its visuals as having more artistic merit than those told beforehand.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline effect rests on a single uncorrected Mann-Whitney test across 29 items in non-randomized, demographically mismatched groups; chance or group composition could explain it.","rationale":"The reader's weakest assumption is exactly the load-bearing issue I identify: the AI/Tell After vs AI/Tell Before comparison is non-randomized, the groups differ in age, and the key p-value would not survive a Bonferroni correction across 29 items. My stress-test adds specificity about the expected false-positive rate under the global null and proposes a concrete permutation-based reanalysis plus a covariate-adjusted model. Since the reader's verdict is already CONDITIONAL with medium correctness risk, and my concern reinforces rather than redirects that assessment, the verdict should remain UNCHANGED. I considered whether a deeper internal inconsistency exists (e.g., the corrupted Table 1 entry or the mismatch between the stated 29 items and the survey appendix) but these are reporting errors, not the central claim's core vulnerability. I did not identify a separate objection stronger than the one the reader already noted, so I agree with the reader's weakest assumption and verdict.","tokens_in":8446,"tokens_out":4021,"duration_ms":46528,"concrete_test":"Reanalyze the AI Tell-Before vs Tell-After contrast for all 29 baseline items using a permutation-based family-wise error correction (e.g., Westfall-Young) that accounts for correlations among items, and fit an ordinal logistic regression for the 'artistic merit' item with age and performance session as covariates. If the adjusted p for the artistic-merit item remains below .05 and the disclosure effect survives the inclusion of age/session, the concern is resolved; if the effect vanishes, the headline claim is attributable to chance or group composition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the AI/Tell After group rating 'The projected visuals demonstrated artistic merit' higher than the AI/Tell Before group (M=4.11 vs 2.83, Z=-2.501, p=.012, Table 1). For this to support the paper's conclusion, the difference must be caused by disclosure timing, not by which participants attended which performance or by statistical noise. Two facts make that assumption fragile. First, the groups were not randomly assigned; they were audiences at different scheduled performances, and the AI/Tell Before group was older on average (29 vs 23.89, Section 2.1). Given n=12 and n=9, a few participants could drive the effect. Second, this p-value comes from one of 29 Likert items tested with no multiple-comparison correction (Section 3). Under the global null, with 29 independent tests, the probability of seeing at least one p<.05 is about 0.77, and the expected number of such findings is 1.45; a single p=.012 is therefore not surprising by chance alone. The paper does not report effect sizes, corrected p-values, or any control for age/session, so the central claim is not yet distinguished from a chance or composition artifact. The corrupted entry in Table 1 ('Being informed...monetary valueThe projected visuals appeared random') also raises reporting-quality concerns, though it is not itself the main issue.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a 2x2 between-subjects field study (N=39) in which audience members watched one of four live dance performances that differed in whether the visuals and sound were created with generative AI or with traditional digital tools, and in whether the technology was disclosed before the performance or after the survey. The central claim is that participants who were told about AI after completing the survey rated the AI-generated visuals as having greater artistic merit than those told before (M = 4.11 vs. 2.83, Z = -2.501, p = .012, Table 1). The authors interpret this as evidence that awareness of AI involvement lowers perceived artistic value and argue that explainable AI research should focus on transparency about AI's presence and capabilities rather than only on algorithmic mechanics.","tokens_in":8663,"tokens_out":4769,"duration_ms":48603,"significance":"The research question is timely and relevant to HCI, XAI, and arts practice. The study's main strength is that it uses real, professionally performed dance pieces with identical choreography, and the authors provide a detailed appendix of prompts, system architecture, and the survey instrument. If the finding were robust, it would have practical implications for artists and curators and theoretical implications for how disclosure timing shapes value judgments. However, the central result currently rests on a small, non-randomized sample and a single uncorrected statistical test among many, so the contribution is best viewed as an exploratory case study rather than a demonstration of the claimed effect.","major_comments":[{"comment":"The headline comparison (AI/Tell After vs. AI/Tell Before) is one of 29 Mann-Whitney tests on individual Likert items, but no multiple-comparison correction is applied. The key p-value of .012 would not survive a Bonferroni threshold of 0.05/29 = .0017. Under the global null, the expected number of p<.05 findings among 29 independent tests is about 1.45, and the probability of at least one is roughly 0.77, so this single significant result is not strong evidence against chance. Please report adjusted p-values (e.g., Holm-Bonferroni or Benjamini-Hochberg) or, preferably, test a pre-specified composite hypothesis about artistic merit.","section":"Section 3 / Table 1"},{"comment":"Participants were not randomly assigned to conditions; they attended one of four scheduled performances. The AI/Tell Before group differs from the AI/Tell After group in mean age (29.0 vs. 23.89) and group size (12 vs. 9). The observed difference in artistic merit could therefore be due to session effects, self-selection, or demographic differences rather than disclosure timing. The paper should report comparability of all groups on age, gender, and prior AI/art experience, or include age as a covariate in a regression/ANCOVA. Random assignment within each performance session would be the most direct fix.","section":"Section 2.1 / Section 2.3"},{"comment":"One row in Table 1 is corrupted: \"Being informed about the production of a piece of art would impact its monetary valueThe projected visuals appeared random (Z = -2.698, p = .007...)\" appears to concatenate two separate survey items. This raises doubts about the accuracy of the table as a whole. The authors should reconstruct the table directly from their analysis output and verify every row, and fix the malformed \"p = <.001\" entry.","section":"Table 1"},{"comment":"No effect sizes are reported for any comparison, despite very small group sizes (n = 12 and n = 9 for the key comparison). With samples this small, a one-point mean difference can be driven by one or two participants. Please report rank-biserial correlation or Cliff's delta with 95% confidence intervals, and consider a leave-one-out sensitivity analysis to assess the stability of the headline result.","section":"Section 3 / Table 1"},{"comment":"The paper states that the Tell After condition added \"six additional Likert scale questions,\" but the survey appendix lists several additional items that are not standard 5-point Likert scales: the two \"Which of the following...\" items use 1-to-5 anchors with qualitatively different endpoint labels, the \"I noticed a mapping...\" item has Yes/No/I don't recall responses, and one item is open-ended. Clarify exactly which six items were included in the Mann-Whitney analyses and how the non-Likert items were handled, since this affects the reproducibility of the statistical tests.","section":"Section 2.2 / Appendix B.2"}],"minor_comments":[{"comment":"The phrase \"we uncovered the mixed opinions\" overstates the contribution of a small exploratory survey; consider phrasing that matches the tentative nature of the evidence.","section":"Abstract"},{"comment":"The procedure says informed consent documents were distributed after the performance, which means audience members did not consent to participate before being exposed to the experimental manipulation. Please clarify the IRB-approved consent procedure, including any debriefing process.","section":"Section 2.3"},{"comment":"The Non-AI comparison row contains a stray \"]).\" at the start of the entry for \"I tried to make a connection...\"; this typo should be corrected.","section":"Table 1"},{"comment":"The \"Used Responses for Scene 3\" bullets appear to describe wave mappings (\"amplitude of the waves,\" \"wave's direction\") rather than stars, and one bullet refers to \"the third dance move's resemblance probability.\" Please verify that these are indeed the responses used for Scene 3 and correct any mislabeling.","section":"Appendix B.1"},{"comment":"Reference [17] contains a URL with spaces and line breaks; clean it up for the camera-ready version.","section":"References"},{"comment":"The paper does not report internal consistency (e.g., Cronbach's alpha) for the survey items used as dependent variables, which is relevant given that items are analyzed individually and many appear to tap similar constructs.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is an exploratory workshop-style case study, and its headline claim currently rests on weak statistical foundations: uncorrected multiple testing and non-random assignment. The corrupted Table 1 entry should be checked against raw data before publication. I do not see evidence of intentional misrepresentation, but the framing in the abstract and discussion should be tempered to match the actual evidence. With a reanalysis that addresses multiplicity and group comparability, or a substantial caveat rewording, the paper could become a useful contribution to the XAIxArts workshop; in its current form, the central claim is not adequately supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: a small but honestly reported case study with a clean 2x2 design; the headline effect is plausible but rests on one uncorrected test across 29 items in non-randomized groups.\n\nWhat's new: the tell-before/tell-after manipulation within the AI condition holds the visual content constant, which is a nice control. The paper ships the full survey and ChatGPT prompts, so the method is reproducible. The writing is direct about limitations.\n\nSoft spots: the key artistic-merit result (p=.012) comes from a single Mann-Whitney on 29 Likert items, no correction. Under the global null you'd expect roughly one p<.05 by chance, so p=.012 isn't surprising. Groups weren't randomized: different performances, and the AI/tell-before group is older (29 vs 23.89). With n=12 and n=9, a couple of participants can move the result. No effect sizes, no corrected p-values, no control for age or session. There's also a corrupted entry in Table 1 ('Being informed...monetary valueThe projected visuals appeared random') that needs fixing. The paper cites prior work on attention but not the existing literature on AI attribution lowering perceived quality, so it underplays how expected this finding is.\n\nWhat holds up: the direction of the effect matches prior findings, and the pattern across items (AI told after rated artistic merit higher, AI told before rated visuals more random) is internally consistent. The paper doesn't overclaim; it calls itself a case study and a call to the XAI community.\n\nWho this is for: people working on XAI for creative contexts, and researchers studying audience perception of AI art. It's a useful data point, not a field-rearranging result.\n\nRecommendation: deserves a serious referee, but only with the expectation of revision. The authors should run corrections, report effect sizes, and ideally analyze the data with a mixed model or permutation test that respects the non-random assignment. If I were handling it, I'd send it out.","headline":"A small but honest case study with a clean 2x2 design; the headline effect is plausible but rests on one uncorrected test in non-randomized groups.","tokens_in":9217,"tokens_out":1554,"would_cite":false,"duration_ms":15388,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that audiences grant more artistic merit to generative-AI visuals when they learn only after the performance that AI was used, and that this disclosure timing effect should shape how AI art is presented and explained.","keywords":["generative AI art","audience perception","disclosure timing","artistic merit","live dance performance","explainable AI","human-computer interaction"],"falsifier":"A preregistered replication with random assignment to disclosure timing and a single pre-specified measure of artistic merit—analyzed with a correction for multiple comparisons—would settle the claim: if the tell-after group still outrates the tell-before group, disclosure timing is the active ingredient; if the gap vanishes, group composition or chance was responsible.","tokens_in":8179,"feed_emoji":"🎭","tokens_out":7549,"duration_ms":74294,"temperature":0.7,"pith_summary":"The paper sets out to show that audiences' knowledge of generative AI's role changes how they value an artwork, independent of the artwork itself. It compares four live dance performances built from the same choreography: the visuals and sound were driven either by generative AI or by traditional digital tools, and viewers were told about the technology either before the show or after filling out the survey. The reported difference is that viewers who learned about the AI after the survey rated the projected visuals' artistic merit higher than those told beforehand (M=4.11 vs M=2.83), while the told-before group was more likely to call the visuals random. The paper takes this as evidence that disclosure timing shapes aesthetic judgment, and argues that explaining AI in the arts should address AI's presence and social context, not just its internal mechanics.","feed_headline":"Viewers rate AI-made visuals more artistic when told afterward","feed_subtitle":"A live dance study shows disclosure timing shifted artistic-merit ratings; how AI is framed may shape its reception.","key_machinery":"The load-bearing device is a 2x2 between-subjects disclosure design: two versions of a twelve-minute dance performance, identical choreography and structure, differ only in whether generative AI or a technologist made the creative visual and sound mapping decisions, and audiences learn which version they saw either before the performance or after the survey. The measurement instrument is a 29-item Likert-scale survey probing whether the projected visuals and sound seemed creative, meaningful, random, distracting, or artistically meritorious, with pairwise group differences analyzed by Mann-Whitney tests. The design's power is that it holds the performed artwork constant and varies only the viewer's knowledge of the technology's role, so any difference between tell-before and tell-after groups is attributable, in the paper's logic, to that knowledge.","core_discovery":"The study's central claim is that the same generative-AI visuals are judged differently depending on when viewers are told AI was involved. In the AI condition, participants who were informed after completing the survey agreed more strongly that 'The projected visuals demonstrated artistic merit' than participants informed before the performance, and the before group agreed more strongly that the visuals 'appeared random.' The paper interprets these differences as a bias against known AI involvement: once viewers know a machine contributed, they search for intention differently and grant less artistic credit. It also reports that among told-after audiences, the AI performance triggered more curiosity about how the visuals were made, while the non-AI version was rated higher on the sound complementing the performance. This is presented as a case study, not a general law, and the paper calls on explainable-AI work to include the viewer's prior knowledge as part of the explanation.","pith_inferences":["Left implicit in the paper is a packaging dilemma: if disclosure after the experience raises artistic-merit ratings, artists and platforms face a transparency trade-off between honest upfront labeling and maximizing perceived value, and the field will need normative guidance on which should win.","The design mixes a large-language model and separately trained neural networks under one 'AI' label, so an obvious extension is to isolate whether the disclosure effect is driven by audience beliefs about AI as a category or by the specific generation mechanism used.","A testable extension would carry the same disclosure manipulation to static media such as images or music to see whether the effect is tied to live embodied performance or generalizes across art forms."],"forward_implications":["Artists and curators who disclose AI involvement before a performance should expect lower artistic-merit ratings from audiences than if the same information comes after the experience, all else equal.","User studies of AI art that announce AI use upfront may systematically understate the aesthetic value audiences would otherwise assign to the work.","Explainable-AI practice in the arts should treat when and how AI's presence is revealed as part of the explanation, alongside any account of model mechanics.","The higher curiosity ratings in the AI/Tell-After group suggest that undisclosed AI can make audiences more inquisitive about the work's creation, not merely more approving."],"supporting_citations":[{"why":"Supplies the learned-value account used to explain why disclosing technology before the performance made Non-AI visuals more attention-grabbing and distracting.","marker":"[1]"},{"why":"Supports the interpretation that viewers can attribute intentional meaning to unintended AI outputs when context is withheld.","marker":"[2]"},{"why":"Describes the machine-learning tool used to train the neural networks that classified dance subroutines in the AI performance.","marker":"[6]"},{"why":"Provides the Likert-scale measurement technique on which the audience survey is built.","marker":"[14]"},{"why":"Defines the Mann-Whitney test used for pairwise comparisons between audience groups.","marker":"[16]"},{"why":"Supplies the generative-AI prompts and adopted responses that produced visuals and data mappings in the AI condition.","marker":"[18]"},{"why":"Provides the ranking-based comparison method that underwrites the Mann-Whitney statistics reported.","marker":"[22]"}],"fun_headline_variants":["Disclosure timing shifts how dance audiences judge AI visuals","Told later, dance audiences see AI visuals as more artistic","Knowing about AI beforehand lowers artistic ratings of dance visuals","Surprise AI in dance: viewers credit it more after the fact"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes the rating gap between the AI/Tell-Before and AI/Tell-After groups was caused by when disclosure happened, rather than by their pre-existing differences—the groups attended different performances, were not randomly assigned, differed in mean age, and 29 survey items were tested without a multiple-comparison correction.","fun_headline_variants_meta":{"raw":{"variants":["Disclosure timing shifts how dance audiences judge AI visuals","Told later, dance audiences see AI visuals as more artistic","Knowing about AI beforehand lowers artistic ratings of dance visuals","Surprise AI in dance: viewers credit it more after the fact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00069,"raw_usage":{"total_tokens":3086,"prompt_tokens":869,"completion_tokens":2217,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":2149}},"tokens_in":485,"tokens_out":2217,"duration_ms":15986,"temperature":1.0,"reasoning_tokens":2149,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:16:55.015628+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A preregistered replication with random assignment to disclosure timing and a single pre-specified measure of artistic merit—analyzed with a correction for multiple comparisons—would settle the claim: if the tell-after group still outrates the tell-before group, disclosure timing is the active ingredient; if the gap vanishes, group composition or chance was responsible.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the learned-value account used to explain why disclosing technology before the performance made Non-AI visuals more attention-grabbing and distracting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the interpretation that viewers can attribute intentional meaning to unintended AI outputs when context is withheld."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the machine-learning tool used to train the neural networks that classified dance subroutines in the AI performance."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Likert-scale measurement technique on which the audience survey is built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the generative-AI prompts and adopted responses that produced visuals and data mappings in the AI condition."},{"cited_title":"If I had to make petals of a painting move according to the respiratory rate of a dancer how should I map this respiratory rate to the petals?","cited_arxiv_id":null,"evidence_quote":"Provides the ranking-based comparison method that underwrites the Mann-Whitney statistics reported."}],"review_version":1}