{"id":"856aeb82-4b8f-4cd7-a12f-92faefdd5fec","arxiv_id":"2502.03253","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"When LLMs rate STEM solution originality, showing example solutions improves accuracy but sharply increases correlations among creativity facets to near 1, unlike human raters.","lead":"This paper compares how human experts and large language models rate the originality of STEM design solutions, and whether showing example solutions changes their ratings. It finds that LLMs collapse the distinct creativity facets of remoteness, uncommonness, and cleverness into a single score, while humans keep them more separate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Facet homogenization may be a prompt/format artifact: LLM originality, uncommon, remote, and clever ratings are generated jointly in one pass, so near-1 correlations may reflect response anchoring rather than evaluation processes.","rationale":"The paper's strongest and most novel claim is the dissociation between predictive and construct validity in LLM creativity ratings. The evidence for the construct-validity part is almost entirely the pairwise facet-originality correlations in Study 2. Those correlations are computed from outputs of a single prompt that asks for all four ratings simultaneously. This design creates a direct threat to the inference: the high correlations could be a property of the measurement instrument, not of the model's evaluation. The reader's weakest assumption concerns LLM-based coding of linguistic markers in explanations. That is relevant to the process claims (e.g., comparative language), but the homogenization claim does not depend on those linguistic markers; it depends on the facet ratings themselves. Thus the prompt-format confound is more load-bearing for the central claim. The paper's own limitation paragraph flags prompt sensitivity, but only as a general caution. A concrete prompt-ablation test would settle whether the 0.99 correlations reflect true facet collapse or output formatting. If the test shows the effect is robust to prompt order and separation, the central claim stands; if not, the conclusion should be limited to the specific joint-prompt protocol. The study otherwise has strengths: two model families, supplementary Kappa analyses, and direct comparison with human ratings under matched conditions.","tokens_in":12415,"tokens_out":3651,"duration_ms":33936,"concrete_test":"Run Study 2 again on a subset of the same responses with each facet rated in a separate API call (or with randomized prompt orders), using identical facet definitions, and compare the resulting inter-facet correlations. If the example-condition correlations drop substantially (e.g., from >0.95 to <0.8), the homogenization effect is attributable to the joint-prompt format rather than to LLM evaluation processes. Also test a prompt with originality rated last instead of first; if the ordering changes the magnitude of the correlations, anchoring is at play.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of weaker construct validity rests on the near-perfect correlations among facet ratings and originality in Study 2. However, the LLM prompt (Figure 10) asks for ORIGINALITY, UNCOMMON, REMOTE, CLEVER, and EXPLANATION in a single response. These ratings are not independent measures: the model may anchor on the first rating or produce coherent values across all scales. The example-condition correlations of 0.99 (GPT-4o-mini; similar in Claude-3.5-Haiku, Figure 11) could therefore reflect a format-induced response style rather than a true homogenization of the constructs. Humans rate originality first and the facets afterwards in separate steps, so their moderate correlations are not directly comparable. The paper's own limitation section acknowledges that results 'may have been driven in part by the structure of the prompt,' but this possibility is not tested. Cohen's Kappa is computed on the same joint ratings, so it does not address output independence. Without a prompt-ablation or separate-call design, the claimed trade-off between predictive and construct validity is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two studies comparing how human experts and LLMs evaluate the originality of responses to engineering design problems (DPT). Raters provide originality scores, brief written explanations, and ratings for three facets of originality—uncommonness, remoteness, and cleverness—either with or without example solutions and ratings. Study 1 (72 Prolific participants with STEM degrees) finds that no-example experts use more comparative and causal/analytical language than example experts, and that the presence of examples changes some facet correlations. Study 2 runs a parallel protocol with GPT-4o-mini and Claude-3.5-Haiku using a structured prompt, finding that examples improve LLM accuracy against ground-truth originality scores but that facet-originality correlations rise to values near 0.99 in the example condition. The paper interprets this as a homogenization of the facets in LLM evaluation and argues that LLM scores have stronger predictive validity but weaker construct validity than human scores.","tokens_in":12604,"tokens_out":4698,"duration_ms":44369,"significance":"If the central claims survive scrutiny, the paper is a useful contribution: it moves beyond accuracy-only benchmarks for LLM creativity evaluation, offers a parallel human–LLM design, uses multiple LLM families, reports both Pearson and Kappa statistics, and makes code and data available. The fine-grained facet protocol and the attention to explanation text are valuable. However, the paper's headline claim about LLM facet homogenization is currently confounded by the joint-format prompt used for LLM ratings, and the linguistic-marker analyses rely on LLM annotations without human validation for the specific categories. The significance is therefore conditional on resolving these issues.","major_comments":[{"comment":"The central evidence for facet homogenization comes from correlations among LLM ratings that are generated in a single API call with a fixed format requiring ORIGINALITY, UNCOMMON, REMOTE, CLEVER, and EXPLANATION in one response. With temperature set to zero, the model is not producing independent measurements of four constructs; the near-1 correlations in the example condition (Figure 2 and Figure 11) may reflect response formatting, anchoring on the first rating, or prompt-induced consistency rather than a collapse of the underlying evaluation constructs. Cohen's Kappa, which is computed on the same jointly generated ratings, does not address output independence. The authors' own Future Work section acknowledges that the results 'may have been driven in part by the structure of the prompt,' but this possibility is not tested. Human participants, by contrast, rated originality first and then rated the facets in a separate step after writing explanations, so the human and LLM procedures are not directly comparable. A prompt-ablation design with separate calls per facet, or at least randomized output order, is needed before the claimed trade-off between predictive and construct validity is established.","section":"Study 2 Methods and Results; Figure 10; Future Work"},{"comment":"The linguistic-marker analyses in both studies use GPT-4o and Claude-3.5-Sonnet as annotators for past/future focus, perceptual details, causal/analytical language, comparative language, and cleverness, with no human-annotated gold standard for these specific prompts and categories. The paper cites Rathje et al. (2024) for the general validity of LLM psycholinguistic rating, but that validation does not cover the precise constructs used here (e.g., 'comparative language,' 'causal/analytical markers') nor the short, domain-specific explanations collected in this study. This assumption is load-bearing for the cognitive-process claims—e.g., that no-example experts rely more on memory retrieval because they produce more comparative language. I request either a human-annotation reliability check on a sample of explanations or a clear statement that the cognitive-process interpretation is provisional pending such validation.","section":"Study 1 Methods (linguistic markers) and Study 2 Results"},{"comment":"The criterion for 'true originality scores' is described as factor scores from Patterson et al. (2025), a manuscript under review. The current paper does not describe how these factor scores were derived, how items were selected or deduplicated, or how the ground-truth ratings relate to the newly collected human ratings. Since the conclusion that LLM originality scores have stronger predictive validity depends on this criterion, readers cannot currently verify the accuracy gain from examples. Please provide a description of the factor-scoring procedure and a direct link to the scoring data, or state clearly which parts of the dataset are available.","section":"Study 1 and Study 2 Methods (ground-truth originality)"}],"minor_comments":[{"comment":"The text reports 'U = 75076.5, p < 0.5' for Claude-3.5-Sonnet's causal/analytical comparison; this appears to be a typo for p > 0.05 or p = 0.5, and the exact p value should be reported.","section":"Study 1 Results"},{"comment":"Model naming is inconsistent: the text refers to 'GPT-4 O3' in Study 1 Methods and 'GPT-4O' elsewhere, and 'CLAUDE-3.5-SONNET' alternates with 'Claude-3.5-Sonnet.' Please unify the model names.","section":"Throughout"},{"comment":"The prompt in Figure 7 says 'Casual / analytical markers' while the text uses 'causal/analytical'; the typo could affect prompt interpretation and should be corrected.","section":"Figure 7 prompt"},{"comment":"Figure numbering is confusing: 'Figure 4' appears to refer to different displays in Study 1 Results and Study 2 Results, and the caption for the linguistic-marker comparison appears late in the supplementary material. Please renumber and cross-check all figure references.","section":"Figures"},{"comment":"The paper reports many pairwise Fisher's z tests without a multiple-comparison correction; because the correlation matrix contains ten pairs per model per condition, I recommend reporting corrected p values or noting which comparisons survive correction.","section":"Study 2 Results"}],"recommendation":"major_revision","confidential_remarks":"This is a well-motivated and readable paper with a clear protocol and useful supplementary materials. The main obstacle to acceptance is the joint-prompt confound in Study 2; a separate-call or prompt-order ablation would address the most serious concern. I would also ask the editor to ensure that the ground-truth dataset from the under-review manuscript is made available or its scoring procedure fully described. The human data collection and the comparison framework are solid enough that the paper is worth revising rather than rejecting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper does something genuinely useful: it runs the same fine-grained creativity-rating protocol (originality plus uncommonness, remoteness, cleverness) on 72 STEM experts and two LLM families, with and without example solutions, and compares not just scores but explanation text. That comparative design is new, and the human data collection is careful (exclusions, attention checks, detail on instructions). The observation that LLM facet correlations with originality rise to 0.99 in the example condition is striking and looks like a real phenomenon at first glance.\n\nBut the stress-test hits the load-bearing wall. The Study 2 prompt asks the LLM to output ORIGINALITY, UNCOMMON, REMOTE, CLEVER, and EXPLANATION in one pass. The near-perfect correlations among the four scales may simply be response anchoring: the model produces coherent, consistent numbers across all scales in a single generation. The paper's own limitation section says the results 'may have been driven in part by the structure of the prompt,' but the authors never test this. A prompt ablation or separate-call design would settle it. Without that, the central claim—that LLMs homogenize the facets and thus trade construct validity for predictive validity—is not established. The human data do show moderate correlations (0.45–0.67), so the contrast is interesting, but the LLM side of the contrast is confounded with format.\n\nThere are other soft spots, smaller ones. The linguistic-marker analysis relies on LLM ratings of explanation text (GPT-4o, Claude-3.5-sonnet) without validating those labels against human annotation for the specific prompts and categories used. A couple of the markers (cleverness, causal/analytical) are not standard LIWC categories. And the ground-truth factor scores come from an under-review manuscript (Patterson et al.). These make the cognitive-process interpretation of the explanation data more provisional than the prose suggests.\n\nNone of this means the paper is junk. The human Study 1 is well executed and the descriptive patterns in Figure 1 are informative. The authors are honest about budget constraints and prompt sensitivity. For the creativity-assessment community, the paper is a useful point of departure: it raises the question of whether LLM creativity scores are measuring the intended constructs, and it gives other researchers a concrete protocol to improve on.\n\nRecommendation: send it to peer review. The empirical design and the data are worth referees' time. But the referee report should require the authors to address the joint-prompt confound, either with an ablation or by changing the experimental design, before the homogenization interpretation is accepted.","headline":"Fine-grained human/LLM creativity comparison with a promising but untested homogenization claim; the joint-prompt confound needs an ablation before that interpretation is justified.","tokens_in":13158,"tokens_out":2254,"would_cite":false,"duration_ms":20110,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Language models can predict human creativity ratings more accurately than human experts can, but they do so by collapsing the three rated facets of creativity into a single dimension, especially when given example solutions with ratings.","keywords":["creativity assessment","large language models","text analysis","STEM","originality facets","remoteness","uncommonness","cleverness"],"falsifier":"Have human annotators label the same set of expert and LLM explanations for the five linguistic markers and compute agreement (e.g., Cohen's kappa) with GPT-4o and Claude-3.5-sonnet ratings. Low kappa (below roughly 0.6) on comparative or causal/analytical markers would undercut the claim that no-example experts rely more on memory retrieval. Alternatively, if an independent LLM evaluation with modified prompts or different examples shows facet–originality correlations below 0.8, then homogenization is not a fixed property of LLM creativity evaluation.","tokens_in":12233,"feed_emoji":"🤖","tokens_out":7848,"duration_ms":61480,"temperature":0.7,"pith_summary":"This paper asks whether human experts and large language models reason about creativity in the same way when they rate solutions to real-world engineering design problems. Using 72 STEM experts and two LLMs, the authors collected originality scores, written justifications, and ratings on three facets of creativity—remoteness, uncommonness, and cleverness. The central finding is that LLMs are more accurate at predicting ground-truth human originality scores, especially when given example solutions with ratings, but they achieve this by collapsing the three facets into a single dimension, with facet–originality correlations rising above 0.99. Humans, by contrast, keep the facets distinct and shift which facet matters most when examples are present. If this holds, automated creativity scoring that agrees with human averages can still be measuring something very different from what human raters mean by originality.","feed_headline":"LLMs predict creativity better than experts, but blur its facets","feed_subtitle":"With examples, LLM facet ratings correlate above 0.99 with originality: more accurate, but conceptually flat.","key_machinery":"The central object is the facet-based originality rating protocol: each solution is rated for originality plus remoteness, uncommonness, and cleverness on five-point Likert scales, and raters write one-to-two-sentence explanations of their originality score. The explanations are then coded by two LLMs (GPT-4o and Claude-3.5-sonnet) for five linguistic markers—comparative, causal/analytical, perceptual, past/future, and cleverness—using zero-shot prompts adapted from Rathje et al. The argument is carried by comparing Pearson correlations among facets and between each facet and originality across conditions (with versus without examples) and across populations (humans versus LLMs), with Fisher's z tests for significance. The key identity at stake is the view of originality as an aggregation of facets: if facet correlations approach 1.0, the facets no longer carry independent information, which the paper terms homogenization.","core_discovery":"The paper claims that LLM evaluators of creativity have stronger predictive validity but weaker construct validity than human experts. In Study 1, experts who rated design-problem solutions without examples used more comparative language and emphasized uncommonness, suggesting memory-based retrieval; those given examples shifted weight toward cleverness. In Study 2, GPT-4o-mini and Claude-3.5-haiku prioritized remoteness and uncommonness, and giving them example solutions raised their correlation with ground-truth originality from roughly 0.6–0.67 to 0.74–0.76. Yet examples also made the three facets nearly perfectly correlated with originality (above 0.99), meaning the models no longer distinguished between 'clever,' 'remote,' and 'uncommon'—the facets homogenized into a single originality judgment. The paper argues these patterns reveal diverging evaluation strategies: humans weigh facets differently depending on context, while LLMs reduce all facets to semantic distance from prior knowledge, especially when shown examples.","pith_inferences":["A testable extension is to probe LLM evaluators with 'trap' responses in which one facet is high and another low; if facet correlations remain near 1.0, the model is likely computing a single semantic-similarity score rather than three distinct constructs.","The contrast between humans shifting to cleverness and LLMs collapsing facets hints that LLM pretraining proxies 'originality' by distance from common responses, not by task-specific insight; this could be tested by fine-tuning on creativity-annotation data and re-running the facet analysis.","If the linguistic-marker labels are accepted, the comparative-language result offers a cheap way to detect memory-based evaluation in human raters, which could be used to calibrate rater training or flag AI-generated justifications."],"forward_implications":["Automated LLM creativity scoring can achieve high agreement with human averages while failing to model the conceptual distinctions human raters make; agreement on scores does not imply agreement on the basis of evaluation.","Low-temperature LLM evaluators produce less diverse explanation styles and more rigid, analytical justifications than human experts, even when their numeric ratings match.","Including few-shot example solutions improves LLM accuracy but amplifies facet homogenization, so benchmark gains in predictive accuracy can conceal a loss in construct validity.","Construct-level evaluation of AI judges—testing whether facets stay separable—is needed alongside accuracy metrics before deploying LLMs as reviewers of STEM proposals or papers."],"supporting_citations":[{"why":"Provides the Design Problems Task dataset and the ground-truth originality ratings that both human and LLM studies evaluate.","marker":"Patterson et al. (2025)"},{"why":"Defines the three facets (remoteness, uncommonness, cleverness) and the factor-scoring method used to compute the ground-truth originality scores.","marker":"Silvia et al. (2008)"},{"why":"Supplies the explanation-based analysis method and the linguistic-category approach that the paper adapts to LLM raters.","marker":"Orwig et al. (2024)"},{"why":"Justifies using LLMs to score psycholinguistic markers zero-shot, which is the basis for the text analysis of explanations.","marker":"Rathje et al. (2024)"},{"why":"Establishes the Consensual Assessment Technique, the framework that motivates using expert ratings as the criterion of creativity.","marker":"Amabile (1982)"},{"why":"Provides the few-shot learning rationale that frames the example condition in Study 2 as analogous to giving exemplars to human raters.","marker":"Brown et al. (2020)"}],"fun_headline_variants":["AI judges creativity accurately but flattens its facets","LLMs predict creativity, but see all facets as one","Better AI creativity scores, but a one-note evaluation","Experts weigh facets separately; LLMs blend them into one","LLMs: accurate creativity prediction, but facet collapse"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that the LLM-generated labels for linguistic markers in the explanations—comparative, causal/analytical, perceptual, past/future, and cleverness—are valid measures of the cognitive processes raters used, yet these labels were not validated against human-annotated gold standards for these specific categories and prompts.","fun_headline_variants_meta":{"raw":{"variants":["AI judges creativity accurately but flattens its facets","LLMs predict creativity, but see all facets as one","Better AI creativity scores, but a one-note evaluation","Experts weigh facets separately; LLMs blend them into one","LLMs: accurate creativity prediction, but facet collapse"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000684,"raw_usage":{"total_tokens":3143,"prompt_tokens":1024,"completion_tokens":2119,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":2040}},"tokens_in":640,"tokens_out":2119,"duration_ms":14598,"temperature":1.0,"reasoning_tokens":2040,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T05:20:26.997825+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human annotators label the same set of expert and LLM explanations for the five linguistic markers and compute agreement (e.g., Cohen's kappa) with GPT-4o and Claude-3.5-sonnet ratings. Low kappa (below roughly 0.6) on comparative or causal/analytical markers would undercut the claim that no-example experts rely more on memory retrieval. Alternatively, if an independent LLM evaluation with modified prompts or different examples shows facet–originality correlations below 0.8, then homogenization is not a fixed property of LLM creativity evaluation.","supporting_citations":[{"cited_title":", Pronchick, J","cited_arxiv_id":null,"evidence_quote":"Provides the Design Problems Task dataset and the ground-truth originality ratings that both human and LLM studies evaluate."},{"cited_title":", Beaty, R E","cited_arxiv_id":null,"evidence_quote":"Supplies the explanation-based analysis method and the linguistic-category approach that the paper adapts to LLM raters."},{"cited_title":", Mirea, D M","cited_arxiv_id":null,"evidence_quote":"Justifies using LLMs to score psycholinguistic markers zero-shot, which is the basis for the text analysis of explanations."},{"cited_title":"APACrefauthors \\ 1982","cited_arxiv_id":null,"evidence_quote":"Establishes the Consensual Assessment Technique, the framework that motivates using expert ratings as the criterion of creativity."}],"review_version":1}