{"id":"09e79da4-b0d2-455a-8aef-a84168ba0c2a","arxiv_id":"2507.15357","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A multi-dataset evaluation finds LLM performance on metaphor inference is driven more by lexical overlap and sentence length than by metaphor understanding.","lead":"This paper tests seven large language models on metaphor interpretation across five datasets, comparing performance on original metaphorical examples against literal paraphrases. It finds accuracy tracks lexical overlap and sentence length more than metaphorical content, suggesting surface features, not metaphor understanding, drive the results.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The -met vs -lit accuracy comparison is scored against original gold labels, but Section 6.2 documents paraphrases that change the entailment relation; without quantifying the label-shift rate, the central claim that surface features explain performance is not established.","rationale":"The paper's central claim is explicitly tied to the adversarial literal paraphrase comparison: the abstract and Section 7 assert that LLM performance is more influenced by lexical overlap and sentence length than by metaphor content, and that emergent metaphor understanding is a combination of surface features, in-context learning, and linguistic knowledge. The key empirical support is the consistent accuracy drop on -lit versions in Tables 2 and 10, interpreted via the increased Levenshtein distance and sentence length in Table 3. This reasoning is only valid if the -lit pair has the same entailment label as the original pair. The authors themselves identify label shifts in Section 6.2 and Table 4, but never quantify how often they occur. In the Meta4XNLI example, the paraphrase 'She worked as a model for Channel' changes the correct label from entailment to not_entailment relative to the unchanged hypothesis 'She was Chanel's muse'; scoring that instance against the original gold label penalizes the model for being correct on the literal version, artificially inflating the -met advantage. The same problem can bias any dataset, and the paper provides no sensitivity analysis. The second line of evidence in Table 3 is also confounded: the paraphrase manipulation varies sentence length, lexical overlap, and label validity simultaneously, so the reported covariance between Levenshtein distance and accuracy across five datasets cannot establish causation. I therefore agree with the reader's weakest assumption: label-shift contamination is the load-bearing concern. A re-annotation of the -lit gold labels is the decisive check. If the label-shift rate is substantial or if the corrected -lit accuracy reaches the -met accuracy, the abstract's strong claim should be rejected or substantially softened. If the corrected gap remains, the surface-feature account is provisionally supported. Thus the reader's CONDITIONAL verdict remains appropriate: the paper is valuable and largely well-executed, but its headline conclusion awaits a reanalysis that re-labels the literal paraphrases.","tokens_in":18379,"tokens_out":6277,"duration_ms":67955,"concrete_test":"Sample 120-150 literal-paraphrase instances from each of the five datasets (Meta4XNLI, Fig-QA, Figurative-NLI, FLUTE, IMPLI) and have two expert annotators (or a validated high-quality LLM with adjudication) judge the correct NLI label for the pair consisting of the literal premise/hypothesis and the unchanged counterpart. Compute (a) the label-shift rate, i.e., the proportion where the corrected label differs from the original gold label, and (b) the accuracy of the best model (Qwen2.5-72B-Instruct with CoT, as in Table 3) evaluated against the corrected labels. If the label-shift rate exceeds about 5%, the uncorrected -lit accuracies in Tables 2 and 10 are unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract, Section 7) is that LLM performance on metaphor interpretation is driven more by surface features (lexical overlap, sentence length) than by metaphorical content. The empirical pivot is the accuracy gap between original metaphorical datasets (-met) and literal paraphrases (-lit) in Tables 2 and 10: the -lit versions achieve lower accuracy, and this drop is attributed to increased Levenshtein distance and sentence length. This interpretation requires that each literal paraphrase preserves the inference label of the original pair. Section 6.2 and Table 4 explicitly document 'paraphrases that result in a label shift' (the Meta4XNLI example 'She was Channel's muse' becoming 'She worked as a model for Channel' changes the entailment to the unchanged hypothesis 'She was Chanel's muse'), yet the -lit accuracies in Tables 2 and 10 are still computed against the original gold labels. The frequency of such shifts is not reported. If label shifts are non-negligible, the -lit accuracy is systematically biased, and the observed gap in Table 2 could reflect changed ground truth rather than loss of metaphorical content or reliance on surface features. The quantitative analysis in Table 3 is likewise confounded: the literal paraphrase manipulation simultaneously changes sentence length, lexical overlap, and label validity. Thus the strongest wording in the abstract, that emergent metaphor understanding is a mirage, rests on an unvalidated measurement of the -lit condition.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper evaluates seven instruction-tuned LLMs on five English datasets for metaphor interpretation framed as Natural Language Inference (NLI) and Question Answering (QA), under zero-shot, few-shot, and chain-of-thought prompting. The authors also generate 'literal' paraphrases of the metaphorical sentences using Command R+ and compare model accuracy on the original metaphorical datasets (-met) and the paraphrased literal versions (-lit). They report that accuracy on -lit is generally lower, that the paraphrases have higher Levenshtein distance and longer average sentence length, and conclude that LLM performance is driven more by surface features such as lexical overlap and sentence length than by metaphorical content, arguing that purported emergent metaphor understanding is actually a combination of surface cues, in-context learning, and linguistic knowledge.","tokens_in":18607,"tokens_out":5303,"duration_ms":55013,"significance":"If the central claim were established, this would be an important cautionary result for the metaphor-processing evaluation community: current benchmarks would be shown to reward shallow pattern matching rather than deep figurative understanding. The paper's strengths include a broad multi-dataset, multi-model, multi-prompt evaluation; public release of code and data; and an honest error analysis that identifies concrete failure modes. The empirical basis, however, is currently not firm because the pivotal -met versus -lit comparison is scored against original gold labels despite documented label shifts, and because the paraphrase manipulation changes several properties at once. The work is useful as a comprehensive evaluation resource, but the headline claim about emergent abilities is stronger than the evidence supports.","major_comments":[{"comment":"The pivotal -met vs. -lit accuracy comparison in Tables 2 and 10 is scored against the original gold labels for both versions, but Section 6.2 explicitly documents paraphrases that change the entailment relation. The Meta4XNLI example makes the point: under the literal paraphrase 'She worked as a model for Channel for seven years', the hypothesis 'She was Chanel's muse' is genuinely not entailed, so the model's 'not_entailment' prediction is correct for the paraphrased pair and is counted as an error only because the original gold label is retained. The frequency of such label shifts is not reported, and any non-negligible rate systematically depresses -lit accuracy, contaminating the gap on which the central claim rests. The authors should quantify the label-shift rate, re-score the -lit sets with re-annotated labels, or exclude shifted instances and show that the conclusions are unchanged.","section":"§6.2, Table 4"},{"comment":"The literal paraphrase manipulation is not a clean removal of metaphor. Section 6.2 acknowledges that some generated paraphrases still contain metaphorical expressions (e.g., 'sharp' in the Fig-QA paraphrase and 'arrive' in the FLUTE paraphrase). The -lit condition therefore changes sentence length, lexical overlap, label validity, and residual metaphoricity simultaneously, so the observed accuracy drop cannot be uniquely attributed to the absence of metaphorical content. A controlled paraphrase set with human verification, or an instance-level analysis that separates paraphrases that are fully literal and label-preserving from those that are not, is required to support the surface-features interpretation.","section":"§4.2, §6.2"},{"comment":"The quantitative evidence for the surface-features claim is a cross-dataset comparison of aggregate Levenshtein distance and mean sentence length against mean accuracy for one model and prompt configuration (Qwen2.5-72B with CoT). With only five datasets, no statistical test or regression is reported, and dataset difficulty, label balance, and other properties are uncontrolled. The statement that performance is 'more influenced' by lexical overlap and sentence length than by metaphorical content is a causal claim that these correlations do not establish. The authors should provide per-instance analyses, a regression that includes both surface features and metaphor-related controls, or a controlled construction that varies one feature at a time.","section":"§6.1, Table 3"},{"comment":"The concluding claim that 'any alleged emergent abilities of LLMs to understand metaphorical language are the result of a combination of surface-level features, in-context learning, and linguistic knowledge' is stronger than the evidence. In Table 2, -lit accuracy is sometimes higher than -met accuracy (e.g., Fig-QA with Llama-3-8B-Instruct under CoT: 81.58 vs. 76.17), and the label-shift and residual-metaphor issues described above prevent the -met/-lit gap from being interpreted as a loss of metaphor understanding. The conclusions should be restricted to the demonstrated sensitivity to lexical overlap and sentence length, with the broader claim about the absence of metaphor understanding presented as a hypothesis requiring further study.","section":"Abstract, §7"}],"minor_comments":[{"comment":"The table contains formatting artifacts that should be cleaned up, including '87.57s' in the IMPLI-lit row, inconsistent column headers ('Qwen-7B' vs. 'Qwen2.5-7B'), and the spacing in 'A vg_met' and 'A vg_lit'.","section":"Table 2"},{"comment":"The FLUTE row does not clearly indicate whether metaphors occur in premises, hypotheses, or both; the 'Met loc.' column should be unambiguous for every dataset, especially since the paraphrase generation targets only sentences containing metaphors.","section":"Table 1"},{"comment":"The text states 'temperature=3' for Mistral-7B-Instruct during paraphrase generation; if this is accurate, it is an unusually high sampling temperature and should be justified, and if it is a typo for 0.3 it should be corrected.","section":"§4.2"},{"comment":"Several references contain author-name encoding or duplication problems, such as 'Coms, a' (should be Comşa), 'Grici¯ut˙e', and duplicate entries for Bollegala and Shutova (2013a/b) and Shutova (2010/2013); these should be normalized.","section":"References"},{"comment":"The legend and axis labels are small and the dual y-axes are not explained in the caption; please label the bars as accuracy, the lines as average sentence length, and identify the datasets more legibly.","section":"Figure 2"},{"comment":"The phrase 'Mistral-7B-Instruct is the worse performing model' should read 'the worst-performing model'.","section":"§5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is an honest empirical study with useful resources, but the central measurement is currently contaminated by label shift and residual metaphoricity. The fix—re-annotating or filtering the -lit sets and re-running the comparison—is clearly within the manuscript's scope and does not require new models or datasets. I would prioritize this over the stylistic issues. The authors' use of their own Meta4XNLI corpus is not a concern because the central claim does not depend on that dataset alone and the code/data release is a genuine asset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis one is worth a look, but the headline claim is bigger than the experiment. The genuinely new thing is the scope: seven models, five datasets, multiple prompt families, and a literal-paraphrase condition. That is a real improvement over the single-dataset studies that dominate this area. The paper also does a reasonable job of situating itself in the artifact literature, and it ships code and data. The finding that few-shot CoT outperforms fine-tuned encoders on most of these benchmarks is a useful data point.\n\nThe soft spot is the -lit comparison. The literal paraphrases are scored against the original gold labels, and the authors themselves show a paraphrase that changes the entailment relation (the 'muse' to 'model' example in Table 4). They never quantify how often that happens. If label shifts are not rare, the -met/-lit accuracy gap is partly an artifact of invalid references, not a measure of lost metaphorical content. The Table 3 correlation analysis is also suggestive rather than definitive: the paraphrase manipulation changes sentence length, lexical overlap, and label validity at once, and the correlation is computed over five datasets without error bars. The abstract's wording that the results 'demonstrate' the absence of emergent ability is too strong.\n\nCredit where due: the paper does not hide these problems. Section 6.2 and the limitations section acknowledge the paraphrase issues, and the manual error analysis is honest. So this is a fixable flaw rather than a fatal one. The authors could re-annotate a sample of paraphrases, report the label-shift rate, and either re-score with corrected labels or restrict the analysis to label-preserving paraphrases.\n\nWho is this for? People working on metaphor processing, NLI pitfalls, and benchmark design. It deserves a serious referee, not because the conclusion is settled, but because the empirical backbone is broad and the flaw is addressable. I would send it to review with a request for the re-analysis before acceptance.","headline":"Worth reading for the scope, not the headline; the literal-paraphrase comparison is contaminated by label shift, so the surface-features claim needs a re-analysis.","tokens_in":19131,"tokens_out":2579,"would_cite":true,"duration_ms":27473,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models appear to understand metaphor on current benchmarks, but the paper argues this is an artifact of lexical overlap and sentence length rather than figurative understanding.","keywords":["metaphor interpretation","large language models","natural language inference","question answering","lexical overlap","surface features","literal paraphrase","chain-of-thought prompting"],"falsifier":"A reader could have two annotators verify whether each automatically generated literal paraphrase preserves the gold entailment label and contains no metaphor, then recompute accuracy on the label-preserving subset; if literal accuracy remains at the level of the metaphorical accuracy on that subset, the central claim would be undermined. Alternatively, constructing literal and metaphorical stimuli matched for lexical overlap and sentence length and showing equal accuracy would falsify the surface-feature explanation.","tokens_in":18147,"feed_emoji":"🧠","tokens_out":6347,"duration_ms":66843,"temperature":0.7,"pith_summary":"Large language models appear to interpret metaphor well on standard benchmarks, but this paper tries to show that appearance is an artifact of surface-level properties. Across five datasets and seven models, accuracy on original metaphorical pairs is higher than on automatically produced literal paraphrases of the same items, and the pattern tracks lexical overlap and sentence length. The authors conclude that any alleged emergent ability to understand metaphorical language is better explained as a combination of surface-feature matching, in-context learning, and linguistic knowledge. This matters because current metaphor-interpretation benchmarks, built largely by lexical replacement, may measure pattern matching rather than figurative understanding.","feed_headline":"LLMs ace metaphor tests on word overlap, not understanding","feed_subtitle":"Across five datasets, literal paraphrases score below metaphors, showing surface cues drive results.","key_machinery":"The central mechanism is an adversarial control condition: every metaphorical sentence is rewritten as a literal paraphrase using an instruction-tuned model, creating a paired literal version of each dataset that is scored against the same gold labels as the original metaphorical version. The argument then uses two surface diagnostics, Levenshtein distance between premise and hypothesis as a proxy for lexical overlap, and average sentence length, to show that accuracy differences align with these features rather than with metaphor content. That is, the literal versions are longer and have lower overlap, and the models do worse on them; the paper takes this as evidence that the metaphor datasets' templatic, high-overlap structure is what carries performance.","core_discovery":"On the paper's own terms, the discovery is that LLMs' performance on metaphor interpretation is more sensitive to lexical overlap and sentence length than to the presence of metaphorical content. When the metaphor-containing sentences are rewritten as literal paraphrases, model accuracy generally falls, even though the paraphrase is meant to remove the figurative difficulty; meanwhile, datasets with higher premise-hypothesis overlap and shorter sentences yield higher accuracy. The paper reads this as evidence against emergent metaphor understanding, attributing the apparent ability to surface-level features, in-context learning, and linguistic knowledge. It also finds that few-shot and chain-of-thought prompting outperform fine-tuned encoder baselines on most of these benchmarks, further undermining the need for metaphor-specific training data.","pith_inferences":["A sharper test would hold length and overlap fixed while toggling metaphoricity, which no current dataset does.","The same lexical-substitution confound likely affects claims about other figurative phenomena, such as idiom and irony benchmarks, not just metaphor.","Because the literal paraphrases were machine-generated and sometimes contain metaphors or shift labels, the true gap between metaphorical and literal performance could be either larger or smaller than reported; manual paraphrase validation would settle which.","If human-authored literal paraphrases that preserve labels also lower accuracy, then the paper's surface-feature explanation would be supported; if they do not, part of the observed drop is an artifact of automatic paraphrase generation."],"forward_implications":["Benchmarks built by lexical replacement overstate LLM metaphor understanding, because high premise-hypothesis overlap alone can drive high accuracy.","Few-shot and chain-of-thought prompting can match or exceed fine-tuned encoder baselines on metaphor-interpretation tasks, so task-specific annotated training sets are less decisive than prompt design.","Metaphor-interpretation evaluations should report or control lexical overlap and sentence length before interpreting differences between models or conditions.","The 'emergent ability' framing should be replaced by a surface-feature-plus-knowledge explanation unless new evidence appears.","Naturally occurring metaphorical text provides a harder and more realistic test than lexicon-substituted data."],"supporting_citations":[{"why":"Supplies Figurative-NLI, the metaphor subset evaluated here, with entailment pairs constructed by lexical replacement.","marker":"Chakrabarty et al. (2021a)"},{"why":"Provides IMPLI, a gold NLI benchmark with metaphorical language, and is the source of the Levenshtein-distance-based overlap analysis.","marker":"Stowe et al. (2022)"},{"why":"Supplies Fig-QA, a Winograd-style figurative QA dataset used for evaluation and the fine-tuned RoBERTa-large baseline.","marker":"Liu et al. (2022)"},{"why":"Supplies FLUTE, the figurative NLI benchmark with entailments and contradictions used in evaluation.","marker":"Chakrabarty et al. (2022)"},{"why":"Supplies Meta4XNLI, the natural-language parallel NLI corpus without lexical substitution, plus its fine-tuned baseline.","marker":"Sanchez-Bayona and Agerri (2024)"},{"why":"Provides the stress-test argument that NLI models exploit lexical overlap artifacts, used to frame the bias analysis.","marker":"Naik et al. (2018)"},{"why":"Defines the edit-distance metric used to quantify lexical overlap between premises and hypotheses.","marker":"Levenshtein (1966)"},{"why":"Supplies RoBERTa-large, the fine-tuned encoder baseline that few-shot and chain-of-thought prompting outperform.","marker":"Liu et al. (2019)"}],"fun_headline_variants":["LLM metaphor scores hinge on lexical overlap, not meaning","Word overlap, not metaphor understanding, drives LLM results","LLMs pass metaphor tests via surface features alone","Metaphor mastery in LLMs is a surface-level illusion","Literal paraphrases trip up LLMs, exposing metaphor shortcuts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the literal paraphrase removes the metaphor while preserving the original inference label, so that a drop in accuracy on the literal version can be attributed to the loss of metaphor rather than to label shift or paraphrase quality.","fun_headline_variants_meta":{"raw":{"variants":["LLM metaphor scores hinge on lexical overlap, not meaning","Word overlap, not metaphor understanding, drives LLM results","LLMs pass metaphor tests via surface features alone","Metaphor mastery in LLMs is a surface-level illusion","Literal paraphrases trip up LLMs, exposing metaphor shortcuts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000492,"raw_usage":{"total_tokens":2372,"prompt_tokens":856,"completion_tokens":1516,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":1436}},"tokens_in":472,"tokens_out":1516,"duration_ms":11794,"temperature":1.0,"reasoning_tokens":1436,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:33:09.424479+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could have two annotators verify whether each automatically generated literal paraphrase preserves the gold entailment label and contains no metaphor, then recompute accuracy on the label-preserving subset; if literal accuracy remains at the level of the metaphorical accuracy on that subset, the central claim would be undermined. Alternatively, constructing literal and metaphorical stimuli matched for lexical overlap and sentence length and showing equal accuracy would falsify the surface-feature explanation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides IMPLI, a gold NLI benchmark with metaphorical language, and is the source of the Levenshtein-distance-based overlap analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies FLUTE, the figurative NLI benchmark with entailments and contradictions used in evaluation."},{"cited_title":"Meta4XNLI: A Crosslingual Parallel Corpus for Metaphor Detection and Interpretation","cited_arxiv_id":"2404.07053","evidence_quote":"Supplies Meta4XNLI, the natural-language parallel NLI corpus without lexical substitution, plus its fine-tuned baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the stress-test argument that NLI models exploit lexical overlap artifacts, used to frame the bias analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the edit-distance metric used to quantify lexical overlap between premises and hypotheses."}],"review_version":1}