{"id":"3bd24b98-fef2-4558-98f8-01a6180c22d4","arxiv_id":"2507.13335","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Models explain puns better than jokes requiring contemporary world knowledge, and none of the tested LLMs explains all joke types reliably.","lead":"A new dataset of 600 jokes across puns, internet humour, and topical jokes is paired with human-written explanations and used to test eight large language models' ability to explain why jokes are funny. The study finds that no model reliably explains all joke types, and that topical jokes are hardest, suggesting that pun-focused humour research may not reflect real-world humour understanding.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The difficulty ordering rests on a single evaluator who also wrote the gold explanations, and the independent agreement check is only moderate; the rankings may reflect one person's explanation standard rather than intrinsic joke difficulty.","rationale":"I read the paper in good faith. It presents a new, balanced dataset of 600 jokes across four types with human-authored explanations, evaluates eight LLMs using human scoring, an LLM judge, and reference-based metrics, and reports a consistent difficulty ordering. These are real contributions, and the dataset release is valuable regardless of the evaluation debate. However, the central claim that no model reliably explains all joke types and that the difficulty ordering holds is only as strong as the evaluation instrument. The weakest point is exactly where the reader placed it: the ground truth is authored by one person, and the primary evaluation is also performed by that person. The third-party check is small and shows only moderate agreement. Although the logistic regression and automatic metrics align with the ordering, they are all referenced to the same gold explanations, so they do not independently validate the gold standard. This is not an accusation of bias or dishonesty; it is a structural concern about what the scores mean. If the golds are unusually demanding about specific details, the completeness scores—especially for topical jokes, where many valid explanations exist—could be deflated across models and inflate the apparent difficulty gap. A focused independent re-annotation study is the most direct way to test whether the reported ordering is an artifact of evaluator calibration. Because the paper already acknowledges some of these limitations and the underlying dataset is a contribution, I do not recommend changing the reader's CONDITIONAL verdict; rather, the concern reinforces the conditions under which the central claim should be accepted.","tokens_in":16788,"tokens_out":3892,"duration_ms":53465,"concrete_test":"Recruit two or three independent annotators who are not authors and have not seen the gold explanations or the primary evaluator's scores. Have them re-score a stratified sample of at least 480 model explanations (e.g., 15 jokes per type x 4 types x 8 models) using only the Section 4.3 rubric. Then compute the joke-type ordering and per-model 'good explanation' proportions from these independent scores and compare them with the reported ordering. If independent annotation reproduces the ordering with Krippendorff's alpha at or above 0.667 on the sampled scores, the central claim survives; if the ordering flips or flattens, the reported difficulty ranking is not robust to evaluator identity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that joke-explanation difficulty follows homographic puns < heterographic puns < non-topical Reddit jokes < topical jokes, and that no tested model reliably explains all types—depends on the validity of its human evaluation. Section 3.6 states that an author wrote all 600 gold explanations, and Section 5 states that the same author scored all 4,800 model explanations. The independent check covers only 320/4,800 explanations (10 jokes per type per model), and Krippendorff's alpha is 0.574 (accuracy) and 0.553 (completeness). These values are below the commonly cited 0.667 threshold for tentative reliability, and the paper's suggestion that 'agreement statistics in the interval between 0.4 and 0.6 are considered good' is not standard for Krippendorff's alpha. The LLM-as-a-judge agreement is similar (alpha about 0.52-0.57), so even the automated check is calibrated to the same gold explanations. If those golds encode one person's standard for 'completeness'—for example, requiring specific named-entity-level detail for topical jokes—then valid alternative explanations can be systematically downgraded. Because the human evaluator is also the gold author, the measured ordering may be amplified by this calibration rather than reflecting an intrinsic property of the joke types.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a dataset of 600 jokes balanced across homographic puns, heterographic puns, non-topical Reddit humour, and topical Reddit humour, each paired with a human-authored gold explanation. Using this dataset, the authors evaluate eight LLMs in a zero-shot joke-explanation setting, scoring 4,800 generated explanations for accuracy and completeness through a single primary evaluator, a third-party reliability subsample, an LLM-as-a-judge, and reference-based automatic metrics. The central finding is a difficulty ordering—homographic puns easiest, then heterographic puns, then non-topical Reddit humour, then topical humour hardest—and that no tested model, including reasoning models, reliably explains all joke types.","tokens_in":17007,"tokens_out":5047,"duration_ms":57779,"significance":"If the difficulty ordering holds, the paper makes a valuable contribution by showing that the field's near-exclusive focus on pun-based humour benchmarks overstates LLM humour understanding. The dataset itself is a useful resource: 600 jokes across four categories with gold explanations, topical-knowledge URLs, and phonetic transcriptions, released with code. The empirical scope is broad (eight models, 4,800 explanations), and the main ordering is corroborated by several complementary evaluation methods. The authors are transparent about their single-evaluator design and include third-party and LLM-judge agreement analyses. However, the validity of the central claim hinges on the objectivity of the gold explanations and the reliability of the human evaluation, which are the weakest points in the current manuscript.","major_comments":[{"comment":"The central difficulty ordering is based on scores assigned by a single evaluator who is also the author of the gold explanations. The third-party reliability check covers only 320 of 4,800 explanations and yields Krippendorff's alpha of 0.574 (accuracy) and 0.553 (completeness), which are below the commonly cited threshold of 0.667 for tentative reliability. The Limitations paragraph's assertion that \"agreement statistics in the interval between 0.4 and 0.6 are considered good\" is not a standard interpretation for Krippendorff's alpha. To make the main claim load-bearing, the paper needs either a larger independent annotation sample or a clearly justified weaker reliability standard, along with per-joke-type agreement figures to confirm that the ordering is not driven by one annotator's idiosyncratic standard.","section":"§4.3/§5.1, Figures 3–4"},{"comment":"The LLM-as-a-judge is selected for best alignment with the primary author's scores, and it is given the author-written gold explanation as a reference when scoring model outputs. This creates a circular validation loop: the automated judge is calibrated to the same ground truth whose validity is in question. The reference-based automatic metrics in Table 1 are similarly aligned to the same golds, so they cannot independently confirm the difficulty ordering. I recommend reporting agreement of the third-party annotators directly against the gold explanations, and using an LLM judge that was not selected post hoc on the same test set, or evaluating the judge on a held-out set of explanations.","section":"§A.6/§6"},{"comment":"The binary \"good\" threshold (scores ≥4 on both accuracy and completeness) is arbitrary, and no sensitivity analysis is provided. Because Figure 4 is the primary visual evidence for the abstract's claim that no model reliably explains all joke types, the paper should either justify this threshold with reference to the task's intended use or demonstrate that the main ordering is robust to reasonable alternative thresholds (e.g., ≥3 on both criteria, or mean score ≥4).","section":"§5.1, Figure 4"},{"comment":"There is a factual inconsistency in the size of the third-party reliability subset: §4.3 states \"a subset of 320 explanations (10 jokes * 4 joke types * 8 models)\", whereas Appendix A.1 describes \"evaluation on a subset of 240 explanations\". Since the agreement statistics are central to the reliability argument, this discrepancy must be resolved and the correct number reported consistently.","section":"§4.3/§A.1"}],"minor_comments":[{"comment":"The title contains a formatting artifact, \"T raditional\", which should be corrected to \"Traditional\".","section":"Title/Abstract"},{"comment":"The model list repeats \"DeepSeek-R1-Distill-Llama-8B\" for both the 8B and 70B variants; the 70B entry should read \"DeepSeek-R1-Distill-Llama-70B\".","section":"Appendix A.1"},{"comment":"The cross-reference \"with the rubric presented in Appendix 4.3\" is confusing; the rubric appears in Section 4.3, so the reference should be to \"§4.3\".","section":"§3.6"},{"comment":"The observation that \"completeness scores are generally lower than accuracy scores across all models\" is based on visual inspection of Figure 3; a paired significance test across the 4,800 explanations would strengthen this claim.","section":"§5"},{"comment":"The automatic-metric differences between joke types are small on some metrics (e.g., BERTScore ranges from 0.87 to 0.89), so the statement that these metrics \"confirm our originally hypothesised ordering\" is stronger than the numbers warrant; report confidence intervals or statistical tests for these comparisons.","section":"Table 1/§6"},{"comment":"The limitation paragraph acknowledges the single-evaluator design but then claims the reliability is \"robustly prove[n]\"; given the moderate Krippendorff's alpha values, this wording should be tempered to match the evidence.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an NLP venue and addresses a genuinely important gap in humour understanding benchmarks. The main risk is the self-referential evaluation design: gold explanations written by an author, all 4,800 model explanations scored by that same author, and the LLM judge selected for best agreement with that same author. The third-party and automatic checks are commendable but insufficient in their current form. If the authors can provide a substantially larger independent annotation sample and a non-circular validation of the judge, the central claim would be much more convincing. I would not reject the paper on the basis of the current limitations, but the reliability argument needs real work before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful new benchmark and the central difficulty ordering (homographic puns easiest, topical hardest) is probably right, but the evaluation is self-calibrated in a way that needs fixing before the numbers can be taken at face value.\n\nThe genuinely new thing is the dataset: 600 jokes balanced across four joke types, with author-written gold explanations, plus a systematic zero-shot comparison of eight LLMs. Prior work was pun-heavy or single-source, so the comparison across puns, Reddit humour, and topical jokes fills a real gap. The finding that no model reliably explains all types, and that topical jokes are hardest, is consistent across the primary human evaluation, the LLM judge, and the automatic metrics. That triangulation gives the main claim real support.\n\nWhat the paper does well: the dataset design is sensible, the rubric is clearly defined, and they release the data and code. The case study on the Tide Pod joke is a nice qualitative illustration. The logistic regression on binary success/failure is a step beyond just reporting averages.\n\nThe soft spots are all in the evaluation loop. The same author who wrote all 600 gold explanations also scored all 4,800 model explanations. The third-party check covers only 320 explanations and Krippendorff's alpha is around 0.55–0.57, which the paper calls 'good' but is usually considered moderate at best. The LLM judge was chosen for best alignment with that same author's scores, so it isn't an independent anchor. On top of that, joke type is confounded with source: puns come from SemEval, the Reddit jokes from r/Jokes, so any source-specific effect (style, era, register) is mixed into the type effect. Model-level comparisons also lack significance tests beyond the aggregate regression.\n\nNone of this breaks the central claim—the third-party annotators and the LLM judge show the same difficulty ordering on the subset they saw, suggesting the ordering is not just one person's idiosyncrasy. But the specific success rates and model rankings are softer than the paper implies. The fix is straightforward: independent annotation of a larger sample, blind to the golds, and significance tests for pairwise model differences.\n\nThis deserves a serious referee. It's a solid empirical contribution for computational humour and LLM evaluation, but the eval section needs work before I'd trust the exact numbers. Tell the authors to strengthen annotation independence and report the stats more carefully.","headline":"A useful new benchmark with a probably-right difficulty ordering, but the eval loop is self-calibrated and needs independent annotation before the exact numbers are trusted.","tokens_in":17560,"tokens_out":2147,"would_cite":true,"duration_ms":23773,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that LLM humour understanding is much weaker on topical and internet jokes than on puns, so pun-only benchmarks overestimate it.","keywords":["humour explanation","joke explanation","LLM humour understanding","topical humour","puns","benchmark dataset","zero-shot evaluation","Reddit humour"],"falsifier":"Independently re-annotate all 4,800 explanations with annotators who never see the authors' gold explanations and grade accuracy and completeness from the joke alone; if topical jokes then score as well as homographic puns, or the model ranking reverses, the central claim fails. A sharper version is to test the models on fresh topical jokes written after their training cutoff, so no memorised explanation can leak into the result.","tokens_in":16549,"feed_emoji":"😂","tokens_out":5775,"duration_ms":66853,"temperature":0.7,"pith_summary":"This paper argues that the standard way of testing machine humour understanding—short, self-contained puns—has created an inflated picture of LLM competence. To test this, the authors built a balanced set of 600 jokes spanning homographic puns, heterographic puns, non-topical Reddit humour, and topical jokes that need outside knowledge, wrote reference explanations for each, and scored 4,800 zero-shot explanations from eight LLMs. Their central finding is a difficulty ordering: homographic puns are easiest, then heterographic puns, then non-topical internet humour, with topical humour hardest. None of the tested models, including reasoning models, reliably produces adequate explanations across all four types, so the narrow pun focus of prior benchmarks does not represent everyday humour.","feed_headline":"Pun-only benchmarks overstate LLM humour understanding","feed_subtitle":"A 600-joke test across four joke styles shows topical humour is hardest for every LLM tested.","key_machinery":"The load-bearing object is the dataset itself: 600 jokes in four balanced categories, each paired with an author-written gold explanation and scored by a six-point accuracy and completeness rubric. The joke-type categories operationalise what the paper means by humour complexity: puns require semantics and phonetics, non-topical Reddit humour requires common-sense and social knowledge, and topical humour requires esoteric real-world knowledge. The binary success criterion (both scores at least 4) combined with logistic regression on model size and joke type is what carries the difficulty-ordering claim.","core_discovery":"The discovery is that joke format, not just joke content, determines LLM explanation quality, and the four-type ordering holds across models and evaluation methods. Homographic puns, where one spelling carries two meanings, produce the highest proportion of explanations scoring at least 4 out of 5 on both accuracy and completeness. Heterographic puns are harder because they require the model to recognise that two differently spelled words sound alike, and phonetic similarity is not visible in orthographic text. Non-topical Reddit humour is harder still, and topical humour, which requires retrieving named entities and events that are not explicitly stated, is hardest. Larger models outperform smaller variants, with the gap widest on topical jokes, but even the best model does not sustain good explanations across all joke types.","pith_inferences":["If the measured gap is mostly about retrieving named-entity knowledge rather than humour reasoning, then giving models access to the URLs the dataset provides for topical jokes should largely close the gap; this is not tested in the paper.","A direct extension would be to evaluate the same models on newly written topical jokes from after their training cutoff, removing any chance that memorised explanations leak into the results.","The author-written gold explanations set the standard for both human and automatic scoring, so a multi-reference or crowd-sourced ground truth would test whether the absolute quality scores are stable even if the difficulty ordering is.","Because topical humour ages, the benchmark's difficulty ordering may shift over time; a living benchmark would need periodically refreshed jokes and explanations."],"forward_implications":["If the ordering holds, benchmarks built only from puns overstate LLM humour understanding relative to the jokes people actually encounter online.","Topical humour is the most discriminating test: differences between large and small models are most visible there, so it can serve as a harder evaluation signal than puns.","Reasoning-specialised models do not automatically beat standard models on joke explanation, suggesting their reasoning style is not tuned to incongruity comprehension.","Heterographic puns expose a phonetic blind spot in models trained on orthographic text, pointing to phonology-aware training as a needed direction."],"supporting_citations":[{"why":"Provides the 300 SemEval-2017 Task 7 puns that form the homographic and heterographic joke subsets.","marker":"Miller et al. (2017)"},{"why":"Provides the r/Jokes corpus from which the non-topical and topical Reddit joke subsets are filtered and selected.","marker":"Weller and Seppi (2020)"},{"why":"Shows prior pun-explanation work with crowdsourced explanations, demonstrating the pun-only scope this paper extends beyond.","marker":"Sun et al. (2022a)"},{"why":"Establishes the humour explanation task on New Yorker caption contests that this work adapts to four joke types with written reference explanations.","marker":"Hessel et al. (2023)"},{"why":"Defines the GPT-4o and GPT-4o Mini models whose zero-shot explanations are evaluated.","marker":"Achiam et al. (2024)"},{"why":"Defines the Gemini 1.5 Pro and Flash models whose zero-shot explanations are evaluated.","marker":"Georgiev et al. (2024)"},{"why":"Defines the Llama 3.1 8B and 70B models whose zero-shot explanations are evaluated.","marker":"Dubey et al. (2024)"},{"why":"Defines the Qwen2.5-72B-Instruct model used as the LLM judge that corroborates the human-rated difficulty ordering.","marker":"Qwen et al. (2025)"}],"fun_headline_variants":["LLMs explain puns better than topical jokes, study finds","Why topical humour stumps every LLM tested","Joke type, not just content, decides LLM humour grasp","600 jokes show LLMs crack puns but not topical wit","Topical jokes are the ultimate LLM humour hurdle"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's conclusions depend on the author-written gold explanations and the 0-5 rubric being an unbiased standard for what counts as explaining a joke; if those explanations have a house style, or the rubric silently rewards matching the author's own reading, the difficulty ordering and model rankings could be artifacts of that standard.","fun_headline_variants_meta":{"raw":{"variants":["LLMs explain puns better than topical jokes, study finds","Why topical humour stumps every LLM tested","Joke type, not just content, decides LLM humour grasp","600 jokes show LLMs crack puns but not topical wit","Topical jokes are the ultimate LLM humour hurdle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00061,"raw_usage":{"total_tokens":2804,"prompt_tokens":878,"completion_tokens":1926,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":1844}},"tokens_in":494,"tokens_out":1926,"duration_ms":14174,"temperature":1.0,"reasoning_tokens":1844,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:24:52.473796+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-annotate all 4,800 explanations with annotators who never see the authors' gold explanations and grade accuracy and completeness from the joke alone; if topical jokes then score as well as homographic puns, or the model ranking reverses, the central claim fails. A sharper version is to test the models on fresh topical jokes written after their training cutoff, so no memorised explanation can leak into the result.","supporting_citations":[],"review_version":1}