Pith. sign in

REVIEW 3 major objections 6 minor 300 references

The paper argues that current AI has largely mastered literal, perception-grounded semantics and is now moving into high-level semantic intelligence: understanding and generating humor, sarcasm, metaphor, empathy, persuasion, and narrative

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-07-31 23:03 UTC pith:WBJXNTF3

load-bearing objection A useful, well-organized survey that is agenda-setting rather than formal; the BLSI/HLSI boundary is asserted and the complexity equations don't yet do work, but the reference map and task taxonomy are genuinely valuable. the 3 major comments →

arxiv 2607.24082 v1 pith:WBJXNTF3 submitted 2026-07-27 cs.AI

Towards High-Level Semantic Intelligence

classification cs.AI
keywords high-level semanticssemantic complexitysemantic intelligencehumor understandingfigurative languagemultimodal reasoningsurvey
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that AI's progress can be read as a shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI): from literal recognition and production to meanings that require inference, context, and social grounding. It argues that six seemingly separate tasks—humor, sarcasm, metaphor, empathy, persuasion, and narrative—should be treated as one coherent research direction, unified by a formal language of semantic clues, chains, causal structure, and semantic effects. If true, the field gains a shared vocabulary for comparing methods and benchmarks across text, speech, vision, and multimodal settings, and a common target for future models. The survey organizes existing data, modeling, and evaluation work for HLS understanding and generation under this umbrella.

Core claim

The central claim is that AI's semantic evolution follows a clear trajectory: early systems handled direct, literal, perception-grounded meanings, while current systems are increasingly expected to handle abstract, context-dependent, socially situated meanings. The paper names these two levels Basic-Level Semantic Intelligence (BLSI) and High-Level Semantic Intelligence (HLSI), and places six task families—humor, sarcasm, metaphor, empathy, persuasion, and narrative—under HLSI. It proposes that all six share underlying structures: semantic clues, semantic chains, causal chains, and semantic effects, with semantic complexity and semantic density as quantitative characterizations. The authors

What carries the argument

The framework centers on four formal elements: Semantic Clues (perceptual evidence), Semantic Chains (inference from clues to derived meaning, called 'Semantic Level-Up'), Causal Chains (causal relations among factors and effects), and Semantic Effects (the holistic cognitive or affective result of composing chains). These feed two proposed metrics: Semantic Complexity Cs, proportional to the complexity of the semantic and causal chains, and Semantic Density Ds = Cs/|R|, the complexity packed into a physical representation. The work that this machinery does is to give the BLSI/HLSI boundary a common descriptive language, and to justify treating humor, sarcasm, metaphor, empathy, persuasion,

Load-bearing premise

The whole BLSI/HLSI stage distinction rests on the claim that semantic complexity is a measurable quantity, but the paper never defines how to compute it; if complexity cannot be measured, the boundary between basic and high-level semantics is a label, not a discovery.

What would settle it

Take paired literal and sarcastic utterances matched in token count: if the paper's metric is real, the sarcastic one should have higher semantic density Ds. Since no procedure for computing Cs is supplied, the prediction cannot be evaluated. A concrete test would be to provide an algorithm for Complexity(SC,CC), compute Ds across a benchmark, and show a threshold that separates BLS from HLS; absent that, the formal apparatus is decorative.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If HLS is one coherent direction, progress on one task—say metaphor—can transfer to others through shared semantic mechanisms and evaluation tools.
  • If the BLSI/HLSI transition is real, benchmarks should measure semantic complexity and density, not just label accuracy, and should separate perceptual clue recognition from higher-level semantic reasoning.
  • If current large models sit near the BLS/HLS boundary, failures in humor, sarcasm, empathy, persuasion, and narrative point to a common reasoning bottleneck rather than task-specific gaps.
  • The framework predicts that training on HLS data can strengthen general reasoning; the paper cites evidence that metaphorical images and humorous content improve complex visual and logical reasoning.
  • Because the same mechanisms can both reveal implicit hate and enable metaphor- or persuasion-based jailbreaks, HLS is a double-edged factor in AI safety.
  • If the BLSI/HLSI distinction is correct, then measuring model capability should include not just correctness but whether the model can explain the mechanism—incongruity, source-target mapping, affective cause—that produces the high-level meaning.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the proposed metrics were ever operationalized, semantic density could become a quantitative ranking of utterances and experiences, connecting HLS to information theory and cognitive-load measures; the paper leaves this connection implicit.
  • A testable extension would be a developmental-style benchmark: paired literal and nonliteral stimuli matched for surface length, presented across modalities, to see whether models acquire BLS skills first and HLS skills later in roughly the order humans do—literal, then metaphor/sarcasm/humor, then persuasion and narrative.
  • Because the paper explicitly excludes formal-rule domains like mathematical reasoning and code generation, HLSI should be read as a social-communicative axis of intelligence, not a complete account of high-level cognition.
  • The taxonomy's real contribution may be to force researchers to define what 'literal' and 'implicit' mean in concrete task terms, since the BLS/HLS boundary is likely to shift as models improve.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that AI is undergoing a transition from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI), defined by a shift from literal, perceptually grounded semantics to context-dependent, socially situated, and cognitively complex meanings. It proposes a conceptual framework built on Semantic Clues, Semantic Chains, Causal Chains, and Semantic Effects, and introduces two quantitative notions, Semantic Complexity (Cs) and Semantic Density (Ds), meant to characterize the BLS/HLS boundary. The survey then organizes recent work on humor, sarcasm, metaphor, empathy, persuasion, and narrative across text, speech, vision, and multimodal settings, covering data construction, modeling/optimization, and evaluation for both understanding and generation. It also discusses applications in AI safety, mental health, social agents, negotiation, creativity, and cultural understanding, and closes with open challenges.

Significance. If the HLS framing is accepted, the paper provides a useful common vocabulary and a broad map of otherwise scattered tasks, spanning modalities and both understanding and generation. Its strengths are the extensive reference coverage, the systematic data/modeling/evaluation taxonomy in Sections IV and V, and its explicit treatment of subjectivity, quality control, and evaluation limitations. The paper also candidly acknowledges in Section VII.B that measuring abstract semantic complexity is particularly challenging and that systematic research is limited. The main risk is that the formal apparatus—especially the complexity metric—is not operational, so the paper's headline theoretical contribution currently rests on stipulation rather than measurement. The survey's organizational value, however, is largely independent of the formal metric, and the claimed trajectory could be supported with more transparent empirical evidence.

major comments (3)
  1. [§II-A2, Eqs. (4)–(5)] Semantic Complexity is defined only as Cs ∝ Complexity(SC, CC), but Complexity(SC, CC) is never defined: there is no function, unit, measurement protocol, or comparison procedure. SC is a semantic inference relation and CC a causal relation, but no account is given of how their 'complexity' is computed. Consequently Ds = Cs/|R| is also non-operational. The BLSI/HLSI boundary in §II-A3 and the stage transition in Fig. 2 are therefore asserted rather than derived from the formal apparatus. The paper itself concedes in §VII.B that measurement is 'particularly challenging' and 'systematic research remains limited.' This does not invalidate the survey, but the formal equations should either be operationalized or repositioned as informal illustrative notation, with the claimed contribution adjusted accordingly.
  2. [Figure 1, footnote 2] The 'clear trajectory' and 'pronounced growth' claims rest on the publication-density statistics, but the methodology is not disclosed. The caption says counts were compiled via the DBLP Search API across 15 top-tier venues, but omits the exact queries, keyword lists, search fields (title/abstract/full text), inclusion and exclusion criteria, normalization procedure, and year groupings. Without this information, the figure cannot be reproduced or checked for query artifacts, duplicate counts, or venue bias. A detailed methodology appendix or released query/count file is needed for this evidence to be load-bearing.
  3. [§II-A3, §III-B] The selection of ChatGPT as the 'rough boundary' between BLSI and HLSI is not connected to the proposed formal quantities. No computation shows that tasks before ChatGPT are below some threshold of Cs or Ds and tasks after are above it; the boundary is chronological. Similarly, the six core HLS phenomena are designated as such by stipulation, and their coherence is supported by shared informal features (§II-B) rather than by the formal metric. These choices may be reasonable as a research proposal, but the paper should present them as an organizing hypothesis, not as a consequence of the semantic-complexity framework.
minor comments (6)
  1. [Footnote 1] 'Working in progress' should be 'Work in progress.'
  2. [§II-A2, Eq. (5)] The notation |R| is explained only by parenthetical examples (token count, pixel area); a formal definition in the definition block would improve precision.
  3. [Figure 2] The row of milestone labels (e.g., 'AlexNetSeq2Seq', 'Deep Speech') is visually compressed and hard to read; adding spacing or separate lines would improve legibility.
  4. [§II-A1d] The abbreviation 'SE' is introduced for Semantic Effect but is rarely used later; consider using it consistently or dropping it to avoid unnecessary notation.
  5. [Figures 5 and 6] The taxonomy diagrams are dense. Several leaf terms (e.g., 'Theory-Guided Semantic Structuring' vs. 'Theory-Guided Semantic Annotation') are close; check that each term is exactly aligned with the section it summarizes.
  6. [References] In-text citations such as [239], [240], [245], and [246] appear in the body; please verify that all cited entries are present in the final reference list. The provided text is truncated before the end, so this is a consistency check rather than a claim of a missing entry.

Circularity Check

1 steps flagged

Mild definitional circularity: HLS is defined through the same cognitive-complexity constructs that Eq. (4) uses to measure complexity, making the BLS/HLS boundary stipulative.

specific steps
  1. self definitional [Section II.A.3(b) (HLS definition); Section II.A.2(a), Eq. (4); Section II.B(c)]
    "High-Level Semantics (HLS) refers to abstract meaning constructed through the integration of basic-level observations and complex cognitive reasoning. ... High semantic complexity implies that multiple layers of inference, causal reasoning, and contextual integration are required. ... HLS typically involves a higher degree of semantic complexity, as its understanding and generation often requires multi-hop reasoning."

    HLS is defined as meaning produced by complex cognitive reasoning over Semantic Chains and Causal Chains, while Eq. (4) defines semantic complexity as Complexity(SC,CC). The claim that HLS has high semantic complexity is therefore true by construction, since the quantity supposed to measure the HLS/BLS distinction is defined in terms of the very chains that constitute HLS. No independent measurement of Cs is given to separate HLS from BLS. The paper concedes in Sec. VII.B that measuring abstract semantic complexity is 'particularly challenging' and systematic research is 'limited,' so the formal boundary is stipulative rather than empirically derived. This does not invalidate the survey's taxonomy, but the central stage-claim rests on definition, not on an operationalized metric.

full rationale

The paper is a survey rather than a prediction-driven study. I found no fitted-parameter circularity, no self-citation chain, and no equation used to generate a numerical prediction from fitted inputs. The only reduction-by-construction appears in the formal framework: HLS is defined as abstract meaning requiring complex cognitive reasoning over Semantic Chains and Causal Chains, and Eq. (4) then states Cs ∝ Complexity(SC,CC). Consequently, asserting that HLS has high semantic complexity is a restatement of the definition rather than an independent empirical finding. The BLS/HLS boundary is therefore stipulative, and the claimed 'clear trajectory' in Fig. 1 is not derived from an operationalized complexity measure; the paper itself acknowledges in Sec. VII.B that measuring abstract semantic complexity is difficult and underdeveloped. This is a mild definitional circularity, not a forced derivation. The survey's extensive literature organization and its characterization of HLS tasks remain independently useful.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 5 invented entities

The framework is entirely conceptual. There are no fitted numeric parameters, but the entire apparatus rests on stipulated constructs (Semantic Clue/Chain/Causal Chain/Effect) and an undefined Complexity function. The human-AI parallelism and the six-task selection are domain assumptions made for the survey's scope. None of the invented entities has an independent falsifiable handle, so the framework cannot be empirically confirmed or refuted as stated.

axioms (5)
  • ad hoc to paper HLS phenomena can be decomposed into Semantic Clues, Semantic Chains, Causal Chains, and Semantic Effects.
    Introduced as the foundational decomposition in Sec. II-A1; no independent evidence that this decomposition is exhaustive or unique.
  • ad hoc to paper Semantic Complexity and Semantic Density are measurable quantities via Cs ∝ Complexity(SC,CC) and Ds = Cs/|R|.
    Sec. II-A2 presents these as formal metrics, but Complexity(SC,CC) is never defined, so the quantities are not computable or falsifiable.
  • domain assumption AI semantic development parallels human cognitive development from literal to abstract, socially embedded semantics.
    Used in Sec. I and Figure 2 to motivate the BLSI-HLSI trajectory, supported only by selected psycholinguistics references, not by a systematic comparative argument.
  • ad hoc to paper The Transformer and ChatGPT mark the boundary between BLSI and HLSI.
    Stated in Sec. I as a rough boundary; this is a modeling choice, not an empirically established threshold.
  • ad hoc to paper Humor, sarcasm, metaphor, empathy, persuasion, and narrative are the core HLS phenomena.
    The survey explicitly excludes math/code and other formal-rule domains (footnote 3) and selects six categories; this scope decision is not derived from a prior theory.
invented entities (5)
  • Semantic Clue (c) no independent evidence
    purpose: Atomic unit of semantic evidence that a system must perceive and reason over.
    No operational definition or measurement procedure is provided; it is a conceptual primitive.
  • Semantic Chain (SC) no independent evidence
    purpose: Represents the inference relation ⊢ from clues to a derived HLS conclusion.
    The ⊢ symbol is used informally in Eq. (1) without a formal semantics or inference rules.
  • Causal Chain (CC) no independent evidence
    purpose: Represents causal factors ⇝ effects within semantic interpretation.
    The ⇝ relation in Eq. (2) is not formally defined or grounded in a causal inference framework.
  • Semantic Effect (E) no independent evidence
    purpose: Composite cognitive/affective/communicative outcome E = Φ({SC},{CC}).
    Φ is unspecified in Eq. (3); no method is given to compute or evaluate E.
  • BLSI / HLSI stage labels no independent evidence
    purpose: Provide a high-level classification of AI semantic capability over time.
    The boundary examples (Transformer, ChatGPT) are illustrative; no quantitative measure is attached to the stages.

pith-pipeline@v1.3.0-alltime-deepseek · 44523 in / 11091 out tokens · 102058 ms · 2026-07-31T23:03:19.071862+00:00 · methodology

0 comments
read the original abstract

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling the understanding and generation of High-Level Semantics (HLS). A similar trajectory can also be observed in human cognitive development. We define this transition as the shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). However, this issue has not yet been systematically and comprehensively examined in prior work. Motivated by this gap, this survey reviews the development of AI semantic intelligence from the perspective of semantic complexity. We systematically survey existing research on HLS tasks, including humor, sarcasm, metaphor, empathy, persuasion, narrative, and other general HLS phenomena, across text, speech, vision, and multimodal scenarios. Specifically, we summarize data construction methods, modeling and optimization strategies, and evaluation methodologies for both understanding and generation. HLS is essential for advancing AI toward genuinely human-like intelligence. By synthesizing existing methods and insights from the perspective of semantic intelligence, this survey aims to support the continued development of AI toward HLSI.

Figures

Figures reproduced from arXiv: 2607.24082 by Gefei Yang, Jiahui Gan, Kai Yu, Mengyue Wu, Qi Jia, Shota Watanabe, Tianxi Wan, Xiujie Song, Yining You.

Figure 1
Figure 1. Figure 1: Evolutionary trajectory and publication density of High-Level Se [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Conceptual taxonomy and evolutionary roadmap from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). Top: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Illustration of Semantic Level-Up through cognitive reasoning across different modalities. By combining lower-level perceptual or contextual semantic [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The cognitive architecture of High-Level Semantic Understanding [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Taxonomy of High-Level Semantic (HLS) understanding methods, covering data construction, modeling approaches, and evaluation protocols. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Taxonomy of High-Level Semantic (HLS) generation methods, covering data construction, generation approaches, and evaluation protocols. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

300 extracted references · 3 canonical work pages

  1. [1]

    Gradient-based learning applied to document recognition,

    Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998

  2. [2]

    Maximum mutual information estimation of hidden markov model parameters for speech recognition,

    L. Bahl, P. Brown, P. de Souza, and R. Mercer, “Maximum mutual information estimation of hidden markov model parameters for speech recognition,” inICASSP ’86. IEEE International Conference on Acous- tics, Speech, and Signal Processing, vol. 11, 1986, pp. 49–52

  3. [3]

    A neural proba- bilistic language model,

    Y . Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A neural proba- bilistic language model,”Journal of machine learning research, vol. 3, no. Feb, pp. 1137–1155, 2003

  4. [4]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication,”The Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948

  5. [5]

    Recent contributions to the mathematical theory of com- munication,

    W. Weaver, “Recent contributions to the mathematical theory of com- munication,”ETC: a review of general semantics, vol. 74, no. 1/2, pp. 136–157, 2017

  6. [6]

    Tomasello, m., constructing a language: a usage-based theory of language acquisition. cambridge, ma: Harvard university press, 2003. pp. 388. hardback, £29.95. isbn 0-674-01030-2

    J. M. PINE, “Tomasello, m., constructing a language: a usage-based theory of language acquisition. cambridge, ma: Harvard university press, 2003. pp. 388. hardback, £29.95. isbn 0-674-01030-2.”Journal of Child Language, vol. 32, no. 3, p. 697–702, 2005

  7. [7]

    Bloom,How children learn the meanings of words

    P. Bloom,How children learn the meanings of words. MIT press, 2002

  8. [8]

    Winner,The point of words: Children’s understanding of metaphor and irony

    E. Winner,The point of words: Children’s understanding of metaphor and irony. Harvard University Press, 1988

  9. [9]

    Playing with expectations: A contextual view of humor development,

    G. Airenti, “Playing with expectations: A contextual view of humor development,”Frontiers in Psychology, vol. 7, p. 1392, 2016

  10. [10]

    Empathy and moral development,

    M. L. Hoffman, “Empathy and moral development,”The annual report of educational psychology in Japan, vol. 35, pp. 157–162, 1996

  11. [11]

    J. S. Bruner,Acts of meaning: Four lectures on mind and culture. Harvard university press, 1990, vol. 3

  12. [12]

    A theory of argumentative understand- ing: Relationships among position preference, judgments of goodness, memory and reasoning,

    N. L. Stein and C. A. Miller, “A theory of argumentative understand- ing: Relationships among position preference, judgments of goodness, memory and reasoning,”Argumentation, vol. 7, no. 2, pp. 183–204, 1993

  13. [13]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017

  14. [14]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkatet al., “Gpt-4 technical report,”arXiv preprint arXiv:2303.08774, 2023. 29

  15. [15]

    B. G. Buchanan and E. H. Shortliffe,Rule Based Expert Systems: The Mycin Experiments of the Stanford Heuristic Programming Project (The Addison-Wesley series in artificial intelligence). USA: Addison- Wesley Longman Publishing Co., Inc., 1984

  16. [16]

    Imagenet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” inAdvances in Neural Information Processing Systems, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., vol. 25. Curran Associates, Inc., 2012. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 2012/file/c399862d3b...

  17. [17]

    Deep speech: Scaling up end-to-end speech recognition,

    A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coateset al., “Deep speech: Scaling up end-to-end speech recognition,”arXiv preprint arXiv:1412.5567, 2014

  18. [18]

    Rich feature hierarchies for accurate object detection and semantic segmentation,

    R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2014

  19. [19]

    Bert: Pre- training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. N. Toutanova, “Bert: Pre- training of deep bidirectional transformers for language understanding,”

  20. [20]

    Improving language understanding by generative pre-training,

    A. Radford, K. Narasimhan, T. Salimans, I. Sutskeveret al., “Improving language understanding by generative pre-training,” 2018

  21. [21]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” inProceedings of the 38th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Ed...

  22. [22]

    Videobert: A joint model for video and language representation learning,

    C. Sun, A. Myers, C. V ondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019

  23. [23]

    Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,

    H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y . Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,” in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y . Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 24 206–24 22...

  24. [24]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynoldset al., “Flamingo: a visual language model for few-shot learning,”Advances in neural information processing systems, vol. 35, pp. 23 716–23 736, 2022

  25. [25]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” inAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 34 892–34 916. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/ 2023/file/6dcf277ea32ce3288914fa...

  26. [26]

    Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,

    S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y . T. Lee, Y . Li, S. Lundberget al., “Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,” 2023

  27. [27]

    Open-sora: Democratizing efficient video production for all,

    Z. Zheng, X. Peng, T. Yang, C. Shen, S. Li, H. Liu, Y . Zhou, T. Li, and Y . You, “Open-sora: Democratizing efficient video production for all,”CoRR, vol. abs/2412.20404, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2412.20404

  28. [28]

    Seedance 2.0: Advancing video generation for world complexity,

    T. Seedance, D. Chen, L. Chen, X. Chen, Y . Chen, Z. Chen, Z. Chen, F. Cheng, T. Cheng, Y . Chenget al., “Seedance 2.0: Advancing video generation for world complexity,”arXiv preprint arXiv:2604.14148, 2026

  29. [29]

    Large lan- guage models for subjective language understanding: A survey,

    C. Song, Y . Zhang, H. Gao, B. Yao, and P. Zhang, “Large lan- guage models for subjective language understanding: A survey,”arXiv preprint arXiv:2508.07959, 2025

  30. [30]

    Looking beyond the obvious: A survey on abstract concept recognition for video understanding: Looking beyond the obvious: A survey on abstract concept

    G. Mago, P. Mettes, and S. Rudinac, “Looking beyond the obvious: A survey on abstract concept recognition for video understanding: Looking beyond the obvious: A survey on abstract concept...”Int. J. Comput. Vision, vol. 134, no. 5, Apr. 2026. [Online]. Available: https://doi.org/10.1007/s11263-026-02784-5

  31. [31]

    Class-basedn-gram models of natural language,

    P. F. Brown, V . J. Della Pietra, P. V . deSouza, J. C. Lai, and R. L. Mercer, “Class-basedn-gram models of natural language,” Computational Linguistics, vol. 18, no. 4, pp. 467–480, 1992. [Online]. Available: https://aclanthology.org/J92-4003/

  32. [32]

    Efficient estimation of word representations in vector space,

    T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,”arXiv preprint arXiv:1301.3781, 2013

  33. [33]

    Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition,

    E. F. Tjong Kim Sang and F. De Meulder, “Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition,” inProceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, 2003, pp. 142–147. [Online]. Available: https://aclanthology.org/W03-0419/

  34. [34]

    Message Understanding Conference- 6: A brief history,

    R. Grishman and B. Sundheim, “Message Understanding Conference- 6: A brief history,” inCOLING 1996 Volume 1: The 16th International Conference on Computational Linguistics, 1996. [Online]. Available: https://aclanthology.org/C96-1079/

  35. [35]

    Statistical phrase-based transla- tion,

    P. Koehn, F. J. Och, and D. Marcu, “Statistical phrase-based transla- tion,” inProceedings of the 2003 human language technology confer- ence of the north american chapter of the association for computational linguistics, 2003, pp. 127–133

  36. [36]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. Cho, and Y . Bengio, “Neural machine translation by jointly learning to align and translate,” in3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Y . Bengio and Y . LeCun, Eds., 2015. [Online]. Available: http://arxiv.org/abs/1409.0473

  37. [37]

    SQuAD: 100,000+ questions for machine comprehension of text,

    P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ questions for machine comprehension of text,” inProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, J. Su, K. Duh, and X. Carreras, Eds. Austin, Texas: Association for Computational Linguistics, Nov. 2016, pp. 2383–2392. [Online]. Available: https://acla...

  38. [38]

    A neural attention model for abstractive sentence summarization,

    A. M. Rush, S. Chopra, and J. Weston, “A neural attention model for abstractive sentence summarization,” inProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, L. M `arquez, C. Callison-Burch, and J. Su, Eds. Lisbon, Portugal: Association for Computational Linguistics, Sep. 2015, pp. 379–389. [Online]. Available: https:/...

  39. [39]

    Automatically constructing a corpus of sentential paraphrases,

    W. B. Dolan and C. Brockett, “Automatically constructing a corpus of sentential paraphrases,” inProceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. [Online]. Available: https://aclanthology.org/I05-5002/

  40. [40]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255

  41. [41]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” inAdvances in Neural Information Processing Systems, C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2015. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2015/...

  42. [42]

    Fully convolutional networks for semantic segmentation,

    J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015

  43. [43]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969

  44. [44]

    Realtime multi-person 2d pose estimation using part affinity fields,

    Z. Cao, T. Simon, S.-E. Wei, and Y . Sheikh, “Realtime multi-person 2d pose estimation using part affinity fields,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7291–7299

  45. [45]

    State of the art in example-based texture synthesis,

    L.-Y . Wei, S. Lefebvre, V . Kwatra, and G. Turk, “State of the art in example-based texture synthesis,”Eurographics 2009, State of the Art Report, EG-STAR, pp. 93–117, 2009

  46. [46]

    Texture synthesis using convo- lutional neural networks,

    L. Gatys, A. S. Ecker, and M. Bethge, “Texture synthesis using convo- lutional neural networks,”Advances in neural information processing systems, vol. 28, 2015

  47. [47]

    Globally and locally consistent image completion,

    S. Iizuka, E. Simo-Serra, and H. Ishikawa, “Globally and locally consistent image completion,”ACM Transactions on Graphics (ToG), vol. 36, no. 4, pp. 1–14, 2017

  48. [48]

    Image completion with structure propagation,

    J. Sun, L. Yuan, J. Jia, and H.-Y . Shum, “Image completion with structure propagation,” inACM Siggraph 2005 Papers, 2005, pp. 861– 868

  49. [49]

    Auto-encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013

  50. [50]

    Generative adversarial nets,

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” Advances in neural information processing systems, vol. 27, 2014

  51. [51]

    A tutorial on hidden markov models and selected applica- tions in speech recognition,

    L. Rabiner, “A tutorial on hidden markov models and selected applica- tions in speech recognition,”Proceedings of the IEEE, vol. 77, no. 2, pp. 257–286, 1989. 30

  52. [52]

    Suppression of acoustic noise in speech using spectral subtraction,

    S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,”IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, 1979

  53. [53]

    Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

    W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 4960–4964

  54. [54]

    X-vectors: Robust dnn embeddings for speaker recognition,

    D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in2018 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), 2018, pp. 5329–5333

  55. [55]

    Emotional speech synthesis: a review

    M. Schr ¨oder, “Emotional speech synthesis: a review.” inInterspeech, vol. 2001, 2001, pp. 561–564

  56. [56]

    Speech parameter generation algorithms for hmm-based speech syn- thesis,

    K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech syn- thesis,” in2000 IEEE international conference on acoustics, speech, and signal processing. Proceedings (Cat. No. 00CH37100), vol. 3. IEEE, 2000, pp. 1315–1318

  57. [57]

    Wavenet: A generative model for raw audio,

    A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuogluet al., “Wavenet: A generative model for raw audio,”arXiv preprint arXiv:1609.03499, vol. 12, no. 1, 2016

  58. [58]

    Tacotron: Towards end- to-end speech synthesis,

    Y . Wang, R. Skerry-Ryan, D. Stanton, Y . Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y . Xiao, Z. Chen, S. Bengioet al., “Tacotron: Towards end- to-end speech synthesis,”arXiv preprint arXiv:1703.10135, 2017

  59. [59]

    Imagebind: One embedding space to bind them all,

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 15 180–15 190

  60. [60]

    Making computers laugh: Inves- tigations in automatic humor recognition,

    R. Mihalcea and C. Strapparava, “Making computers laugh: Inves- tigations in automatic humor recognition,” inProceedings of human language technology conference and conference on empirical methods in natural language processing, 2005, pp. 531–538

  61. [61]

    SemEval-2017 task 6: #HashtagWars: Learning a sense of humor,

    P. Potash, A. Romanov, and A. Rumshisky, “SemEval-2017 task 6: #HashtagWars: Learning a sense of humor,” inProceedings of the 11th International Workshop on Semantic Evaluation (SemEval- 2017), S. Bethard, M. Carpuat, M. Apidianaki, S. M. Mohammad, D. Cer, and D. Jurgens, Eds. Vancouver, Canada: Association for Computational Linguistics, Aug. 2017, pp. 49...

  62. [62]

    “president vows to cut <taxes>hair

    N. Hossain, J. Krumm, and M. Gamon, ““president vows to cut <taxes>hair”: Dataset and analysis of creative text editing for humorous headlines,” inProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), J. Burstein, C. Doran, and T. So...

  63. [63]

    Cards against AI: Predicting humor in a fill-in-the-blank party game,

    D. Ofer and D. Shahaf, “Cards against AI: Predicting humor in a fill-in-the-blank party game,” inFindings of the Association for Computational Linguistics: EMNLP 2022, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 5397–5403. [Online]. Available: https://aclantholog...

  64. [64]

    Can language models make fun? a case study in Chinese comical crosstalk,

    J. Li, X. Wu, X. Liu, Q. Xie, P. Tiwari, and B. Wang, “Can language models make fun? a case study in Chinese comical crosstalk,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Jul....

  65. [65]

    Talk funny! a large-scale humor response dataset with chain-of-humor interpretation,

    Y . Chen, Y . Yuan, P. Liu, D. Liu, Q. Guan, M. Guo, H. Peng, B. Liu, Z. Li, and Y . Xiao, “Talk funny! a large-scale humor response dataset with chain-of-humor interpretation,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, p. 17826–17834, Mar. 2024. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/ article/view/29736

  66. [66]

    Chumor 2.0: Towards better benchmarking Chinese humor understanding from (ruo zhi ba),

    R. He, Y . He, L. Bai, J. Liu, Z. Sun, Z. Tang, H. Wang, H. Xia, R. Mihalcea, and N. Deng, “Chumor 2.0: Towards better benchmarking Chinese humor understanding from (ruo zhi ba),” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: Association for Computational Li...

  67. [67]

    “what do you call a dog that is incontrovertibly true? dogma

    A. Cocchieri, L. Ragazzi, P. Italiani, G. Tagliavini, and G. Moro, ““what do you call a dog that is incontrovertibly true? dogma”: Testing LLM generalization through humor,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Aus...

  68. [68]

    Cfunmodel: A

    Z. Yu, X. Hu, and X. Wan, “Cfunmodel: A” funny” language model capable of chinese humor generation and processing,”arXiv preprint arXiv:2503.20417, 2025

  69. [69]

    Comparing apples to oranges: A dataset & analysis of LLM humour understanding from traditional puns to topical jokes,

    T. Loakman, W. Thorne, and C. Lin, “Comparing apples to oranges: A dataset & analysis of LLM humour understanding from traditional puns to topical jokes,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov....

  70. [70]

    Drivel-ology: Challenging llms with interpreting nonsense with depth,

    Y . Wang, C. Xiao, C.-Y . Hsiao, Z. Y . Chang, C.-L. Chen, T. Loakman, and C. Lin, “Drivel-ology: Challenging llms with interpreting nonsense with depth,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 23 085–23 107

  71. [71]

    Do androids laugh at electric sheep? humor “understanding

    J. Hessel, A. Marasovic, J. D. Hwang, L. Lee, J. Da, R. Zellers, R. Mankoff, and Y . Choi, “Do androids laugh at electric sheep? humor “understanding” benchmarks from the new yorker caption contest,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki,...

  72. [72]

    Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon caption- ing,

    J. Zhang, L. Jain, Y . Guo, J. Chen, K. L. Zhou, S. Suresh, A. Wagen- maker, S. Sievert, T. Rogers, K. Jamiesonet al., “Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon caption- ing,”Advances in Neural Information Processing Systems, vol. 37, pp. 125 264–125 286, 2024

  73. [73]

    Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,

    S. Zhong, Z. Huang, S. Gao, W. Wen, L. Lin, M. Zitnik, and P. Zhou, “Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,” inProceed- ings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, ser. Proceedings of the IEEE Computer Society Conference on Computer Vision a...

  74. [74]

    Cracking the code of juxtaposition: can ai models understand the humorous contradictions,

    Z. Hu, T. Liang, J. Li, Y . Lu, Y . Zhou, Y . Qiao, J. Ma, and Y . Yin, “Cracking the code of juxtaposition: can ai models understand the humorous contradictions,” inProceedings of the 38th International Conference on Neural Information Processing Systems, ser. NIPS ’24. Red Hook, NY , USA: Curran Associates Inc., 2024

  75. [75]

    Visionarena: 230k real world user-vlm conversations with preference labels,

    C. Chou, L. Dunlap, K. Mashita, K. Mandal, T. Darrell, I. Stoica, J. E. Gonzalez, and W.-L. Chiang, “Visionarena: 230k real world user-vlm conversations with preference labels,” in2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 3877– 3887

  76. [76]

    D-humor: Dark humor understanding via multimodal open-ended reasoning,

    S. K. R. Kasu, M. Z. Ur Rehman, S. S. Dar, R. B. Junghare, D. S. Namboodiri, and N. Kumar, “D-humor: Dark humor understanding via multimodal open-ended reasoning,” in2025 IEEE International Conference on Data Mining (ICDM), 2025, pp. 377–386

  77. [77]

    Humor in pixels: Benchmarking large multimodal models understanding of online comics,

    Y . Ryan, R. Y . Tan, K. T. W. Choo, and R. K.-W. Lee, “Humor in pixels: Benchmarking large multimodal models understanding of online comics,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, Eds. Suzhou, China: Association for Computational Linguistics, Nov. 2025, pp. 1...

  78. [78]

    HumorDB: Can AI understand graphical humor?

    V . V . Jain, F. d. S. A. Feitosa, and G. Kreiman, “HumorDB: Can AI understand graphical humor?” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2025/papers/Jain HumorDB Can AI understand graphical humor ICCV 2025 paper.pdf

  79. [79]

    V-hub: A visual-centric humor understanding benchmark for video llms,

    Z. Shi, H. Li, Y . Zhao, J. Zhou, Y . Wang, Q. Cui, W. Bi, S. Zhu, B. Zhao, and Z. Zheng, “V-hub: A visual-centric humor understanding benchmark for video llms,”arXiv preprint arXiv:2509.25773, 2025

  80. [80]

    GODBench: A benchmark for multimodal large language models in video comment art,

    Y . Lei, C. Zhang, Z. Liu, H. Leng, S. Liu, T. Gao, Q. Liu, and Y . Wang, “GODBench: A benchmark for multimodal large language models in video comment art,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 31 (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: Associat...

Showing first 80 references.