Pith. sign in

REVIEW 4 cited by

Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12725 v2 pith:7Y6QRBN7 submitted 2024-07-17 cs.CL

classification cs.CL
keywords cuesllmsfourframeworkmodelssarcasmcomprehensivehuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Elaborating a series of intermediate reasoning steps significantly improves the ability of large language models (LLMs) to solve complex problems, as such steps would evoke LLMs to think sequentially. However, human sarcasm understanding is often considered an intuitive and holistic cognitive process, in which various linguistic, contextual, and emotional cues are integrated to form a comprehensive understanding, in a way that does not necessarily follow a step-by-step fashion. To verify the validity of this argument, we introduce a new prompting framework (called SarcasmCue) containing four sub-methods, viz. chain of contradiction (CoC), graph of cues (GoC), bagging of cues (BoC) and tensor of cues (ToC), which elicits LLMs to detect human sarcasm by considering sequential and non-sequential prompting methods. Through a comprehensive empirical comparison on four benchmarks, we highlight three key findings: (1) CoC and GoC show superior performance with more advanced models like GPT-4 and Claude 3.5, with an improvement of 3.5%. (2) ToC significantly outperforms other methods when smaller LLMs are evaluated, boosting the F1 score by 29.7% over the best baseline. (3) Our proposed framework consistently pushes the state-of-the-art (i.e., ToT) by 4.2%, 2.0%, 29.7%, and 58.2% in F1 scores across four datasets. This demonstrates the effectiveness and stability of the proposed framework.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Introduces Sarc7, a seven-type sarcasm benchmark on MUStARD, and shows an emotion-based prompting method improves sarcasm type macro-F1 (0.3664) and generation success (72 vs 52 of 100) over zero-shot prompting.

  2. CAF-I: A Collaborative Multi-Agent Framework for Enhanced Irony Detection with Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    CAF-I, a multi-agent LLM framework with context, semantic, and rhetorical agents plus a refinement evaluator, reports state-of-the-art zero-shot irony detection, averaging 76.31 Macro-F1 across four benchmarks.

  3. Pragmatic Metacognitive Prompting Improves LLM Performance on Sarcasm Detection

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A two-call prompt that adds pragmatic analysis and reflection improves GPT-4o's sarcasm detection on MUStARD and SemEval2018, but the effect is not consistent across models and lacks statistical validation.

  4. LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts

    cs.CL 2025-01 reject novelty 4.0 of 10

    LLMQuoter uses a distilled 3B model to extract quotes for RAG; the paper shows gold quotes greatly improve QA, but does not test its own model's quotes end-to-end.

Pith tools