Pith. sign in

REVIEW 4 major objections 6 minor 59 references

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

T0 review · 4 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Stance flips in multimodal chat can be forecasted by extracting sextuples and attributing triggers with role-based reasoning.

desk verdict Solid task-and-resource package for multimodal stance flips, with real ablations—but the core Flip-Trig numbers sit on an inconsistent trigger taxonomy that needs fixing before anyone trusts the headline metric. read the letter →

arxiv 2607.24191 v1 pith:3FZCSZ5S submitted 2026-07-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords stancedetectionmultimodallearningconversationalAIflippinginformationextractionrationalegenerationThought-of-Stance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that stance in conversation is not a one-shot label but a persistent state that can reverse when new multimodal evidence lands. Existing benchmarks miss belief reversals, mix emotion with argumentative stance, and underuse images, audio, video, and stickers that resolve sarcasm. StanceFlip is a bilingual multi-turn dataset of thousands of dialogues with five modalities and over 20% flip density, defined by two tasks: extract a holder–target–emotion–sentiment–stance–rationale sextuple per turn, and locate flips with their socio-cognitive trigger type. ConStaFF, built on a multimodal LLM, runs a Thought-of-Stance pipeline of specialized personas plus a self-reflective check on rationales. On both subtasks it beats strong multimodal LLM baselines by large margins, including joint flip-and-trigger scores that remain hard for all systems.

What carries the argument

Thought-of-Stance (ToS): a four-step persona chain (Cartographer for target proposition, Psychologist for multimodal affect, Discourse Analyst for stance-state transitions, Synthesizer–Critic for rationale and trigger) plus self-reflective verification that revises draft rationales against target, stance, and evidence consistency.

What would settle it

Hold out human-only dialogues with independently annotated flips and triggers (no LLM synthesis or LLM-as-judge), retrain or evaluate ConStaFF, and check whether Micro-F1, Iden., and Flip-Trig margins over untuned multimodal LLMs collapse.

Watch

Extended reading notes

Core claim

Modeling conversational stance as evolving state snapshots plus causal flip attribution, with a role-decomposed Thought-of-Stance reasoner and self-reflective rationale verification, yields state-of-the-art Multimodal Stance Sextuple Extraction and Dynamic Stance Flip Attribution on the StanceFlip benchmark, substantially above strong MLLM baselines in English and Chinese.

Load-bearing premise

That GPT-assisted synthesis, retrieval, and labeling—even after filters and expert review—produce gold stance flips and media alignments faithful enough that model gains measure real flip understanding rather than matching the generator’s style.

Editorial extensions

If this is right

  • Stance systems can report structured sextuples and flip causes instead of isolated Support/Oppose tags.
  • Multimodal cues (especially video and audio) become first-class evidence for sarcasm and reversal triggers, not optional features.
  • Role-ordered stance reasoning plus one-pass rationale critique becomes a reusable pattern for other discourse state-tracking tasks.
  • High flip-density bilingual data forces models to track persistence and true reversals rather than majority static labels.
  • Zero-shot transfer to unseen debate targets is framed as learning target-invariant evolution patterns via ToS.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Joint Flip-Trig scores staying low for every system suggests trigger attribution may need explicit causal or counterfactual supervision beyond persona prompting.
  • Heavy reliance on the same LLM family for data, training, and soft matching invites a future human-only test split as the real stress test of the claim.
  • Separating emotion from stance in the sextuple could transfer to moderation and negotiation tools that must distinguish hostile agreement from topical opposition.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces StanceFlip, a bilingual (EN/ZH) benchmark of 7,710 multi-turn dialogues (37,836 turns) augmented across five modalities for a new task, Multimodal Conversational Stance Flipping Forecasting, split into (I) panoptic stance sextuple extraction (holder, target, emotion, sentiment, stance, rationale) and (II) dynamic stance-flip attribution with a four-way trigger taxonomy. Data are built by a GPT-4o-driven simulation–retrieval–labeling pipeline over DailyDialog/MELD/ZS-CSD seeds, with dual-consistency filtering, a 17-expert full review, and a 100-dialogue double-blind audit (κ = 0.84 stance, 0.87 media relevance). The authors also propose ConStaFF, a Vicuna/Qwen2.5-7B model with ImageBind encoding, a four-persona "Thought-of-Stance" reasoning scaffold, self-reflective verification, and three-stage LoRA tuning, reporting state-of-the-art results over prompted LLM baselines (e.g., EN Micro-F1 25.84, Iden. 46.77, Flip-Trig 26.38), plus ablations for reasoning strategy, modality drops, and self-reflection, and implicit-stance and zero-shot transfer probes.

Significance. If the trigger-taxonomy inconsistency is repaired and the evaluation re-run, StanceFlip would be a useful community resource: a bilingual, five-modality conversational benchmark that operationalizes stance as a persistent state with explicit flip-trigger attribution, a genuinely under-explored formulation. Strengths worth naming: full prompt disclosure for construction and labeling (Tables 5–7), documented three-stage training with hyperparameters (Table 8), five-seed averaging throughout, ToS/CoS/ToT and modality-drop ablations, implicit-stance and zero-shot transfer probes, and a double-blind inter-annotator study (κ = 0.84/0.87). The ConStaFF framework itself is a reasonable strong baseline rather than a conceptual advance.

major comments (4)
  1. [§3.1, Table 7, App. A.2, App. D.2] The trigger taxonomy is stated three inconsistent ways, and the version that generated the gold labels does not match the version used for evaluation. §3.1 and D.2 use {Factual & Logical; Emotional & Value-based; Personal Experience & Anecdote; Social Influence}; the Table 7 labeling prompt (which produced gold shift_reason_category) uses {Introduction of New Info, Logical Argument, Emotional Appeal, Social Pressure}; Appendix A.2 uses {Logical Argumentation, Emotional Resonance, New Information, Social Interaction}. D.2's keyword buckets contain no entry that can match a gold 'New Information'/'Introduction of New Info' label, so every such gold trigger is a guaranteed false negative for every model, systematically depressing and distorting Trig and Flip-Trig — the paper's headline metric (Table 4, Fig. 3d). The authors must harmonize the taxonomy across §3.1, Table 7, A.2, and D.2, and
  2. [§5.1 vs App. D.4] Identification F1 is defined inconsistently. §5.1 defines Iden. F1 on the quadruple (Holder, Target, Stance, Flip), while App. D.4 defines it as the triple (Holder, Target, Rationale). These are materially different criteria, and Iden. = 46.77% (EN) is one of the two headline numbers supporting the state-of-the-art claim. Please state which definition was actually computed, correct the other, and if necessary re-report Table 3 and Figs. 5–7.
  3. [§5.1–5.2, Tables 3–4] The state-of-the-art claim rests on an asymmetric comparison. In Tables 3–4, ConStaFF is three-stage fine-tuned on StanceFlip, while all listed baselines (Vicuna, Llama2/3, Qwen2.5, Mistral, Flan-T5-XXL, with or without ToS prompting) appear to be zero/few-shot prompted. The margin (e.g., Micro-F1 25.84 vs 14.85) may reflect fine-tuning on the benchmark rather than the ToS architecture per se. Fig. 5 (ToS vs CoS/ToT 'under the same setting') partially addresses this, but a fine-tuned non-ToS ablation of the identical backbone (Vicuna-7B EN / Qwen2.5-7B ZH) trained end-to-end on the same data is needed to isolate the contribution of the ToS/self-reflection design.
  4. [§3.2–3.3, App. D.3] Pipeline coupling needs a direct test. GPT-4o plans media injection, generates irony-aware queries, and produces gold labels (Tables 5–7); ConStaFF is instruction-tuned on this distribution; and Target/Rationale scoring uses GPT-4o-mini as judge with a lenient containment fallback (App. D.3). The 100-dialogue κ audit (App. A.1) validates stance labels but not the judge's equivalence decisions. A concrete, cheap fix: report human adjudication of a sample of judge YES/NO decisions (or re-score a test subset with human Target/Rationale matching) to bound how much of the ConStaFF-vs-baseline gap is judge agreement with the generator's style rather than true extraction quality.
minor comments (6)
  1. [§5.1 vs App. C.1] §5.1 states experiments used 8×A6000 GPUs; App. C.1 states 10×A6000. Please reconcile.
  2. [App. C.2] Vicuna-7B-v1.5 is cited as [?] — missing reference.
  3. [§5.5, Fig. 5] Fig. 5 compares 'ToS, CoS, ToT' but CoS is never defined in the text; presumably Chain-of-Stance [31]. Please define and describe how CoS/ToT were adapted to this task.
  4. [§5.4] The zero-shot transfer experiment (§5.4) evaluates on MT-CSD, which is text-only per Table 1's modality column for comparable corpora; how ConStaFF handles absent modalities at inference is not described.
  5. [Throughout] Venue template artifacts remain: 'Conference acronym ’XX, June 03–05, 2018, Woodstock, NY' header and 'Received 20 February 2007' footer; three authors list the same email (hychai@szu.edu.cn). Typo: 'Motivation by [32]' should be 'Motivated by [32]'.
  6. [Table 2] Table 2 reports 'Av. Sext.' exactly equal to turn counts (21,534/16,302), implying one sextuple per turn with no multi-target turns; if so, state this explicitly, as it affects the difficulty of the sextuple claim relative to panoptic sentiment benchmarks like PanoSent [30].

Circularity Check

1 steps flagged · score 2.0 of 10

No derivation-by-construction circularity; only mild generator–judge family coupling on open-ended fields, with central claims still independently tested.

  1. other [§3.2 Evolution-aware Automated Labeling; Appendix D.3 Model-based Semantic Evaluation]
    "GPT-4o performs a panoptic scan based on our evolution logic, prioritizing multimodal evidence to decode latent conviction under irony. This yields per-turn logic-based rationales... For open-ended fields (Target and Rationale), rigid string matching yields high false negatives. We adopt an LLM-as-a-Judge approach using GPT-4o-mini."

    Gold Target/Rationale strings are produced in the GPT-4o construction pipeline; the same open-ended fields are then scored by GPT-4o-mini semantic match. That couples generator and judge families, so elevated Rationale/St-R/Micro scores can partly reward reproducing GPT-4o phrasing rather than independent causal understanding. This is evaluation-distribution coupling, not a mathematical identity, and is mitigated by human audit plus closed-set Flip/Stance metrics and external MT-CSD—hence only a minor step.

full rationale

StanceFlip/ConStaFF is an empirical benchmark-and-model paper, not a first-principles derivation. Task definitions, ToS role decomposition, and multi-stage tuning are design choices evaluated by held-out F1 against labeled test splits and external MT-CSD zero-shot transfer; nothing reduces identity-wise to a fitted parameter renamed as prediction, a self-cited uniqueness theorem, or an ansatz smuggled in as necessity. Author self-citations (prior Chai et al. stance work) appear only in related-work positioning and are not load-bearing for the SOTA claim. The sole mild circularity-adjacent issue is pipeline coupling: GPT-4o synthesizes queries/labels/rationales in construction, while Appendix D.3 scores open-ended Target/Rationale via GPT-4o-mini semantic equivalence—so part of rationale/St-R gains can reflect style match to the generator family rather than fully independent understanding. Expert review, κ, categorical Flip/stance metrics, ablations, and external MT-CSD results keep the central comparative claim from collapsing by construction. Taxonomy inconsistencies noted by the skeptic are correctness bugs, not circular reductions. Score 2.

Assumptions & free parameters 6 free parameters · 6 assumptions · 3 invented entities

Load-bearing premises are dataset-construction and modeling conventions rather than physical constants: stance as a persistent four-way state with a specific flip definition; four trigger categories; GPT-4o-mediated multimodal injection and labeling validity after human audit; ImageBind+linear projector sufficiency; ToS role order; N_max=1 reflection; and LLM-as-judge semantic equivalence for open fields.

free parameters (6)
  • Multimodal retrieval cosine threshold = ≥ 0.8
    Candidates kept only if similarity ≥ 0.8; directly controls media relevance and task difficulty.
  • Stance confidence gate for auto-labels = ≥ 0.9
    Self-assessed score ≥ 0.9 used in dual-consistency filtering during construction.
  • Stance significance score threshold for media injection = > 0.7
    Dialogues proceed to multimodal injection only if GPT-4o significance > 0.7.
  • Self-reflection max iterations N_max = 1
    Critic–synthesizer loop capped at 1 iteration by empirical diminishing returns.
  • LoRA rank/alpha and stage learning rates = r=16, α=32; stage LRs as Table 8
    PEFT and optimizer settings (r=16, α=32; LRs 2e-5 / 1e-4 / 5e-5) chosen for training stability, not derived.
  • Train/val/test split ratio = 8:1:1
    8:1:1 stratified split defines all reported metrics.
assumptions (6)
  • domain assumption A holder’s stance is a continuous discourse state with Establishment, Persistence (inherit prior definitive stance on off-topic turns), and Flip only between established Support/Neutral/Oppose labels—not Unknown→definitive.
    Definition 3.1 and Stance Evolution Logic; labels and Subtask-II depend on this state machine.
  • domain assumption Emotion/sentiment and stance are separable dimensions that should be bound in one sextuple (h,g,e,s,st,r).
    Motivated via PanoSent-style panoptic analysis; structures Subtask-I metrics.
  • ad hoc to paper Flip triggers fall into exactly four socio-cognitive categories: Factual & Logical; Emotional & Value-based; Personal Experience & Anecdote; Social Influence.
    Subtask-II taxonomy; evaluation maps free text to these buckets via keywords.
  • domain assumption When text and non-text conflict (e.g., sarcasm), multimodal evidence should dominate pragmatic interpretation.
    Stated in construction labeling rules and Psychologist step of ToS.
  • domain assumption Frozen ImageBind embeddings plus a learned linear/MLP projector adequately ground image/audio/video/sticker cues for LLM stance reasoning.
    Stage-1 multimodal alignment architecture (§4.3, Appendix C).
  • ad hoc to paper GPT-4o-mini semantic YES/NO judgments (with containment fallback) are valid soft matches for Target and Rationale.
    Appendix D.3 evaluation protocol for open-ended fields.
invented entities (3)
  • StanceFlip benchmark (sextuple + flip-trigger schema)
    purpose: Provide bilingual multimodal multi-turn supervision for static cognitive snapshots and dynamic flip attribution.
    New resource defined in §3; independent use by others would constitute external evidence, but release is not demonstrated in-text.
  • Thought-of-Stance (ToS) cognitive personas (Cartographer, Psychologist, Discourse Analyst, Synthesizer–Critic)
    purpose: Decompose MCSFF into ordered target grounding, affect resolution, state tracking, and verified explanation.
    Prompt/role scaffold over standard CoT/ToT; no existence claim beyond a reasoning procedure.
  • ConStaFF multi-stage instruction-tuned system
    purpose: End-to-end baseline for MCSFF combining alignment, ToS tuning, and self-reflection tuning.
    Method artifact; performance is internal to StanceFlip plus limited MT-CSD transfer.

how reviews work

0 comments
Cite this review

Pith. "Pith review of StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting." pith.science (2026). https://pith.science/paper/3FZCSZ5S

@misc{pith2026260724191,
  author       = {Pith},
  title        = {Pith review of: StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3FZCSZ5S}},
  note         = {Machine review of arXiv:2607.24191}
}
read the original abstract

Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key limitations: failure to capture the dynamic evolution of beliefs, particularly during stance reversals; difficulty in disentangling affective states from logical reasoning; and neglect of the critical role of multimodal cues in resolving pragmatic ambiguities such as sarcasm. To address these limitations, we propose StanceFlip, a benchmark designed for multimodal conversational stance flipping forecasting over multi-turn dialogues across five modalities and multi-scenarios, which includes two novel subtasks: 1) Multimodal Stance Sextuple Extraction, extracting holder, target, emotion, sentiment, stance, and rationale as static state snapshots of dialogue to capture fine-grained cognitive structures. 2) Dynamic Stance Flip Attribution, tracking stance reversals across the conversation and identifying their underlying triggers. Alongside the dataset, we propose a dedicated framework, named ConStaFF, for Multimodal Conversational Stance Flipping Forecasting (MCSFF). Built upon a large language model, ConStaFF performs end-to-end stance reasoning, with a Thought-of-Stance (ToS) reasoning framework and a self-reflective verification mechanism integrated for structured stance modeling and faithful flip attribution. Specifically, ToS decomposes the reasoning process into specialized cognitive personas to formulate target propositions, resolve cross-modal conflicts, and infer historical stance trajectories. Extensive experiments show that our approach achieves state-of-the-art performance on both sextuple extraction and flip-trigger attribution, outperforming strong multimodal large language model baselines by substantial margins.

Figures

Figures reproduced from arXiv: 2607.24191 by the authors.

Figure 1
Figure 1. Illustration of the StanceFlip benchmark. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The overall architecture of our ConStaFF model. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 2
Figure 2. Our framework is specifically designed to tackle the key [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figures from the paper (4 more)
Figure 3
Figure 3. Figure 3: Performance evaluation on implicit stance datasets. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png]
Figure 6
Figure 6. Figure 6: Evaluation of the contribution of each modality. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Five-Dimensional Comparison: ToS (w/o Self-Reflection) [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Distribution of principal domains in the dataset. [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

59 extracted references · 9 linked inside Pith

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report.arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Abeer AlDayel and Walid Magdy. 2020. Stance Detection on Social Media: State of the Art and Trends.CoRRabs/2006.03644 (2020). arXiv:2006.03644

  3. [3]

    Emily Allaway and Kathleen R. McKeown. 2020. Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic Representations. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 8913–8931

  4. [4]

    Emily Allaway, Malavika Srikanth, and Kathleen McKeown. 2021. Adversarial Learning for Zero-Shot Stance Detection on Social Media. InProceedings of the 2021 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies. Association for Computational Linguistics, 4756–4767

  5. [5]

    Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021. Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval. In2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021. IEEE, Montreal, QC, Canada, 1708–1718

  6. [6]

    Heyan Chai, Jinhao Cui, Siyu Tang, Ye Ding, Xinwang Liu, Binxing Fang, and Qing Liao. 2025. MG-SIN: Multigraph Sparse Interaction Network for Multitask Stance Detection.IEEE Trans. Neural Networks Learn. Syst.36, 2 (2025), 3111–3125

  7. [7]

    Heyan Chai, Siyu Tang, Jinhao Cui, Ye Ding, Binxing Fang, and Qing Liao. 2022. Improving Multi-task Stance Detection with Multi-task Interaction Network. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP. Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 2990–3000

  8. [8]

    Feiyu Chen, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, and Jie Shao. 2021. Learning What and When to Drop: Adaptive Multimodal and Contextual Dy- namics for Emotion Recognition in Conversation. InProceedings of the ACM Multimedia Conference, MM. ACM, Virtual Event, China, 1064–1073

Show all 59 references
  1. [9]

    Zhanpeng Chen, Zhihong Zhu, Wanshi Xu, Yunyan Zhang, Xian Wu, and Yefeng Zheng. 2024. Aspects are Anchors: Towards Multimodal Aspect-based Sentiment Analysis via Aspect-driven Alignment and Refinement. InProceedings of the 32nd ACM International Conference on Multimedia, MM 20...

  2. [10]

    Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2024. Scaling instruction-finetuned language models.Journal of Machine Learning Research25, 70 (2024), 1–53

  3. [11]

    Yuzhe Ding, Kang He, Bobo Li, Li Zheng, Haijun He, Fei Li, Chong Teng, and Donghong Ji. 2025. Zero-Shot Conversational Stance Detection: Dataset and Approaches. InFindings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025 (Fi...

  4. [12]

    Gemmeke, Daniel P

    Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Au- dio Set: An ontology and human-labeled dataset for audio events. In2017 IEEE International Conference on Acoustics, Speech and Signal ...

  5. [13]

    Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. ImageBind One Embedding Space to Bind Them All. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17...

  6. [14]

    Dong Guo, Faming Wu, Feida Zhu, Fuxing Leng, Guang Shi, Haobin Chen, Haoqi Fan, Jian Wang, Jianyu Jiang, Jiawei Wang, et al . 2025. Seed1. 5-vl technical report.arXiv preprint arXiv:2505.07062(2025)

  7. [15]

    Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria. 2020. MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Anal- ysis. InProceedings of The 28th ACM International Conference on Multimedia,. ACM, Virtual Event / Seattle, WA, USA, 1122–1131

  8. [16]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. InThe Tenth International Conference on Learning Representa- tions, ICLR 2022, Virtual Event, April 25-29, 2...

  9. [17]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thoma...

  10. [18]

    Xincheng Ju, Dong Zhang, Suyang Zhu, Junhui Li, Shoushan Li, and Guodong Zhou. 2024. ECFCON: Emotion Consequence Forecasting in Conversations. In Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024. ACM, Melbourne, VIC, Australia, 2233–2241

  11. [19]

    J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data.Biometrics(1977), 159–174

  12. [20]

    Woojin Lee, Jaewook Lee, and Harksoo Kim. 2024. LOGIC: LLM-originated guidance for internal cognitive improvement of small language models in stance detection.PeerJ Comput. Sci.10 (2024), e2585

  13. [21]

    Bobo Li, Hao Fei, Lizi Liao, Yu Zhao, Chong Teng, Tat-Seng Chua, Donghong Ji, and Fei Li. 2023. Revisiting Disentanglement and Fusion on Modality and Context in Conversational Multimodal Emotion Recognition. InProceedings of the 31st ACM International Conference on Multimedia,...

  14. [22]

    Yingjie Li, Krishna Garg, and Cornelia Caragea. 2023. A New Direction in Stance Detection: Target-Stance Extraction in the Wild. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, July 9-14, 2023. Associ...

  15. [23]

    Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset. InProceedings of the Eighth International Joint Conference on Natural Language Processing, IJCNLP 2017, Taipei, Taiwan, November 27 - Dece...

  16. [24]

    Yupeng Li, Dacheng Wen, Haorui He, Jianxiong Guo, Xuan Ning, and Francis C. M. Lau. 2023. Contextual Target-Specific Stance Detection on Twitter: Dataset and Method. InIEEE International Conference on Data Mining, ICDM, December 1-4, 2023. IEEE, Shanghai, China, 359–367

  17. [25]

    Yingjie Li, Chenye Zhao, and Cornelia Caragea. 2023. TTS: A Target-based Teacher-Student Framework for Zero-Shot Stance Detection. InProceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023. ACM, Austin, TX, USA, 1500–1509

  18. [26]

    Bin Liang, Zixiao Chen, Lin Gui, Yulan He, Min Yang, and Ruifeng Xu. 2022. Zero-Shot Stance Detection via Contrastive Learning. InWWW ’22: The ACM Web Conference 2022, Virtual Event, Lyon, France, April 25 - 29, 2022. ACM, Virtual Event, Lyon, France, 2738–2747

  19. [27]

    Bin Liang, Ang Li, Jingqian Zhao, Lin Gui, Min Yang, Yue Yu, Kam-Fai Wong, and Ruifeng Xu. 2024. Multi-modal Stance Detection: New Datasets and Model. InFindings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 20...

  20. [28]

    Bin Liang, Chenwei Lou, Xiang Li, Lin Gui, Min Yang, and Ruifeng Xu. 2021. Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal Graphs. InProceedings of the ACM Multimedia Conference, October 20 - 24, 2021. ACM, Virtual Event, China, 4707–4715

  21. [29]

    Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C

    Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. InComputer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014,...

  22. [30]

    Meng Luo, Hao Fei, Bobo Li, Shengqiong Wu, Qian Liu, Soujanya Poria, Erik Cambria, Mong-Li Lee, and Wynne Hsu. 2024. PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis. InProceedings of the 32nd ACM International Co...

  23. [31]

    Junxia Ma, Changjiang Wang, Hanwen Xing, Dongming Zhao, and Yazhou Zhang

  24. [32]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. Self-Refine: It...

  25. [33]

    Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry

    Saif M. Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu, and Colin Cherry. 2016. SemEval-2016 Task 6: Detecting Stance in Tweets. InPro- ceedings of the 10th International Workshop on Semantic Evaluation (SemEval- 2016). Association for Computational Linguistics, ...

  26. [34]

    Fuqiang Niu, Zebang Cheng, Xianghua Fu, Xiaojiang Peng, Genan Dai, Yin Chen, Hu Huang, and Bowen Zhang. 2024. Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model. InProceedings of the 32nd ACM International Conference on Multimedia, MM ...

  27. [35]

    Fuqiang Niu, Genan Dai, Yisha Lu, Jiayu Liao, Xiang Li, Hu Huang, and Bowen Zhang. 2025. MT2-CSD: A New Dataset and Multi-Semantic Knowledge Fu- sion Method for Conversational Stance Detection.CoRRabs/2506.21053 (2025). arXiv:2506.21053

  28. [36]

    Fuqiang Niu, Min Yang, Ang Li, Baoquan Zhang, Xiaojiang Peng, and Bowen Zhang. 2024. A Challenge Dataset and Effective Models for Conversational Stance Detection. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Eval...

  29. [37]

    Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. InProceedings of the 57th Conference of the Association for Computational Linguistics, ACL...

  30. [38]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCN...

  31. [39]

    Parinaz Sobhani, Diana Inkpen, and Xiaodan Zhu. 2017. A Dataset for Multi- Target Stance Detection. InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, April 3-7, 2017, Volume 2: Short Papers. Association fo...

  32. [40]

    Foulds, Bert Huang, Lise Getoor, and Marilyn A

    Dhanya Sridhar, James R. Foulds, Bert Huang, Lise Getoor, and Marilyn A. Walker

  33. [41]

    Teng Sun, Juntong Ni, Wenjie Wang, Liqiang Jing, Yinwei Wei, and Liqiang Nie

  34. [42]

    Teng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui, Xuemeng Song, and Liqiang Nie

  35. [43]

    Llama Team. 2024. The Llama 3 Herd of Models.CoRRabs/2407.21783 (2024). arXiv:2407.21783 doi:10.48550/ARXIV.2407.21783

  36. [44]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucu- rull, David Esiobu, Jude Fernandes, Jeremy...

  37. [45]

    Fanfan Wang, Heqing Ma, Xiangqing Shen, Jianfei Yu, and Rui Xia. 2024. Observe before Generate: Emotion-Cause aware Video Caption for Multimodal Emotion Cause Generation in Conversations. InProceedings of the 32nd ACM International Conference on Multimedia, MM. ACM, Melbourne,...

  38. [46]

    Jiawen Wang, Longfei Zuo, Siyao Peng, and Barbara Plank. 2024. MultiClimate: Multimodal Stance Detection on Climate Change Videos. InProceedings of the Third Workshop on NLP for Positive Impact. Association for Computational Lin- guistics, Miami, Florida, USA, 315–326

  39. [47]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. InAdvances in Neural Information Processing Systems 35: Annual Conference on Neura...

  40. [48]

    Penghui Wei, Junjie Lin, and Wenji Mao. 2018. Multi-Target Stance Detection via a Dynamic Memory-Augmented Network. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, SIGIR 2018, July 08-12, 2018. ACM, Ann Arbor, MI, USA, 1229–1232

  41. [49]

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...

  42. [50]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Pro...

  43. [51]

    Jiaqing Yuan, Ruijie Xi, and Munindar P. Singh. 2025. Reasoner Outperforms: Generative Stance Detection with Rationalization for Social Media. InProceedings of the 36th ACM Conference on Hypertext and Social Media, HT, Chicago, IL, USA, September 15-18. ACM, Chicago, IL, USA, 28–32

  44. [52]

    Guido Zarrella and Amy Marsh. 2016. MITRE at SemEval-2016 Task 6: Transfer Learning for Stance Detection. InProceedings of the 10th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT , June 16-17, 2016. The Association for Computer Linguistics, San Diego, CA, USA...

  45. [53]

    Changmeng Zheng, Junhao Feng, Ze Fu, Yi Cai, Qing Li, and Tao Wang. 2021. Multimodal Relation Extraction with Efficient Graph Alignment. InProceedings of the ACM Multimedia Conference, MM. ACM, Virtual Event, China, 5298–5306

  46. [54]

    Great idea

    Arkaitz Zubiaga, Elena Kochkina, Maria Liakata, Rob Procter, Michal Lukasik, Kalina Bontcheva, Trevor Cohn, and Isabelle Augenstein. 2017. Discourse-Aware Rumour Stance Classification in Social Media Using Sequential Classifiers.CoRR abs/1712.02223 (2017). arXiv:1712.02223 A M...

  47. [59]

    No". 2) Stance Establishment: Moving from ’Unknown’ to the FIRST declared stance is NOT a shift. It is Establishment. shift_occurred =

    Stance Persistence (Default): If the current turn is a filler, ques- tion, or ambiguous remark, the core stance remains UNCHANGED (inherits previous). shift_occurred = "No". 2) Stance Establishment: Moving from ’Unknown’ to the FIRST declared stance is NOT a shift. It is Estab...

  48. [2015]

    Joint Models of Disagreement and Stance in Online Debate. InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing,...

  49. [2022]

    InProceedings of The 30th ACM International Conference on Multimedia, October 10 - 14, 2022

    Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment Analysis. InProceedings of The 30th ACM International Conference on Multimedia, October 10 - 14, 2022. ACM, Lisboa, Portugal, 15–23

  50. [2023]

    InProceedings of the 31st ACM International Conference on Multimedia, MM

    General Debiasing for Multimodal Sentiment Analysis. InProceedings of the 31st ACM International Conference on Multimedia, MM. ACM, Ottawa, ON, Canada, 5861–5869

  51. [2024]

    InNatural Language Processing and Chinese Computing - 13th National CCF Conference, NLPCC (Lecture Notes in Computer Science, Vol

    Chain of Stance: Stance Detection with Large Language Models. InNatural Language Processing and Chinese Computing - 13th National CCF Conference, NLPCC (Lecture Notes in Computer Science, Vol. 15363). Springer, Hangzhou, China, 82–94

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.