Pith. sign in

REVIEW 2 major objections 5 minor 32 references

The first English multimodal multi-party dialogue discourse dataset shows audio cues improve structure and relation recovery over text alone.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 22:15 UTC pith:RD6PZWZ7

load-bearing objection First English multimodal multi-party discourse-parsing resource; solid annotation and modality ablations, modest audio gains, sitcom-scale limits already owned by the authors. the 2 major comments →

arxiv 2606.00012 v1 pith:RD6PZWZ7 submitted 2026-04-13 cs.CL cs.AI

DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset

classification cs.CL cs.AI
keywords dialogue discourse parsingmultimodalmulti-party dialogueSDRTdatasetaudio cuesTV drama
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Multi-party dialogue discourse parsing finds which utterances depend on which others and what relation holds between them (question-answer, comment, contrast, and so on). Prior public resources were either text-only or restricted to two-party Chinese conversations, so they miss the visual and acoustic signals that often explain topic jumps and parallel side-conversations. This paper releases DraDDP, 495 dialogue segments drawn from Friends Season 1, with 6,374 utterances aligned to 9.1 hours of video and audio, annotated under the 16 SDRT relation labels used by STAC. Benchmarks with both classical parsers and multimodal large models show that adding audio consistently lifts Link&Rel-F1, while video helps mainly in two-party settings and can introduce noise when many speakers are present. The resource therefore supplies the missing English multi-party multimodal testbed and quantifies when each modality actually helps.

Core claim

DraDDP is the first publicly available English multimodal multi-party dialogue discourse parsing dataset; on it, multimodal (especially audio) information improves recovery of both dependency edges and relation types relative to text-only models, with the gain growing as speaker count rises.

What carries the argument

DraDDP itself: 495 Friends Season 1 segments whose subtitle lines serve as elementary discourse units, annotated with directed SDRT dependency graphs and the 16 STAC relation labels, plus time-aligned video and audio.

Load-bearing premise

That Friends Season 1 subtitle lines, treated as discourse units and labelled under STAC’s 16 SDRT relations, are a sufficiently representative and reliable proxy for real multimodal multi-party conversation.

What would settle it

A controlled human annotation of a spontaneous multi-party English corpus (meetings or podcasts) of comparable size, using the same relation inventory, that yields substantially lower inter-annotator agreement or that shows no audio gain for multimodal models trained and tested the same way.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • English multi-party multimodal discourse parsers can now be trained and compared on a single public resource.
  • Audio features (intonation, intensity, emotion) should be prioritised over raw video frames for multi-speaker scenes.
  • Models that fuse all three modalities without speaker-aware gating risk degrading relation-type accuracy.
  • Downstream tasks such as meeting summarisation and multiparty emotion recognition gain a shared multimodal discourse backbone.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The modest relation-type Kappa (0.60) and sitcom domain bias imply that gains measured on DraDDP may shrink on unscripted speech; a second spontaneous English set is the natural next measurement.
  • Because most long-distance edges appear only when speaker count exceeds two, any model that ignores speaker identity will under-use the very signal audio is providing.
  • The same annotation protocol applied to non-English multi-party video would immediately test whether the audio advantage is language-independent.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces DraDDP, claimed as the first publicly available English multimodal multi-party dialogue discourse parsing dataset, built from Friends Season 1 (495 segments, 6,374 utterances, 9.1 hours of aligned video/audio). Annotation follows SDRT with STAC’s 16 relation labels; EDUs are official subtitle lines. A four-stage protocol with LLaMA3 pre-annotation (STAC-fine-tuned) and human correction yields structure Kappa 0.91 and relation Kappa 0.60. The authors benchmark RLTST, BERTLine, MODDP, and LLaMIPa variants (including Qwen2.5 / VL / Audio / Omni backbones) under Link and Link&Rel F1, report modest gains from audio (e.g., LLaMIPa‡ Qwen2-Audio Link&Rel 55.09 vs text 53.55 on DraDDP), and analyze speaker-count, error patterns, and modality ablations. Dataset, guidelines, and code are to be released.

Significance. If the resource claim and release hold, DraDDP fills a genuine gap: existing multimodal discourse-parsing sets (JDDC 2.1, MODDP) are Chinese and two-party, while English multi-party sets (STAC, Molweni) are text-only. The controlled pre-annotation bias check, speaker-stratified results (Table 4), modality ablations (Table 7 / Appendix E), and error-type analysis (Figure 5) give the community a usable English multi-party multimodal benchmark and a clear empirical signal that audio helps relation typing more than video under current MLLM pipelines. Public release of data, guidelines, and code is a concrete contribution that can seed follow-on work on fusion and long-range multi-party structure.

major comments (2)
  1. §3.3 and §8: Relation-type inter-annotator Kappa is only 0.60 (structure 0.91). While the paper correctly notes this is slightly above STAC’s 0.58 and supplies confusable-pair guidance (Appendix B / Table 9), Link&Rel-F1 is the primary metric used to claim multimodal value (Table 2: 55.09 vs 53.55). With moderate label reliability, the absolute Link&Rel numbers and the small absolute gains should be accompanied by uncertainty estimates (e.g., bootstrap CIs over dialogues or annotator-adjudicated subsets) so readers can judge whether the audio improvement is robust to residual label noise.
  2. §3.1 and Limitations §8: All EDUs are official subtitle lines from a single sitcom season. The paper acknowledges domain bias and small scale, yet the central novelty claim (“first English multimodal multi-party…”) and the modality conclusions rest on this proxy. A short quantitative comparison of dependency-distance / relation distributions against STAC or a spontaneous multi-party corpus (or an explicit statement that conclusions are sitcom-conditioned) would better bound how far the audio-gain result is expected to transfer.
minor comments (5)
  1. Table 2: Report standard deviations or multiple-run averages for the LLM fine-tunes; single-run F1 differences of ~1.5 points are hard to interpret without variance.
  2. Figure 1 / case study (Figure 6): The gold graph and model outputs would be clearer if utterance IDs and speaker labels were aligned consistently across panels.
  3. §6.1: Video sampling (1 fps, max 16 frames) and Mel spectrogram settings are stated; a one-sentence note on whether audio is speaker-diarized or raw mixture would help reproducibility of the Qwen2-Audio / Omni runs.
  4. Appendix A Table 5: Percentages for DraDDP sum cleanly; a brief note that they are computed over gold arcs (not utterances) would avoid ambiguity.
  5. Typos / polish: “Clafi” abbreviation is used before full expansion in some figure captions; expand on first use in main text for readers unfamiliar with STAC shorthand.

Circularity Check

0 steps flagged

No circularity: DraDDP is an empirical dataset+benchmark paper whose claims rest on independent human gold labels and held-out model evaluation, not on self-definitional or fitted-as-prediction steps.

full rationale

The paper's central claims are (1) construction of the first English multimodal multi-party dialogue discourse parsing dataset and (2) empirical observation that multimodal (especially audio) features improve Link and Link&Rel F1 under fixed model families on that resource. Neither claim is obtained by redefining an input as an output. Pre-annotation (LLaMA3 fine-tuned on STAC) is explicitly non-binding: annotators watch video and correct independently; a controlled 50-sample experiment yields inter-group Kappa 0.90/0.58 nearly identical to overall agreement, so gold labels are not forced by the pre-annotator. Relation and link metrics are computed against held-out human annotations on a random train/dev/test split; no parameter fitted on the test set is re-labeled as a prediction. Self-citations (authors' prior discourse-parsing work) appear only as baselines or related work and are not load-bearing uniqueness theorems that force the present results. The STAC 16-label inventory and SDRT graph formalism are external standards, not circular redefinitions. Domain bias and modest relation Kappa (0.60) are limitations of representativeness, not circularity. Score 0 is therefore the correct finding.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

Load-bearing premises are standard discourse-theory choices and dataset-construction decisions, not free physical constants. Model hyperparameters affect benchmark numbers but not the existence claim of the dataset. No new physical entities are postulated.

free parameters (3)
  • LoRA rank / alpha
    Set to 8 / 16 for all LLM fine-tunes; affects reported F1 but is a training choice, not a scientific constant.
  • Video frame rate and max frames
    1 fps, max 16 frames; directly shapes the weak video results and is chosen by hand.
  • Learning rate and training schedule
    AdamW 1e-4, 3 epochs, batch 1 with grad accum 8; standard but free experimental knobs for the benchmark numbers.
axioms (4)
  • domain assumption SDRT directed-graph discourse structure with STAC’s 16 relation inventory is the correct annotation target for multimodal multi-party dialogue.
    Section 3.2 adopts Asher & Lascarides SDRT and STAC labels without re-deriving suitability for TV multimodal data.
  • ad hoc to paper Each official subtitle line is an elementary discourse unit (EDU).
    Section 3.1 equates subtitle lines with EDUs for alignment convenience; alternative segmentations could change graphs.
  • domain assumption Friends Season 1 multi-party scenes are representative enough of multi-party multimodal interaction for a public benchmark.
    Section 3.1 and Limitations §8; sitcom scriptedness is acknowledged but still underwrites generalization claims.
  • domain assumption Human gold after dual annotation / adjudication is the evaluation ground truth despite relation Kappa 0.60.
    Section 3.3; moderate relation agreement is reported but still used as the sole target for Link&Rel-F1.

pith-pipeline@v1.1.0-grok45 · 20261 in / 2956 out tokens · 35821 ms · 2026-07-12T22:15:05.815444+00:00 · methodology

0 comments
read the original abstract

Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations. Previous studies are mostly limited to textual modality or two-party dialogue, failing to meet the multimodal and multi-party settings. In this paper, we construct the first publicly available English multimodal dataset DraDDP for multi-party dialogue discourse parsing, based on American TV dramas. DraDDP contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, covering rich multi-party interaction scenarios. Moreover, we establish comprehensive benchmarks by evaluating this task on DraDDP and conducting in-depth analysis on the impact of different modalities. Experimental results demonstrate the value of multimodal information in capturing dialogue structures and relation types. We will publicly release the dataset, annotation guidelines, and code to promote future research in multimodal dialogue understanding.

Figures

Figures reproduced from arXiv: 2606.00012 by Peifeng Li, Qiaoming Zhu, Shannan Liu, Yaxin Fan.

Figure 1
Figure 1. Figure 1: An example of multimodal dialogue discourse parsing. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Error statistics from pre-annotation where [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: Data analysis of the DraDDP dataset. ticipants expands from “dyadic” to “multi-party”, dialogue topics are more prone to branching and jumping, making the semantic structure of multi￾party dialogues more “non-local”. 4.2 Speaker Statistics We analyze the speakers in each dialogue seg￾ment. As shown in [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Error pattern statistics of the DraDDP test set under different modalities, where [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: A case study on unimodal (T) and multimodal (T+V+A) dialogue parsing using the DraDDP dataset. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Controversial annotation cases in the DraDDP dataset. [PITH_FULL_IMAGE:figures/full_fig_p013_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

32 extracted references · 4 linked inside Pith

  1. [1]

    Nicholas Asher, Julie Hunter, Mathieu Morey, Farah Benamara, and Stergos Afantenos. 2016. Discourse structure and dialogue acts in multiparty dialogue: the stac corpus. In Proceedings of the 10th International Conference on Language Resources and Evaluation, pages 2721--2727

  2. [2]

    Nicholas Asher and Alex Lascarides. 2003. Logics of conversation. Cambridge University Press

  3. [3]

    Zineb Bennis, Julie Hunter, and Nicholas Asher. 2023. A simple but effective model for attachment in discourse parsing with multi-task learning for relation labeling. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 3412--3417

  4. [4]

    Chunkit Chan, Jiayang Cheng, Weiqi Wang, Yuxin Jiang, Tianqing Fang, Xin Liu, and Yangqiu Song. 2023. Chatgpt evaluation on sentence level relations: A focus on temporal, causal, and discourse relations. arXiv preprint arXiv:2304.14827

  5. [5]

    Yaxin Fan, Feng Jiang, Peifeng Li, Fang Kong, and Qiaoming Zhu. 2023. Improving dialogue discourse parsing via reply-to structures of addressee recognition. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8484--8495

  6. [6]

    Yaxin Fan, Feng Jiang, Peifeng Li, and Haizhou Li. 2024 a . Uncovering the potential of chatgpt for discourse analysis in dialogue: An empirical study. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, pages 16998--17010

  7. [7]

    Yaxin Fan, Peifeng Li, Fang Kong, and Qiaoming Zhu. 2025 a . Enhancing multiparty dialog discourse parsing with dynamic task-adaptive graph transformer and difficulty-aware task scheduling. IEEE Transactions on Neural Networks and Learning Systems, 36(9):16492--16506

  8. [8]

    Yaxin Fan, Peifeng Li, and Qiaoming Zhu. 2024 b . Improving multi-party dialogue generation via topic and rhetorical coherence. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 3240--3253

  9. [9]

    Yaxin Fan, Peifeng Li, and Qiaoming Zhu. 2025 b . Improving dialogue discourse parsing through discourse-aware utterance clarification. arXiv preprint arXiv:2506.15081

  10. [10]

    Xiachong Feng, Xiaocheng Feng, Bing Qin, and Xinwei Geng. 2021. Dialogue discourse-aware graph model and data augmentation for meeting summarization. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 3808--3814

  11. [11]

    Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378

  12. [12]

    Shen Gao, Xin Cheng, Mingzhe Li, Xiuying Chen, Jinpeng Li, Dongyan Zhao, and Rui Yan. 2023. Dialogue summarization with static-dynamic structure fusion graph. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 13858--13873

  13. [13]

    Chen Gong, DeXin Kong, Suxian Zhao, Xingyu Li, and Guohong Fu. 2024. Moddp: A multi-modal open-domain chinese dataset for dialogue discourse parsing. In Findings of the Association for Computational Linguistics, pages 10561--10573

  14. [14]

    Xiulan Hao, Shaohua Wei, Qian Cao, and Xiongtao Zhang. 2024. Emotion recognition in conversations based on discourse parsing and graph attention network. Telecommunications Science, 40(5):100--111

  15. [15]

    Yuru Jiang, Yu Li, Weikai He, Jie Chen, Yanchao Yu, and Yangsen Zhang. 2023. A new dataset and parsing model for chinese multiparty dialogue discourse structure. In Proceedings of the 2023 International Conference on Asian Language Processing, pages 221--227

  16. [16]

    Xincheng Ju, Dong Zhang, Suyang Zhu, Junhui Li, Shoushan Li, and Guodong Zhou. 2024. Ecfcon: Emotion consequence forecasting in conversations. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 2233--2241

  17. [17]

    Chuyuan Li, Chlo \'e Braud, Maxime Amblard, and Giuseppe Carenini. 2024 a . Discourse relation prediction and discourse parsing in dialogues with minimal supervision. In Proceedings of the 5th Workshop on Computational Approaches to Discourse, pages 161--176

  18. [18]

    Jiaqi Li, Ming Liu, Min-Yen Kan, Zihao Zheng, Zekun Wang, Wenqiang Lei, Ting Liu, and Bing Qin. 2020. Molweni: A challenge multiparty dialogues-based machine reading comprehension dataset with discourse structure. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2642--2652

  19. [19]

    Jingyang Li, Shengli Song, Yixin Li, Hanxiao Zhang, and Guangneng Hu. 2024 b . Chatmdg: A discourse parsing graph fusion based approach for multi-party dialogue generation. Information Fusion, 110:102469

  20. [20]

    Wei Li, Luyao Zhu, Wei Shao, Zonglin Yang, and Erik Cambria. 2023. Task-aware self-supervised framework for dialogue discourse parsing. In Findings of the Association for Computational Linguistics, pages 14162--14173

  21. [21]

    Shannan Liu, Peifeng Li, Yaxin Fan, and Qiaoming Zhu. 2025. Enhancing multi-party dialogue discourse parsing with explanation generation. In Proceedings of the 31st International Conference on Computational Linguistics, pages 1531--1544

  22. [22]

    Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2023. Enhanced speaker-aware multi-party multi-turn dialogue comprehension. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2410--2423

  23. [23]

    Virgile Rennard, Guokan Shang, Michalis Vazirgiannis, and Julie Hunter. 2024. Leveraging discourse structure for extractive meeting summarization. arXiv preprint arXiv:2405.11055

  24. [24]

    Kate Thompson, Akshay Chaturvedi, Julie Hunter, and Nicholas Asher. 2024 a . Llamipa: an incremental discourse parser. In Findings of the Association for Computational Linguistics, pages 6418--6430

  25. [25]

    Kate Thompson, Julie Hunter, and Nicholas Asher. 2024 b . Discourse structure for the minecraft corpus. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, pages 4957--4967

  26. [26]

    Chengrui Wang, Shaoming Ji, and Fang Kong. 2024. Local or global optimization for dialogue discourse parsing. In CCF International Conference on Natural Language Processing and Chinese Computing, pages 149--161. Springer

  27. [27]

    Jiahui Xu, Feng Jiang, Anningzhe Gao, Luis Fernando D'Haro, and Haizhou Li. 2024. Unsupervised mutual learning of discourse parsing and topic segmentation in dialogue. arXiv preprint arXiv:2405.19799

  28. [28]

    Duzhen Zhang, Feilong Chen, and Xiuyi Chen. 2023. Dualgats: Dual graph attention networks for emotion recognition in conversations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 7395--7408

  29. [29]

    Hanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou, Shaojie Zhao, and Jiayan Teng. 2022. Mintrec: A new dataset for multimodal intent recognition. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1688--1697

  30. [30]

    Nan Zhao, Haoran Li, Youzheng Wu, and Xiaodong He. 2022. Jddc 2.1: A multimodal chinese dialogue dataset with joint tasks of query rewriting, response generation, discourse parsing, and summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 12037--12051

  31. [31]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  32. [32]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...