REVIEW 2 major objections 5 minor 32 references
The first English multimodal multi-party dialogue discourse dataset shows audio cues improve structure and relation recovery over text alone.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 22:15 UTC pith:RD6PZWZ7
load-bearing objection First English multimodal multi-party discourse-parsing resource; solid annotation and modality ablations, modest audio gains, sitcom-scale limits already owned by the authors. the 2 major comments →
DraDDP: A Multimodal Multi-Party Dialogue Discourse Parsing Dataset
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
DraDDP is the first publicly available English multimodal multi-party dialogue discourse parsing dataset; on it, multimodal (especially audio) information improves recovery of both dependency edges and relation types relative to text-only models, with the gain growing as speaker count rises.
What carries the argument
DraDDP itself: 495 Friends Season 1 segments whose subtitle lines serve as elementary discourse units, annotated with directed SDRT dependency graphs and the 16 STAC relation labels, plus time-aligned video and audio.
Load-bearing premise
That Friends Season 1 subtitle lines, treated as discourse units and labelled under STAC’s 16 SDRT relations, are a sufficiently representative and reliable proxy for real multimodal multi-party conversation.
What would settle it
A controlled human annotation of a spontaneous multi-party English corpus (meetings or podcasts) of comparable size, using the same relation inventory, that yields substantially lower inter-annotator agreement or that shows no audio gain for multimodal models trained and tested the same way.
If this is right
- English multi-party multimodal discourse parsers can now be trained and compared on a single public resource.
- Audio features (intonation, intensity, emotion) should be prioritised over raw video frames for multi-speaker scenes.
- Models that fuse all three modalities without speaker-aware gating risk degrading relation-type accuracy.
- Downstream tasks such as meeting summarisation and multiparty emotion recognition gain a shared multimodal discourse backbone.
Where Pith is reading between the lines
- The modest relation-type Kappa (0.60) and sitcom domain bias imply that gains measured on DraDDP may shrink on unscripted speech; a second spontaneous English set is the natural next measurement.
- Because most long-distance edges appear only when speaker count exceeds two, any model that ignores speaker identity will under-use the very signal audio is providing.
- The same annotation protocol applied to non-English multi-party video would immediately test whether the audio advantage is language-independent.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DraDDP, claimed as the first publicly available English multimodal multi-party dialogue discourse parsing dataset, built from Friends Season 1 (495 segments, 6,374 utterances, 9.1 hours of aligned video/audio). Annotation follows SDRT with STAC’s 16 relation labels; EDUs are official subtitle lines. A four-stage protocol with LLaMA3 pre-annotation (STAC-fine-tuned) and human correction yields structure Kappa 0.91 and relation Kappa 0.60. The authors benchmark RLTST, BERTLine, MODDP, and LLaMIPa variants (including Qwen2.5 / VL / Audio / Omni backbones) under Link and Link&Rel F1, report modest gains from audio (e.g., LLaMIPa‡ Qwen2-Audio Link&Rel 55.09 vs text 53.55 on DraDDP), and analyze speaker-count, error patterns, and modality ablations. Dataset, guidelines, and code are to be released.
Significance. If the resource claim and release hold, DraDDP fills a genuine gap: existing multimodal discourse-parsing sets (JDDC 2.1, MODDP) are Chinese and two-party, while English multi-party sets (STAC, Molweni) are text-only. The controlled pre-annotation bias check, speaker-stratified results (Table 4), modality ablations (Table 7 / Appendix E), and error-type analysis (Figure 5) give the community a usable English multi-party multimodal benchmark and a clear empirical signal that audio helps relation typing more than video under current MLLM pipelines. Public release of data, guidelines, and code is a concrete contribution that can seed follow-on work on fusion and long-range multi-party structure.
major comments (2)
- §3.3 and §8: Relation-type inter-annotator Kappa is only 0.60 (structure 0.91). While the paper correctly notes this is slightly above STAC’s 0.58 and supplies confusable-pair guidance (Appendix B / Table 9), Link&Rel-F1 is the primary metric used to claim multimodal value (Table 2: 55.09 vs 53.55). With moderate label reliability, the absolute Link&Rel numbers and the small absolute gains should be accompanied by uncertainty estimates (e.g., bootstrap CIs over dialogues or annotator-adjudicated subsets) so readers can judge whether the audio improvement is robust to residual label noise.
- §3.1 and Limitations §8: All EDUs are official subtitle lines from a single sitcom season. The paper acknowledges domain bias and small scale, yet the central novelty claim (“first English multimodal multi-party…”) and the modality conclusions rest on this proxy. A short quantitative comparison of dependency-distance / relation distributions against STAC or a spontaneous multi-party corpus (or an explicit statement that conclusions are sitcom-conditioned) would better bound how far the audio-gain result is expected to transfer.
minor comments (5)
- Table 2: Report standard deviations or multiple-run averages for the LLM fine-tunes; single-run F1 differences of ~1.5 points are hard to interpret without variance.
- Figure 1 / case study (Figure 6): The gold graph and model outputs would be clearer if utterance IDs and speaker labels were aligned consistently across panels.
- §6.1: Video sampling (1 fps, max 16 frames) and Mel spectrogram settings are stated; a one-sentence note on whether audio is speaker-diarized or raw mixture would help reproducibility of the Qwen2-Audio / Omni runs.
- Appendix A Table 5: Percentages for DraDDP sum cleanly; a brief note that they are computed over gold arcs (not utterances) would avoid ambiguity.
- Typos / polish: “Clafi” abbreviation is used before full expansion in some figure captions; expand on first use in main text for readers unfamiliar with STAC shorthand.
Circularity Check
No circularity: DraDDP is an empirical dataset+benchmark paper whose claims rest on independent human gold labels and held-out model evaluation, not on self-definitional or fitted-as-prediction steps.
full rationale
The paper's central claims are (1) construction of the first English multimodal multi-party dialogue discourse parsing dataset and (2) empirical observation that multimodal (especially audio) features improve Link and Link&Rel F1 under fixed model families on that resource. Neither claim is obtained by redefining an input as an output. Pre-annotation (LLaMA3 fine-tuned on STAC) is explicitly non-binding: annotators watch video and correct independently; a controlled 50-sample experiment yields inter-group Kappa 0.90/0.58 nearly identical to overall agreement, so gold labels are not forced by the pre-annotator. Relation and link metrics are computed against held-out human annotations on a random train/dev/test split; no parameter fitted on the test set is re-labeled as a prediction. Self-citations (authors' prior discourse-parsing work) appear only as baselines or related work and are not load-bearing uniqueness theorems that force the present results. The STAC 16-label inventory and SDRT graph formalism are external standards, not circular redefinitions. Domain bias and modest relation Kappa (0.60) are limitations of representativeness, not circularity. Score 0 is therefore the correct finding.
Axiom & Free-Parameter Ledger
free parameters (3)
- LoRA rank / alpha
- Video frame rate and max frames
- Learning rate and training schedule
axioms (4)
- domain assumption SDRT directed-graph discourse structure with STAC’s 16 relation inventory is the correct annotation target for multimodal multi-party dialogue.
- ad hoc to paper Each official subtitle line is an elementary discourse unit (EDU).
- domain assumption Friends Season 1 multi-party scenes are representative enough of multi-party multimodal interaction for a public benchmark.
- domain assumption Human gold after dual annotation / adjudication is the evaluation ground truth despite relation Kappa 0.60.
read the original abstract
Multi-party dialogue discourse parsing aims to identify dependency structures and relation types between utterances in conversations. Previous studies are mostly limited to textual modality or two-party dialogue, failing to meet the multimodal and multi-party settings. In this paper, we construct the first publicly available English multimodal dataset DraDDP for multi-party dialogue discourse parsing, based on American TV dramas. DraDDP contains 495 dialogue segments with 6,374 utterances and 9.1 hours of parallel video content, covering rich multi-party interaction scenarios. Moreover, we establish comprehensive benchmarks by evaluating this task on DraDDP and conducting in-depth analysis on the impact of different modalities. Experimental results demonstrate the value of multimodal information in capturing dialogue structures and relation types. We will publicly release the dataset, annotation guidelines, and code to promote future research in multimodal dialogue understanding.
Figures
Reference graph
Works this paper leans on
-
[1]
Nicholas Asher, Julie Hunter, Mathieu Morey, Farah Benamara, and Stergos Afantenos. 2016. Discourse structure and dialogue acts in multiparty dialogue: the stac corpus. In Proceedings of the 10th International Conference on Language Resources and Evaluation, pages 2721--2727
2016
-
[2]
Nicholas Asher and Alex Lascarides. 2003. Logics of conversation. Cambridge University Press
2003
-
[3]
Zineb Bennis, Julie Hunter, and Nicholas Asher. 2023. A simple but effective model for attachment in discourse parsing with multi-task learning for relation labeling. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 3412--3417
2023
-
[4]
Chunkit Chan, Jiayang Cheng, Weiqi Wang, Yuxin Jiang, Tianqing Fang, Xin Liu, and Yangqiu Song. 2023. Chatgpt evaluation on sentence level relations: A focus on temporal, causal, and discourse relations. arXiv preprint arXiv:2304.14827
Pith/arXiv arXiv 2023
-
[5]
Yaxin Fan, Feng Jiang, Peifeng Li, Fang Kong, and Qiaoming Zhu. 2023. Improving dialogue discourse parsing via reply-to structures of addressee recognition. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8484--8495
2023
-
[6]
Yaxin Fan, Feng Jiang, Peifeng Li, and Haizhou Li. 2024 a . Uncovering the potential of chatgpt for discourse analysis in dialogue: An empirical study. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, pages 16998--17010
2024
-
[7]
Yaxin Fan, Peifeng Li, Fang Kong, and Qiaoming Zhu. 2025 a . Enhancing multiparty dialog discourse parsing with dynamic task-adaptive graph transformer and difficulty-aware task scheduling. IEEE Transactions on Neural Networks and Learning Systems, 36(9):16492--16506
2025
-
[8]
Yaxin Fan, Peifeng Li, and Qiaoming Zhu. 2024 b . Improving multi-party dialogue generation via topic and rhetorical coherence. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 3240--3253
2024
-
[9]
Yaxin Fan, Peifeng Li, and Qiaoming Zhu. 2025 b . Improving dialogue discourse parsing through discourse-aware utterance clarification. arXiv preprint arXiv:2506.15081
Pith/arXiv arXiv 2025
-
[10]
Xiachong Feng, Xiaocheng Feng, Bing Qin, and Xinwei Geng. 2021. Dialogue discourse-aware graph model and data augmentation for meeting summarization. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 3808--3814
2021
-
[11]
Joseph L Fleiss. 1971. Measuring nominal scale agreement among many raters. Psychological bulletin, 76(5):378
1971
-
[12]
Shen Gao, Xin Cheng, Mingzhe Li, Xiuying Chen, Jinpeng Li, Dongyan Zhao, and Rui Yan. 2023. Dialogue summarization with static-dynamic structure fusion graph. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 13858--13873
2023
-
[13]
Chen Gong, DeXin Kong, Suxian Zhao, Xingyu Li, and Guohong Fu. 2024. Moddp: A multi-modal open-domain chinese dataset for dialogue discourse parsing. In Findings of the Association for Computational Linguistics, pages 10561--10573
2024
-
[14]
Xiulan Hao, Shaohua Wei, Qian Cao, and Xiongtao Zhang. 2024. Emotion recognition in conversations based on discourse parsing and graph attention network. Telecommunications Science, 40(5):100--111
2024
-
[15]
Yuru Jiang, Yu Li, Weikai He, Jie Chen, Yanchao Yu, and Yangsen Zhang. 2023. A new dataset and parsing model for chinese multiparty dialogue discourse structure. In Proceedings of the 2023 International Conference on Asian Language Processing, pages 221--227
2023
-
[16]
Xincheng Ju, Dong Zhang, Suyang Zhu, Junhui Li, Shoushan Li, and Guodong Zhou. 2024. Ecfcon: Emotion consequence forecasting in conversations. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 2233--2241
2024
-
[17]
Chuyuan Li, Chlo \'e Braud, Maxime Amblard, and Giuseppe Carenini. 2024 a . Discourse relation prediction and discourse parsing in dialogues with minimal supervision. In Proceedings of the 5th Workshop on Computational Approaches to Discourse, pages 161--176
2024
-
[18]
Jiaqi Li, Ming Liu, Min-Yen Kan, Zihao Zheng, Zekun Wang, Wenqiang Lei, Ting Liu, and Bing Qin. 2020. Molweni: A challenge multiparty dialogues-based machine reading comprehension dataset with discourse structure. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2642--2652
2020
-
[19]
Jingyang Li, Shengli Song, Yixin Li, Hanxiao Zhang, and Guangneng Hu. 2024 b . Chatmdg: A discourse parsing graph fusion based approach for multi-party dialogue generation. Information Fusion, 110:102469
2024
-
[20]
Wei Li, Luyao Zhu, Wei Shao, Zonglin Yang, and Erik Cambria. 2023. Task-aware self-supervised framework for dialogue discourse parsing. In Findings of the Association for Computational Linguistics, pages 14162--14173
2023
-
[21]
Shannan Liu, Peifeng Li, Yaxin Fan, and Qiaoming Zhu. 2025. Enhancing multi-party dialogue discourse parsing with explanation generation. In Proceedings of the 31st International Conference on Computational Linguistics, pages 1531--1544
2025
-
[22]
Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2023. Enhanced speaker-aware multi-party multi-turn dialogue comprehension. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2410--2423
2023
-
[23]
Virgile Rennard, Guokan Shang, Michalis Vazirgiannis, and Julie Hunter. 2024. Leveraging discourse structure for extractive meeting summarization. arXiv preprint arXiv:2405.11055
Pith/arXiv arXiv 2024
-
[24]
Kate Thompson, Akshay Chaturvedi, Julie Hunter, and Nicholas Asher. 2024 a . Llamipa: an incremental discourse parser. In Findings of the Association for Computational Linguistics, pages 6418--6430
2024
-
[25]
Kate Thompson, Julie Hunter, and Nicholas Asher. 2024 b . Discourse structure for the minecraft corpus. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, pages 4957--4967
2024
-
[26]
Chengrui Wang, Shaoming Ji, and Fang Kong. 2024. Local or global optimization for dialogue discourse parsing. In CCF International Conference on Natural Language Processing and Chinese Computing, pages 149--161. Springer
2024
-
[27]
Jiahui Xu, Feng Jiang, Anningzhe Gao, Luis Fernando D'Haro, and Haizhou Li. 2024. Unsupervised mutual learning of discourse parsing and topic segmentation in dialogue. arXiv preprint arXiv:2405.19799
Pith/arXiv arXiv 2024
-
[28]
Duzhen Zhang, Feilong Chen, and Xiuyi Chen. 2023. Dualgats: Dual graph attention networks for emotion recognition in conversations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 7395--7408
2023
-
[29]
Hanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou, Shaojie Zhao, and Jiayan Teng. 2022. Mintrec: A new dataset for multimodal intent recognition. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1688--1697
2022
-
[30]
Nan Zhao, Haoran Li, Youzheng Wu, and Xiaodong He. 2022. Jddc 2.1: A multimodal chinese dialogue dataset with joint tasks of query rewriting, response generation, discourse parsing, and summarization. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 12037--12051
2022
-
[31]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[32]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.