{"id":"c8bc3923-81cc-468b-8fe9-98d6286d7a77","arxiv_id":"2412.18489","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Existing SLU speech datasets are insufficient for training ML models on collaborative problem solving because they lack multimodal, longitudinal, ambiguous, and team-dynamics data.","lead":"This report reviews 22 speech and text datasets used for Spoken Language Understanding and rates their suitability for training machine learning models on team-based collaborative problem solving. It concludes that current datasets lack the multimodal, longitudinal, ambiguous, and team-dynamics data needed for that goal.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quantitative suitability ratings in Section V are the only direct evidence for the insufficiency claim, but they rest on undocumented manual inspection of 100–300 utterances per category with no sampling protocol, annotator count, or inter-rater reliability; the central conclusion is…","rationale":"The paper is a position report, and the central claim is conditional: if the characterization is correct, existing SLU datasets are insufficient and new datasets should be created with certain features. The qualitative part of the argument—that SLU annotation schemas (intents, slots, speaker IDs) do not directly encode team-level constructs like consensus, role shifts, or psychological safety—is plausible and supported by Table I and the dataset descriptions. The paper also deserves credit for deriving dataset requirements from a cognitive/social/emotional framework. However, the load-bearing support for the quantitative 'suitability' characterization is the manual rating procedure in Section V. The text is explicit that most metrics were manually evaluated on 100–300 utterances from 2–4 datasets per category, but it never specifies how utterances were sampled, who did the rating, what rubric mapped observations to the reported ranges, or whether ratings were stable across raters. This makes the percentage estimates in Tables II–IX non-reproducible. In addition, 'Expected Accuracy' is a model-dependent quantity rather than a dataset attribute, and no benchmark citations support the numbers. The reader's weakest assumption identifies exactly this gap. My proposed check—an independent re-scoring with reliability statistics—would settle whether the numbers are stable. If they are not, the paper should be read as a qualitative position paper, and the quantitative tables should be revised or removed. That does not change the conditional verdict: the qualitative conclusion remains plausible, but the quantitative evidence should not be treated as established.","tokens_in":28770,"tokens_out":7916,"duration_ms":77237,"concrete_test":"Release the exact sampling script and per-utterance rating sheets for the metrics in Section V, then have two independent annotators re-score a stratified random sample of 50 utterances per dataset category and compute Krippendorff's alpha for each metric. If alpha < 0.6 for metrics such as Difficulty in Representation (%) or Presence of Ambiguities (%), or if category-level medians shift by more than one rating band, the quantitative suitability tables are not reproducible and should be presented as preliminary expert judgment rather than quantitative evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim that current SLU datasets are inadequate for CPS and that new datasets must be created with multimodal, longitudinal, ambiguous, and conflict-laden speech rests on the metric tables in Section V (Tables II, IV, VI, VII, VIII, and IX). Sections V.A–V.H state repeatedly that these metrics were obtained by 'manually analyzing samples of 100–300 utterances for 2–4 datasets per category' (e.g., V.A), and Section V.B says that only utterance size and presence of ambiguity words were computed automatically. No sampling frame, random seed, annotator qualifications, coding rubric, or reliability statistic is reported. Appendix A lists per-dataset values but not raw ratings or sampling details. 'Expected Accuracy (%)' is treated as a dataset property even though it is a model-dependent outcome; no citations to published benchmarks support values like 'Text: 90–95' or 'Sound: 60–70' for Task-Oriented Dialogue. If these numbers are not repeatable, the quantitative evidence for 'insufficiency' collapses to expert opinion, and specific claims such as sound representation difficulty being 40–50% or Multi-Speaker Interaction having 30–40% ambiguity are unverifiable. The qualitative taxonomy (Table I) still suggests low CPS coverage, but the paper's quantitative characterization cannot be treated as established evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a set of metrics for characterizing whether existing speech datasets capture the cognitive, social, and emotional activities involved in Collaborative Problem Solving (CPS), and it applies these metrics to four categories of Spoken Language Understanding (SLU) datasets: Task-Oriented Dialogue, Multi-Speaker Interaction, Text Understanding, and Speech Recognition. Based on the resulting quantitative ratings, the paper concludes that current SLU datasets are poorly suited for training ML models to improve CPS, and it lists requirements for future datasets, including multimodal data, longitudinal tracking, ambiguous and ill-defined utterances, and situations involving disruptions and conflicts. The qualitative taxonomy and the enumerated dataset features are the paper's main contributions, while the quantitative suitability ratings in Section V and the appendix are presented as evidence for the insufficiency claim.","tokens_in":29029,"tokens_out":3231,"duration_ms":33430,"significance":"If the quantitative characterization were reproducible, the paper would provide a useful mapping between SLU dataset categories and CPS-relevant constructs, and its requirements list would be a practical starting point for dataset design. The qualitative comparison is genuinely informative: the SLU annotation schemes (intents, slots, dialogue states, transcriptions) do not naturally encode team-level constructs such as Team Agreement, Team Synchronization, conflict resolution, or longitudinal team dynamics, and the paper makes this mismatch visible through a structured taxonomy. A further strength is that the paper states its assumptions explicitly, including the team size of about three or four members and the use of speech as the primary channel, which makes the scope of the claim clear. The quantitative tables, however, are not currently established evidence: they are based on undocumented manual inspection and are internally inconsistent, so the paper's value at present is largely conceptual and agenda-setting rather than an empirical measurement study.","major_comments":[{"comment":"The quantitative suitability ratings are load-bearing for the paper's central insufficiency claim, but the manuscript reports no sampling protocol, annotator count, coding rubric, or inter-rater reliability for the manual inspection of 100-300 utterances from 2-4 datasets per category. Values such as 'Difficulty in Representation (%) Sound: 40-50' (Table II) and 'Presence and Amount of Ambiguities/Unknowns (% of data affected): 30-40%' (Table III) are presented as precise ranges without any evidence that they are representative or repeatable. I request either a documented methodology for the manual analysis (including how datasets and utterances were sampled, how many annotators participated, and a reliability statistic such as Cohen's kappa) or an explicit reframing of these numbers as the authors' expert estimates, with all inference in the conclusions downgraded accordingly.","section":"Section V, Tables II-IX and XIII-XVI"},{"comment":"The metric 'Expected Accuracy (%)' is treated as a property of a dataset category, but accuracy is a property of a model trained and evaluated on a dataset, not of the dataset itself. The reported values such as 'Text: 90-95' and 'Sound: 70-80' for Task-Oriented Dialogue are not accompanied by citations to published benchmark results or by a specification of the model family, train/test split, or evaluation metric. As written, these entries are unverifiable and should either be replaced with documented benchmark results or removed from the dataset-characterization scheme.","section":"Section V.A, Table II"},{"comment":"The appendix tables duplicate the main metric tables but assign different values to the same constructs. For example, Table II reports 'Expected Accuracy (%) Text: 90-95' for Task-Oriented Dialogue, while Table X reports 'MultiWOZ: 85-90' and 'SGD: 70-80' for the same quantity; similarly, 'Difficulty in Representation (%) Text: 10-20' in Table II becomes 'MultiWOZ: 10-20; SGD: 40-50' in Table X. This internal inconsistency means a reader cannot tell which numbers are the authors' final estimates, and it materially undermines the quantitative basis for the conclusion that current SLU datasets are inadequate. The authors should reconcile the two sets of tables or clearly designate one as the reported result.","section":"Appendix A, Tables X-XVI vs Section V, Tables II-IX"},{"comment":"The 'Presence of Ambiguities' metric is said to be computed automatically by identifying ambiguous words such as 'maybe', 'probably', or 'unsure' in utterances, but no details are given about tokenization, normalization, the exact word list, or which datasets and splits were processed. This operationalization conflates lexical hedges with semantic ambiguity and has no validation against human judgments or downstream ambiguity-related tasks. Since ambiguity plays a central role in the paper's recommendation that future datasets include 'short, ambiguous, and ill-defined speech utterances', this metric needs either a rigorous validation study or a more cautious interpretation as a proxy for one type of uncertainty.","section":"Section V.B, ambiguity metric"}],"minor_comments":[{"comment":"The text contains an unresolved placeholder '[ ?]' in the description of semantic parsing metrics; this should be completed or removed.","section":"Section V.B"},{"comment":"Table references are inconsistent: Section V.C says 'Table XIII summarizes' for the problem-solving metrics that appear as Table IV in the main text, and Section V.E refers to 'Table XIV' for the reactivity metrics that appear as Table VI. The table numbering should be made consistent throughout.","section":"Section V.C and Section V.E"},{"comment":"The appendix table XIII introduces 'ICSI' as a Multi-Speaker Interaction dataset, but ICSI is neither described in Table I nor defined in the text; every dataset used in the quantitative analysis should be listed and described.","section":"Table XIII and Table I"},{"comment":"The abbreviation 'TEB' is defined as 'Team Emotional Behavior', but the defining sentence says 'TEM refers to psychological safety'; this appears to be a typo for TEB.","section":"Section II.E.2"},{"comment":"Several definitions of key SLU activities rely on non-archival blog or vendor pages rather than primary literature; for a research paper the definitions should be anchored in peer-reviewed sources or standard textbooks.","section":"References [60], [61], [97], [98], [99]"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads as a technical report or position paper rather than a traditional journal article. The conceptual framework and dataset-needs list are useful, but the quantitative evidence is currently presented with a precision that the methodology cannot support. I would advise the editor that a revision should either add the missing measurement documentation or explicitly change the epistemic status of the tables to 'expert judgments' and rewrite the conclusions so that they do not rest on unverifiable numeric ranges. The paper also cites several of the authors' own earlier works to motivate the activity taxonomy; this is acceptable if those works substantiate the specific constructs, but the citation pattern should be checked for self-citation inflation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is a position report that maps CPS cognitive/social/emotional activities onto SLU datasets and concludes existing corpora don't cover team dynamics. That framing is genuinely new, and the conclusion is plausible. The problem is Section V: the metric tables (Tables II–XVI) look quantitative but rest on undocumented manual inspection of 100–300 utterances per category. No sampling protocol, no annotator count, no inter-rater reliability. 'Expected Accuracy (%)' is treated as a dataset property even though it is model-dependent. So the numbers should not be treated as evidence.\n\nWhat's good: the taxonomy in Table I is a useful catalog, the list of CPS activities (framing, reframing, consensus, fixation) is well grounded in the cognitive science literature, and the requirements in Section VI (multimodal, longitudinal, ambiguous, conflict-laden data) are sensible. The qualitative mismatch argument is solid: SLU annotations are designed for intents and slots, not for team-level constructs like bridging knowledge gaps, team synchronization, or psychological safety. That central claim holds up without the numbers.\n\nSoft spots: the quantitative characterization. Even the internal consistency is shaky—Appendix tables assign metric values to individual datasets, but the main tables only give category-level ranges, and there is no explanation of how per-dataset values were aggregated. Also 'Difficulty in Representation (%)' and 'Expected Accuracy (%)' are just asserted. The paper would be better if it dropped the percentages and kept the qualitative ordinal comparisons, or if it published the rubric and raw ratings.\n\nMinor: some references are weak (GeeksforGeeks, blog posts) for a survey. Self-citation in the intro is fine; the cited work is the authors' own lab experiments, which is relevant.\n\nWho this is for: researchers building datasets for collaborative AI or team cognition. They will get a useful requirements checklist and a structured way to think about what is missing from SLU benchmarks. It deserves a serious referee, mostly because it fills a real gap and makes a falsifiable claim about dataset insufficiency. I would send it to review but ask for the quantitative tables to be either removed, tied to a released rubric, or replaced with existing benchmark numbers.\n\nFor my own work: I would cite the requirements list, not the percentages.","headline":"A useful requirements list for CPS datasets, undermined by unsupported quantitative ratings; the qualitative claim is plausible but the numbers should not be trusted.","tokens_in":29537,"tokens_out":1561,"would_cite":true,"duration_ms":15548,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that existing spoken language understanding datasets are not suitable for training machine learning models to support collaborative problem solving, and it identifies the specific features such datasets would need.","keywords":["collaborative problem solving","spoken language understanding","dataset suitability","speech datasets","machine learning training","team dynamics","multi-modal data","ambiguity and conflict"],"falsifier":"A concrete audit: take a random sample of at least 1000 utterances from each of the four dataset categories, have multiple independent annotators score the same metrics (ambiguity rate, abstraction levels, reframing types, social interaction scores) using a pre-registered coding manual, and compute inter-rater agreement. If the resulting scores differ substantially from the paper's ranges—for instance, if task-oriented dialogue shows ambiguity rates above 40% or multi-speaker interaction shows low social interaction scores—the paper's conclusion that existing datasets lack CPS-relevant features would be undercut. Additionally, if a model trained on an existing multi-speaker dataset (e.g., AMI Meeting Corpus) achieved human-level performance on a held-out CPS task involving ambiguous, disrupted, and longitudinally tracked team dialogues, that would directly contradict the claim that current data are insufficient.","tokens_in":28545,"feed_emoji":"🗣️","tokens_out":3821,"duration_ms":34872,"temperature":0.7,"pith_summary":"The paper tries to establish that current speech datasets, built for spoken language understanding tasks like intent detection and slot filling, lack the properties needed to train machine learning models for collaborative problem solving in small teams. It analyzes a broad set of popular SLU datasets using a battery of metrics that capture cognitive, social, and emotional dimensions of problem solving. The central finding is that these datasets do not represent the multi-modal, longitudinal, ambiguous, and conflict-laden interactions that characterize real team problem solving. If this is right, then models trained on existing data will not transfer well to collaborative settings, and new dataset collection efforts should prioritize those missing features.","feed_headline":"Existing speech data can't train team problem-solving AI","feed_subtitle":"Analysis of 22 SLU datasets finds they lack the ambiguity, conflict, and multi-modal interaction that team problem solving needs.","key_machinery":"The central mechanism is a two-part characterization framework: a taxonomy of SLU datasets grouped by purpose (task-oriented dialogue, multi-speaker interaction, text understanding, speech recognition), and a set of CPS-based metrics that operationalize cognitive, social, and emotional activities. Metrics include kinds of modalities and expected accuracy, size and abstraction levels of utterances, presence of ambiguities and unknowns, number of reframing steps, iteration counts for consensus, scores for social interaction and behavioral adaptability, and measures of goal changes, emotion tracking, and response to critique. These metrics are applied to samples of 100-300 utterances from 2-4 datasets per category, yielding quantitative profiles that expose where each dataset category falls short of CPS needs.","core_discovery":"The paper's central claim is that no existing SLU dataset adequately represents collaborative problem solving as it occurs in teams of about four members talking to each other. The analysis organizes datasets into four categories—task-oriented dialogue, multi-speaker interaction, text understanding, and speech recognition—and scores them on metrics for multi-modal tracking, semantic parsing, solution elaboration, reactivity to unexpected situations, social and emotional feature management, individual-in-team issues, and problem solving process. The conclusion is that the datasets are weakest exactly where CPS is most demanding: they contain little ambiguous or ill-defined speech, few sudden disruptions or conflicts, no longitudinal tracking of team dynamics, and no integrated multi-modal signals beyond speech and text. The paper therefore specifies that new datasets should include multi-modal data capturing diverse team interactions, longitudinal data for tracking dynamics over time, short ambiguous and ill-defined utterances, and situations of sudden disruptions and conflicts.","pith_inferences":["An implication the paper leaves implicit is that the same metric battery could be applied prospectively to newly collected multimodal team-interaction data, giving dataset builders a pre-hoc suitability score rather than a post-hoc justification.","The paper's emphasis on sudden disruptions and conflicts suggests a testable extension: adding scripted 'disruption events' to an existing multi-speaker corpus and measuring whether models trained on that augmented data handle unexpected turns better than models trained on the original corpus.","If the quantitative ratings are representative, a practical consequence is that 'general-purpose' SLU pretraining may not transfer to CPS even with fine-tuning, because the missing phenomena (ambiguity, role shifts, long-horizon team dynamics) are not just rare but absent from the pretraining distribution.","The reliance on manual inspection of small samples points to a concrete next step: a larger-scale annotation study with multiple raters could quantify inter-rater reliability and produce confidence intervals for each metric, converting the current ordinal scores into statistically grounded estimates."],"forward_implications":["If the paper's analysis is correct, anyone training an ML model for collaborative problem solving on existing SLU datasets should expect poor performance in real team settings, because the training data will not contain the ambiguous, conflicting, and dynamic interactions that CPS requires.","The proposed list of needed dataset features gives concrete guidance for new data collection efforts: multimodal recordings (not just speech), longitudinal sessions, deliberately ambiguous and ill-defined utterances, and scripted or naturally occurring disruptions and conflicts.","The metric battery itself can serve as a checklist for evaluating any future speech-based dataset's fitness for CPS research, allowing comparisons across datasets and categories on a common scale.","The finding that multi-speaker interaction datasets like the AMI Meeting Corpus come closest to CPS conditions suggests that extending such corpora with ambiguity, conflict, and longitudinal structure would be a high-value direction.","If these conclusions hold, benchmarks for collaborative problem solving should not rely on existing SLU test sets without augmentation, since those test sets will not reflect the target task's true difficulty."],"supporting_citations":[{"why":"ATIS is the canonical task-oriented SLU dataset used to characterize the task-oriented dialogue category in the metric evaluations.","marker":"[72]"},{"why":"SNIPS provides multi-domain intent and slot data for the task-oriented dialogue category, contributing to the modality and ambiguity metrics.","marker":"[73]"},{"why":"MultiWOZ is the large multi-domain Wizard-of-Oz dataset whose utterance sizes, abstraction levels, and ambiguity rates feed the semantic parsing metrics.","marker":"[76]"},{"why":"SGD (Schema-Guided Dialogue) supplies domain and dialogue state annotations used to score the problem solving and consensus metrics for task-oriented dialogue.","marker":"[79]"},{"why":"AMI Meeting Corpus is the primary multi-speaker interaction dataset; its speaker turns, interruptions, and discourse annotations shape the social and emotional feature metrics.","marker":"[93]"},{"why":"CoNLL-2003 provides the named entity recognition data that anchors the text understanding category's representation and ambiguity metrics.","marker":"[80]"},{"why":"LibriSpeech is the base ASR dataset whose phoneme-level transcriptions and acoustic features underlie the speech recognition category's simplicity and adaptability scores.","marker":"[82]"},{"why":"Common Voice contributes crowd-sourced speech variety and noise, informing the difficulty-in-representation and trial-and-error metrics for speech recognition.","marker":"[83]"}],"fun_headline_variants":["Speech datasets lack the messiness team AI needs","No existing speech dataset fits collaborative problem solving","Team problem-solving AI needs data that doesn't exist yet","Clean speech data is why team AI fails","Missing: speech data with ambiguity and conflict for team AI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative suitability ratings in Section V, such as 'Expected Accuracy (%) Text: 90-95' and 'Difficulty in Representation (%) Sound: 40-50', are derived from manual inspection of 100-300 utterances from 2-4 datasets per category, with no described sampling procedure, inter-rater reliability check, or raw data release, so the paper's quantitative evidence of inadequacy depends on those ratings being representative and repeatable.","fun_headline_variants_meta":{"raw":{"variants":["Speech datasets lack the messiness team AI needs","No existing speech dataset fits collaborative problem solving","Team problem-solving AI needs data that doesn't exist yet","Clean speech data is why team AI fails","Missing: speech data with ambiguity and conflict for team AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000259,"raw_usage":{"total_tokens":1525,"prompt_tokens":821,"completion_tokens":704,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":631}},"tokens_in":437,"tokens_out":704,"duration_ms":7282,"temperature":1.0,"reasoning_tokens":631,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:41:26.582664+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete audit: take a random sample of at least 1000 utterances from each of the four dataset categories, have multiple independent annotators score the same metrics (ambiguity rate, abstraction levels, reframing types, social interaction scores) using a pre-registered coding manual, and compute inter-rater agreement. If the resulting scores differ substantially from the paper's ranges—for instance, if task-oriented dialogue shows ambiguity rates above 40% or multi-speaker interaction shows low social interaction scores—the paper's conclusion that existing datasets lack CPS-relevant features would be undercut. Additionally, if a model trained on an existing multi-speaker dataset (e.g., AMI Meeting Corpus) achieved human-level performance on a held-out CPS task involving ambiguous, disrupted, and longitudinally tracked team dialogues, that would directly contradict the claim that current data are insufficient.","supporting_citations":[{"cited_title":"T., Godfrey, J","cited_arxiv_id":null,"evidence_quote":"ATIS is the canonical task-oriented SLU dataset used to characterize the task-oriented dialogue category in the metric evaluations."},{"cited_title":"H., Tseng, B","cited_arxiv_id":null,"evidence_quote":"MultiWOZ is the large multi-domain Wizard-of-Oz dataset whose utterance sizes, abstraction levels, and ambiguity rates feed the semantic parsing metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SGD (Schema-Guided Dialogue) supplies domain and dialogue state annotations used to score the problem solving and consensus metrics for task-oriented dialogue."},{"cited_title":"& Wellner, P","cited_arxiv_id":null,"evidence_quote":"AMI Meeting Corpus is the primary multi-speaker interaction dataset; its speaker turns, interruptions, and discourse annotations shape the social and emotional feature metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"LibriSpeech is the base ASR dataset whose phoneme-level transcriptions and acoustic features underlie the speech recognition category's simplicity and adaptability scores."},{"cited_title":"& Weber, F","cited_arxiv_id":null,"evidence_quote":"Common Voice contributes crowd-sourced speech variety and noise, informing the difficulty-in-representation and trial-and-error metrics for speech recognition."}],"review_version":1}