Pith. sign in

REVIEW 4 major objections 5 minor 99 references

An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that existing spoken language understanding datasets are not suitable for training machine learning models to support collaborative problem solving, and it identifies the specific features such datasets would need.

desk verdict A useful requirements list for CPS datasets, undermined by unsupported quantitative ratings; the qualitative claim is plausible but the numbers should not be trusted. read the letter →

arxiv 2412.18489 v1 pith:JEMWYHG5 submitted 2024-12-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords collaborativeproblemsolvingspokenlanguageunderstandingdatasetsuitabilityspeechdatasetsmachinelearningtrainingteamdynamicsmulti-modaldataambiguityandconflict
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that current speech datasets, built for spoken language understanding tasks like intent detection and slot filling, lack the properties needed to train machine learning models for collaborative problem solving in small teams. It analyzes a broad set of popular SLU datasets using a battery of metrics that capture cognitive, social, and emotional dimensions of problem solving. The central finding is that these datasets do not represent the multi-modal, longitudinal, ambiguous, and conflict-laden interactions that characterize real team problem solving. If this is right, then models trained on existing data will not transfer well to collaborative settings, and new dataset collection efforts should prioritize those missing features.

What carries the argument

The central mechanism is a two-part characterization framework: a taxonomy of SLU datasets grouped by purpose (task-oriented dialogue, multi-speaker interaction, text understanding, speech recognition), and a set of CPS-based metrics that operationalize cognitive, social, and emotional activities. Metrics include kinds of modalities and expected accuracy, size and abstraction levels of utterances, presence of ambiguities and unknowns, number of reframing steps, iteration counts for consensus, scores for social interaction and behavioral adaptability, and measures of goal changes, emotion tracking, and response to critique. These metrics are applied to samples of 100-300 utterances from 2-4 datasets per category, yielding quantitative profiles that expose where each dataset category falls short of CPS needs.

What would settle it

A concrete audit: take a random sample of at least 1000 utterances from each of the four dataset categories, have multiple independent annotators score the same metrics (ambiguity rate, abstraction levels, reframing types, social interaction scores) using a pre-registered coding manual, and compute inter-rater agreement. If the resulting scores differ substantially from the paper's ranges—for instance, if task-oriented dialogue shows ambiguity rates above 40% or multi-speaker interaction shows low social interaction scores—the paper's conclusion that existing datasets lack CPS-relevant features would be undercut. Additionally, if a model trained on an existing multi-speaker dataset (e.g., AMI Meeting Corpus) achieved human-level performance on a held-out CPS task involving ambiguous, disrupted, and longitudinally tracked team dialogues, that would directly contradict the claim that current data are insufficient.

Watch

Extended reading notes

Core claim

The paper's central claim is that no existing SLU dataset adequately represents collaborative problem solving as it occurs in teams of about four members talking to each other. The analysis organizes datasets into four categories—task-oriented dialogue, multi-speaker interaction, text understanding, and speech recognition—and scores them on metrics for multi-modal tracking, semantic parsing, solution elaboration, reactivity to unexpected situations, social and emotional feature management, individual-in-team issues, and problem solving process. The conclusion is that the datasets are weakest exactly where CPS is most demanding: they contain little ambiguous or ill-defined speech, few sudden disruptions or conflicts, no longitudinal tracking of team dynamics, and no integrated multi-modal signals beyond speech and text. The paper therefore specifies that new datasets should include multi-modal data capturing diverse team interactions, longitudinal data for tracking dynamics over time, short ambiguous and ill-defined utterances, and situations of sudden disruptions and conflicts.

Load-bearing premise

The quantitative suitability ratings in Section V, such as 'Expected Accuracy (%) Text: 90-95' and 'Difficulty in Representation (%) Sound: 40-50', are derived from manual inspection of 100-300 utterances from 2-4 datasets per category, with no described sampling procedure, inter-rater reliability check, or raw data release, so the paper's quantitative evidence of inadequacy depends on those ratings being representative and repeatable.

Editorial extensions

If this is right

  • If the paper's analysis is correct, anyone training an ML model for collaborative problem solving on existing SLU datasets should expect poor performance in real team settings, because the training data will not contain the ambiguous, conflicting, and dynamic interactions that CPS requires.
  • The proposed list of needed dataset features gives concrete guidance for new data collection efforts: multimodal recordings (not just speech), longitudinal sessions, deliberately ambiguous and ill-defined utterances, and scripted or naturally occurring disruptions and conflicts.
  • The metric battery itself can serve as a checklist for evaluating any future speech-based dataset's fitness for CPS research, allowing comparisons across datasets and categories on a common scale.
  • The finding that multi-speaker interaction datasets like the AMI Meeting Corpus come closest to CPS conditions suggests that extending such corpora with ambiguity, conflict, and longitudinal structure would be a high-value direction.
  • If these conclusions hold, benchmarks for collaborative problem solving should not rely on existing SLU test sets without augmentation, since those test sets will not reflect the target task's true difficulty.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit is that the same metric battery could be applied prospectively to newly collected multimodal team-interaction data, giving dataset builders a pre-hoc suitability score rather than a post-hoc justification.
  • The paper's emphasis on sudden disruptions and conflicts suggests a testable extension: adding scripted 'disruption events' to an existing multi-speaker corpus and measuring whether models trained on that augmented data handle unexpected turns better than models trained on the original corpus.
  • If the quantitative ratings are representative, a practical consequence is that 'general-purpose' SLU pretraining may not transfer to CPS even with fine-tuning, because the missing phenomena (ambiguity, role shifts, long-horizon team dynamics) are not just rare but absent from the pretraining distribution.
  • The reliance on manual inspection of small samples points to a concrete next step: a larger-scale annotation study with multiple raters could quantify inter-rater reliability and produce confidence intervals for each metric, converting the current ordinal scores into statistically grounded estimates.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a set of metrics for characterizing whether existing speech datasets capture the cognitive, social, and emotional activities involved in Collaborative Problem Solving (CPS), and it applies these metrics to four categories of Spoken Language Understanding (SLU) datasets: Task-Oriented Dialogue, Multi-Speaker Interaction, Text Understanding, and Speech Recognition. Based on the resulting quantitative ratings, the paper concludes that current SLU datasets are poorly suited for training ML models to improve CPS, and it lists requirements for future datasets, including multimodal data, longitudinal tracking, ambiguous and ill-defined utterances, and situations involving disruptions and conflicts. The qualitative taxonomy and the enumerated dataset features are the paper's main contributions, while the quantitative suitability ratings in Section V and the appendix are presented as evidence for the insufficiency claim.

Significance. If the quantitative characterization were reproducible, the paper would provide a useful mapping between SLU dataset categories and CPS-relevant constructs, and its requirements list would be a practical starting point for dataset design. The qualitative comparison is genuinely informative: the SLU annotation schemes (intents, slots, dialogue states, transcriptions) do not naturally encode team-level constructs such as Team Agreement, Team Synchronization, conflict resolution, or longitudinal team dynamics, and the paper makes this mismatch visible through a structured taxonomy. A further strength is that the paper states its assumptions explicitly, including the team size of about three or four members and the use of speech as the primary channel, which makes the scope of the claim clear. The quantitative tables, however, are not currently established evidence: they are based on undocumented manual inspection and are internally inconsistent, so the paper's value at present is largely conceptual and agenda-setting rather than an empirical measurement study.

major comments (4)
  1. [Section V, Tables II-IX and XIII-XVI] The quantitative suitability ratings are load-bearing for the paper's central insufficiency claim, but the manuscript reports no sampling protocol, annotator count, coding rubric, or inter-rater reliability for the manual inspection of 100-300 utterances from 2-4 datasets per category. Values such as 'Difficulty in Representation (%) Sound: 40-50' (Table II) and 'Presence and Amount of Ambiguities/Unknowns (% of data affected): 30-40%' (Table III) are presented as precise ranges without any evidence that they are representative or repeatable. I request either a documented methodology for the manual analysis (including how datasets and utterances were sampled, how many annotators participated, and a reliability statistic such as Cohen's kappa) or an explicit reframing of these numbers as the authors' expert estimates, with all inference in the conclusions downgraded accordingly.
  2. [Section V.A, Table II] The metric 'Expected Accuracy (%)' is treated as a property of a dataset category, but accuracy is a property of a model trained and evaluated on a dataset, not of the dataset itself. The reported values such as 'Text: 90-95' and 'Sound: 70-80' for Task-Oriented Dialogue are not accompanied by citations to published benchmark results or by a specification of the model family, train/test split, or evaluation metric. As written, these entries are unverifiable and should either be replaced with documented benchmark results or removed from the dataset-characterization scheme.
  3. [Appendix A, Tables X-XVI vs Section V, Tables II-IX] The appendix tables duplicate the main metric tables but assign different values to the same constructs. For example, Table II reports 'Expected Accuracy (%) Text: 90-95' for Task-Oriented Dialogue, while Table X reports 'MultiWOZ: 85-90' and 'SGD: 70-80' for the same quantity; similarly, 'Difficulty in Representation (%) Text: 10-20' in Table II becomes 'MultiWOZ: 10-20; SGD: 40-50' in Table X. This internal inconsistency means a reader cannot tell which numbers are the authors' final estimates, and it materially undermines the quantitative basis for the conclusion that current SLU datasets are inadequate. The authors should reconcile the two sets of tables or clearly designate one as the reported result.
  4. [Section V.B, ambiguity metric] The 'Presence of Ambiguities' metric is said to be computed automatically by identifying ambiguous words such as 'maybe', 'probably', or 'unsure' in utterances, but no details are given about tokenization, normalization, the exact word list, or which datasets and splits were processed. This operationalization conflates lexical hedges with semantic ambiguity and has no validation against human judgments or downstream ambiguity-related tasks. Since ambiguity plays a central role in the paper's recommendation that future datasets include 'short, ambiguous, and ill-defined speech utterances', this metric needs either a rigorous validation study or a more cautious interpretation as a proxy for one type of uncertainty.
minor comments (5)
  1. [Section V.B] The text contains an unresolved placeholder '[ ?]' in the description of semantic parsing metrics; this should be completed or removed.
  2. [Section V.C and Section V.E] Table references are inconsistent: Section V.C says 'Table XIII summarizes' for the problem-solving metrics that appear as Table IV in the main text, and Section V.E refers to 'Table XIV' for the reactivity metrics that appear as Table VI. The table numbering should be made consistent throughout.
  3. [Table XIII and Table I] The appendix table XIII introduces 'ICSI' as a Multi-Speaker Interaction dataset, but ICSI is neither described in Table I nor defined in the text; every dataset used in the quantitative analysis should be listed and described.
  4. [Section II.E.2] The abbreviation 'TEB' is defined as 'Team Emotional Behavior', but the defining sentence says 'TEM refers to psychological safety'; this appears to be a typo for TEB.
  5. [References [60], [61], [97], [98], [99]] Several definitions of key SLU activities rely on non-archival blog or vendor pages rather than primary literature; for a research paper the definitions should be anchored in peer-reviewed sources or standard textbooks.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: the dataset-suitability ratings are manual assessments, not fitted outputs, and the central insufficiency claim does not reduce to its inputs.

full rationale

The paper is a survey and qualitative/quantitative characterization, not a derivation of predictions from fitted parameters. The central claim that existing SLU datasets are insufficient for CPS is supported by metrics in Section V that were obtained by manually analyzing 100–300 utterances per category. Those ratings are expert judgments with no reported sampling protocol or inter-rater reliability, which is a reproducibility concern rather than a circularity concern. The metrics are not defined in terms of the conclusion, and the conclusion is not fed back into the metric values. There are self-citations in Section I (references [1]–[8]) used to motivate the importance of datasets, but these are not load-bearing for the final analysis. The taxonomy and metric framework are the paper's own constructs, and applying them to datasets does not make the outcome true by construction. No equation or fitted parameter is reused as a prediction, and no uniqueness theorem or prior result by the same authors is invoked to force the conclusion. Therefore no significant circularity is present.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The report relies on domain assumptions about what constitutes CPS and which data modalities matter. It introduces no new physical or mathematical entities. The main epistemic risk is the untested sampling assumption underlying all quantitative ratings.

free parameters (1)
  • CPS metric ratings for SLU categories = Various ranges (e.g., 2-3 abstraction levels, 10-20% ambiguity, expected accuracy 90-95%)
    These values are hand-assigned after informal inspection of 100-300 utterances per dataset. They are not fitted through a reproducible procedure and are treated as factual evidence for the paper's conclusions.
assumptions (4)
  • domain assumption Speech dialogue in small teams is the primary modality for capturing CPS processes.
    The entire report limits datasets to speech recordings of 3-4 member teams, excluding other modalities like gestures, shared artifacts, and physiological signals from the suitability assessment.
  • domain assumption The cognitive, social, and emotional activity taxonomy from Section II is complete and sufficient for characterizing CPS.
    All metrics are derived from this taxonomy. If important CPS activities are missing, the suitability judgment is incomplete.
  • ad hoc to paper Manual inspection of 100-300 utterances from 2-4 datasets per category yields representative estimates for the whole category.
    This sampling assumption is load-bearing for every quantitative table; without it the numbers cannot be generalized.
  • domain assumption SLU tasks are sufficiently similar to CPS that their datasets can be evaluated as proxies.
    The paper argues similarity but does not validate the transfer empirically.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving." pith.science (2026). https://pith.science/paper/JEMWYHG5

@misc{pith2026241218489,
  author       = {Pith},
  title        = {Pith review of: An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JEMWYHG5}},
  note         = {Machine review of arXiv:2412.18489}
}
read the original abstract

This report characterized the suitability of existing datasets for devising new Machine Learning models, decision making methods, and analysis algorithms to improve Collaborative Problem Solving and then enumerated requirements for future datasets to be devised. Problem solving was assumed to be performed in teams of about three, four members, which talked to each other. A dataset consists of the speech recordings of such teams. The characterization methodology was based on metrics that capture cognitive, social, and emotional activities and situations. The report presented the analysis of a large group of datasets developed for Spoken Language Understanding, a research area with some similarity to Collaborative Problem Solving.

Figures

Figures reproduced from arXiv: 2412.18489 by the authors.

Figure 1
Figure 1. SLU Datasets Classification the dataset is designed, such as intent detection, dialogue systems, or speech recognition. • Data Collection: The method used to gather data, such as real-world scenarios, Wizard-of-Oz experiments, or crowdsourced contributions. • Sentence Characteristics: Specific features of the sen￾tences or utterances in the dataset, including length, formality, and content diversity. • Annotations: … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

99 extracted references · 74 canonical work pages

  1. [1]

    Applications of diaLogic System in Individual and Team-based Problem Solving Applications

    R. Duke and A. Doboli, “Applications of diaLogic System in Individual and Team-based Problem Solving Applications”, Proc. IEEE Interna- tional Symposium on Smart Electronic Systems (iSES), 2022

  2. [2]

    Using Speech Data to Automatically Charac- terize Team Effectiveness to Optimize Power Distribution in Internet- of-Things Applications

    G.Villuri and A. Doboli, “Using Speech Data to Automatically Charac- terize Team Effectiveness to Optimize Power Distribution in Internet- of-Things Applications”, Proc. IEEE 3rd Conference on Information Technology and Data Science (CITDS), 2024

  3. [3]

    Studying Consensus and Disagreement during Problem Solving in Teams through Learning and Response Generation Agents Model

    A. Doboli and D. Curiac, “Studying Consensus and Disagreement during Problem Solving in Teams through Learning and Response Generation Agents Model”, Mathematics, MDPI, June 2023

  4. [4]

    A novel agent-based, evolutionary model for expressing the dynamics of creative open-problem solving in small groups

    A., Doboli and S., Doboli, “A novel agent-based, evolutionary model for expressing the dynamics of creative open-problem solving in small groups”, Applied Intelligence, 51, 2094–2127, 2021

  5. [5]

    Modeling Group Creativity as the Evolution of Community-level

    A. Doboli, X. Liu, H. Li and S. Doboli, “Modeling Group Creativity as the Evolution of Community-level”, Creative Problem Solving, In “The Oxford Handbook”, Oxford University Press, 2019

  6. [6]

    Understanding the Significance of Mid-Tier Research Teams in Idea Flow through a Community

    X.Liu, A., Doboli and S., Doboli, “Understanding the Significance of Mid-Tier Research Teams in Idea Flow through a Community”, IEEE Transactions on Computational Social Systems, pp. 1–21, 2022

  7. [7]

    Combining Informetrics and Trend Analysis to Understand Past and Current Directions in Electronic Design Automa- tion

    C. Curiac and A. Doboli, “Combining Informetrics and Trend Analysis to Understand Past and Current Directions in Electronic Design Automa- tion”, Scientometrics, Springer, August 2022, DOI: 10.1007/s11192- 022-04481-9

  8. [8]

    Towards Insightful Automated Dialog for Therapy through Top-down/Bottom-up Response Generation

    A. Doboli, “Towards Insightful Automated Dialog for Therapy through Top-down/Bottom-up Response Generation”, IEEE International Sym- posium on Smart Electronic Systems (iSES), 2022

Show all 99 references
  1. [9]

    Toward an understanding of macrocognition in teams: Pre- dicting processes in complex collaborative contexts

    S. Fiore, M. Rosen, K. Smith-Jentsch, E. Salas, M. Letsky and N. Warner, “Toward an understanding of macrocognition in teams: Pre- dicting processes in complex collaborative contexts”, Human Factors, 52(2), pp. 203–224, 2010

  2. [10]

    Towards a generalized competency model of collaborative problem solving

    C. Sun, V .J. Shute, A. Stewart, J. Yonehiro, N. Duran, and S. D’Mello, “Towards a generalized competency model of collaborative problem solving”, Computers & Education, 143:103672, 2020

  3. [11]

    How the group affects the mind: A cognitive model of idea generation in groups

    B. Nijstad and W. Stroebe, “How the group affects the mind: A cognitive model of idea generation in groups”, Personality and Social Psychology Review, 10(3), pp. 186–213, 2006

  4. [12]

    Understanding team learning dynamics over time

    C. Wiese and S. Burke, “Understanding team learning dynamics over time”, Frontiers in Psychology, 10:1417, 2019

  5. [13]

    Problem-solving phase transitions during team collaboration

    T. Wiltshire, J. Butner, and S. Fiore, “Problem-solving phase transitions during team collaboration”, Cognitive science, 42(1), pp. 129–167, 2018

  6. [14]

    Team learning: Collectively connecting the dots

    A. Ellis, J. Hollenbeck, D. Ilgen, C. Porter, B. West and H. Moon, “Team learning: Collectively connecting the dots”, Journal of applied Psychology, 88(5):821, 2003

  7. [15]

    Cognitive processes in well- defined and ill-defined problem solving

    G. Schraw, M. Dunkle and L. Bendixen, “Cognitive processes in well- defined and ill-defined problem solving”, Applied Cognitive Psychology, 9(6), pp. 523–538, 1995

  8. [16]

    Modeling semantic knowledge structures for creative problem solving: Studies on expressing concepts, categories, associations, goals and context

    A. Doboli, A. Umbarkar, S. Doboli and J. Betz, “Modeling semantic knowledge structures for creative problem solving: Studies on expressing concepts, categories, associations, goals and context”, Knowledge-based Systems, 78, pp. 34–50, 2015

  9. [17]

    The role of precedents in increasing creativity during iterative design of electronic embedded systems

    A. Doboli and A. Umbarkar, “The role of precedents in increasing creativity during iterative design of electronic embedded systems”, Design Studies, 35(3), pp. 298–326, 2014

  10. [18]

    Not too much, not too little: The influence of constraints on creative problem solving

    K. Medeiros, P. Partlow and M. Mumford, “Not too much, not too little: The influence of constraints on creative problem solving”, Psychology of Aesthetics, Creativity, and the Arts, 8(2), pp. 198–210, 2014

  11. [19]

    Temporal construal effects on abstract and concrete thinking: Consequences for insight and creative cognition

    J. Forster, R. Friedman and N. Liberman, “Temporal construal effects on abstract and concrete thinking: Consequences for insight and creative cognition”, Journal of Personality and Social Psychology, 87(2), pp. 177–189, 2004

  12. [20]

    Psychological safety and learning behavior in work teams

    A. Edmondson, “Psychological safety and learning behavior in work teams”, Administrative Science Quarterly, 44(2), pp. 350–383, 1999

  13. [21]

    What do you mean by collaborative learning?

    P. Dillenbourg, “What do you mean by collaborative learning?”, In P. Dillenbourg, “Collaborative learning: Cognitive and Computational Approaches”, Oxford: Elsevier, 1999, pp.1–19

  14. [22]

    The use of environmental clues during incubation

    R. A. Dodds, S. M. Smith and T. B. Ward, “The use of environmental clues during incubation”, Creativity Research Journal, 14, pp. 287–304, 2002

  15. [23]

    Climates and cultures for innovation and creativity at work

    M. West and A. Richter, “Climates and cultures for innovation and creativity at work”, In C. E. J.Zhou, J. Shalley, editor, “Handbook of organizational creativity”, pp. 211–236, New York: Taylor Francis Group., 2008

  16. [24]

    Problem Framing Activities Carried Out by Student Design Teams to Enhance Creativity: Comparative Analysis of High and Low Creative Teams

    K. Suk and H. Lee, “Problem Framing Activities Carried Out by Student Design Teams to Enhance Creativity: Comparative Analysis of High and Low Creative Teams”, Arch. Des. Res., 34, pp. 23–38, 2021

  17. [25]

    All Frames Are Not Created Equal: A Typology and Critical Analysis of Framing Effects

    I. Levin, S. Schneider and G. Gaeth, “All Frames Are Not Created Equal: A Typology and Critical Analysis of Framing Effects”, Organizational Behavior and Human Decision Processes, V olume 76, Issue 2, 1998, pp. 149–188

  18. [26]

    The framing effect and risky decisions: Examining cognitive functions with fMRI

    C. Gonzalez, J. Dana, H. Koshino and M. Just, “The framing effect and risky decisions: Examining cognitive functions with fMRI”, Journal of Economic Psychology, V ol. 26, Issue 1, 2005, pp. 1–20

  19. [27]

    Problem frame patterns: an exploration of patterns in the problem space

    R. Wirfs-Brock, P. Taylor and J. Noble, James, “Problem frame patterns: an exploration of patterns in the problem space”, Proc. Conference on Pattern Languages of Programs, 2006

  20. [28]

    Two Minds, One Dialog: Coordinating Speaking and Understanding

    S. Brennan, A. Galati and A. Kuhlen, “Two Minds, One Dialog: Coordinating Speaking and Understanding”, In The Psychology of Learning and Motivation: Advances in Research and Theory, B. Ross (Ed.), V ol. 53: Psychology of Learning and Motivation, Academic Press: Cambridge, 2010...

  21. [29]

    Systematic Methodology for Design- ing Reconfigurable Delta Sigma Modulator Topologies for Multimode Communication Systems

    Y . Wei, H. Tang and A. Doboli, “Systematic Methodology for Design- ing Reconfigurable Delta Sigma Modulator Topologies for Multimode Communication Systems”, IEEE Transactions on CADICS, V ol. 26, No. 3, pp. 480–496, 2007

  22. [30]

    High-Level Synthesis of Delta-Sigma Modu- lators Optimized for Complexity, Sensitivity and Power Consumption

    H. Tang and A. Doboli, “High-Level Synthesis of Delta-Sigma Modu- lators Optimized for Complexity, Sensitivity and Power Consumption”, IEEE Transactions on CADICS, V ol. 25, No. 3, pp. 597–607, 2006

  23. [31]

    Learning and Memory. An Integrated Approach

    J. Anderson, “Learning and Memory. An Integrated Approach”, Wiley, 2000

  24. [32]

    Pragmatics in Analogical Mapping

    B. Spellman and K. Holyoak, “Pragmatics in Analogical Mapping”, Cognitive Psychology, V ol. 31, Issue 3, 1996, pp. 307–346

  25. [33]

    Design, analogy, and creativity

    A. Goel, “Design, analogy, and creativity”, IEEE Expert, vol. 12, no. 3, pp. 62-70, May-June 1997

  26. [34]

    Cap- turing scientists’ insight from dddas

    P. Reynolds, D. Brogan, J. Carnahan, Y . Loitiere and M. Spiegel, “Cap- turing scientists’ insight from dddas”, Proc. International Conference on Computational Science (ICCS) - V olume Part III, pp. 570–577, 2006

  27. [35]

    How scientists think: On-line creativity and conceptual change in science. Conceptual Structures and Processes: Emergence, discovery, and change

    K. Dunbar, “How scientists think: On-line creativity and conceptual change in science. Conceptual Structures and Processes: Emergence, discovery, and change”, in T. Ward, S. Smith, and J. Vaid, eds., American Psychological Association Press, 1997

  28. [36]

    Creative foraging: An experimental paradigm for study- ing exploration and discovery

    Y . Hart, A. Mayo, R. Mayo, L. Rozenkrantz, A. Tendler, U. Alon and et al, “Creative foraging: An experimental paradigm for study- ing exploration and discovery”, PLoS ONE 12(8): e0182133. https:// doi.org/10.1371/journal.pone.0182133, 2017

  29. [37]

    The design of divide and conquer algorithms

    D. Smith, “The design of divide and conquer algorithms”, Science of Computer Programming, 5, pp. 37–58, 1985

  30. [38]

    Evocation and elaboration of solutions: Different types of problem-solving actions. An empirical study on the design of an aerospace artifact

    W. Visser, “Evocation and elaboration of solutions: Different types of problem-solving actions. An empirical study on the design of an aerospace artifact”, in T. Kohonen & F. Fogelman-Souli ´e (Eds.), “At the crossroads of Artificial Intelligence, Cognitive science, and Neuro-...

  31. [39]

    Designing web sites: opportunistic actions and cognitive effort of lay-designers

    N. Bonnardel, L. Lanzone, and S. Sumner, “Designing web sites: opportunistic actions and cognitive effort of lay-designers” Cognitive Science Quarterly, 3(1), pp. 25–56, 2003

  32. [40]

    Efficient creativity: Constraint-guided con- ceptual combination

    F. Costello and M. Keane, “Efficient creativity: Constraint-guided con- ceptual combination”, Cognitive Science, 24(2), pp. 299–349, 2000

  33. [41]

    Relations versus properties in concept combination. Journal of Memory and Language

    E. J. Wisniewski and B. C. Love, “Relations versus properties in concept combination. Journal of Memory and Language”, 38, pp. 177–202, 1998

  34. [42]

    Effects of problem scope and creativity instructions on idea generation and selection

    E. F. Rietzschel, B. A. Nijstad and W. Stroebe, “Effects of problem scope and creativity instructions on idea generation and selection”, Creativity Research Journal, 26, pp. 185–191, 2014

  35. [43]

    Methodologies for examining problem solving success and failure

    M. DeCaro, M. Wieth and S. Beilock, “Methodologies for examining problem solving success and failure”, Methods, V olume 42, Issue 1, 2007, pp. 58–67

  36. [44]

    Critical Thinking Assessment in Engineer- ing Education: A Scopus-Based Literature Review

    S. Deo and K. Holtta-Otto, “Critical Thinking Assessment in Engineer- ing Education: A Scopus-Based Literature Review”, ASME. J. Mech. Des., July 2024, 146(7): 072301

  37. [45]

    MCD: A Model-Agnostic Counterfactual Search Method For Multi-modal Design Modifications

    L. Regenwetter, Y . Obaideh and F. Ahmed, “MCD: A Model-Agnostic Counterfactual Search Method For Multi-modal Design Modifications”, arXiv, 2305.11308, 2024, https://arxiv.org/abs/2305.11308

  38. [46]

    VisiFit: Structuring Iterative Improvement for Novice Designers

    L. Chilton, E. Ozmen, S. Ross and V . Liu, “VisiFit: Structuring Iterative Improvement for Novice Designers”, Proc. CHI Conference on Human Factors in Computing Systems, 2021

  39. [47]

    Mental fixation and metacognitive predic- tions of insight in creative problem solving

    B. Storm and M. Hickman, “Mental fixation and metacognitive predic- tions of insight in creative problem solving”, The Quarterly Journal of Experimental Psychology, 68:4, pp. 802–813, 2015

  40. [48]

    Effects of task instructions and brief breaks on brainstorming

    P. Paulus, T. Nakui, V . L. Putman and V . R. Brown, “Effects of task instructions and brief breaks on brainstorming”, Group Dynamics Theory Research and Practice, 10(3), pp. 206–219, 2006

  41. [49]

    Categorization and representation of physics problems by experts and novices

    M. Chi, P. Feltovich and R. Glaser, “Categorization and representation of physics problems by experts and novices”, Cognitive Science, 3, pp. 121–152, 1981

  42. [50]

    Social Neuroscience: People Thinking about Thinking People

    J. Cacioppo, P. Visser and C. Pickett (Eds.), “Social Neuroscience: People Thinking about Thinking People”, MIT Press,2006

  43. [51]

    Conflict across representational gaps: Threats to and opportunities for improved communication

    M. A. Cronin and L. R. Weingart, “Conflict across representational gaps: Threats to and opportunities for improved communication”, Proceedings of the National Academy of Sciences, 116(16), pp. 7642–7649, 2019

  44. [52]

    Joint Action: Mental Representations, Shared Information and General Mechanisms for Coordinating with Others

    C, Vesper, E. Sangati, J. Butepage, F. Ciardo, B. Crossey, A. Effen- berg, D. Hristova, A. Karlinsky, L. McEllin, S. Nijssen and et al., “Joint Action: Mental Representations, Shared Information and General Mechanisms for Coordinating with Others”, Front. Psychol., 2016, 7, pp. 2039

  45. [53]

    Social yet creative: The role of social relationships in facilitating individual creativity

    J. Perry-Smith, “Social yet creative: The role of social relationships in facilitating individual creativity”, The Academy of Management Journal, 49(1), pp. 85–101, 2006

  46. [54]

    Making group brainstorming more effective: Recommendations from an associative memory perspective

    V . Brown and P. Paulus, “Making group brainstorming more effective: Recommendations from an associative memory perspective”, Current Directions in Psychological Science, 11, pp. 208–212, 2002

  47. [55]

    Psychological safety, trust, and learning in organizations: A group-level lens. Trust and distrust in organizations: Dilemmas and approaches

    A. Edmondson, R. Kramer and K. Cook, “Psychological safety, trust, and learning in organizations: A group-level lens. Trust and distrust in organizations: Dilemmas and approaches”, 12, pp. 239–272, 2004

  48. [56]

    Recognizing devel- opers’ emotions while programming

    D. Girardi, N. Novielli, D. Fucci and F. Lanubile, “Recognizing devel- opers’ emotions while programming”, Proc. ACM/IEEE International Conference on Software Engineering, 2020, pp. 666–677

  49. [57]

    Negative affective environments improve com- plex solving performance

    C. Barth and J. Funke, “Negative affective environments improve com- plex solving performance”, Cognition and Emotion, 24(7), pp. 1259– 1268, 2010

  50. [58]

    Exploring Causes of Frustration for Software Developers

    D. Ford and C. Parnin, “Exploring Causes of Frustration for Software Developers”, Proc. IEEE/ACM International Workshop on Cooperative and Human Aspects of Software Engineering, 2015, pp. 115–116

  51. [59]

    Quinlan, T. (2004). Speech recognition technology and students with writing difficulties: Improving fluency. Journal of Educational Psychol- ogy, 96(2), 337–346

  52. [60]

    TAUS. (2021). Domain Classification with Natural Language Pro- cessing. Retrieved from https://www.taus.net/resources/blog/domain- classification-with-natural-language-processing

  53. [61]

    GeeksforGeeks. (2024). Intent Recognition using TensorFlow. Re- trieved from https://www.geeksforgeeks.org/intent-recognition-using- tensorflow

  54. [63]

    Veyseh, A. P. B., Dernoncourt, F., & Nguyen, T. H. (2020). Improving Slot Filling by Utilizing Contextual Information. In Proceedings of the 2nd Workshop on Natural Language Processing for Conversational AI, pages 90–95. Association for Computational Linguistics

  55. [64]

    E, H., et al. (2020). Efficient Context and Schema Fusion Networks for Multi-Domain Dialogue State Tracking. In Findings of the Association for Computational Linguistics: EMNLP 2020

  56. [65]

    Dehghan, M., et al. (2024). EWEK-QA : Enhanced Web and Efficient Knowledge Graph Retrieval for Citation-based Question Answering Systems. In Proceedings of the 62nd Annual Meeting of the Associ- ation for Computational Linguistics (V olume 1: Long Papers), pages 14169–14187, B...

  57. [66]

    Yu, J., Bohnet, B., & Poesio, M. (2020). Named Entity Recognition as Dependency Parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 6470–6476, Online. Association for Computational Linguistics

  58. [67]

    Henderson, M., Thomson, B., & Young, S. (2014). Word-Based Dialog State Tracking with Recurrent Neural Networks. In Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), pages 292–299, Philadelphia, PA, USA

  59. [68]

    Multimodal expressive embodied conversa- tional agents

    C. Pelachaud, Catherine, “Multimodal expressive embodied conversa- tional agents”, Proc. ACM International Conference on Multimedia (MULTIMEDIA ’05), 2005, pp. 683–689

  60. [69]

    Zooming on Multimodality and Attuning. A Multilayer Model for the Analysis of the V ocal Act in Conversational In- teractions

    R.Ciceri and F. Biassoni, “Zooming on Multimodality and Attuning. A Multilayer Model for the Analysis of the V ocal Act in Conversational In- teractions” In G. Riva, M.T. Anguera, B.K. Wiederhold and F. Mantovani (Eds.), “From Communication to Presence: Cognition, Emotions and...

  61. [70]

    A Stacked Multi- Layered Perceptron - LLM Model for Extracting the Relations in Textual Descriptions

    G. Villuri, H. Shaik, S. Doboli and A. Doboli, “A Stacked Multi- Layered Perceptron - LLM Model for Extracting the Relations in Textual Descriptions”, Proc. IEEE Symposium on Computational Intelligence in Natural Language Processing and Social Media Companion, 2025

  62. [71]

    Towards Semantic Classification: An Experimental Study on Automated Understanding of the Meaning of Verbal Utterances

    G. Villuri, H. Pallapu, S. Doboli and A. Doboli, “Towards Semantic Classification: An Experimental Study on Automated Understanding of the Meaning of Verbal Utterances”, Proc. IEEE CCWC, 2025

  63. [72]

    T., Godfrey, J

    Hemphill, C. T., Godfrey, J. J., & Doddington, G. R. (1990). The ATIS spoken language systems pilot corpus. In Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990

  64. [73]

    & Dureau, J

    Coucke, A., Saade, A., Ball, A., Bluche, T., Caulier, A., Leroy, D., ... & Dureau, J. (2018). Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces. arXiv preprint arXiv:1805.10190

  65. [74]

    & Suleman, K

    El Asri, L., Schulz, H., Sharma, S., Zumer, J., Harris, J., Fine, E., ... & Suleman, K. (2017). Frames: a corpus for adding memory to goal- oriented dialogue systems. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue

  66. [75]

    Henderson, M., Thomson, B., & Williams, J. D. (2014). The second dialog state tracking challenge. In Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL)

  67. [76]

    H., Tseng, B

    Budzianowski, P., Wen, T. H., Tseng, B. H., Casanueva, I., Ultes, S., Ramadan, O., & Ga ˇsi´c, M. (2018). MultiWOZ-A Large-Scale Multi- Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Langu...

  68. [77]

    S., Hoi, S

    Wu, C. S., Hoi, S. C., Socher, R., & Xiong, C. (2020). TOD-BERT: Pre- trained Natural Language Understanding for Task-Oriented Dialogue. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)

  69. [78]

    [Online]

    Google Cloud, ”Dialogflow API,” Google Cloud Documentation, 2024. [Online]. Available: https://cloud.google.com/dialogflow/es/docs/reference/rest/v2-overview

  70. [79]

    Rastogi, A., Zang, X., Sunkara, S., Gupta, R., & Khaitan, P. (2020). Towards scalable multi-domain conversational agents: The schema- guided dialogue dataset. In Proceedings of the AAAI Conference on Artificial Intelligence

  71. [80]

    F., & De Meulder, F

    Sang, E. F., & De Meulder, F. (2003). Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050

  72. [81]

    & Xue, N

    Weischedel, R., Palmer, M., Marcus, M., Hovy, E., Pradhan, S., Ramshaw, L., ... & Xue, N. (2013). OntoNotes Release 5.0 LDC2013T19. Linguistic Data Consortium, Philadelphia, PA

  73. [82]

    Panayotov, V ., Chen, G., Povey, D., & Khudanpur, S. (2015). Lib- rispeech: an ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

  74. [83]

    & Weber, F

    Ardila, R., Branson, M., Davis, K., Kohler, M., Meyer, J., Henretty, M., ... & Weber, F. (2020). Common voice: A massively-multilingual speech corpus. In Proceedings of the 12th Language Resources and Evaluation Conference

  75. [84]

    H., Wu, S

    Lee, C. H., Wu, S. L., Liu, C. L., & Lee, H. Y . (2018). Spoken SQuAD: A Study of Mitigating the Impact of Speech Recognition Errors on Listening Comprehension. In Interspeech

  76. [85]

    & Toutanova, K

    Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., ... & Toutanova, K. (2019). Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7, 453-466

  77. [86]

    S., & Bengio, Y

    Lugosch, L., Ravanelli, M., Ignoto, P., Tomar, V . S., & Bengio, Y . (2019). Speech model pre-training for end-to-end spoken language understanding. arXiv preprint arXiv:1904.03670

  78. [87]

    S., & Zisserman, A

    Nagrani, A., Chung, J. S., & Zisserman, A. (2017). V oxceleb: a large- scale speaker identification dataset. arXiv preprint arXiv:1706.08612

  79. [88]

    Kahn, J., Rivi `ere, M., Zheng, W., Kharitonov, E., Xu, Q., Mazar ´e, P. E., ... & Dupoux, E. (2020). Libri-light: A benchmark for ASR with limited or no supervision. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

  80. [89]

    Bastianelli, E., Vanzo, A., Swietojanski, P., & Rieser, V . (2020). SLURP: A Spoken Language Understanding Resource Package. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)

  81. [90]

    Warden, P. (2018). Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209

  82. [91]

    & Dupoux, E

    Wang, C., Rivi `ere, M., Lee, A., Wu, A., Talnikar, C., Haziza, D., ... & Dupoux, E. (2021). V oxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. In Proceedings of the 59th Annual Meeting of the Associat...

  83. [92]

    Rousseau, A., Del ´eglise, P., & Est `eve, Y . (2014). Enhancing the TED- LIUM corpus with selected data for language modeling and more TED talks. In LREC

  84. [93]

    & Wellner, P

    Carletta, J., Ashby, S., Bourban, S., Flynn, M., Guillemot, M., Hain, T., ... & Wellner, P. (2005). The AMI meeting corpus: A pre-announcement. In International workshop on machine learning for multimodal interac- tion

  85. [94]

    Tur, G., & De Mori, R. (2011). Spoken language understanding: Systems for extracting semantic information from speech. John Wiley & Sons

  86. [95]

    McTear, M. (2016). The Dialogue Manager: Coordinating the Dialogue. In Spoken Dialogue Systems (pp. 89-117). Springer, Cham

  87. [96]

    Xing, B., Liao, L., Huang, M., & Tsang, I. (2024). DC-Instruct: An Effective Framework for Generative Multi-intent Spoken Language Understanding. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 14520-14534). Associa- tion for Comp...

  88. [97]

    AltexSoft. (2023). What Is Named Entity Recognition (NER) and How It Works? AltexSoft Blog

  89. [98]

    Transkriptor. (2023). How Does V oice-to-Text Work? Transkriptor Blog

  90. [99]

    AltexSoft. (2023). Quality Assurance (QA), Quality Control and Testing - The Basics of Software Quality Management. AltexSoft Whitepaper

  91. [100]

    Shu, L., Xu, H., Liu, B., & Molino, P. (2019). Modeling Multi- Action Policy for Task-Oriented Dialogues. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.