REVIEW 3 major objections 6 minor 41 references
An Exploration of Internal States in Collaborative Problem Solving
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Retrospective video-recall monologues reveal a twelve-label map of internal states during collaborative problem solving.
desk verdict Honest, useful exploratory dataset, but the 61.4% positive-prevalence headline is not yet supported: the label mapping is post hoc, overlapping, and unreported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the stimulated-recall protocol: each participant, alone after the task, watches a video of their team and narrates their thoughts and feelings moment-to-moment, with the video allowing those states to be mapped back to task time. On top of that transcript, the analysis builds a small emotion taxonomy by hand: 29 frequent keywords and phrases are assigned to twelve labels, and the labels are validated by within-category and between-category cosine similarities computed from pretrained transformer word embeddings. The taxonomy is what turns free-form monologue into a quantitative distribution of internal states.
What would settle it
Collect concurrent in-task measures of emotion, such as physiological arousal or button-press emotion ratings, during the collaborative task and compare them with the retrospective label distribution from the video-recall monologues; if the two show little or no correspondence, the claim that retrospective speech reveals actual internal states would be falsified.
Extended reading notes
Core claim
The central claim is that internal states during collaborative problem solving leave reliable traces in retrospective speech, and that a hand-derived set of twelve emotion labels—Engaged, Disengaged, Conflicted, Confident, Reserved, Frustrated, Optimistic, Anxious, Disappointed, Satisfied, Confused, Surprised—captures the distribution of those states. The authors show that frequent n-grams are mostly filler, so they manually select 29 content-bearing keywords and phrases, assign them to labels using discrete-emotion literature, and then use cosine similarity of word embeddings to show within-label words are semantically close and labels are reasonably distinct. The resulting frequency distribution is skewed positive: engagement, optimism, and satisfaction dominate, while frustration, disengagement, and reservedness are rare.
Load-bearing premise
The entire dataset is retrospective speech, so the paper must assume that what participants say while watching the video accurately reconstructs what they actually felt during the task; if memory or social desirability distorts those recollections, the emotion-label frequencies describe the retelling rather than the experience.
Editorial extensions
If this is right
- Future collaborative problem solving studies can use prompted video recall to collect internal-state data at scale without interrupting collaboration.
- The twelve labels and their associated keywords give later researchers a starting vocabulary for automatic emotion annotation.
- The distribution suggests positive states dominate in this kind of cooperative construction task, so interventions aimed at reducing frustration may target a minority of experience.
- Because confusion is relatively frequent and known to accompany learning, its presence can be treated as a sign of engagement rather than failure.
- The within-category and between-category semantic similarity method offers a template for checking that hand-built emotion categories are coherent.
Reading between the lines
- Inference: if the retrospective protocol is validated against concurrent measures, the same method could be extended to role-level analysis, since directors and builders may show different emotion profiles.
- Inference: the label set is likely task- and context-specific; the Lego construction task may induce more positive states than competitive or high-stakes collaboration, so the 61% positive figure should not be generalized without replication.
- Inference: the semantic-similarity validation could be turned into a testable prediction that automatic classifiers trained on the twelve labels should predict task-relevant events, such as a builder making an error or a director giving confusing instructions, better than chance.
- Inference: one could test whether the act of narrating changes the experience by comparing first-phase monologues with second-phase task behavior, a check for reactivity of the method.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an exploratory linguistic analysis of retrospective self-reports collected from 29 participants immediately after a four-person Lego collaborative problem-solving task. Participants watched a video replay of the task and narrated their moment-to-moment internal states; the authors transcribed these monologues, applied frequency analyses of unigrams, bigrams, and trigrams, hand-selected 29 keywords, grouped them into twelve emotion labels, and report that "Optimistic," "Engaged," and "Satisfied" account for 333 (61.4%) of emotion-label word occurrences. A BERT-based semantic-similarity analysis (within- and between-category cosine similarities) is presented as validation that the labels are coherent and distinct.
Significance. The study addresses a real gap: internal states during CPS are usually inferred from observable behavior, and retrospective replay is a relatively underexplored method for accessing individual experience. The paper is commendable for making its full keyword tables and label mappings visible, for transparently reporting demographic and exclusion information, and for acknowledging the retrospective-recall limitation in Section 6. If the label-based prevalence claims were supported by a reproducible matching rule and independent validation, the dataset would be a useful descriptive baseline for future CPS emotion research. As it stands, the exploratory n-gram findings are credible, but the headline distributional claim is not yet supported.
major comments (3)
- [Section 4.3, Table 8, Figure 2] The 333-occurrence/61.4% figure is not well-defined because Table 8 lists the same surface keywords under multiple labels without a stated counting rule. For instance, "figure out" appears under Engaged and Optimistic; "easier"/"easi" under Optimistic and Satisfied; "vagu" under Conflicted and Confused; "assum" under Conflicted and Reserved; "mistak" under Anxious and Disappointed; and "crazi" under Frustrated and Anxious. If each occurrence is counted once per label, positive labels with duplicated keywords are systematically inflated; if occurrences are assigned to one label only, the adjudication rule is missing. Table 8 also contains words not in Table 7 (e.g., "zoning out," "conflict," "sure," "stress," "worri," "fault," "lost," "surpris"), so the source of the occurrence counts in Figure 2 is unclear. Please report a complete, single-label mapping from each keyword/phrase occurrence to one label, or a precise multi-label counting rule, and provide a sensitivity analysis of the 61.4% figure under alternative assignments.
- [Section 4.4, Table 9] The semantic-similarity validation is circular. The labels and their member words were constructed post hoc from the same keyword-frequency data by the same authors, so high within-category BERT cosine similarity is expected and does not independently confirm that the categories are coherent. No random baseline, shuffled-label comparison, or inter-rater reliability is reported. Please add a baseline such as average similarity of random word sets of matched size, a second coder re-applying the codebook to transcripts, or a hold-out coding exercise, so that Table 9 can be interpreted as evidence for label coherence.
- [Sections 3.1 and 6] The entire dataset consists of retrospective verbal reports produced during video replay, and the paper itself states in Section 6 that "participants may forget details or not accurately report their emotional experiences after the task is completed." Because the paper's central claims concern internal states during the task, the absence of any concurrent or independent check on the fidelity of recall leaves the construct validity of the frequency tables open. Either provide evidence for the fidelity of the replay procedure (e.g., alignment with known task events, or a subset with in-task measures), or revise the wording throughout Sections 4 and 5 so that the results are explicitly framed as characterizations of retrospective accounts rather than directly measured in-task states.
minor comments (6)
- [Section 3.3] "When mapping, we accounted for the the context used by participants" contains a duplicated "the."
- [Section 3.3] The frequency cutoffs for unigrams (>=30), bigrams (>=6), and key terms (>=5) are introduced without a rationale; a sentence on why these thresholds were chosen would help.
- [Section 3.1, Table 1] The text says the age range was 20 to 32, while the table groups ages as 18-24, 25-31, and 32+; please reconcile these descriptions.
- [Section 4.1] The trigram analysis reports phrases with frequencies between 3 and 5, but no trigram cutoff is specified in Section 3.3; please state the criterion used.
- [Section 4.4, Figure 4] The discussion of highest and lowest between-category similarities would be easier to check if the specific numeric values or a full similarity matrix were provided in the text.
- [General] The paper does not include a data or code availability statement; if the de-identified transcripts can be shared, adding such a statement would substantially improve reproducibility and would directly help reviewers evaluate the label-mapping concerns.
Circularity Check
The 61.4% positive-prevalence claim and the within-category semantic-similarity 'validation' reduce to the same hand-built keyword-to-label mapping used to define the labels, making the central quantitative result circular.
-
fitted input called prediction
[Section 3.3 (Data Analysis), Section 4.3 (Emotion Labels), Table 8]
"We developed a set of emotion labels based on the distribution of frequent key words and phrases, drawing on work on discrete emotions and their corresponding emotional expressions [12]. Words and phrases were then mapped to these labels."
The reported prevalence in Section 4.3 — "Optimistic", "Engaged", and "Satisfied" labels were the most prevalent, collectively accounting for 333 (61.4%) of the total word occurrences — is the sum of frequencies of the keywords assigned to each label in Table 8. Because the labels were constructed from that same keyword-frequency distribution in Section 3.3, the prevalence is the input frequency relabeled, not an independent measurement. The mapping is non-exclusive: "figure out" appears under Engaged and Optimistic, "easier"/"easi" under Optimistic and Satisfied, "vagu" under Conflicted and Confused, "assum" under Conflicted and Reserved, and "mistak" under Anxious and Disappointed.
-
self definitional
[Section 4.4 (Semantic Similarity), Table 9]
"The within-category represents average semantic similarity between word-pairs that we associate with the same label."
This is presented as a validation: "We used semantic similarity to validate that the words within each label were appropriately similar." But the labels in Table 8 were created by the authors choosing keywords that are semantically related (e.g., "frustrat", "stupid", "stress" for Frustrated). Computing BERT cosine similarity within each such hand-grouped set and reporting high scores is therefore expected by construction; no random baseline, inter-rater reliability, or held-out coding is provided. The high within-category similarity confirms only that the labels were defined using similar words, not that the labels independently capture the participants' internal states.
full rationale
The paper's raw n-gram and keyword-frequency analyses are self-contained and not circular: Table 5, Table 6, and Table 7 report counts directly from transcripts. The circularity enters when these same frequencies are used to construct the twelve emotion labels (Section 3.3), after which Section 4.3 presents the label distribution (including the headline 61.4% positive prevalence) as a finding about internal states. Since each label is defined by the very keywords whose frequencies are then counted, the distribution is forced by the construction. The Section 4.4 semantic-similarity check is similarly self-confirming because it measures average similarity within word sets that were grouped on the basis of semantic relatedness; there is no independent validation that a second coder would assign the same labels or that the overlapping keywords would resolve to the same categories. Self-citations to the authors' prior multimodal-CPS work appear only in related work and are not load-bearing. The acknowledged retrospective-recall limitation is a construct-validity concern rather than a circularity concern. Overall, the exploratory n-gram observations remain credible, but the central quantitative prevalence claim reduces in part to the hand-built mapping, so a score of 6 is appropriate.
Assumptions & free parameters
free parameters (3)
- Frequency cutoffs (unigram >=30, bigram >=6, keyword >=5) =
30/6/5
- Emotion label-to-keyword mapping (Table 8) =
12 labels, 69 keyword entries
- Pretrained BERT (bert-base-uncased) embeddings =
768-dimensional vectors
assumptions (4)
- domain assumption Retrospective stimulated recall accurately reproduces in-task internal states
- domain assumption Frequency of emotion words reflects prevalence of internal states
- domain assumption Cosine similarity between BERT embeddings is a valid measure of semantic relatedness for emotion categories
- domain assumption The first phase of the task is representative of the second phase
Cite this review
Pith. "Pith review of An Exploration of Internal States in Collaborative Problem Solving." pith.science (2026). https://pith.science/paper/NSLUGSQG
@misc{pith2026250702229,
author = {Pith},
title = {Pith review of: An Exploration of Internal States in Collaborative Problem Solving},
year = {2026},
howpublished = {\url{https://pith.science/paper/NSLUGSQG}},
note = {Machine review of arXiv:2507.02229}
}
read the original abstract
Collaborative problem solving (CPS) is a complex cognitive, social, and emotional process that is increasingly prevalent in educational and professional settings. This study investigates the emotional states of individuals during CPS using a mixed-methods approach. Teams of four first completed a novel CPS task. Immediately after, each individual was placed in an isolated room where they reviewed the video of their group performing the task and self-reported their internal experiences throughout the task. We performed a linguistic analysis of these internal monologues, providing insights into the range of emotions individuals experience during CPS. Our analysis showed distinct patterns in language use, including characteristic unigrams and bigrams, key words and phrases, emotion labels, and semantic similarity between emotion-related words.
Figures
Reference graph
Works this paper leans on
-
[1]
Educational psychologist 50(1), 84–94 (2015)
Azevedo, R.: Defining and measuring engagement and learning in science: Concep- tual, theoretical, methodological, and analytical issues. Educational psychologist 50(1), 84–94 (2015)
work page 2015
-
[2]
Science329(5995), 1081–1085 (2010)
Bahrami, B., Olsen, K., Latham, P.E., Roepstorff, A., Rees, G., Frith, C.D.: Opti- mally interacting minds. Science329(5995), 1081–1085 (2010)
work page 2010
-
[3]
Cambridge University Press (2011)
Bakeman, R., Quera, V.: Sequential analysis and observational methods for the behavioral sciences. Cambridge University Press (2011)
work page 2011
-
[4]
In: Pro- ceedings of the sixth international conference on learning analytics & knowledge
Beheshitha, S.S., Hatala, M., Gašević, D., Joksimović, S.: The role of achievement goal orientations when studying effect of learning analytics visualizations. In: Pro- ceedings of the sixth international conference on learning analytics & knowledge. pp. 54–63 (2016)
work page 2016
-
[5]
In: International Conference on Artificial Intelligence in Education
Bradford, M., Khebour, I., Blanchard, N., Krishnaswamy, N.: Automatic detection of collaborative states in small groups using multimodal features. In: International Conference on Artificial Intelligence in Education. pp. 767–773. Springer (2023)
2023
-
[6]
Computational linguistics32(1), 13–47 (2006)
Budanitsky, A., Hirst, G.: Evaluating wordnet-based measures of lexical semantic relatedness. Computational linguistics32(1), 13–47 (2006)
work page 2006
-
[7]
Dillenbourg, P.: What do you mean by collaborative learning? Collaborative- learning: Cognitive and computational approaches. pp. 1–19 (1999)
work page 1999
-
[8]
Contemporary Educational Psychology 69, 102050 (2022)
Dindar, M., Järvelä, S., Nguyen, A., Haataja, E., Çini, A.: Detecting shared physio- logical arousal events in collaborative problem solving. Contemporary Educational Psychology 69, 102050 (2022)
work page 2022
Show all 41 references
-
[9]
Journal of business and Psychology17, 245–260 (2002)
Donaldson, S.I., Grant-Vallone, E.J.: Understanding self-report bias in organiza- tional behavior research. Journal of business and Psychology17, 245–260 (2002)
2002
-
[10]
Learning and Instruction22(2), 145–157 (2012)
D’Mello, S., Graesser, A.: Dynamics of affective states during complex learning. Learning and Instruction22(2), 145–157 (2012)
2012
-
[11]
Learning and Instruction29, 153–170 (2014)
D’Mello, S., Lehman, B., Pekrun, R., Graesser, A.: Confusion can be beneficial for learning. Learning and Instruction29, 153–170 (2014)
2014
-
[12]
Cognition & emotion6(3-4), 169–200 (1992)
Ekman, P.: An argument for basic emotions. Cognition & emotion6(3-4), 169–200 (1992)
1992
-
[13]
International Journal of Organizational Analysis10, 343–362 (12 2002)
Feyerherm, A., Rice, C.: Emotional intelligence and team performance: The good, the bad and the ugly. International Journal of Organizational Analysis10, 343–362 (12 2002). https://doi.org/10.1108/eb028957
2002 doi
-
[14]
TechTrends59, 64–71 (2015) An Exploration of Internal States in Collaborative Problem Solving 15
Gašević, D., Dawson, S., Siemens, G.: Let’s not forget: Learning analytics are about learning. TechTrends59, 64–71 (2015) An Exploration of Internal States in Collaborative Problem Solving 15
2015
-
[15]
Jour- nal of Learning Analytics4(2), 113–128 (2017)
Gasevic, D., Jovanovic, J., Pardo, A., Dawson, S.: Detecting learning strategies with analytics: Links with self-reported measures and academic performance. Jour- nal of Learning Analytics4(2), 113–128 (2017)
2017
-
[16]
International Journal of Intel- ligent Networks2, 64–69 (2021)
Geetha, M., Renuka, D.K.: Improving the performance of aspect based sentiment analysis using fine-tuned bert base uncased model. International Journal of Intel- ligent Networks2, 64–69 (2021)
2021
-
[17]
In: Applied natural language processing: Identification, investigation and resolution, pp
Graesser, A.C., D’Mello, S., Hu, X., Cai, Z., Olney, A., Morgan, B.: Autotutor. In: Applied natural language processing: Identification, investigation and resolution, pp. 169–187. IGI Global (2012)
2012
-
[18]
Hampton, J.A., Passanisi, A.: When intensions do not map onto extensions: Indi- vidualdifferencesinconceptualization.JournalofExperimentalPsychology:Learn- ing, Memory, and Cognition42(4), 505 (2016)
2016
-
[19]
Jour- nal of veterinary medical education40(4), 333–341 (2013)
Hazel,S.J.,Heberle,N.,McEwen,M.M.,Adams,K.:Team-basedlearningincreases active engagement and enhances development of teamwork and communication skills in a first-year course for veterinary and animal science undergraduates. Jour- nal of veterinary medical education40(4), 333–3...
2013
-
[20]
Organizational Research Methods 12(3), 554–566 (2009)
Highhouse, S.: Designing experiments that generalize. Organizational Research Methods 12(3), 554–566 (2009)
2009
-
[21]
Huang, X., Lajoie, S.P.: Social emotional interaction in collaborative learning: Why it matters and how can we measure it? Social Sciences & Humanities Open7(1), 100447 (2023)
2023
-
[22]
arXiv preprint arXiv:1607.01759 (2016)
Joulin, A., Grave, E., Bojanowski, P., Mikolov, T.: Bag of tricks for efficient text classification. arXiv preprint arXiv:1607.01759 (2016)
2016 arXiv
-
[23]
In: Proceedings of naacL-HLT
Kenton, J.D.M.W.C., Toutanova, L.K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of naacL-HLT. vol. 1. Minneapolis, Minnesota (2019)
2019
-
[24]
Kerr, N.L., Tindale, R.S.: Group performance and decision making. Annu. Rev. Psychol. 55(1), 623–655 (2004)
2004
-
[25]
Khebour, I., Lai, K., Bradford, M., Zhu, Y., Brutti, R., Tam, C., Tu, J., Ibarra, B., Blanchard, N., Krishnaswamy, N., Pustejovsky, J.: Common ground tracking in multimodal dialogue (2024), https://arxiv.org/abs/2403.17284
2024 arXiv
-
[26]
In: Mining text data, pp
Liu, B., Zhang, L.: A survey of opinion mining and sentiment analysis. In: Mining text data, pp. 415–463. Springer (2012)
2012
-
[27]
arXiv preprint cs/0205028 (2002)
Loper, E., Bird, S.: Nltk: The natural language toolkit. arXiv preprint cs/0205028 (2002)
2002 arXiv
-
[28]
The MIT Press (1999)
Manning, C.: Foundations of statistical natural language processing. The MIT Press (1999)
1999
-
[29]
Advances in neural information processing systems26 (2013)
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed repre- sentations of words and phrases and their compositionality. Advances in neural information processing systems26 (2013)
2013
-
[30]
Nath, A., Venkatesha, V., Bradford, M., Chelle, A., Youngren, A., Mabrey, C., Blanchard, N., Krishnaswamy, N.: Any other thoughts, hedgehog? linking deliber- ation chains in collaborative dialogues (2024), https://arxiv.org/abs/2410.19301
2024 arXiv
-
[31]
Handbook of research methods in personality psychology1(2007), 224–239 (2007)
Paulhus, D.L., Vazire, S., et al.: The self-report method. Handbook of research methods in personality psychology1(2007), 224–239 (2007)
2007
-
[32]
Program14(3), 130–137 (1980)
Porter, M.F.: An algorithm for suffix stripping. Program14(3), 130–137 (1980)
1980
-
[33]
In: International conference on machine learning
Radford,A.,Kim,J.W.,Xu,T.,Brockman,G.,McLeavey,C.,Sutskever,I.:Robust speech recognition via large-scale weak supervision. In: International conference on machine learning. pp. 28492–28518. PMLR (2023) 16 S. Anindho et al
2023
-
[34]
In: The 7th international student conference on advanced science and technology ICAST
Rahutomo, F., Kitasuka, T., Aritsugi, M., et al.: Semantic cosine similarity. In: The 7th international student conference on advanced science and technology ICAST. vol. 4, p. 1. University of Seoul South Korea (2012)
2012
-
[35]
DICE Discussion Paper (2020)
Riener, G., Schneider, S., Wagner, V.: Addressing validity and generalizability concerns in field experiments. DICE Discussion Paper (2020)
2020
-
[36]
International Journal of Computer-Supported Collaborative Learning 9, 365–370 (2014)
Stahl, G., Law, N., Cress, U., Ludvigsen, S.: Analyzing roles of individuals in small-group collaboration processes. International Journal of Computer-Supported Collaborative Learning 9, 365–370 (2014)
2014
-
[37]
Computers & Education 143, 103672 (2020)
Sun, C., Shute, V.J., Stewart, A., Yonehiro, J., Duran, N., D’Mello, S.: Towards a generalized competency model of collab- orative problem solving. Computers & Education 143, 103672 (2020). https://doi.org/https://doi.org/10.1016/j.compedu.2019.103672, https://www.sciencedirec...
2020
-
[38]
Frontiers in psychology8, 235933 (2017)
Tyng, C.M., Amin, H.U., Saad, M.N., Malik, A.S.: The influences of emotion on learning and memory. Frontiers in psychology8, 235933 (2017)
2017
-
[39]
In: International Conference on Human-Computer In- teraction
VanderHoeven, H., Blanchard, N., Krishnaswamy, N.: Point target detection for multimodal communication. In: International Conference on Human-Computer In- teraction. pp. 356–373. Springer (2024)
2024
-
[40]
In: International Conference on Human-Computer Inter- action
VanderHoeven, H., Bradford, M., Jung, C., Khebour, I., Lai, K., Pustejovsky, J., Krishnaswamy, N., Blanchard, N.: Multimodal design for interactive collaborative problem-solving support. In: International Conference on Human-Computer Inter- action. pp. 60–80. Springer (2024)
2024
-
[41]
In: Paaßen, B., Epp, C.D
Venkatesha, V., Nath, A., Khebour, I., Chelle, A., Bradford, M., Tu, J., Puste- jovsky, J., Blanchard, N., Krishnaswamy, N.: Propositional extraction from natu- ral speech in small group collaborative tasks. In: Paaßen, B., Epp, C.D. (eds.) Proceedings of the 17th Internation...
2024 doi
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.