{"id":"9470b653-6f81-4a9c-9b52-e1aff42e7bbd","arxiv_id":"2607.13614","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A paired within-subject multimodal corpus of 32 groups recorded in-person and via videoconference reveals setting-specific shifts in turn-taking, language complexity, facial expression, and rated enjoyment.","lead":"VIP-MINGLE is a new 59-hour multimodal dataset of 105 people in 32 groups, each recorded both in-person and over videoconference doing the same collaborative task. The paper reports behavioral shifts between settings — slower turn-taking, simpler speech, different facial expressions, and lower enjoyment ratings — arguing that models built for one mode may not transfer to the other.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Diarization quality is the load-bearing risk for the behavioral-shift claim, but the transparent open corpus makes it testable rather than fatal.","rationale":"The reader identified exactly the load-bearing concern: speaker diarization validity for in-person audio (Section 2.3.1) underpins the speech-temporal results in Section 3.1 that constitute the headline behavioral-shift evidence. I agree that this is the correct weak point: the approach applies a state-of-the-art diarization model in a low inter-microphone distance, high-crosstalk setup with no DER, no overlap analysis, and no manual validation, while the videoconference condition is naturally clean. Because the turn-level statistics are directly derived from diarization boundaries, any systematic diarization bias (e.g., truncating or merging turns during overlap, or inserting false transitions in crosstalk) would produce exactly the reported pattern of longer gaps and shorter utterances. However, the corpus contribution itself remains valuable and independently verifiable: raw audio/video are released with a DOI and scripts, the within-subject design is real, and the main effects replicate prior work. Therefore the concern does not move the verdict to REJECT—it makes the manuscript's strong behavioral conclusions conditional on an audit that the released data can support. I would keep the reader's CONDITIONAL verdict and agree with the assessment that this is the weakest assumption; if the audit passes, the claim is substantially strengthened; if it fails, only the corpus-release framing survives and the behavioral-shift claim becomes unsupported.","tokens_in":7080,"tokens_out":1525,"duration_ms":15898,"concrete_test":"Run the released pyannote diarization on a random sample (e.g., 10% of in-person sessions) and compute DER against a fresh manual speaker-turn annotation, reporting separate error rates for speech vs. silence and for overlap regions. Then recompute the Section 3.1 mixed-effects models using only turns whose boundaries are confirmed within ±100 ms by manual annotation; if the β coefficients for turn-taking gap and utterance duration lose significance or change direction, the behavioral-shift claim is a diarization artifact rather than a medium effect.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central behavioral-shift claims in Section 3.1 (longer turn-taking gaps, shorter utterance durations in videoconference) rest on per-participant speech timing derived from pyannote speaker diarization applied to in-person recordings where the target-to-interfering distance ratio was below the 3:1 rule of thumb and crosstalk was only mitigated by manual gain reduction (Section 2.3.1). No diarization error rate (DER), no per-speaker overlap analysis, and no comparison against a held-out manual segmentation is reported. In multiparty recordings with substantial crosstalk, speaker diarization errors—especially missed short utterances and merged/split turns—directly bias utterance-duration and gap distributions. Since the videoconference audio is intrinsically clean per-room, any diarization-error asymmetry between settings would masquerade as a setting effect. The paper's own limitations do not flag diarization accuracy, and the reader's verdict already identified this as the weakest assumption. The corpus is released with raw audio, so this is an audit problem, not an unfixable design flaw. The claim that the corpus is a controlled bridge is independently supported by the within-subject design and raw-data release, and the reported effects replicate known directions (e.g., Boland 2022), so the concern is measurement validity, not internal contradiction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents VIP-MINGLE, a multimodal corpus of 32 groups and 105 participants who each took part in paired in-person and videoconference sessions of a standardized collaborative game. The dataset release includes raw audio/video, psychometric baselines, processed features (pyannote diarization, Whisper transcripts, OpenFace/DeepFace facial features), and crowd-sourced segment-level annotations. The authors report exploratory comparisons showing longer turn-taking gaps, shorter utterance durations, lower syntactic complexity, altered facial-expression patterns, and lower enjoyment in videoconference sessions, using these differences to argue that videoconferencing is a qualitatively different communicative medium rather than a degraded version of in-person interaction.","tokens_in":7371,"tokens_out":5854,"duration_ms":55421,"significance":"If the corpus and its processing pipeline are validated, VIP-MINGLE is a valuable community resource. Its main strength is the within-subject design: the same groups perform the same task in both settings, which removes many speaker- and task-level confounds present in unpaired cross-setting comparisons. The decision to release raw recordings in addition to processed features is also commendable, because it makes the automatic-pipeline outputs auditable and re-computable. The key caveat is that the central behavioral-shift claims rest on automatic diarization, transcription, and annotation whose measurement quality is not documented in the manuscript. Since the raw data are available, this is a fixable validation gap rather than a fatal design flaw, but it must be addressed before the reported comparisons can be interpreted as evidence about communication-medium effects.","major_comments":[{"comment":"The central speech-timing effects (longer turn-taking gaps and shorter utterance durations in videoconference, Fig. 2A–B) are computed from pyannote diarization of in-person recordings under conditions explicitly described as violating the 3:1 target-to-interferer rule. No diarization error rate (DER), precision/recall against manual segmentation, or per-speaker overlap statistics are reported. Because the videoconference condition is intrinsically room-separated, any systematic diarization error in the in-person condition—such as missed short utterances or split/merged turns—will shift the gap and duration distributions and can masquerade as a setting effect. The authors should provide DER on a held-out subset, overlap statistics, and a comparison of automatically derived turn boundaries against manual annotations. This is necessary to support the §3.1 claim.","section":"§2.3.1, §3.1"},{"comment":"The language-complexity analysis in §3.2 uses Whisper transcripts as the basis for MDD and other text metrics, but no word error rate (WER) or transcript-quality assessment is reported. Moreover, the speaker-attributed transcripts are obtained by aligning Whisper segments to diarized speaker labels; misalignment or diarization errors will add noise or bias to per-speaker text. The authors should report WER on a sampled subset and describe how overlapping speech and alignment failures were handled. This is particularly relevant to the MDD difference in Fig. 2C, which is the only significant language-complexity result.","section":"§2.3.2, §3.2"},{"comment":"The annotation analysis relies on 7,077 clips rated by 192 annotators, with only a reference to prior work [10] for excluding low-reliability annotators. No inter-annotator agreement coefficients are reported for the enjoyment ratings or for the multi-label event annotations (interruptions, gaps), and the number/percentage of excluded annotators is not given. Without these, the enjoyment and event-rate comparisons in Fig. 3 are not independently verifiable. Please report Krippendorff's alpha or ICC for the Likert ratings and per-label agreement for the event annotations, along with the exclusion criteria and counts.","section":"§4"},{"comment":"The headline corpus-scale numbers are internally inconsistent. Thirty-two groups × two settings × an average session duration of 21 minutes amounts to approximately 22.4 hours of wall-clock session recordings, not the stated ≈59 hours. If 'recording hours' counts individual audio/video tracks or streams rather than session time, that should be stated explicitly, and the relationship between 'average session duration' and 'recording hours' should be defined. This statistic is central to the corpus's contribution and must be unambiguous.","section":"§2"},{"comment":"The facial-expression comparisons test a large number of action units and emotion categories with uncorrected Wilcoxon signed-rank tests. Given the multiple comparisons, the few significant differences could include false positives. The exploratory framing in the text is helpful, but the abstract's claim of 'significant behavioral distribution shifts' would be better supported by reporting multiple-comparison correction (or explicitly acknowledging the uncorrected exploratory nature) and effect sizes for the significant AUs/emotions.","section":"§3.3, Fig. 2D–E"}],"minor_comments":[{"comment":"Minor formatting: '16 kHZ' should be '16 kHz.' Also, §2.3.1 mentions encoding diarized audio as 'mono, 32 kHz audio-only MP4 files,' while §2.3.2 says audio was standardized to 16 kHz mono WAV; please clarify the final format(s) and sampling rates shipped in the corpus.","section":"§2.3.2"},{"comment":"The mixed-effects models are described only by β, SE, and p. Please report the model specification explicitly—random effects (group/participant/session), fixed effects, and how order/counterbalancing was entered—so the coefficients are interpretable and reproducible.","section":"§3.1"},{"comment":"The annotation procedure says annotators viewed 120 clips each, but it is unclear how this maps to 7,077 rated clips across 192 annotators. A sentence describing the annotation design (how many clips per segment, overlap, quality-control clips) would improve transparency.","section":"§4"},{"comment":"The event-label term 'gaps' in the annotation task should be defined consistently with the Heldner–Edlund gap measure used in §3.1; it is currently unclear whether they refer to the same construct or to a coarser subjective judgment.","section":"Fig. 3"},{"comment":"Reference [22] is listed as an 'Advanced Online Publication' from 2026; please verify the publication year and provide the final DOI if available.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The corpus design is genuinely useful, and I do not see a circularity or fatal design flaw. My recommendation is driven by missing validation metrics for the automatic pipelines and annotations. If the authors supply DER, WER, and inter-annotator agreement numbers—and clarify the recording-hour accounting—the paper is likely ready for acceptance. The open release of raw data makes this an audit-and-revise situation rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a corpus paper, and the corpus is the contribution. VIP-MINGLE looks like the first open, paired within-subject dataset of group conversation in both in-person and videoconference settings, with ~59 hours of raw audio/video, processed features, psychometrics, and 7,077 annotated segments. If the release matches the description, that fills a real gap: AMI/ICSI are in-person only, CANDOR/RoomReader are remote only, and cross-setting comparisons out there use unpaired or unavailable data.\n\nThe design is genuinely good: same groups do the same task in both settings, order counterbalanced, individual laptop audio plus lavalier mics, raw data and scripts on Zenodo. The self-citations aren't circular; they supply annotation criteria and modeling context rather than the outcome. The behavioral findings also replicate known directions—longer gaps, shorter utterances, simpler syntax in videoconference—so the paper is consistent with prior work rather than inventing a new effect.\n\nWhere it gets soft is measurement quality. Section 3.1's central claims come from pyannote speaker diarization applied to in-person recordings where the target-to-interfering distance ratio was below the 3:1 rule of thumb and crosstalk was only mitigated by manual gain reduction. No diarization error rate is reported. In multiparty recordings with overlap, diarization errors systematically bias utterance duration and gap distributions, and since the videoconference audio is cleanly separated by room, any error asymmetry between settings will masquerade as a setting effect. The authors don't flag this in the limitations, which is a real omission. The good news is they release raw audio, so this is an audit problem rather than an unfixable one—reviewers can demand a DER estimate on a held-out subset, or the authors can add it.\n\nThere are smaller issues in the same spirit: no WER for Whisper transcripts, no inter-annotator reliability for the annotations despite mention of excluding low-reliability annotators, uncorrected multiple comparisons over the AU/emotion tests, and a causal flourish in Section 4 that connects gaps/interruptions to 'Zoom fatigue'—the design can't support that. None of these undercut the corpus's value; they undercut the strength of the exploratory claims.\n\nWho's this for? People building or evaluating conversational AI across settings, and researchers studying videoconferencing effects. A serious referee should see it, with the expectation that measurement-quality audits are added; the release-with-data means those audits can actually be done.","headline":"The corpus is the contribution; the behavioral-shift claims need a data audit before they carry weight.","tokens_in":7885,"tokens_out":2517,"would_cite":true,"duration_ms":23458,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Videoconferencing reshapes group conversation, not merely degrades it","keywords":["multimodal corpus","group conversation","videoconferencing","in-person interaction","turn-taking","speaker diarization","facial expression","domain shift"],"falsifier":"Take a random sample of, say, 30 in-person and 30 videoconference sessions from the released corpus, have human annotators mark true speaker turns on the raw audio, and compute the diarization error rate separately for each setting. If the in-person error rate is dramatically higher and re-estimating standard turn-taking gaps and utterance durations from the human labels eliminates the significant differences, the central temporal claim is refuted. A second falsifier: check whether the higher in-person interruption rate survives human annotation, since the annotation study already measured int","tokens_in":7006,"feed_emoji":"🎥","tokens_out":5891,"duration_ms":55728,"temperature":0.7,"pith_summary":"VIP-MINGLE is a new corpus built to answer a deceptively simple question: when the same people do the same collaborative task face-to-face and over a video link, what actually changes? The paper's claim is that the change is substantial and qualitative. Videoconference sessions show longer turn-taking gaps, shorter utterances, syntactically simpler speech, and less salient facial expressions overall, though a few localized expressions increase. Human raters also find in-person sessions more enjoyable, even though the two settings are rated equally fluid. If right, this means communication-medium effects can be isolated from speaker and task confounds for the first time in an open dataset, and it warns that models built on one setting will not silently transfer to the other.","feed_headline":"Video calls yield shorter turns, longer pauses, less joy","feed_subtitle":"Same people, same task: switching to video changes speech, expressions, and enjoyment","key_machinery":"The engine of the design is within-subject pairing: the same group, the same task, two settings. Because speaker identity, group chemistry, and task demands are held constant, any systematic difference between the two sessions can be attributed to the communication medium. Supporting this are per-participant signal-processing pipelines — speaker diarization to separate individual voices, speech transcription with word-level timestamps, facial-action-unit and emotion extraction from video, and human annotation of 10-second clips — all aligned to a common timeline. The turn-taking metrics specifically come from a standard model of conversation timing applied to diarized speech, so the quality","core_discovery":"The central discovery is a corpus plus a set of paired comparisons. The corpus records 32 groups (105 participants) in both settings on the same trivia-style collaborative task, with order counterbalanced, yielding about 59 hours of raw audio/video, per-participant processed features, psychometric baselines, and 7,077 segment-level human ratings. The analysis shows that switching from in-person to videoconferencing shifts entire distributions: turn-taking gaps lengthen, utterance durations shrink, dependency distance — a syntactic-complexity measure — drops, and facial-expression profiles change qualitatively, not just in amplitude. Ratings show lower enjoyment but unchanged fluidity in vide","pith_inferences":["A direct validation step the paper leaves implicit: compute diarization error rates on a hand-labeled subset of the in-person recordings. If errors are concentrated in overlapping speech, the reported gap/utterance differences could be an artifact of speaker-boundary estimation rather than a medium effect.","The 360° in-person video is a hidden asset: it can be used to test whether the lower enjoyment in video is driven by reduced gaze alignment or mutual attention, an explanation the paper does not pursue.","The syntactic-complexity drop without a change in lexical diversity suggests the effect may be about on-line production under video latency rather than vocabulary; a follow-up could examine disfluencies and filler words in the transcripts.","The equal-fluidity, lower-enjoyment dissociation implies that enjoyment is carried by non-temporal channels — expression intensity, gaze, or interruption patterns — which the corpus is well suited to disentangle in a mediation analysis."],"forward_implications":["Models trained on traditional in-person meeting corpora will likely need adaptation before they work on videoconference data; the corpus gives a benchmark for measuring that domain shift.","The longer gaps and shorter turns in video suggest that videoconferencing systems could improve perceived engagement by adjusting for turn-taking delays or making interruption cues more visible.","Because in-person conversation was rated more enjoyable despite more interruptions, 'clean' remote exchanges may be missing interactional dynamics that users value; interface design should consider restoring them.","The released raw recordings let others extract novel features (e.g., spatial cues from the 360° camera) to test whether the observed shifts replicate with newer models.","The significant behavioral differences across modalities motivate building multimodal, domain-aware models rather than assuming one-size-fits-all conversation understanding."],"fun_headline_variants":["Video calls lengthen pauses, shorten turns, cut joy","In-person to video: longer gaps, shorter speech, less joy","Same groups, video calls: distinct speech and emotion shifts","VIP-MINGLE corpus reveals video call behavior shifts","Video conferencing reshapes group talk and mood"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The in-person individual audio tracks are assumed to be correctly separated by automated speaker diarization even though the microphones were closer to interfering speakers than the usual 3:1 guideline recommends, and no error rate for that separation is reported; if the separation is unreliable during overlapping speech, the measured differences in turn-taking gaps and utterance durations could be artifacts of the separation software rather than real effects of the communica","fun_headline_variants_meta":{"raw":{"variants":["Video calls lengthen pauses, shorten turns, cut joy","In-person to video: longer gaps, shorter speech, less joy","Same groups, video calls: distinct speech and emotion shifts","VIP-MINGLE corpus reveals video call behavior shifts","Video conferencing reshapes group talk and mood"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1109,"prompt_tokens":666,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":410,"completion_tokens_details":{"reasoning_tokens":363}},"tokens_in":410,"tokens_out":443,"duration_ms":4707,"temperature":1.0,"reasoning_tokens":363,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T04:38:27.051388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 30 in-person and 30 videoconference sessions from the released corpus, have human annotators mark true speaker turns on the raw audio, and compute the diarization error rate separately for each setting. If the in-person error rate is dramatically higher and re-estimating standard turn-taking gaps and utterance durations from the human labels eliminates the significant differences, the central temporal claim is refuted. A second falsifier: check whether the higher in-person interruption rate survives human annotation, since the annotation study already measured int","supporting_citations":[],"review_version":1}