{"id":"65654fd1-95b6-496f-ba18-219f65bd5cc2","arxiv_id":"2501.15328","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A triadic VR study reports that how well one teammate's future joystick actions can be predicted from other teammates' physiology and behavior correlates with rings passed, while conventional synchrony measures do not.","lead":"This study reports that in a three-person virtual reality spacecraft task, how well a team's future actions can be predicted from teammates' physiology and behavior correlates with team performance, while simple synchrony measures do not. The authors propose this predictability as a new quantitative biomarker for team dynamics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported predictability score is not shown to be out-of-sample; without a train/test split, Section 2.5's beta = 3.20 may reflect memorization rather than genuine forecasting.","rationale":"The reader's REJECT is well aligned with the technical state of the paper. The central result depends on a quantity, predictability, that is never shown to be a valid out-of-sample forecast. Section 4.10 is the only evaluation description, and it lacks a split. The paper contains no code release or formal verification that would independently corroborate the result, so the burden falls entirely on the described protocol. I considered whether the strongest concern should instead be the missing trajectory-only baseline; that is real, but it is secondary because the first-order question is whether any reported r is a generalizable prediction at all. If the model is trained and scored on the same epochs, the GLMM in Fig. 3d could be significant even if no true predictability exists. The leave-one-team-out check directly settles this: if the significant beta survives only in-sample scoring, the main claim fails. If it survives, the paper needs a trajectory ablation to support the physiological claim, but the core correlation would be credible. Thus my read does not change the verdict: REJECT remains appropriate, with revision contingent on a proper cross-validated evaluation and ablations.","tokens_in":11915,"tokens_out":6430,"duration_ms":62711,"concrete_test":"Retrain the transformer under leave-one-team-out cross-validation: for each of the 10 teams, train on the other 9 teams' epochs, predict the held-out team's actions, and recompute the Section 4.10 r and the Section 2.5 GLMM (beta and P) using only these held-out predictions. Repeat with a trajectory-only input (no pupil, EEG, or controller data from teammates). If the GLMM beta becomes non-significant, or if the trajectory-only model matches the full model's held-out r, the reported predictability-performance link is not supported as a physiologically-informed, out-of-sample biomarker.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The decisive assumption is that the predictability score in Section 4.10 is a measure of out-of-sample predictive accuracy. The paper does not establish this. Section 4.9 states only that a validation dataset was used to monitor loss and accuracy during training; Section 4.10 defines r as the Pearson correlation between concatenated target actions and concatenated model predictions, with no statement that those predictions were made on data held out from training. If the same epochs are scored that were used to fit the transformer, r is inflated by memorization, and the team-level correlation in Fig. 3d and Section 2.5 (beta = 3.20, P < 0.001) may be a correlation with a training artifact rather than with genuine predictability. This is especially serious because each epoch is short (1 s input, 0.5 s output), each team contributes at most three sessions, and the final predictability analysis uses only n = 10 teams with no reconciliation of why 7 of the 17 usable teams drop out. A trajectory-only control is also missing from Sections 4.9-4.10, so even a properly held-out r could reflect the deterministic mapping from ring position to corrective action rather than teammates' physiology. Both omissions point to the same load-bearing gap: the metric has not been shown to measure what the title claims.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a virtual-reality triadic collaborative task (Apollo Distributed Control Task, ADCT) in which 17 teams of three performed three sessions each. The authors collected EEG, pupillometry, eye gaze, speech, and controller inputs and computed inter-subject correlations (ISC) across modalities, finding that speech synchrony is positively associated with team performance while physiological synchronies are not. The main contribution is a 'predictability' metric: a multi-head attention transformer is trained to predict one co-pilot's future controller action from the other two co-pilots' multimodal data and spacecraft trajectory, and the team-level predictability (average Pearson r between predicted and true actions) is reported to correlate positively with team performance (β=3.20, P<0.001). The paper argues that this predictability is a biomarker of team performance, challenging the emphasis on synchrony.","tokens_in":12205,"tokens_out":7438,"duration_ms":59580,"significance":"If the predictability metric were validated as an out-of-sample forecast, it would provide a novel, quantitative biomarker for team coordination and a template for multimodal team analysis; the ADCT dataset with synchronized multimodal recordings across repeated sessions is a valuable resource. The study also contributes a useful negative result on physiological synchrony. However, the central claim is not yet established because the paper does not show that the predictability scores are out-of-sample, does not compare against trajectory-only or self-history baselines, uses a small sample with unexplained exclusions, and reports the headline estimate without confidence intervals or a full model specification. These issues are load-bearing for the main conclusion.","major_comments":[{"comment":"The model evaluation in Section 4.10 does not demonstrate that the reported Pearson r is based on out-of-sample predictions. Section 4.9 mentions only that a validation dataset was used to monitor loss during training, and Section 4.10 computes the correlation between concatenated target actions and model predictions without stating that the predicted epochs were held out from training. If the scored epochs are the same ones used for fitting, the predictability scores are inflated by memorization, and the team-level correlation in Fig. 3d (β=3.20, P<0.001 in Section 2.5) may reflect a training artifact rather than genuine predictability. Please specify the train/validation/test split, the number of epochs in each, and report the correlation on held-out data.","section":"4.9–4.10"},{"comment":"No baseline or ablation is reported for the forecasting model. The input combines spacecraft trajectory with the two co-pilots' multimodal data, but the paper never tests whether the physiological and behavioral inputs add predictive power beyond the trajectory alone, or whether the target's own recent actions suffice. Without such controls, the predictability metric could be driven by the deterministic relationship between ring position and required corrective action, which is available from the environment rather than from teammates. The claim that this is a 'physiologically-informed' biomarker requires showing that the multimodal teammate data improves prediction over trajectory-only and self-history baselines.","section":"4.9–4.10"},{"comment":"The predictability analysis is based on n=10 teams, yet the Methods section reports only one team excluded for incomplete sessions and lists modality-specific exclusions (pupil: 4 teams, EEG: 9 teams, speech: 10 teams). The paper does not explain how 17 usable teams reduce to 10 for the predictability analysis, nor whether these exclusions are related to team performance. If data-quality issues are more common in low- or high-performing teams, the correlation between predictability and performance could be biased. Please provide a participant-flow diagram showing the sample size for each analysis and a sensitivity analysis with alternative exclusion criteria.","section":"4.1, Fig. 3c–d"},{"comment":"The headline association (β=3.20, P<0.001) is reported without confidence intervals, standard errors, or the number of observations used in the GLMM. With only 10 teams, the apparent correlation in Fig. 3d may be sensitive to a few influential points. The paper should report the GLMM specification explicitly (fixed and random effects as used, not just the general form in Eq. 6), include confidence intervals, and show team-level scatters with per-session points to assess robustness. A nonparametric or robust correlation at the team level would also be helpful.","section":"2.5"},{"comment":"The rationale for excluding speech data from the model input is not convincing: the authors state that speech event synchrony significantly correlates with team performance and therefore they excluded speech to 'avoid potential bias.' Excluding a modality because it correlates with the outcome removes a potentially informative predictor and may itself bias the predictability metric. The supplementary comparison with speech included should be summarized in the main text, or the exclusion should be justified on other grounds (e.g., technical artifacts or preprocessing failures).","section":"2.4"}],"minor_comments":[{"comment":"The repeated-measures ANOVA results report degrees of freedom such as F(2,32) for N=17 teams, but the error degrees of freedom are not consistently derived; please clarify the within-subject design and report effect sizes or at least mean squared errors.","section":"2.2"},{"comment":"The text states that 8-head attention is used, but Eq. (4) shows Concat(head1,...,head4); please correct the inconsistency.","section":"4.9, Eq. (4)"},{"comment":"Hyperparameters for training (loss function, optimizer, learning rate, batch size, number of epochs, early stopping criterion) are not reported; without these, the forecasting model is not reproducible.","section":"4.9"},{"comment":"The Pearson correlation is computed on target and predicted controller actions that are largely discrete ternary values (-1, 0, 1); a rank-based metric or classification accuracy would be more appropriate and less sensitive to the action baseline rate.","section":"4.10"},{"comment":"Equation (1) is typeset with garbled fragments ('qPn', '\\sqrt{}'); please fix the formatting so the formula is readable.","section":"4.7, Eq. (1)"},{"comment":"The claim that 'the total number of rings passed by each team increased monotonically' is based on group means; individual teams may not all show monotonic improvement, so please soften to 'on average' and present individual trajectories.","section":"2.1"},{"comment":"The paper states that co-pilots make about 0.3 remote controller actions in the 0.5-second output window; report the exact distribution and how this threshold was chosen, since the rarity of actions affects the interpretability of the correlation.","section":"2.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from code and data availability statements; the current preprint does not link to any. The title's 'Forecasts' overstates the correlational analysis. The paper is primarily a methods/biomarker study and should be evaluated accordingly; the missing out-of-sample validation is the key technical risk that the revision must address."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result here—that a teammate-predictability score computed from others' physiology and behavior correlates with team performance—is interesting, and the VR triad task is a good vehicle for it. But the paper never establishes that the score is actually out-of-sample, so the key correlation may be inflated by memorization or by the environment rather than by genuine interpersonal predictability.\n\nWhat is new: applying a transformer to predict one co-pilot's controller action from the other two co-pilots' multimodal data and linking that predictability to performance is a novel measurement. The synchrony null results are also a useful contribution; the data collection is careful and the decision to exclude speech from the model input is defensible.\n\nWhere it falls short: Sections 4.9–4.10 define the model and the correlation between predicted and true actions, but they never state that the scored predictions came from data held out from training. A validation dataset is mentioned only for monitoring loss and accuracy during training. Without a train/test split or cross-validation, the reported beta = 3.20 could be a correlation with a training artifact. Equally important, there is no control model trained on trajectory alone. Since the prediction target is the action in the 0.5 s before a ring, a model could learn the deterministic mapping from ring position to corrective action and achieve high 'predictability' without using teammate data at all. That would make the biomarker a property of the task, not of the team.\n\nThe sample also shrinks from 17 usable teams to n=10 in the predictability analysis with no explanation, and the key coefficients lack confidence intervals. The title says 'Forecasts Team Performance,' but the analysis is a contemporaneous GLMM correlation across sessions, not a prediction of future performance. These are load-bearing omissions, not minor quibbles.\n\nWho this is for: the team-neuroscience and human-AI teaming community. The idea is worth pursuing, and the experimental design is solid enough that the flaws can be fixed. This deserves a serious referee. I would send it out, but with the clear expectation of major revision: add proper held-out evaluation, trajectory-only and no-physiology ablations, confidence intervals, and reconcile the sample sizes. If those analyses come back and the correlation survives, it would be a credible biomarker.","headline":"Interesting idea, solid VR task, but the predictability metric is not shown to be out-of-sample, so the main correlation claim outruns the evidence.","tokens_in":12715,"tokens_out":4101,"would_cite":false,"duration_ms":34835,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A team-level predictability score—how well two teammates' physiology and behavior forecast the third's next control action—is positively correlated with team performance in a three-person virtual-reality spacecraft task.","keywords":["team performance","predictability","multi-modal physiology","virtual reality","transformer","inter-subject correlation","human teaming","biomarker"],"falsifier":"The claim would be settled by re-estimating the team predictability scores with an explicit held-out protocol—for example, training on trials from two sessions and predicting the third session, or leaving entire teams out—and checking whether the per-person Pearson $r$ values and the team-level $\\beta = 3.20$ correlation survive; if out-of-sample predictability drops to near zero, the biomarker reflects the model memorizing its training data rather than real predictability.","tokens_in":11685,"feed_emoji":"🚀","tokens_out":11073,"duration_ms":89094,"temperature":0.7,"pith_summary":"Teams that perform better in a three-person virtual-reality spacecraft task are teams whose members' next control actions can be predicted from the other two members' physiology and behavior. The paper builds a transformer model that, for each ring-passing event, takes one second of two co-pilots' multi-modal physiological and behavioral signals plus the spacecraft trajectory and predicts the third co-pilot's controller action over the next half-second. Averaging the Pearson correlation between predicted and actual actions across the three co-pilots gives a team-level predictability score that is positively linked to rings passed ($\\beta = 3.20$, $P < 0.001$). The same data show only limited or inconsistent correlations between synchrony and performance, which the authors read as evidence that predictability, not synchronization, is the more telling team biomarker.","feed_headline":"How well teammates predict a pilot's next move forecasts team success","feed_subtitle":"Two teammates' signals predicting a pilot's next action track rings passed; synchrony doesn't.","key_machinery":"The machinery is a leave-one-out predictive model built on a multi-head attention transformer. For each ring event, the encoder receives one second of spacecraft trajectory and the behavioral and physiological signals of two co-pilots; cross-modal attention layers fuse those modalities, self-attention layers extract intra-modal patterns, and a decoder with masked attention generates the third co-pilot's controller action over the next 0.5 seconds. Predictability is defined as the Pearson correlation $r$ between the predicted and actual actions for each co-pilot, then averaged over the three co-pilots to form the team biomarker; speech data are excluded from the input so that the result is not driven by the one modality whose synchrony already correlates with performance.","core_discovery":"On the paper's own terms, the central discovery is that a team-level predictability score—the average, across the three co-pilots, of the Pearson correlation between a co-pilot's true controller action and a transformer model's prediction of that action from the other two co-pilots' multi-modal physiology and behavior—predicts team performance. In the Apollo Distributed Control Task, teams with higher predictability passed more rings ($\\beta = 3.20$, $P < 0.001$, across 10 teams), while inter-subject synchrony in pupil size, EEG, controller actions, and speech was mostly not significantly or only weakly related to performance. The paper interprets this as evidence that high-performing team members become mutually anticipatable, and that this anticipatability is visible in their physiology and behavior before the team outcome is known.","pith_inferences":["A direct extension the paper does not pursue is an intervention test: if raising mutual predictability through practice, shared displays, or communication protocols also raises rings passed, that would turn the observed correlation into a causal handle on team performance.","The same leave-one-out predictability metric could be ported to human-AI teams, treating an AI teammate's actions as the target to measure whether an artificial agent behaves in a way human teammates can anticipate.","A modality-ablation study—removing EEG, pupil, or controller inputs one at a time—would reveal which signals carry the predictability effect, a question the paper leaves open."],"forward_implications":["In this task, team performance is indexed by how forecastable each member's next controller action is from teammates' combined physiology and behavior, rather than by how synchronized teammates are.","Synchrony metrics alone—EEG, pupil, controller-action, and speech-event inter-subject correlations—are not reliable performance biomarkers in this setting.","The predictability score offers a continuous, quantitative outcome for team training, role assignment, or feedback systems in collaborative sensorimotor tasks.","Because the main result is computed without speech input, the predictability-performance link is not explained by the one modality whose synchrony already predicted performance."],"supporting_citations":[{"why":"Supplies the multi-head attention transformer architecture used to forecast controller actions.","marker":"(44)"},{"why":"Supplies the boundary avoidance task design that the Apollo Distributed Control Task extends to a three-person team.","marker":"(12)"},{"why":"Supplies the premise that high-performing teams are characterized by predictable interactions, which the predictability metric operationalizes.","marker":"(3)"},{"why":"Supplies CorrCA, the method used to compute EEG inter-subject synchrony for comparison with predictability.","marker":"(29)"},{"why":"Documents brain-to-brain synchrony in group settings, a basis for the synchrony-performance expectations tested here.","marker":"(8)"},{"why":"Reports that physiological and behavioral synchrony predict group cohesion and performance, the conventional result the paper compares against.","marker":"(13)"},{"why":"Shows synchronization of brains, hearts, and eyes during shared stimuli, supporting the synchrony hypothesis.","marker":"(23)"},{"why":"Links inter-brain phase synchronization to team performance, a baseline the predictability result is contrasted with.","marker":"(41)"}],"fun_headline_variants":["Teammate predictability, not synchrony, forecasts team performance","How well you predict your teammate's next move predicts success","In VR teams, anticipating teammates forecasts performance more than syncing","Physiology-based predictability of teammate actions forecasts team success","Predictability of teammate moves forecasts rings, synchrony doesn't"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the reported predictability scores measure genuine prediction of new data; the paper does not say how the transformer model's training and test data were separated, so the correlation with team performance is only trustworthy if the model generalizes beyond the examples it was fit on.","fun_headline_variants_meta":{"raw":{"variants":["Teammate predictability, not synchrony, forecasts team performance","How well you predict your teammate's next move predicts success","In VR teams, anticipating teammates forecasts performance more than syncing","Physiology-based predictability of teammate actions forecasts team success","Predictability of teammate moves forecasts rings, synchrony doesn't"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000862,"raw_usage":{"total_tokens":3690,"prompt_tokens":849,"completion_tokens":2841,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":2758}},"tokens_in":465,"tokens_out":2841,"duration_ms":20650,"temperature":1.0,"reasoning_tokens":2758,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:23:06.614417+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The claim would be settled by re-estimating the team predictability scores with an explicit held-out protocol—for example, training on trials from two sessions and predicting the third session, or leaving entire teams out—and checking whether the per-person Pearson $r$ values and the team-level $\\beta = 3.20$ correlation survive; if out-of-sample predictability drops to near zero, the biomarker reflects the model memorizing its training data rather than real predictability.","supporting_citations":[],"review_version":1}