{"id":"96c952c9-a6b8-48c5-b60e-27d6405a09a8","arxiv_id":"2607.27296","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SKY-Piano pairs real hand and body motion capture with synchronized audio, MIDI, video, and scores for 19 pianists, plus pseudo-fingering labels.","lead":"SKY-Piano is a new 11-hour multimodal dataset of piano performances from 19 pianists, with directly measured hand and body motion synced to audio, MIDI, video, and scores. It gives MIR researchers a common testbed for linking how pianists move with what they play.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc alignment and SAITS imputation are not validated end-to-end; if imputed fingertip z error at key-press exceeds the 5 mm Tier-B threshold, the 'directly measured' hand-motion claim and pseudo-fingering labels degrade.","rationale":"I read the paper as a dataset contribution whose strongest claim is the first public corpus combining real hand and body mocap with MIDI and structured professional/amateur comparison. That claim depends on hand-marker trajectories being reliable at key-press frames. The reader's weakest assumption identifies exactly this bottleneck: post-hoc alignment with centimeter-scale residuals plus substantial marker dropout filled by an unevaluated imputation model. I agree that this is the most load-bearing concern. It is not fatal, because the flagged raw data is released and the pipeline is transparent, but it justifies a conditional acceptance rather than full acceptance, and it should be resolved by an end-to-end validation against hardware-synced ground truth. This does not change the reader's CONDITIONAL verdict, so I recommend UNCHANGED.","tokens_in":10916,"tokens_out":4755,"duration_ms":42613,"concrete_test":"Use hardware-timecode-synchronized professional sessions as ground truth. Artificially mask marker samples using the reported dropout distribution (14% median, 44% thumb-carpal), run the released SAITS imputer, and re-apply the post-hoc wrist-velocity cross-correlation and weighted Procrustes alignment to the subsampled data. Compare the resulting trajectories against the untouched hardware-timecode-aligned trajectories, focusing on per-marker Euclidean error at MIDI onset frames. If median fingertip error at onsets exceeds 5 mm (or the 90th percentile exceeds 15 mm), the imputed hand stream and Tier B/C fingering labels cannot sustain the current accuracy claims, and the paper should report quantitative accuracy bounds for the imputed and aligned streams rather than relying on the 'directly measured' label unqualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the released 120 Hz hand-marker trajectories being accurate enough at key-press instants to support the dataset's novelty as 'directly measured' hand/body mocap and to ground the pseudo-fingering pipeline. Section 3.3 reports that earlier sessions are temporally synchronized post hoc by wrist-velocity cross-correlation and spatially aligned via weighted Procrustes, with median per-trial landmark residuals 'at the centimeter scale.' Section 4 reports a median per-marker missing-cell rate of 14%, rising to 44% on thumb-carpal pairs, filled by SAITS. No held-out accuracy for SAITS imputation is reported, and no temporal-alignment error for the post-hoc sessions is given. The only reported validation, 91.4% keystroke detection within ±50 ms and below 15 mm above the key surface, uses tolerances much looser than the ±25 ms onset window and ≥5 mm z-gap used by the Tier B fingering rule. Since the pseudo-fingering labels and the fine-tuning target (MPJPE against imputed mocap) inherit this stream, unquantified errors here propagate into downstream claims. The flagged release mitigates raw-data concerns, but the imputed form is the one used for generation and full-coverage fingering annotations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SKY-Piano, an 11-hour multimodal piano performance dataset recorded from 19 pianists (7 professional, 12 amateur). The dataset combines optical hand and body motion capture with audio, MIDI, multi-view video, MusicXML scores, and Visual3D kinematics, and is structured so that technique, difficulty, and performer expertise can be compared on shared repertoire. The authors also contribute a geometry-based pseudo-fingering pipeline with an imputation tier, an interactive web explorer, and a fine-tuning experiment with the Tipiano MIDI-to-motion model as a use case. The central claim is that this is the first public corpus combining directly measured hand and body motion capture with synchronized MIDI, score, and a controlled amateur/professional design.","tokens_in":11261,"tokens_out":5081,"duration_ms":46878,"significance":"If the validation gaps are closed, this would be a substantial community resource. The combination of real hand and body optical mocap with MIDI, multi-view video, and structured repertoire is genuinely novel in the MIR dataset landscape. The release of both flagged and imputed motion, per-note source tags for fingering, an interactive explorer, and reproducible pipeline code are concrete strengths. The 2,269-note human audit of fingering labels on the exact released trials is a good reproducibility practice, and the CQT-based audio-to-MIDI alignment plus per-stage manual review add credibility. The fine-tuning case study usefully demonstrates interoperability with an existing MIDI-to-motion model, although the reported key-contact F1 drop needs to be addressed more convincingly. Overall, the dataset has high potential value, but the paper's load-bearing accuracy claims — particularly 'directly measured' sub-centimeter motion and the fingering validation — are not yet fully supported.","major_comments":[{"comment":"The Introduction motivates the dataset by requiring 'sub-centimeter motion data,' and the paper repeatedly describes the motion as 'directly measured.' However, §3.3 states that earlier sessions are temporally synchronized via wrist-velocity cross-correlation and spatially aligned via weighted Procrustes with 'median per-trial landmark residuals at the centimeter scale.' This is not sub-centimeter, and it is larger than the 5 mm z-gap used by the Tier B fingering rule in §5. The manuscript does not report how many sessions used the post-hoc alignment, nor the distribution of alignment errors at the fingertips and at key-press instants. Please report these quantities separately for hardware-synced and post-hoc sessions, and either restrict the sub-centimeter claim to the subset that supports it or provide evidence that key-press fingertip z accuracy is sufficient despite centimeter-scale","section":"§3.3 (Coordinate alignment)"},{"comment":"The released motion has a median per-marker missing-cell rate of 14%, rising to 44% on thumb-carpal pairs, and these gaps are filled by SAITS. No held-out validation of the imputed trajectories is reported. Since the imputed form is the full-coverage version that users will consume and likely the target of the fine-tuning experiment in §6.2, the absence of a measured imputation error — especially for fingertip z at key-press instants — leaves the downstream claims unquantified. Please report SAITS accuracy on held-out occlusions, ideally stratified by marker and by proximity to key-press events, and clarify explicitly whether the fine-tuning MPJPE is evaluated against imputed or flagged trajectories.","section":"§4 and §3.3 (Imputation)"},{"comment":"The fingering evaluation has two gaps. First, the 84.3% Tier C accuracy is a self-consistency statistic: the BiLSTM is trained and evaluated on labels produced by the same algorithmic pipeline. The external 2,269-note human audit evaluates only the geometry-based committed subset (94.0% coverage), so the imputed 6% of notes are not externally validated. Second, the keystroke-detection validation uses tolerances of ±50 ms and 15 mm above the key surface, which are much looser than the Tier B assignment thresholds of ±25 ms and ≥5 mm z-gap; therefore the 91.4% figure does not validate the finger-assignment stage at the thresholds actually used. Please report accuracy separately by source tag (algorithm/imputed) on the audited notes, provide confidence intervals for the 94.5% precision, and either audit the imputed notes or explicitly state that their accuracy is unknown.","section":"§5 (Fingering validation)"}],"minor_comments":[{"comment":"The Video column shows a checkmark without noting that video is only available for the professional cohort. Table 3 indicates amateur and mixed cohorts have no video; please qualify the Table 1 entry (e.g., 'up to 4-view, professional only').","section":"Table 1"},{"comment":"The key-contact F1 drops from 0.93 to 0.66 after fine-tuning while MPJPE improves. The fingertip-versus-MANO-joint explanation is plausible but unverified. Consider adding an evaluation that maps the released markers to the MANO convention, or reporting both conventions explicitly, so users can interpret the fine-tuning result as a compatibility demonstration rather than a regression.","section":"§6.2 / Table 6"},{"comment":"The text uses both 'Set1 graded' and 'Set 1 graded' (also 'Set 1' in §3.1). Please unify the terminology.","section":"§5 / Table 5"},{"comment":"Minor typo: 'Y AMAHA' should be 'YAMAHA'.","section":"§8 Acknowledgments"}],"recommendation":"major_revision","confidential_remarks":"This is a promising dataset paper with a strong and timely contribution. The main risk is that the 'directly measured' and 'sub-centimeter' claims outrun the validation, particularly for post-hoc aligned sessions and SAITS-imputed markers. The fingering evaluation also needs to clearly distinguish external validation from algorithmic self-consistency, and the 6% imputed-cover notes should be acknowledged as unvalidated. The case-study F1 drop should be addressed head-on in revision. I believe these issues are fixable within the manuscript's scope and do not require rejecting the dataset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a real dataset contribution, not a repackaging. SKY-Piano is the first corpus I've seen that combines directly measured hand and body optical mocap with synchronized audio/MIDI/video/score, includes amateurs alongside professionals, and structures repertoire by technique, difficulty, and expertise. That's a genuine gap in MIR, and the paper documents acquisition, preprocessing, and validation thoroughly. The release of both flagged and imputed motion, plus an interactive browser and explicit ambiguity flags for fingering, shows care.\n\nThe fingering pipeline is a sensible adaptation of PianoVAM to measured mocap. The keystroke validation against MIDI is strong: 91.4% within ±50 ms and below 15 mm over 113k notes. The 2,269-note human audit on three exact shipped pieces is reproducible and gives 94.5% precision over the committed subset. Credit where due.\n\nNow soft spots. Tier C validation is circular: the BiLSTM is trained on algorithm-generated labels and evaluated against the same labels, so the 84.3% figure measures self-consistency, not correctness. The human audit doesn't cover Tier C, since it only rates the geometry-based committed subset. That's not fatal—the explicit source tags let users filter—but it should be stated plainly, and a small sample of Tier C should be audited.\n\nThe stress-test concern about alignment and imputation is partially valid. Post-hoc sessions use wrist-velocity cross-correlation and Procrustes with median residuals at the centimeter scale, and the imputed form has up to 44% missing cells per marker. No held-out imputation accuracy is given. That matters for users treating the imputed stream as ground truth in generation tasks. However, the paper says pseudo-fingering uses the flagged motion, not the imputed stream, so the fingering labels are less exposed than the stress-test note suggests. Still, the imputed stream is the default for full-coverage uses, and the paper should add an end-to-end validation or stronger caveats.\n\nThe fine-tuning case study shows an F1 drop after adaptation; the explanation is plausible (marker vs MANO joint convention) but it undercuts the use-case claim slightly. That's a minor issue.\n\nWho is this for: MIR researchers working on motion-conditioned generation, expressive performance analysis, fingering estimation, and multimodal transcription. It deserves a serious referee and likely publication after revision. I'd cite it if I worked on piano hand motion.\n\nRecommendation: send to peer review.","headline":"A genuinely new and useful dataset—first to combine real hand and body mocap with amateurs, structured repertoire, and synchronized audio/MIDI/video—published as-is it would be a solid resource, but the fingering Tier C validation is circular and the imputed motion stream needs end-to-end validation.","tokens_in":11769,"tokens_out":2385,"would_cite":true,"duration_ms":19598,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SKY-Piano is an 11-hour multimodal piano dataset pairing directly measured hand and body motion capture with audio, MIDI, multi-view video, scores, and pseudo-fingering labels for 19 pianists, structured across technique, difficulty, and ex","keywords":["multimodal piano dataset","motion capture","hand mocap","piano fingering","MIDI-to-motion generation","performance analysis","optical motion capture","amateur vs professional"],"falsifier":"Compute the per-note fingering agreement on the three hand-audited pieces using only the flagged (non-imputed) motion, then compare with the released audit labels: if the subset of notes whose labels came from the imputation tier has an agreement far below the reported 94.5%, or if median fingertip z at MIDI onsets shifts by more than a few millimeters when occlusion flags are removed, the central precision claim fails.","tokens_in":10835,"feed_emoji":"🎹","tokens_out":4866,"duration_ms":40089,"temperature":0.7,"pith_summary":"This paper introduces SKY-Piano, an 11-hour multimodal piano performance dataset that synchronizes directly measured hand and body motion capture with audio, MIDI, multi-view video, MusicXML scores, and body-segment kinematics for 19 pianists, 7 professional and 12 amateur. The central claim is that no prior corpus combined real hand and body motion with a repertoire structured to allow controlled comparisons across playing technique, difficulty, and performer expertise. To make the motion useful, the authors include a fingering-labeling pipeline that uses measured fingertip depth to assign each MIDI note to the finger that pressed it, with explicit ambiguity flags and an imputation tier for unresolved cases. They also show the dataset works with an existing MIDI-to-motion generation model by fine-tuning it on their data. If the release holds up, it offers a reference point for piano research that needs true motion rather than motion estimated from video.","feed_headline":"Piano corpus records true hand motion across 19 pianists","feed_subtitle":"First public multimodal set with directly measured hand and body motion, plus fingering labels and MIDI.","key_machinery":"The central mechanism is a hardware-synchronized optical motion-capture setup: reflective markers on each hand plus a full-body marker suit, driven by a common timecode and cross-checked with audio-to-MIDI alignment. On top of that, the fingering pipeline operates on measured fingertip z-depth: for each MIDI note it scores every fingertip by whether it falls within the key's 3D rectangle and below a press-depth threshold, resolves ties with z-onset evidence, and imputes leftovers with a bidirectional LSTM, releasing per-note source tags. The dual flagged/imputed release of the motion streams is part of the machinery: users can choose between raw markers with validity flags and self-attention","core_discovery":"The paper claims to be the first public dataset to pair real optical motion capture of both hands and body with frame-synchronized audio, MIDI, multi-view video, and score data, across a shared core repertoire that crosses two technique categories, three difficulty levels, and two expertise tiers. The measured fingertip trajectories, rather than video-estimated landmarks, are the enabling element: they let a geometry-based fingering labeler assign each note to a finger by detecting which fingertip sits inside the key and lowest in depth at the onset, reaching 94.5% strict precision on a hand-audited subset while covering all notes via a learned imputation tier. A fine-tuning experiment confi","pith_inferences":["The audio-MIDI and motion alignment precision could make SKY-Piano a benchmark for evaluating marker-occlusion imputation: because flagged and imputed versions ship together, anyone can measure how much downstream tasks degrade when synthetic samples are removed.","If the centimeter-scale coordinate-alignment residuals for pre-timecode sessions are representative, users working on finger-level analyses may want to restrict to the hardware-synchronized subset; the paper does not yet quantify this subset's share of the 11 hours.","The tiered confidence of fingering labels suggests a natural curriculum for fingering-estimation models: train on Tier A/B labels, then fine-tune on Tier C, exploiting the source tags rather than treating all labels as equally reliable.","Because the repertoire avoids extended techniques and recent music, a natural extension would add contemporary or jazz material to test whether the measured-motion advantage persists in non-classical idioms."],"forward_implications":["Piano motion research can train and evaluate generation models on measured hand and body trajectories instead of video-derived pseudo motion, avoiding centimeter-scale per-joint noise.","The shared scale and technique exercises across professionals and amateurs make it possible to isolate expertise effects on keystroke and movement kinematics under matched musical conditions.","The fingering labels, released per note with algorithm/imputed/null source tags and explicit ambiguity flags, provide training signal for automatic fingering estimation with known confidence.","The synchronized audio, MIDI, video, and score streams enable joint audio-visual transcription and score-following experiments within one corpus.","The fine-tuning result indicates that models trained on estimated-motion corpora can adapt to measured motion, but that fingertip-definition mismatches require convention-aware evaluation."],"fun_headline_variants":["First piano dataset with directly measured hand and body motion","Piano corpus pairs measured fingertip motion with audio, MIDI, and video","Fingering labels at 94.5% precision from optical motion capture","Real fingertip tracking dataset for piano with synchronized modalities"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the released hand-marker trajectories, after imputation of occluded frames and post-hoc alignment of older sessions, place fingertips accurately enough at the moment of keystrokes that the pseudo-fingering labels and the claim of 'directly measured' motion hold; the paper itself reports median per-trial landmark residuals at the centimeter scale and a 14% median missing-marker rate, so that premise is not guaranteed.","fun_headline_variants_meta":{"raw":{"variants":["First piano dataset with directly measured hand and body motion","Piano corpus pairs measured fingertip motion with audio, MIDI, and video","Fingering labels at 94.5% precision from optical motion capture","Real fingertip tracking dataset for piano with synchronized modalities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1394,"prompt_tokens":705,"completion_tokens":689,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":615}},"tokens_in":449,"tokens_out":689,"duration_ms":7072,"temperature":1.0,"reasoning_tokens":615,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:01:17.207457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the per-note fingering agreement on the three hand-audited pieces using only the flagged (non-imputed) motion, then compare with the released audit labels: if the subset of notes whose labels came from the imputation tier has an agreement far below the reported 94.5%, or if median fingertip z at MIDI onsets shifts by more than a few millimeters when occlusion flags are removed, the central precision claim fails.","supporting_citations":[],"review_version":1}