{"id":"932a835d-5dd9-40c6-abb5-72435333850c","arxiv_id":"2504.19514","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new figure skating dataset and benchmark, FSAnno and FSBench, evaluates and improves multimodal large language models on technical and artistic understanding of figure skating.","lead":"This paper introduces FSAnno, a large multimodal dataset for figure skating with per-element and whole-performance annotations, and FSBench, a benchmark that tests AI models on figure skating knowledge, action recognition, assessment, and commentary. Initial tests show current multimodal large language models understand figure skating poorly, and instruction-tuning on FSAnno improves their motion descriptions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SkateLLM's reported gain rests on GPT-4-generated captions and a GPT-3.5-based metric with no human validation; the improvement may reflect alignment with the annotation pipeline rather than genuine figure-skating understanding.","rationale":"The reader's weakest assumption is that the automatic annotation pipeline (4DHumans, HRNet, Whisper, GPT-4 synthesis) is unvalidated. I agree that this is a serious weakness, but I would sharpen it: the load-bearing failure mode is not only that annotations may contain errors, but that the central demonstration of improvement is circular in a specific way. Instruction-tuning data are generated by GPT-4; the evaluation metric is built on GPT-3.5-turbo event extraction and matching; no human ground truth or preference judgment is used anywhere in Table 4. Therefore the observed F1 improvement does not distinguish genuine technical understanding from stylistic mimicry of the model family that produced both the training targets and the scoring rubric. The paper's own text supports this concern: §4.1 says GPT-4 generated data in batches according to templates, and §3.5 introduces AutoDQ as the open-ended evaluation metric. The result is that the strongest empirical claim, 'our fine-grained, multi-modal annotations significantly enhance the LLMs' capabilities,' is underdetermined by the evidence. This does not invalidate the dataset or the benchmark as a resource; it means acceptance should be conditional on validation of both the annotations and the evaluation metric against human experts, plus release of the data and code for independent reproduction. I therefore keep the reader's CONDITIONAL verdict, with partial agreement because the reader identified the pipeline-validation issue but did not foreground the GPT-to-GPT circularity in the improvement claim.","tokens_in":13126,"tokens_out":2314,"duration_ms":26641,"concrete_test":"Take a random sample of 100 element-motion pairs from FSBench. Have two ISU-trained technical specialists independently write gold captions and then blind-rate the correctness of captions from Motion-GPT and SkateLLM. Compare (a) human preference rates against the AutoDQ F1 gap in Table 4, and (b) AutoDQ scores when the ground truth is a human expert caption instead of the GPT-4-generated caption. If human raters do not prefer SkateLLM by a margin commensurate with the reported F1 gain, or if AutoDQ rates expert-written gold captions lower than GPT-4-generated captions, the enhancement claim is a pipeline artifact rather than evidence of improved figure-skating understanding.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim has two parts: existing models underperform on artistic-sports understanding, and fine-grained FSAnno annotations significantly enhance LLM capabilities. The first part is reasonably supported by the prior-knowledge tests. The second part is not. Training captions in §4.1 are synthesized by GPT-4 from manually crafted templates, with no expert verification of technical accuracy. Evaluation in §3.5 uses AutoDQ, an LLM-based metric whose event extraction and cross-checking are performed by GPT-3.5-turbo. The only reported improvement, Table 4, compares Motion-GPT (F1 7.1) with SkateLLM (F1 38.0) on a single task, element description, with no error bars, no statistical testing, and no human evaluation. The paper itself concedes in §5.2 that it 'focus[es] temporarily on the most fundamental task' and that existing video-based and motion-based LLMs were only partially tested. Consequently, the 30-point F1 gain could arise from SkateLLM learning to imitate GPT-4's template language, which the same LLM family's evaluation metric would reward regardless of whether the caption is technically accurate for the actual motion. This is a correctness risk in the argument, not a disagreement with community consensus: the benchmark may still be useful, but the causal claim that fine-grained multimodal annotations 'significantly enhance' LLM capabilities is not established by the presented evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FSAnno, a large-scale figure skating dataset with multi-level annotations spanning prior knowledge, individual actions, and whole performances, and FSBench, a benchmark for evaluating LLMs on figure skating understanding. The dataset is constructed from 783 competition performances with RGB, motion, skeleton, audio, and text modalities, and includes official judge reports as scoring ground truth. The authors also present SkateLLM, an instruction-tuned variant of Motion-GPT trained on FSAnno, and report that it substantially outperforms the base model on a motion-captioning task. The paper concludes that fine-grained multi-modal annotations significantly enhance LLM capabilities in artistic sports understanding.","tokens_in":13400,"tokens_out":5861,"duration_ms":53613,"significance":"If the annotations are validated, FSAnno/FSBench would fill a clear gap in sports-understanding benchmarks, which currently focus on ball sports and single-task figure skating recognition or scoring. The inclusion of multiple modalities, multiple task levels, and official judge reports is a strength, and the preliminary result that current LLMs score poorly on figure skating knowledge is an interesting finding. However, the central causal claim that fine-grained annotations 'significantly enhance' LLM capabilities is not established by the presented evidence, because the evaluation pipeline is circular and the empirical demonstration is limited to a single task. The benchmark has potential value, but the paper needs stronger validation and more thorough experiments before the claim can be accepted.","major_comments":[{"comment":"Section 3.3 (Data Annotation) and Section 3.4: the entire annotation pipeline is automatic—4DHumans for motion extraction, HRNet for skeletons, Whisper for speech-to-text, and GPT-4 for generating summary comments—with no reported human validation, inter-annotator agreement, or error analysis. Because FSBench and FSAnno use these annotations as ground truth for all tasks (e.g., action segmentation aligned to video cue boxes, GOE scores, commentary), the reliability of the benchmark under this pipeline is not established. The authors should include a human verification study on a subset (with agreement statistics) or explicitly quantify the error modes of each automatic stage.","section":"§3.3–§3.4"},{"comment":"Sections 3.5 and 4.1: the evaluation metric AutoDQ uses GPT-3.5-turbo for both event extraction and cross-checking, while the SkateLLM training captions are generated by GPT-4 from manually crafted templates (Section 4.1). Consequently, SkateLLM is trained to imitate the annotation pipeline's language, and AutoDQ measures similarity to that same language. The large F1 gain in Table 4 (7.1 to 38.0) may therefore reflect pipeline alignment rather than technically accurate understanding of figure skating; the claim in the abstract that annotations 'significantly enhance the LLMs' capabilities' requires either a human expert evaluation of the generated captions or an evaluation metric independent of the annotation generation process.","section":"§3.5, §4.1, Table 4"},{"comment":"Section 5.2 and Table 4: the reported evidence for the enhancement claim is limited to a single task (element description) and a single baseline (Motion-GPT), with no error bars, no statistical testing, and no human evaluation. The paper itself acknowledges in Section 5.2 that existing video- and motion-based LLMs were only partially tested and that it 'focus[es] temporarily on the most fundamental task.' To support the abstract's general conclusion, results on additional FSBench tasks (e.g., action recognition, performance commentary) and comparisons with at least one video-based MLLM are needed.","section":"§5.2, Table 4"},{"comment":"Introduction (third contribution bullet) states that 'FSBench-Text includes multiple-choice questions with human-annotated explanations,' but Section 3.3 describes using a large language model to process collected commentary and does not describe any human annotation of the explanations. The authors should either correct the description or provide the annotation procedure for these explanations; this discrepancy currently undermines the paper's claim of fine-grained human-quality annotations.","section":"Introduction vs. §3.3"}],"minor_comments":[{"comment":"The model names are inconsistent (e.g., 'GPT3.5-turbo', 'GPT4', 'LLaV A 13B'); please standardize the notation.","section":"Table 3"},{"comment":"The '5 shot strategy' is not described; please clarify how the few-shot examples are selected.","section":"§5.1"},{"comment":"Training details are too sparse (e.g., learning rate, batch size, optimizer, number of steps) to reproduce SkateLLM; please provide these hyperparameters.","section":"§4.2"},{"comment":"The pie chart and stacked bar chart are small and difficult to read; please increase the font sizes.","section":"Figure 2"},{"comment":"The category abbreviations (e.g., 'LS' for long-video scoring) are not defined in the caption; please add a full list of abbreviations.","section":"Table 2"},{"comment":"The paper claims that FSAnno is 'open-access' and points to an official repository, but no download link is provided anywhere in the manuscript; please add the link in the final version.","section":"Abstract/§3.3"}],"recommendation":"major_revision","confidential_remarks":"The dataset has potential to be a useful community resource, but the empirical section falls short of the standard needed to support the paper's central claim of significant capability enhancement. I recommend requiring human validation of the annotation pipeline, a non-circular evaluation, and a broader demonstration across FSBench tasks before acceptance. The contradiction between the claimed human-annotated explanations and the described LLM-based annotation should also be resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi [Name],\n\nRead FSBench. The useful part is real: it is the first figure skating dataset that strings together prior-knowledge questions, element recognition, single-action assessment, performance commentary, and multi-modal inputs (RGB, motion, skeleton, audio, text), with official judge reports as the backbone. That gives it a plausible claim to fill a genuine gap, since sports MLLM benchmarks are almost all ball sports. The prior-knowledge tests (Table 3) also support the modest claim that current LLMs lack figure-skating expertise. If the dataset ships, it will get used.\n\nThe soft spots are mostly in the second half. The annotation pipeline is almost fully automatic (4DHumans, HRNet, Whisper, GPT-4) and there is no human validation, no inter-annotator agreement, no error analysis. For a benchmark, that matters: systematic errors in the automatic annotations become systematic biases in the scores. The SkateLLM experiment is the weakest link. The training captions are GPT-4-generated from hand-made templates, and the evaluation metric, AutoDQ, is itself an LLM (GPT-3.5-turbo) that extracts events from the same kind of language. So the 30-point F1 gain in Table 4 may largely reflect the model learning to imitate GPT-4's template style, which the evaluation metric is predisposed to reward. There are also no error bars or significance tests, and only one task is fully tested; the paper concedes the rest of FSBench is only partially evaluated.\n\nThat said, the benchmark itself is a contribution. It just needs grounding: release the data and code; add human validation on a sample, with agreement scores; show that AutoDQ correlates with human judgment; and add baselines and variance. If the evaluation is tightened, the paper's first claim (models are weak on artistic sports) likely survives. The second claim (fine-grained annotations significantly enhance LLMs) is not yet established.\n\nMy take: worth a serious referee, not a desk reject. The reviewer should ask for the validation work before acceptance. I'd bring it to the reading group mainly to discuss the circularity problem in LLM-based evaluation.\n\nRecommended verdict: major revision.","headline":"A genuinely useful multi-modal figure skating benchmark, but the reported instruction-tuning gain is not convincingly separated from the GPT-4-generated training and evaluation pipeline.","tokens_in":13942,"tokens_out":1578,"would_cite":false,"duration_ms":15870,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that current large language models understand figure skating poorly, and that instruction-tuning on FSAnno's fine-grained, multi-modal annotations substantially improves their performance, with the resulting model…","keywords":["figure skating","artistic sports","benchmark","multimodal LLM","instruction tuning","motion understanding","action quality assessment","dataset"],"falsifier":"Have a panel of human figure-skating officials re-annotate a random sample of FSBench elements: the element category, GOE score, segmentation boundaries, and the alignment of commentary to the correct action. If agreement between the human labels and FSAnno's automatically generated labels is no better than chance or is systematically biased, then model scores on FSBench and SkateLLM's gains could reflect the annotation pipeline rather than genuine understanding of figure skating.","tokens_in":12920,"feed_emoji":"⛸️","tokens_out":6823,"duration_ms":64914,"temperature":0.7,"pith_summary":"The paper introduces FSAnno, a large-scale figure skating dataset with fine-grained annotations at three levels—prior knowledge, individual actions, and whole performances—and FSBench, an evaluation benchmark built from it. The authors' central claim is that current large language models understand artistic sports such as figure skating poorly, and that instruction-tuning on FSAnno's fine-grained, multi-modal annotations substantially improves their ability to describe and analyze skating elements. They support this claim with initial tests showing low accuracy on figure-skating rules and event questions, and with an instruction-tuned model whose captioning quality, measured by the AutoDQ metric, rises sharply compared with its base motion-language model. The intended contribution is a reusable resource: a dataset and benchmark that treat technical execution and artistic expression as connected rather than separate tasks.","feed_headline":"AI stumbles on figure skating's artistry, new benchmark shows","feed_subtitle":"FSBench tests technical and artistic moves; instruction-tuning lifts model caption F1 from 7 to 38.","key_machinery":"The central object is FSAnno combined with FSBench: a dataset and benchmark built from 783 complete performances across eleven Grand Prix and Junior Grand Prix events, spanning four program types and including negative samples from younger and less experienced skaters. The load-bearing mechanism is the multi-source annotation pipeline that fuses official judging reports (element categories, GOE scores, TES/PCS), transcribed commentator speech aligned to action timestamps, and LLM-synthesized summaries into a single multi-level, multi-modal resource. The motion and skeleton modalities, extracted from raw video and stripped of identifying appearance, are what allow the benchmark to evaluate understanding of movement quality while protecting athlete privacy and reducing reliance on prior knowledge of specific competitions. Instruction-tuning a motion-language model on this resource yields SkateLLM, the model whose improved captions demonstrate the dataset's value.","core_discovery":"FSBench is presented as the first benchmark for figure skating that spans technical and artistic understanding at multiple granularities, with FSAnno providing the underlying annotations. The benchmark contains FSBench-Text, roughly 4,200 multiple-choice questions on rules and event facts, and FSBench-Motion, which pairs motion and skeleton data with question-answer sets for six tasks: action recognition, single-action assessment, single-action commentary, action segmentation, whole-performance scoring across seven dimensions, and whole-performance commentary. The paper reports that existing LLMs score far below expert level on the knowledge questions, and that motion-based models without figure-skating instruction tuning produce poor captions. The paper's constructive claim is that fine-grained, multi-modal annotations—grounded in official judge reports, transcribed commentator audio, and LLM-synthesized performance summaries—are sufficient to teach a motion-language model noticeably more professional and technically grounded skating descriptions.","pith_inferences":["Beyond the paper, the same three-source recipe—official score sheets, timestamp-aligned commentator audio, and LLM-synthesized summaries—looks directly transferable to other judged artistic sports such as gymnastics, dance, and synchronized swimming, where the missing ingredient is also multi-level technical plus artistic annotation.","Editorial inference: since the same kind of LLM used to synthesize training templates is also used to score model outputs through AutoDQ, part of SkateLLM's improvement may reflect stylistic alignment with the evaluator rather than deeper skating knowledge; the paper does not test this.","Editorial inference: a human study in which judges rate artistry from motion/skeleton data alone would test whether the identity-free representations preserve the artistic signal FSBench claims to evaluate.","Editorial inference: the paper's headline conclusion is currently demonstrated mainly on the captioning task; the claim about artistic understanding would be better supported if full results on assessment and commentary tasks appear in the released repository."],"forward_implications":["FSBench provides a six-task evaluation protocol that other researchers can use to compare models on figure skating, with a public training/test split and an official benchmark split.","The benchmark's results quantify a gap: current LLMs perform limitedly on figure skating rules and event facts, so progress in artistic-sports knowledge is directly measurable.","Instruction-tuning on FSAnno improves captioning F1 from about 7% to about 38% under AutoDQ, indicating that fine-grained multi-modal annotations, not just more video data, drive the gain.","Because FSBench includes identity-free motion and skeleton data, it enables experiments on whether models can judge artistry without seeing the athlete's face or costume.","The annotation design is extensible to anomaly detection and motion generation tasks, which the paper lists as supported future uses."],"supporting_citations":[{"why":"supplies speech-to-text transcription of commentator audio used for action-level commentary annotations","marker":"[30]"},{"why":"extracts the motion modality from raw competition video","marker":"[9]"},{"why":"estimates the skeleton modality used alongside motion in FSBench-Motion","marker":"[33, 39]"},{"why":"provides the base motion-language model that instruction tuning turns into SkateLLM","marker":"[15]"},{"why":"defines the AutoDQ metric used to evaluate open-ended caption and commentary quality","marker":"[34]"},{"why":"the low-rank adaptation technique used during instruction tuning to mitigate catastrophic forgetting","marker":"[13]"},{"why":"supplies the data-centric training rationale behind the two-phase multimodal tuning strategy","marker":"[12]"},{"why":"the prior temporal-segmentation dataset whose granularity and annotation standards FSBench aligns with for segmentation tasks","marker":"[24]"}],"fun_headline_variants":["Figure skating benchmark puts AI's artistic sense on thin ice","FSBench: AI stumbles on skating's blend of art and skill","New benchmark shows AI still can't grasp figure skating's artistry","Figure skating's artistic complexity trips up AI in new test","Benchmark: AI lags on technical and artistic figure skating tasks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's scores and the reported improvement of SkateLLM depend on the assumption that the automatically produced labels—extracted movement data, transcribed commentary, and machine-written summaries—agree with what human judges and commentators actually said and scored.","fun_headline_variants_meta":{"raw":{"variants":["Figure skating benchmark puts AI's artistic sense on thin ice","FSBench: AI stumbles on skating's blend of art and skill","New benchmark shows AI still can't grasp figure skating's artistry","Figure skating's artistic complexity trips up AI in new test","Benchmark: AI lags on technical and artistic figure skating tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000953,"raw_usage":{"total_tokens":4044,"prompt_tokens":908,"completion_tokens":3136,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":3059}},"tokens_in":524,"tokens_out":3136,"duration_ms":24714,"temperature":1.0,"reasoning_tokens":3059,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:50:14.022597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a panel of human figure-skating officials re-annotate a random sample of FSBench elements: the element category, GOE score, segmentation boundaries, and the alignment of commentary to the correct action. If agreement between the human labels and FSAnno's automatically generated labels is no better than chance or is systematically biased, then model scores on FSBench and SkateLLM's gains could reflect the annotation pipeline rather than genuine understanding of figure skating.","supporting_citations":[{"cited_title":"Ro- bust speech recognition via large-scale weak supervi- sion","cited_arxiv_id":null,"evidence_quote":"supplies speech-to-text transcription of commentator audio used for action-level commentary annotations"},{"cited_title":"Motiongpt: Human motion as a foreign language","cited_arxiv_id":null,"evidence_quote":"provides the base motion-language model that instruction tuning turns into SkateLLM"},{"cited_title":"Tempo- ral segmentation of fine-gained semantic action: A motion-centered figure skating dataset","cited_arxiv_id":null,"evidence_quote":"the prior temporal-segmentation dataset whose granularity and annotation standards FSBench aligns with for segmentation tasks"}],"review_version":1}