{"id":"4901624c-3e1f-4453-b193-30a170b9a0e5","arxiv_id":"2505.00525","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of AI-based Parkinson's disease detection across MRI, gait, handwriting, speech, EEG, and multimodal data, undermined by citation errors and recycled content.","lead":"This paper is a survey of machine learning and deep learning methods that detect Parkinson's disease from data like MRI scans, gait, handwriting, speech, and brain waves. It claims to be the first review covering all six data types together, but it is poorly organized and contains substantial citation and content errors.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's benchmark tables are not trustworthy: a handwriting study is cited as an MRI preprocessing method, and the discussion contains sign-language content, so the claimed cross-modal survey fails its central purpose.","rationale":"The paper's stated purpose is to be a reliable, comprehensive reference organizing 347 papers by modality. For that claim to hold, each cited paper must actually concern PD and must be placed in the correct modality table. The Pereira [88] example is directly verifiable from the manuscript alone: the same reference is simultaneously presented as an MRI preprocessing study in Table 2 and as a NewHandPD handwriting study in Table 10, and the cited source is a handwriting paper. This is not a matter of contested scientific judgment; it is a structural failure in the survey's core product. The sign-language material in the discussion and abbreviations reinforces that the writing and selection process were not consistently PD-specific, rather than being an isolated formatting slip. Because the deliverable is the organization and transcription of other works, pervasive citation and modality errors require redoing the survey, not patching individual cells. The reader's REJECT verdict therefore stands, and no adjustment is needed.","tokens_in":53179,"tokens_out":4100,"duration_ms":44062,"concrete_test":"Audit every row of Tables 1-19: fetch the cited reference and verify (a) it is a Parkinson's disease study and (b) its modality matches the table/section where it is placed. Start with Table 2 row 'Pereira et al. [88]': if reference [88] is the SIBGRAPI 2016 handwriting paper, that row is a confirmed misattribution. If more than a handful of audited rows fail, or if any non-PD content remains in the discussion or abbreviations, the benchmark tables cannot be considered reliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II-B's MRI preprocessing table (Table 2) includes the row: Pereira et al. [88], 'Visual Rhythm approaches', dataset HandPD. Reference [88] is Pereira et al., 'Deep learning-aided Parkinson's disease diagnosis from handwritten dynamics' (SIBGRAPI 2016); the same reference reappears in the handwriting section as 'Perai [88]', NewHandPD. A handwriting study cannot be an MRI preprocessing method. This is not a one-cell typo: it shows the modality-classification pipeline can place a paper in the wrong section and cite it under a different author name elsewhere. The same unreliability appears in Section X (Discussion), which discusses 'gesture localization within realistic, uncut, and extended videos', 'gesture captioning', and 'life log devices', and the abbreviations list contains ADDSL (Danish Sign Language), HSL (Hong Kong Sign Language), and FPHA (First Person Hand Action), none of which are PD datasets. These items violate the paper's own exclusion criteria in Section I-C, which exclude papers that 'only mention PD briefly or indirectly'. The central claim is to provide a comprehensive, accurate cross-modal survey with benchmark tables and citations for each modality; if citations are assigned to the wrong modality and non-PD content remains in the final text, the selection and attribution process is unreliable, and the tables and gap analysis cannot be used as a reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of Parkinson's disease (PD) recognition systems across multiple data modalities: MRI, video/pose-based gait analysis, sensor-based gait, handwriting, speech, EEG, other single modalities, and multimodal fusion. Based on 347 articles from 2014–2024, the paper proposes to be the first comprehensive review of PD detection across these modalities, and it provides benchmark tables of datasets and accuracy figures for each modality, along with per-modality discussions of preprocessing, state-of-the-art methods, limitations, and future trends. The central claim is that the paper is a reliable and comprehensive cross-modal reference for researchers developing PD diagnostic systems.","tokens_in":53326,"tokens_out":2747,"duration_ms":27932,"significance":"If the claims were accurate, a cross-modal survey with consolidated benchmark tables would be a useful resource for a community that currently lacks such a reference. The paper is organized by modality, includes a large number of references, and attempts to cover data collection, preprocessing, classifiers, and accuracy. However, the value of a survey rests entirely on the correctness of its paper selection, modality classification, and citation attributions. The manuscript contains several concrete errors that undermine this foundation: a handwriting-dynamics paper is listed as an MRI preprocessing method, non-PD sign-language and gesture-recognition content appears in the discussion and abbreviations, and the stated contribution list promises coverage of modalities (RGB, depth, audio, EMG) that are not actually treated as separate sections. These are not cosmetic issues; they directly invalidate the central claim of a comprehensive, accurate, and reliable review. There is no code, proof, or other machine-checkable artifact to offset these attribution problems.","major_comments":[{"comment":"Table 2 lists Pereira et al. [88] as an MRI preprocessing method with 'Visual Rhythm approaches' on dataset HandPD and CNN classification. Reference [88] is C.R. Pereira et al., 'Deep learning-aided Parkinson's disease diagnosis from handwritten dynamics' (SIBGRAPI 2016), which is a handwriting-dynamics study, not an MRI study. The same reference appears in Table 10 (handwriting datasets) as 'Perai [88]' with NewHandPD. A handwriting study cannot be an MRI preprocessing method, and the inconsistent author-name rendering compounds the problem. This error is load-bearing because the paper's stated contribution includes 'Creation of Benchmark Dataset Tables' with accurate citations for each modality; if a table intended for MRI preprocessing is populated with non-MRI papers, the tables cannot be trusted as a reference.","section":"Section II-B, Table 2"},{"comment":"The Discussion chapter discusses 'gesture localization within realistic, uncut, and extended videos', 'gesture captioning', and 'life log devices', none of which are PD-specific topics. The Abbreviations list contains ADDSL (Annotated Dataset for Danish Sign Language), HSL (Hong Kong Sign Language), and FPHA (First Person Hand Action), none of which are PD datasets or PD-related terms. This content is inconsistent with the paper's own inclusion/exclusion criteria in Section I-C, which exclude papers that 'only mention PD briefly or indirectly'. The presence of sign-language and general gesture-recognition material indicates that the manuscript has not been properly filtered for PD relevance, and it directly contradicts the claim of a comprehensive and accurate PD-focused survey.","section":"Section X (Discussion) and Abbreviations"},{"comment":"The Contributions section states that 'For the first time, this study systematically examines the advancements in multiple data modalities used in PD detection systems, including RGB, skeleton, depth, audio, EMG, EEG, and multimodal fusion.' The abstract similarly promises coverage of these modalities. However, the actual body of the paper does not contain dedicated sections for RGB, depth, audio, or EMG as standalone modalities; the modality sections are MRI, video/pose, sensor, handwriting, speech, EEG, other single modalities, and multimodal fusion. The claim of what is covered is therefore not supported by the manuscript's content, and the promised scope is not delivered.","section":"Section I-E (Contribution) and Abstract"}],"minor_comments":[{"comment":"The organization paragraph lists sections I, III, IV, V, VI, VII, IX, and XI, but omits Section VIII ('Other Single Modalities'), which does appear in the body. The reading of the paper's structure would be easier if all sections were listed.","section":"Section I-G (Organization)"},{"comment":"Several rows in Table 2 list datasets that are not MRI datasets for PD, such as Noor et al. [89] with ADNI, OASIS, and MIRIAD (Alzheimer's datasets). If these are included as MRI-based PD preprocessing work, the relevance to PD should be explicitly justified, and if not, the rows should be removed.","section":"Table 2"},{"comment":"Many abbreviations listed (e.g., RTDPDS, SMKD, SSC-DNN, HDCAM) do not appear to be used in the text, while other abbreviations used in the text (e.g., MDS-UPDRS) are inconsistently expanded. The list should be pruned and aligned with the actual content.","section":"Abbreviations"},{"comment":"The description of Pereira et al. [59] as a review of 'sleep dysfunction in PD' does not match the reference list entry, which is titled 'A survey on computer-assisted Parkinson's disease diagnosis'. Please verify and correct this description.","section":"Section I-B (Existing PD Detection Survey Papers)"},{"comment":"The tables contain inconsistent formatting and incomplete entries, such as missing years (e.g., Table 1 rows for 'DS et al [21]' and 'PPMI et al. [78],[79]'), inconsistent author name spellings (e.g., 'Perai' vs. 'Pereira'), and unexpanded abbreviations (e.g., 'DMFEN'). A careful proofreading pass is needed throughout.","section":"Various tables"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to contain substantial recycled content from a different survey topic (sign-language and gesture recognition), as evidenced by the ADDSL, HSL, and FPHA abbreviations and the gesture-localization discussion in Section X. This is a scope and provenance concern that goes beyond ordinary presentation issues; the editor may wish to verify whether the manuscript text was properly assembled from materials intended for this submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: nice idea, poor execution. The paper compiles a wide range of AI-based PD detection literature into per-modality tables, which would be a useful starting point if the entries were reliable. But they are not, and there are structural signs that the article-selection and attribution pipeline is broken.\n\nWhat it does well: it pulls together roughly 347 references across MRI, video/pose, sensor, handwriting, speech, EEG, and multimodal fusion, with tables listing datasets, methods, and reported accuracies. Coverage extends to 2024-2025, and the authors' self-citations are not egregious for a survey.\n\nWhere it falls apart: in Section II-B, Table 2 lists Pereira et al. [88] as an MRI preprocessing method using the HandPD dataset. Reference [88] is a handwriting-dynamics paper (SIBGRAPI 2016), and the same reference later appears as \"Perai\" in the handwriting section. A handwriting study cannot be an MRI preprocessing method; this is not a one-cell typo but evidence that papers are being assigned to categories without sufficient checking.\n\nThe Discussion (Section X) wanders into \"gesture localization within realistic, uncut, and extended videos\", \"gesture captioning\", and \"life log devices\", none of which are about PD. The abbreviations list includes ADDSL (Danish Sign Language), HSL (Hong Kong Sign Language), and FPHA (First Person Hand Action) — all unrelated to Parkinson's disease. This suggests the text may have been recycled from a gesture-recognition survey without proper adaptation, violating the paper's own exclusion criteria.\n\nThe novelty claim is also thin. The paper says it is the first to cover six modalities, but prior surveys (Loh 2021, Rana 2022, Shaban 2023, Zhao 2025) already span multiple overlapping modalities. Adding one or two modalities to a multimodal survey is incremental, not a first.\n\nThe reader's REJECT verdict is fair. The problems are pervasive, not isolated; if the categorization system can put a handwriting study in the MRI section, the benchmark tables cannot be trusted as a reference. The unrelated discussion further suggests the paper needs substantial reworking, not light revision.\n\nWho might still use it: someone doing a quick, shallow scan for dataset names, with the understanding that every entry must be verified against the original source. As a citable reference for a research paper, it is not safe.\n\nMy recommendation: desk reject. This is not ready for peer review as a reliable survey; the authors would need to redo the article selection and attribution process from the ground up.","headline":"Broad PD-detection survey whose benchmark tables are not trustworthy: a handwriting paper is cited as MRI preprocessing and the discussion contains sign-language content, so the paper does not deliver its promised comprehensive reference.","tokens_in":53994,"tokens_out":3985,"would_cite":false,"duration_ms":40837,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review maps Parkinson's disease detection research across six data modalities and claims to be the first to compare them side by side, with benchmark tables of datasets and reported accuracies.","keywords":["Parkinson's disease detection","machine learning","deep learning","multimodal fusion","benchmark datasets","gait analysis","speech analysis","EEG"],"falsifier":"Pick one table, such as the MRI preprocessing table, and look up each cited source: a handwriting-study citation appearing in that section would fail this check, and finding a comparable share of misplaced references would show the tables cannot be trusted as a modality guide.","tokens_in":52868,"feed_emoji":"🧠","tokens_out":4994,"duration_ms":48924,"temperature":0.7,"pith_summary":"This paper is a survey that tries to establish a complete map of machine-learning and deep-learning approaches to detecting Parkinson's disease across six data modalities: MRI, video-based pose and gait, wearable sensors, handwriting, speech, and EEG, plus multimodal combinations. It claims to be the first review to cover all six modalities side by side, and it provides benchmark tables that list datasets, methods, and reported accuracies for each modality. If the map is reliable, a researcher could use it to choose a modality, a dataset, and a starting model without re-searching the literature. The paper's own conclusion is that single-modality systems are mature but fragmented, and that multimodal fusion and continuous, real-time recognition are the open problems.","feed_headline":"Six modalities, 347 papers, one Parkinson's detection map","feed_subtitle":"Benchmark tables for MRI, gait, sensor, handwriting, speech, and EEG show where the field stands and where fusion is needed.","key_machinery":"The central object is the modality-structured benchmark table: each table gathers datasets, class counts, sample sizes, feature-extraction method, classifier, reported accuracy, and stated limitations for one modality. These tables do the paper's main work by turning a scattered literature into rows that can be scanned across modalities, and they are also where the paper's reliability claims live.","core_discovery":"The central claim is that Parkinson's disease detection research is best understood modality by modality, and that a side-by-side view across MRI, gait pose, gait sensors, handwriting, speech, EEG, and multimodal fusion reveals both a common trajectory and modality-specific bottlenecks. The authors argue that this is the first systematic review to cover six data modalities at once, and they support the claim with benchmark tables that compile datasets, sample sizes, classifiers, feature-extraction approaches, and reported accuracies for each track. Read on its own terms, the paper establishes a working reference map of the field, identifies the shift from handcrafted features to deep learning as the dominant trend, and points to small datasets and the absence of robust multimodal frameworks as the main barriers to clinical use.","pith_inferences":["The 'first to cover six modalities' claim is conditional on the 2014-2024 window and on the chosen databases; a different search window could yield earlier or different multimodal surveys.","If the benchmark tables were cleaned and released as structured data, they could become a living registry that researchers update when new Parkinson's results appear.","The presence of sign-language and gesture-recognition material in the discussion and abbreviations suggests part of the text was adapted from another survey; that makes the surrounding synthesis less reliable even where the tables are accurate.","A reader who wants to use this map should re-verify any specific accuracy figure against the original paper before citing it in a comparison."],"forward_implications":["A new researcher can use the modality tables to select a dataset and a baseline model without repeating the literature search.","The reported accuracies should be read as study-specific, not comparable across rows, because datasets and protocols differ.","Multimodal fusion is consistently named as the route to higher diagnostic accuracy, but the paper finds no robust benchmark framework for it yet.","Across every modality, the dominant pattern is a shift from handcrafted features to deep learning, with small datasets the main brake on progress.","Continuous, real-time recognition of Parkinson's symptoms remains an open problem across all modalities."],"supporting_citations":[{"why":"A prior review that supplies many of the 'latest accuracy' figures printed in the benchmark tables.","marker":"[77]"},{"why":"A prior deep-learning review of Parkinson's diagnosis that the paper positions itself as extending to more modalities.","marker":"[8]"},{"why":"A recent survey of machine and deep learning for Parkinson's recognition whose narrower modality coverage motivates the new review.","marker":"[65]"},{"why":"A gait-analysis survey that defines the baseline for the video and pose section.","marker":"[63]"},{"why":"The flagship longitudinal Parkinson's dataset used across the MRI and multimodal tables.","marker":"[78]"},{"why":"The standard gait force-plate dataset that underlies many sensor-modality performance entries.","marker":"[155]"},{"why":"The handwriting database used as the core benchmark in the handwriting section.","marker":"[197]"},{"why":"The voice dataset that anchors the speech-modality benchmark tables.","marker":"[242]"},{"why":"A multimodal clinical dataset cited for the highest reported accuracy in the multimodal fusion section.","marker":"[338]"},{"why":"An EEG biomarker study whose spectral-amplitude features anchor the EEG section's machine-learning discussion.","marker":"[47]"}],"fun_headline_variants":["347 papers, 6 modalities: Parkinson's detection benchmarked","Parkinson's detection: 6 modalities, 1 review, many gaps","Multimodal PD review: where MRI, gait, and speech stand","The PD detection landscape across six data modalities","From MRI to EEG: 347 studies map Parkinson's detection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the paper correctly sorted each of its 347 sources into the right modality and faithfully transcribed their datasets and accuracies; if that sorting or transcription is wrong, the cross-modal map misleads rather than guides.","fun_headline_variants_meta":{"raw":{"variants":["347 papers, 6 modalities: Parkinson's detection benchmarked","Parkinson's detection: 6 modalities, 1 review, many gaps","Multimodal PD review: where MRI, gait, and speech stand","The PD detection landscape across six data modalities","From MRI to EEG: 347 studies map Parkinson's detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000669,"raw_usage":{"total_tokens":3054,"prompt_tokens":950,"completion_tokens":2104,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":2017}},"tokens_in":566,"tokens_out":2104,"duration_ms":14731,"temperature":1.0,"reasoning_tokens":2017,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:39:08.305251+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pick one table, such as the MRI preprocessing table, and look up each cited source: a handwriting-study citation appearing in that section would fail this check, and finding a comparable share of misplaced references would show the tables cannot be trusted as a modality guide.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The voice dataset that anchors the speech-modality benchmark tables."}],"review_version":1}