{"id":"9c9ee5ae-0b2d-41af-aaf5-bcfee543744f","arxiv_id":"2607.01838","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces Learnable Gate (LG) methods that improve group-level sparsity in counterfactual explanations for rehabilitation IMU data while preserving validity and smoothness compared to channel-level baselines.","lead":"The paper proposes a two-stage framework with Shapley-Adaptive ranking followed by Learnable Gate methods to create group-based counterfactual explanations for IMU time-series data in rehabilitation. This targets making AI suggestions more aligned with how clinicians reason about muscle groups rather than individual sensors.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flags the group-choice assumption as the key external validity point but the central empirical claim (LG improves sparsity without harming other metrics) has no visible internal defect from the abstract. Full-text verification of the concrete_test would be the natural next step; no adjustment to UNVERDICTED is warranted on the basis of an identifiable flaw in the argument structure.","tokens_in":1693,"tokens_out":284,"duration_ms":15486,"concrete_test":"Re-run the LG vs M-CELS comparison on KneE-PAD using the exact reported metrics (sparsity, validity, smoothness, efficiency) after confirming the group definitions and optimization procedure match the manuscript; if the sparsity delta disappears or validity drops below baseline, the headline claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract presents an empirical claim of improved group sparsity via LG gates while preserving validity/smoothness/efficiency on the KneE-PAD dataset, with the two-stage SA-then-LG structure explicitly motivated by the failure of ranking alone. No internal inconsistency, circularity, or unstated assumption that would invalidate the reported comparison is visible from the given description. The reader's weakest assumption (predefined groups being appropriate) is a domain-validity question rather than a load-bearing flaw in the reported experimental result.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces a two-stage framework for generating group-based counterfactual explanations for multivariate time-series data in rehabilitation, specifically for IMU sensor data. The first stage uses Shapley-Adaptive (SA) group ranking to select relevant groups, which maintains counterfactual validity but does not enforce sparsity. The second stage introduces Learnable Gate (LG) methods with trainable per-group gates jointly optimized with perturbation masks. Experiments on the KneE-PAD dataset show that LG improves modality-group sparsity over the channel-level M-CELS baseline while maintaining or improving validity, temporal smoothness, and generation efficiency. Exercise-specific analyses indicate that the group-structured counterfactuals provide concise, muscle-level guidance aligned with clinical reasoning.","tokens_in":1762,"tokens_out":453,"duration_ms":28583,"significance":"If the results hold, this work contributes to making counterfactual explanations more interpretable and actionable in clinical domains by aligning with how experts reason about semantic groups rather than individual channels. This could be particularly valuable in rehabilitation movement analysis, where biomechanically coherent explanations are needed. The motivation for the two-stage approach based on the limitations of ranking alone is a strength, as is the focus on a real-world dataset.","major_comments":[{"comment":"Abstract and Experiments section: The central claim that LG 'substantially improves modality-group sparsity compared to the channel-level M-CELS baseline while maintaining or improving validity, temporal smoothness, and generation efficiency' is presented without any quantitative metrics (e.g., sparsity ratios or validity scores), statistical tests, ablation details, or error analysis. This absence makes it impossible to evaluate the magnitude or reliability of the reported gains, which is load-bearing for the paper's primary empirical contribution.","section":"Abstract and Experiments"}],"minor_comments":[{"comment":"The abstract would be strengthened by including at least one or two key numerical results to support the improvement claims.","section":"Abstract"},{"comment":"Ensure the joint optimization procedure for the LG gates (including any sparsity-inducing terms) is specified with sufficient mathematical detail to support reproducibility.","section":"Methods"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback highlighting the need for quantitative support of our central empirical claims. We address the major comment below and commit to revisions that improve transparency without altering the core contributions.","responses":[{"response":"We agree that the abstract would benefit from explicit quantitative metrics to allow readers to assess the magnitude of improvements. The experiments section reports results via tables comparing LG against M-CELS on sparsity (modality-group level), validity, smoothness, and efficiency metrics, with exercise-specific breakdowns. To directly address the concern, we will revise the abstract to include key numerical results (e.g., sparsity ratios and validity scores) and add statistical tests, expanded ablation details, and error analysis to the experiments section where they strengthen the presentation.","revision_made":"yes","referee_comment":"[Abstract and Experiments] Abstract and Experiments section: The central claim that LG 'substantially improves modality-group sparsity compared to the channel-level M-CELS baseline while maintaining or improving validity, temporal smoothness, and generation efficiency' is presented without any quantitative metrics (e.g., sparsity ratios or validity scores), statistical tests, ablation details, or error analysis. This absence makes it impossible to evaluate the magnitude or reliability of the reported gains, which is load-bearing for the paper's primary empirical contribution."}],"tokens_in":1363,"tokens_out":286,"duration_ms":14078,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper's core move is to replace channel-by-channel counterfactuals with ones that act on predefined semantic groups in multivariate IMU data. They first try Shapley-Adaptive ranking to pick groups but see it fails to enforce sparsity, then add trainable per-group gates that are optimized jointly with the perturbation masks. On the KneE-PAD dataset the LG version produces sparser group-level outputs than the M-CELS baseline while holding validity, smoothness, and speed roughly constant, and the group-structured results line up better with how clinicians describe corrective movements.\n\nThe approach is straightforward and directly targets the mismatch between how most counterfactual methods work and how domain experts reason. The motivation for moving beyond ranking is clear from their description.\n\nThe main limitation is that the abstract gives no numbers, confidence intervals, or ablation tables, so the size and reliability of the gains are hard to judge from what is shown. Everything also depends on the groups being fixed in advance; if those semantic partitions do not match the clinically relevant structure in other settings, the sparsity benefit may not transfer. Only one dataset is mentioned.\n\nThis is useful reading for people working on explainable models for sensor-based rehab or similar high-dimensional time-series where experts already think in grouped terms. It is not a general theoretical advance but a practical adaptation that could matter in applied settings.\n\nI would send it to peer review. The idea is coherent and the problem is real; a referee can check the quantitative claims and the sensitivity to group definitions.","headline":"The paper adds learnable per-group gates to counterfactual generation for IMU time-series so explanations align with muscle-group thinking in rehab, and the two-stage SA-then-LG setup looks like a reasonable fix for sparsity.","tokens_in":2229,"tokens_out":390,"would_cite":false,"duration_ms":20765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Learnable per-group gates generate sparser counterfactual explanations for time-series rehabilitation data while preserving validity and smoothness.","keywords":["counterfactual explanations","time-series classification","rehabilitation data","group-based sparsity","IMU sensors","explainable AI"],"falsifier":"An experiment on the same KneE-PAD dataset or a comparable rehabilitation IMU collection that finds no statistically significant gain in group-level sparsity metrics for the learnable-gate method over the channel-level baseline would falsify the performance claim.","tokens_in":2612,"feed_emoji":"🏃","tokens_out":689,"duration_ms":22306,"temperature":0.7,"pith_summary":"The paper develops a two-stage method to produce counterfactual explanations for multivariate time-series models that operate on semantic feature groups rather than individual sensor channels. In rehabilitation IMU data, clinicians reason about movements through muscle groups and joint segments, yet standard counterfactual methods scatter changes across channels and produce hard-to-interpret outputs. The first stage ranks groups with Shapley values; the second stage introduces trainable per-group gates that are optimized together with the perturbation mask to select entire groups at once. Experiments on the KneE-PAD dataset show the resulting explanations are substantially sparser at the group level than channel-wise baselines, with no loss in validity, temporal smoothness, or speed. The group-level outputs supply concise, muscle-specific corrective suggestions that match the way clinicians actually analyze motion.","feed_headline":"Learnable gates sparsify group counterfactuals in rehab time series","feed_subtitle":"The method keeps validity and smoothness while producing explanations that match how clinicians group muscles and joints.","key_machinery":"Learnable Gate (LG) methods that incorporate trainable per-group relevance gates jointly optimized with perturbation masks to enforce group-level sparsity during counterfactual generation.","core_discovery":"The central claim is that Learnable Gate methods, which add trainable per-group relevance gates jointly optimized with perturbation masks, achieve markedly higher modality-group sparsity than channel-level baselines on rehabilitation IMU time series while maintaining or improving validity, temporal smoothness, and generation efficiency; exercise-specific results further indicate that the resulting group-structured counterfactuals supply concise muscle-level guidance aligned with clinical reasoning.","pith_inferences":["The same gating approach could transfer to other multi-channel time-series settings such as activity recognition or physiological monitoring where domain experts use grouped features.","A follow-up study could measure whether the generated group explanations actually change how clinicians prescribe corrections in a real rehabilitation session.","If the predefined groups prove suboptimal, the framework could be extended to learn the groupings themselves rather than treating them as fixed inputs."],"forward_implications":["Counterfactuals become concise enough to supply muscle-level corrective instructions for specific rehabilitation exercises.","Interpretability improves in any domain where experts already think in terms of semantic feature groups rather than raw sensor channels.","The two-stage process first ranks groups and then applies learnable selection, separating ranking from sparsity enforcement.","Generation remains efficient while group sparsity rises, so the method scales to high-dimensional multi-sensor recordings."],"fun_headline_variants":["Learnable gates boost group sparsity in rehab IMU counterfactuals","Per-group gates increase sparsity in rehabilitation time-series CEs","LG methods sparsify modality groups in rehab counterfactual generation","Group-structured gates yield concise muscle-level rehab explanations"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That the predefined semantic groupings of features into muscle groups and joint segments are the right granularity at which to enforce sparsity so that explanations become biomechanically coherent and clinically useful.","fun_headline_variants_meta":{"raw":{"variants":["Learnable gates boost group sparsity in rehab IMU counterfactuals","Per-group gates increase sparsity in rehabilitation time-series CEs","LG methods sparsify modality groups in rehab counterfactual generation","Group-structured gates yield concise muscle-level rehab explanations"]},"model":"grok-4.3","cost_usd":0.006529,"raw_usage":{"total_tokens":3045,"prompt_tokens":652,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":65287000,"prompt_tokens_details":{"text_tokens":652,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2331,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":652,"tokens_out":62,"duration_ms":15843,"temperature":1.0,"reasoning_tokens":2331,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-03T17:51:49.553394+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment on the same KneE-PAD dataset or a comparable rehabilitation IMU collection that finds no statistically significant gain in group-level sparsity metrics for the learnable-gate method over the channel-level baseline would falsify the performance claim.","supporting_citations":[],"review_version":1}