REVIEW 3 major objections 6 minor 74 references
A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proposes Rehab-Pile, an archive of 60 skeleton-motion datasets, and reports that the lightweight LITEMV model wins the benchmark on both accuracy and efficiency.
desk verdict Useful benchmark archive with a reproducible protocol, but the model ranking needs re-running with validation-based checkpoint selection before trusting the LITEMV claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The instrument that carries the argument is the Rehab-Pile benchmark protocol. Nine public motion repositories are converted into 60 per-exercise datasets (39 classification, 21 regression), each resampled to its average sequence length, normalized with min-max statistics computed only on training folds, and split by cross-subject folds that keep patients out of the training set whenever the repository marks them unhealthy. All nine models share one training configuration (1,500 epochs, batch size 64, learning-rate reduction on plateau, checkpoint chosen by training loss) and are compared with the Multi-Comparison Matrix, which reports average performance differences, win-tie-loss counts, and paired signed-rank p-values. The winning candidate, LITEMV, is a three-block convolutional network built on depthwise separable convolutions, multiplexed filters, hand-crafted trend and peak filters, and exponentially increasing dilation, which together keep its parameter and FLOP counts low while preserving both temporal and spatial information.
What would settle it
Re-run the full protocol on a subset of Rehab-Pile (or the whole archive) selecting checkpoints by a held-out validation fold instead of training loss, and compare the resulting average ranks; if LITEMV's regression wins and classification edge vanish or flip, the paper's conclusion is an artifact of checkpoint selection.
Extended reading notes
Core claim
The paper claims that a carefully unified benchmark changes how rehabilitation-motion models should be compared and which architecture should be the default. Across 60 datasets built from nine public repositories, the LITEMV model—a lightweight convolutional architecture with depthwise separable convolutions, hand-crafted filters, and increasing dilation—achieves the best average rank on both regression (average MAE 5.35 and RMSE 6.43 across 21 datasets) and classification (77.37% average accuracy across 39 datasets), while also having the lowest FLOPs and lowest parameter counts across all 60 tasks. On regression, the advantage over every competitor is statistically significant by the paper's paired signed-rank comparisons; on classification, LITEMV edges VanTran (77.37% vs 77.22% average accuracy) without a statistically significant difference. The paper also reports that STGCN, a graph-convolution model widely used in rehabilitation regression, is the weakest classifier of the nine in this protocol.
Load-bearing premise
The rankings assume that choosing the checkpoint with the lowest training loss gives the best test performance, since no validation set is used in model selection.
Editorial extensions
If this is right
- New rehabilitation-motion studies can report results on Rehab-Pile under a known protocol, making accuracy and MAE figures comparable across papers for the first time.
- LITEMV, especially as a five-run ensemble, becomes a credible default architecture for skeleton-based quality assessment, requiring little compute.
- Researchers using graph-convolution models for rehabilitation should not assume they are the best choice for classification, since STGCN ranked last on the 39 classification datasets in this setup.
- The 60-dataset archive gives the community a single testbed to stress future models across heterogeneous exercise types, joint formats, and label distributions.
- Because the best average classification accuracy is only 77.37%, the benchmark still leaves clear headroom for architectures that handle class imbalance better.
Reading between the lines
- The protocol's checkpoint selection by training loss is a candidate source of bias; re-running the leaderboard with a held-out validation set for model selection would test whether LITEMV's margins persist.
- A leave-one-repository-out analysis would tell whether LITEMV's win is driven by particular source repositories or generalizes across the archive.
- The benchmark protocol itself could be reused to evaluate new inputs, such as 2D pose estimates from consumer video rather than depth-camera skeletons, without changing the models.
- Because accuracy trails balanced accuracy by roughly 7 percentage points, methods that explicitly target class imbalance may move the leaderboard more than architectural changes alone.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Rehab-Pile, a unified archive of 60 skeleton-based rehabilitation datasets (39 classification, 21 extrinsic regression) aggregated from nine public repositories. It proposes a benchmarking protocol with cross-subject train/test folds, a unified training configuration, and public release of code and data. Nine deep learning architectures (FCN, H-Inception, LITEMV, DisjointCNN, ConvLSTM, MotionGRU, VanTran, ConvTran, STGCN) are compared on classification and regression tasks using accuracy, balanced accuracy, MAE, RMSE, and efficiency metrics. The authors conclude that LITEMV offers the best trade-off between performance and efficiency, particularly for regression, while classification results are close between LITEMV and VanTran.
Significance. If the benchmark protocol is valid, this is a valuable community resource: it provides a standardized evaluation suite for skeleton-based rehabilitation assessment, with publicly available datasets and code, and it compares architectures under a common protocol across a substantially larger collection of datasets than prior work. The efficiency analysis (FLOPs, parameter counts, runtimes) is a useful contribution. The central comparative claim, however, rests on the integrity of the model-selection rule and the exact data-preprocessing choices; the current protocol does not yet support the headline ranking with sufficient confidence.
major comments (3)
- [Section 5.3.1] The protocol states: 'Model selection is based on training loss, with the best performing model checkpoint used for final evaluation on the test set.' No validation set is used. Under 1500 epochs with ReduceLROnPlateau, training loss generally decreases monotonically, so the selected checkpoint will typically be the final or most overfit epoch. Because architectures differ in capacity and built-in regularization (e.g., depthwise separable convolutions in LITEMV versus the high-capacity VanTran), this rule can differentially favor models that memorize the training set rather than generalize. This affects every reported test accuracy, MAE, and the MCM p-values in Section 6, and therefore the headline ranking of LITEMV as best-performing. I recommend holding out a subject-disjoint validation split within each training fold and selecting checkpoints by validation loss or accuracy before evaluating on the test set.
- [Section 5.2.1] The description of regression-label normalization is ambiguous. The text says labels are normalized 'by dividing them by the maximum value provided in Table 1.' If this maximum is the observed maximum over the full dataset (including test samples), then target information from the test set is used during training, which is a form of leakage and would invalidate the regression comparisons. If, instead, the maximum is a fixed clinical or nominal range from the original repositories (e.g., 100 for KIMORE, 10 for EHE, 1 for UI-PRMD), this should be stated explicitly and the actual values used per dataset should be reported. Because the regression results are a major part of the paper's claims, this point must be resolved.
- [Section 5.1.2] The fold-creation rule 'When applicable, we further ensure that the test set contains only samples from unhealthy subjects' is not specified precisely. It is unclear which datasets use this rule, how many folds result in each case, and how healthy subjects are distributed across training folds. For datasets with both healthy and unhealthy subjects (e.g., IRDS, KIMORE, KERAAL), this choice creates a systematic distribution shift (training includes healthy subjects, test contains only unhealthy subjects) that can affect model rankings differently across architectures. Please provide a per-dataset table of the number of folds, the number of subjects in train/test, and the test-set composition, and discuss the implications of this protocol choice.
minor comments (6)
- [Tables 1 and 2] The entries 'L T' (Table 1) and 'R TK' (Table 2) appear to be missing spaces and should be 'LT' and 'RTK' respectively; please fix the typography.
- [Section 6.1.1] The sentence 'other CNN-based models do not perform as strongly, as ConvTran and MotionGRU outperform them' is confusing because ConvTran is self-attention-based and MotionGRU is recurrent-based; please rephrase to clarify the intended comparison.
- [Figure 24] The caption describes a 'one-vs-one scatter plot' but the figure is a scatter of performance differences (LITEMV minus VanTran) against dataset characteristics; please adjust the terminology for accuracy.
- [Section 5.3.1] For reproducibility, please report the optimizer, the loss function, and the learning rate schedule hyperparameters (initial learning rate, patience, etc.) in addition to the number of epochs and batch size.
- [Section 5.1.3] Equation (13) defines the min/max statistics over all samples with notation that could be read as using the full dataset; since the text later states these statistics are computed exclusively on the training set, please make this explicit in the equation or its surrounding text.
- [Section 7] The conclusion states that LITEMV 'consistently outperforms others in both accuracy and efficiency,' but Section 6.1.2 reports no statistically significant accuracy difference between LITEMV and VanTran; please soften the wording to match the reported statistics.
Circularity Check
No circularity: LITEMV's benchmark win is an empirical, externally-checked result; the training-loss checkpoint rule is a protocol validity concern, not a circular reduction.
full rationale
This paper makes no theoretical derivation whose conclusion is equivalent to its premises. Its claims are empirical: that Rehab-Pile is a usable standardized archive and that LITEMV ranks best in the benchmark. Those claims are supported by running nine architectures on 60 datasets aggregated from external repositories, with subject-disjoint train/test folds, and by reporting test-set accuracy, MAE, RMSE, ranks, and Wilcoxon p-values. No equation defines a predicted quantity in terms of a fitted input; the benchmark outcomes are not forced by construction. The fact that LITEMV and H-Inception are the authors' own architectures is a self-citation, but the comparison does not reduce to that fact: LITEMV's measured performance comes from the benchmark protocol, not from citation. The checkpoint-selection rule in Section 5.3.1 (choosing the best model by training loss) is a methodological threat to the validity of the model ranking, and the statement that LITEMV's channel-count trend 'aligns with the original findings in the LITEMV paper' is a corroborative self-citation; however, neither is a circular derivation. I therefore find no circularity in the paper's derivation chain.
Assumptions & free parameters
free parameters (3)
- KIMORE binary threshold =
50
- KINECAL retention threshold =
5 at-risk participants
- Ensemble size =
5
assumptions (4)
- domain assumption Each exercise can be treated as an independent dataset and task.
- domain assumption A cross-subject split that restricts the test set to unhealthy subjects when both groups exist is fair.
- ad hoc to paper Training loss is a valid model-selection criterion.
- domain assumption Datasets derived from the same repository are statistically independent for the Wilcoxon signed-rank test.
Cite this review
Pith. "Pith review of A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning." pith.science (2026). https://pith.science/paper/Y763M4YE
@misc{pith2026250721018,
author = {Pith},
title = {Pith review of: A Standardized Benchmark for Skeleton-Based Rehabilitation Assessment Using Deep Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y763M4YE}},
note = {Machine review of arXiv:2507.21018}
}
read the original abstract
Automated assessment of human motion plays a vital role in rehabilitation, enabling objective evaluation of patient performance and progress. Unlike general human activity recognition, rehabilitation motion assessment focuses on analyzing the quality of movement within the same action class, requiring the detection of subtle deviations from ideal motion. Recent advances in deep learning and video-based skeleton extraction have opened new possibilities for accessible, scalable motion assessment using affordable devices such as smartphones or webcams. However, the field lacks standardized benchmarks, consistent evaluation protocols, and reproducible methodologies, limiting progress and comparability across studies. In this work, we address these gaps by (i) aggregating existing rehabilitation datasets into a unified archive called Rehab-Pile, (ii) proposing a general benchmarking framework for evaluating deep learning methods in this domain, and (iii) conducting extensive benchmarking of multiple architectures across classification and regression tasks. All datasets and implementations are released to the community to support transparency and reproducibility. This paper aims to establish a solid foundation for future research in automated rehabilitation assessment and foster the development of reliable, accessible, and personalized rehabilitation solutions. The datasets, source-code and results of this article are all publicly available.
Figures
Figures from the paper (21 more)
Reference graph
Works this paper leans on
-
[1]
M. Devanne, S. M. Nguyen, Multi-level motion analysis for physical exercises assessment in kinaesthetic rehabilitation, in: IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids), IEEE, 2017, pp. 529–534. URL https://doi.org/10.1109/HUMANOIDS.2017.8246923 42
-
[2]
J. Li, J. Xue, R. Cao, X. Du, S. Mo, K. Ran, Z. Zhang, Finerehab: A multi-modality and multi-task dataset for rehabilitation analysis, in: International Workshop on Computer Vision in Sports (CVsports) at CVPR 2024, 2024
work page 2024
-
[3]
Deep Learning For Time Series Analysis With Application On Human Motion
A. Ismail-Fawaz, Deep learning for time series analysis with application on human motion, arXiv preprint arXiv:2502.19364 (2025)
work page Pith review arXiv 2025
- [4]
-
[5]
S. García-de Villa, A. Jiménez-Martín, J. J. García-Domínguez, A database of physical therapy exercises with variability of execution col- lected by wearable sensors, Scientific Data 9 (1) (2022) 266
work page 2022
-
[6]
S. Sardari, S. Sharifzadeh, A. Daneshkhah, B. Nakisa, S. W. Loke, V. Palade, M. J. Duncan, Artificial intelligence for skeleton-based phys- ical rehabilitation action evaluation: A systematic review, Computers in Biology and Medicine 158 (2023) 106835
work page 2023
-
[7]
Y. Liao, A. Vakanski, M. Xian, D. Paul, R. Baker, A review of computa- tional approaches for evaluation of rehabilitation exercises, Computers in biology and medicine 119 (2020) 103687
work page 2020
-
[8]
A. Ismail-Fawaz, M. Devanne, S. Berretti, J. Weber, G. Forestier, Weighted average of human motion sequences for improving rehabili- tation assessment, in: International Workshop on Advanced Analytics and Learning on Temporal Data, Springer, 2024, pp. 131–146
work page 2024
Show all 74 references
-
[9]
Mahmood, N
N. Mahmood, N. Ghorbani, N. F. Troje, G. Pons-Moll, M. J. Black, Amass: Archive of motion capture as surface shapes, in: Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 5442–5451
2019
-
[10]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, S. Berretti, J. Weber, G. Forestier, A supervised variational auto-encoder for human motion generation using convolutional neural networks, in: International Conference on Pattern Recognition and Artificial Intelligence, Springer, 2024, pp. 166–181. 43
2024
-
[11]
Mennella, U
C. Mennella, U. Maniscalco, G. De Pietro, M. Esposito, A deep learning system to monitor and assess rehabilitation exercises in home-based re- mote and unsupervised conditions, Computers in Biology and Medicine 166 (2023) 107485
2023
-
[12]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, S. Berretti, J. Weber, G. Forestier, Look into the lite in deep learning for time series classification, International Journal of Data Science and Analytics (2025) 1–21
2025
-
[13]
Canton-Ferrer, J
C. Canton-Ferrer, J. R. Casas, M. Pardas, Marker-based human mo- tion capture in multiview sequences, EURASIP Journal on Advances in Signal Processing 2010 (1) (2010) 105476
2010
-
[14]
Asteriadis, A
S. Asteriadis, A. Chatzitofis, D. Zarpalas, D. S. Alexiadis, P. Daras, Es- timating human motion from multiple kinect sensors, in: Proceedings of the 6th international conference on computer vision/computer graphics collaboration techniques and applications, 2013, pp. 1–6
2013
-
[15]
Z. Cao, G. Hidalgo Martinez, T. Simon, S. Wei, Y. A. Sheikh, Openpose: Realtimemulti-person2dposeestimationusingpartaffinityfields, IEEE Transactions on Pattern Analysis and Machine Intelligence (2019)
2019
-
[16]
Ismail Fawaz, G
H. Ismail Fawaz, G. Forestier, J. Weber, L. Idoumghar, P.-A. Muller, Deep learning for time series classification: a review, Data mining and knowledge discovery 33 (4) (2019) 917–963
2019
-
[17]
Middlehurst, P
M. Middlehurst, P. Schäfer, A. Bagnall, Bake off redux: a review and experimental evaluation of recent time series classification algorithms, Data Mining and Knowledge Discovery 38 (4) (2024) 1958–2031
2024
-
[18]
Ismail-Fawaz, S
A. Ismail-Fawaz, S. Berretti, M. Devanne, J. Weber, G. Forestier, Re- framing time series augmentation through the lens of generative models, in: International Workshop on Advanced Analytics and Learning on Temporal Data, Springer, 2025
2025
-
[19]
Meyer, A
C. Meyer, A. Ismail-Fawaz, M. Devanne, J. Weber, G. Forestier, A deep diveintoalternativestotheglobalaveragepoolingfortimeseriesclassifi- cation, in: International Workshop on Advanced Analytics and Learning on Temporal Data, Springer, 2025. 44
2025
-
[20]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition, Ieee, 2009, pp. 248–255
2009
-
[21]
H. A. Dau, A. Bagnall, K. Kamgar, C.-C. M. Yeh, Y. Zhu, S. Gharghabi, C. A. Ratanamahatana, E. Keogh, The ucr time series archive, IEEE/CAA Journal of Automatica Sinica 6 (6) (2019) 1293–1305
2019
-
[22]
Bagnall, H
A. Bagnall, H. A. Dau, J. Lines, M. Flynn, J. Large, A. Bostrom, P. Southam, E. Keogh, The uea multivariate time series classification archive, 2018, arXiv preprint arXiv:1811.00075 (2018)
2018 arXiv
-
[23]
Dempster, N
A. Dempster, N. M. Foumani, C. W. Tan, L. Miller, A. Mishra, M. Salehi, C. Pelletier, D. F. Schmidt, G. I. Webb, Monster: Monash scalable time series evaluation repository, arXiv preprint arXiv:2502.15122 (2025)
2025 arXiv
-
[24]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, J. Weber, G. Forestier, Enhancing time series classification with self-supervised learning, in: International Con- ference on Agents and Artificial Intelligence (ICAART), SCITEPRESS- Science and Technology Publications, 2023, pp. 40–47
2023
-
[25]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, S. Berretti, J. Weber, G. Forestier, Find- ing foundation models for time series classification with a pretext task, in: Pacific-Asia Conference on Knowledge Discovery and Data Mining, Springer, 2024, pp. 123–135
2024
-
[26]
O. Badi, M. Devanne, A. Ismail-Fawaz, J. Abdullayev, V. Lemaire, S. Berretti, J. Weber, G. Forestier, Cocalite: A hybrid model combining catch22 and lite for time series classification, in: 2024 IEEE Interna- tional Conference on Big Data (BigData), IEEE, 2024, pp. 1229–1236
2024
-
[27]
Ismail-Fawaz, H
A. Ismail-Fawaz, H. Ismail Fawaz, F. Petitjean, M. Devanne, J. Weber, S. Berretti, G. I. Webb, G. Forestier, Shapedba: Generating effective time series prototypes using shapedtw barycenter averaging, in: Inter- national Workshop on Advanced Analytics and Learning on Temporal D...
2023
-
[28]
Holder, M
C. Holder, M. Middlehurst, A. Bagnall, A review and evaluation of elas- tic distance functions for time series clustering, Knowledge and Infor- mation Systems 66 (2) (2024) 765–809. 45
2024
-
[29]
C. W. Tan, C. Bergmeir, F. Petitjean, G. I. Webb, Monash univer- sity, uea, ucr time series extrinsic regression archive, arXiv preprint arXiv:2006.10996 (2020)
2020 arXiv
-
[30]
Bagnall, M
A. Bagnall, M. Middlehurst, G. Forestier, A. Ismail-Fawaz, A. Guil- laume, D. Guijo-Rubio, C. W. Tan, A. Dempster, G. I. Webb, A hands- on introduction to time series classification and regression, in: Proceed- ings of the 30th ACM SIGKDD Conference on Knowledge Discovery and ...
2024
-
[31]
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, S. R. Bowman, Glue: A multi-task benchmark and analysis platform for natural language un- derstanding, arXiv preprint arXiv:1804.07461 (2018)
2018 arXiv
-
[32]
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, S. Bowman, Superglue: A stickier benchmark for general- purpose language understanding systems, Advances in neural informa- tion processing systems 32 (2019)
2019
-
[33]
Rajpurkar, J
P. Rajpurkar, J. Zhang, K. Lopyrev, P. Liang, Squad: 100,000+ questions for machine comprehension of text, arXiv preprint arXiv:1606.05250 (2016)
2016 arXiv
-
[34]
Capecci, M
M. Capecci, M. G. Ceravolo, F. Ferracuti, S. Iarlori, A. Monteriu, L. Romeo, F. Verdini, The kimore dataset: Kinematic assessment of movement and clinical scores for remote monitoring of physical rehabil- itation, IEEE Transactions on Neural Systems and Rehabilitation En- gine...
2019
-
[35]
Vakanski, H.-p
A. Vakanski, H.-p. Jun, D. Paul, R. Baker, A data set of human body movements for physical rehabilitation exercises, Data 3 (1) (2018) 2
2018
-
[36]
Bruce, Y
X. Bruce, Y. Liu, K. C. Chan, Q. Yang, X. Wang, Skeleton-based hu- man action evaluation using graph convolutional network for monitoring alzheimer’s progression, Pattern Recognition 119 (2021) 108095
2021
-
[37]
Miron, N
A. Miron, N. Sadawi, W. Ismail, H. Hussain, C. Grosan, Intellirehabds (irds)—adatasetofphysicalrehabilitationmovements, Data6(5)(2021) 46. 46
2021
-
[38]
S. M. Nguyen, M. Devanne, O. Remy-Neris, M. Lempereur, A. Thepaut, A medical low-back pain physical rehabilitation database for human body movement analysis, in: International Joint Conference on Neural Networks, 2024
2024
-
[39]
Blanchard, S
A. Blanchard, S. M. Nguyen, M. Devanne, M. Simonnet, M. L. Goff- Pronost, O.Rémy-Néris, Technicalfeasibilityofsupervisionofstretching exercises by a humanoid robot coach for chronic low back pain: The r- cool randomized trial, BioMed Research International 2022 (2022) 1–10
2022
-
[40]
Maudsley-Barton, M
S. Maudsley-Barton, M. H. Yap, Kinecal: a dataset for falls-risk as- sessment and balance impairment analysis, Scientific data 10 (1) (2023) 633
2023
-
[41]
A. T. Paiement, L. Tao, S. L. Hannuna, M. Camplani, D. Damen, M.Mirmehdi, Onlinequalityassessmentofhumanmovementfromskele- ton data, in: British Machine Vision Conference, 2014
2014
-
[42]
Singh, A
A. Singh, A. Bevilacqua, T. B. Aderinola, T. L. Nguyen, D. Whelan, M. O’Reilly, B. Caulfield, G. Ifrim, An examination of wearable sensors and video data capture for human exercise classification, in: Joint Eu- ropean Conference on Machine Learning and Knowledge Discovery in D...
2023
-
[43]
Singh, B
A. Singh, B. T. Le, T. L. Nguyen, D. Whelan, M. O’Reilly, B. Caulfield, G. Ifrim, Interpretable classification of human exercise videos through pose estimation and multivariate time series analysis, in: International Workshop on Health Intelligence, Springer, 2021, pp. 181–199
2021
-
[44]
Singh, A
A. Singh, A. Bevilacqua, T. L. Nguyen, F. Hu, K. McGuinness, M. O’Reilly, D. Whelan, B. Caulfield, G. Ifrim, Fast and robust video- based exercise classification via body pose tracking and scalable mul- tivariate time series classifiers, Data Mining and Knowledge Discovery 37 ...
2023
-
[45]
LeCun, B
Y. LeCun, B. Boser, J. Denker, D. Henderson, R. Howard, W. Hub- bard, L. Jackel, Handwritten digit recognition with a back-propagation network, Advances in neural information processing systems 2 (1989). 47
1989
-
[46]
Zheng, Q
Y. Zheng, Q. Liu, E. Chen, Y. Ge, J. L. Zhao, Time series classification using multi-channels deep convolutional neural networks, in: Interna- tional conference on web-age information management, Springer, 2014, pp. 298–310
2014
-
[47]
J. L. Elman, Finding structure in time, Cognitive science 14 (2) (1990) 179–211
1990
-
[48]
Hochreiter, J
S. Hochreiter, J. Schmidhuber, Long short-term memory, Neural com- putation 9 (8) (1997) 1735–1780
1997
-
[49]
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, Y. Bengio, Learning phrase representations using rnn encoder-decoder for statistical machine translation, arXiv preprint arXiv:1406.1078 (2014)
2014 arXiv
-
[50]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)
2017
-
[51]
Scarselli, M
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, G. Monfardini, The graph neural network model, IEEE transactions on neural networks 20 (1) (2008) 61–80
2008
-
[52]
T. N. Kipf, M. Welling, Semi-supervised classification with graph con- volutional networks, arXiv preprint arXiv:1609.02907 (2016)
2016 arXiv
-
[53]
Z. Wang, W. Yan, T. Oates, Time series classification from scratch with deep neural networks: A strong baseline, in: 2017 International joint conference on neural networks (IJCNN), IEEE, 2017, pp. 1578–1585
2017
-
[54]
Szegedy, W
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Er- han, V. Vanhoucke, A. Rabinovich, Going deeper with convolutions, in: Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1–9
2015
-
[55]
Szegedy, S
C. Szegedy, S. Ioffe, V. Vanhoucke, A. Alemi, Inception-v4, inception- resnet and the impact of residual connections on learning, in: Proceed- ings of the AAAI conference on artificial intelligence, Vol. 31, 2017. 48
2017
-
[56]
Ismail Fawaz, B
H. Ismail Fawaz, B. Lucas, G. Forestier, C. Pelletier, D. F. Schmidt, J. Weber, G. I. Webb, L. Idoumghar, P.-A. Muller, F. Petitjean, Incep- tiontime: Finding alexnet for time series classification, Data Mining and Knowledge Discovery 34 (6) (2020) 1936–1962
2020
-
[57]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, J. Weber, G. Forestier, Deep learning for time series classification using new hand-crafted convolution filters, in: 2022 IEEE International Conference on Big Data (Big Data), IEEE, 2022, pp. 972–981
2022
-
[58]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, S. Berretti, J. Weber, G. Forestier, Lite: Light inception with boosting techniques for time series classification, in: 2023IEEE10thInternationalConferenceonDataScienceandAdvanced Analytics (DSAA), IEEE, 2023, pp. 1–10
2023
-
[59]
S. N. M. Foumani, C. W. Tan, M. Salehi, Disjoint-cnn for multivariate time series classification, in: 2021 International Conference on Data Mining Workshops (ICDMW), IEEE, 2021, pp. 760–769
2021
-
[60]
F. J. Ordóñez, D. Roggen, Deep convolutional and lstm recurrent neural networks for multimodal wearable activity recognition, Sensors 16 (1) (2016) 115
2016
-
[61]
C. Guo, X. Zuo, S. Wang, S. Zou, Q. Sun, A. Deng, M. Gong, L. Cheng, Action2motion: Conditioned generation of 3d human motions, in: Pro- ceedings of the 28th ACM International Conference on Multimedia, 2020, pp. 2021–2029
2020
-
[62]
Ismail-Fawaz, M
A. Ismail-Fawaz, M. Devanne, S. Berretti, J. Weber, G. Forestier, Es- tablishing a unified evaluation framework for human motion generation: A comparative analysis of metrics, Computer Vision and Image Under- standing 254 (2025) 104337
2025
-
[63]
Petrovich, M
M. Petrovich, M. J. Black, G. Varol, Action-conditioned 3d human mo- tion synthesis with transformer vae, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 10985–10995
2021
-
[64]
N. M. Foumani, C. W. Tan, G. I. Webb, M. Salehi, Improving position encoding of transformers for multivariate time series classification, Data mining and knowledge discovery 38 (1) (2024) 22–48. 49
2024
-
[65]
B. Yu, H. Yin, Z. Zhu, Spatio-temporal graph convolutional networks: A deep learning framework for traffic forecasting, in: International Joint Conference on Artificial Intelligence (IJCAI), 2017, pp. 3634–3640
2017
-
[66]
S. Deb, M. F. Islam, S. Rahman, S. Rahman, Graph convolutional net- works for assessment of physical rehabilitation exercises, IEEE Trans- actions on Neural Systems and Rehabilitation Engineering 30 (2022) 410–419
2022
-
[67]
Virtanen, R
P. Virtanen, R. Gommers, T. E. Oliphant, M. Haberland, T. Reddy, D. Cournapeau, E. Burovski, P. Peterson, W. Weckesser, J. Bright, S. J. van der Walt, M. Brett, J. Wilson, K. J. Millman, N. Mayorov, A. R. J. Nelson, E. Jones, R. Kern, E. Larson, C. J. Carey, İ. Polat, Y. Feng,...
2020
-
[68]
Abadi, A
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kud- lur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M....
2015
-
[69]
Chollet, et al., Keras,https://keras.io (2015)
F. Chollet, et al., Keras,https://keras.io (2015)
2015
-
[70]
Middlehurst, A
M. Middlehurst, A. Ismail-Fawaz, A. Guillaume, C. Holder, D. Guijo- Rubio, G. Bulatova, L. Tsaprounis, L. Mentel, M. Walter, P. Schäfer, et al., aeon: a python toolkit for learning from time series, Journal of Machine Learning Research 25 (289) (2024) 1–10
2024
-
[71]
Abdullayev, M
J. Abdullayev, M. Devanne, C. Meyer, A. Ismail-Fawaz, J. Weber, G. Forestier, Enhancing time series classification with diversity-driven 50 neural network ensembles, in: IEEE International Joint Conference on Neural Networks (IJCNN), Rome, Italy, 2025
2025
-
[72]
H. I. Fawaz, G. Forestier, J. Weber, L. Idoumghar, P.-A. Muller, Deep neural network ensembles for time series classification, in: 2019 Interna- tional Joint Conference on Neural Networks (IJCNN), IEEE, 2019, pp. 1–6
2019
-
[73]
Ismail-Fawaz, A
A. Ismail-Fawaz, A. Dempster, C. W. Tan, M. Herrmann, L. Miller, D. F. Schmidt, S. Berretti, J. Weber, M. Devanne, G. Forestier, et al., An approach to multiple comparison benchmark evaluations that is stable under manipulation of the comparate set, arXiv preprint arXiv:2305.1...
2023 arXiv
-
[74]
Wilcoxon, Individual comparisons by ranking methods, in: Break- throughs in statistics: Methodology and distribution, Springer, 1992, pp
F. Wilcoxon, Individual comparisons by ranking methods, in: Break- throughs in statistics: Methodology and distribution, Springer, 1992, pp. 196–202. 51 Appendix A. Regression and Classification Datasets Label Distri- bution 0 25 50 75 100 125 150 175 200 Sorted label index 0 ...
1992
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.