REVIEW 4 major objections 5 minor 42 references
Comparison of Visual Trackers for Biomechanical Analysis of Running
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Joint-based pose trackers, after SVR outlier correction and late fusion, reproduce expert-labeled sprint knee, hip, and trunk angles with RMSE between 3.88 and 6.99 degrees, making markerless video useful but not high-precision for…
desk verdict A useful, honest benchmark of six trackers for sprint kinematics, with a real post-processing recipe; the numbers are only as good as the unquantified manual labels, which the paper should have addressed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a post-processing module attached to the tracker output. For each joint, the module smooths the raw x and y trajectories with support vector regression, computes the residual error curve between raw and smoothed trajectories, and flags times where the residual exceeds three times its standard deviation. Flagged outliers are first corrected by swapping left and right joint labels, the most common confusion in side-view sprints, and if that fails, replaced by the smoothed value. A late-fusion step then averages joint positions from MediaPipe, MoveNet, and RTMPose before angles are computed. This mechanism is what converts noisy tracker outputs into angle curves that visually match expert ground truth.
What would settle it
Re-annotate a subset of the same frames with a second independent expert team and measure the inter-annotator difference on occluded-side joints; if that difference is comparable to the reported RMSE values or to the fusion improvement of about $1^\circ$, the claimed accuracy is not separable from label noise. Alternatively, record the same sprints with a marker-based 3D motion-capture system and recompute all RMSE values against that reference.
Extended reading notes
Core claim
The paper's central claim is that joint-based pose estimators, once their output is smoothed and corrected, reproduce expert-labeled sprint kinematics closely enough for practical biomechanical analysis. Using the VideoRun2D dataset of 40 sprints, the authors show that point-tracking models (CoTracker2, CoTracker3) are unsuited to this task, with mean RMSE of $73.78^\circ$ and $24.83^\circ$, while the best skeleton-based trackers produce angle curves similar to ground truth. Their post-processing module, which smooths joint trajectories with support vector regression, detects outliers at three standard deviations of the residual error, and swaps left-right labels or replaces the outlier with the smoothed value, reduces hip and knee RMSE, for example from $10.43^\circ$ to $8.27^\circ$ for the hip angles of MediaPipe and MoveNet. Late fusion of the three best trackers further improves the occluded left-side joints, lowering left hip RMSE from $8.00^\circ$ to $6.99^\circ$. The best single tracker, RTMPose, already achieves $3.88^\circ$ on the right trunk and $4.27^\circ$ on the right knee, leading the authors to conclude that markerless tracking is a valuable but not high-precision resource for sprint biomechanics.
Load-bearing premise
The argument treats the manual expert annotations as error-free ground truth, even though occluded joints were filled in by anatomical symmetry and no inter-annotator agreement is measured.
Editorial extensions
If this is right
- Joint-based pose trackers with the proposed post-processing can serve as a non-invasive, cost-effective alternative to marker-based and IMU systems for routine sprint monitoring, since their angle curves resemble those reported for marker-based systems.
- Near-side joints visible to the side-view camera can already be measured with errors around $4^\circ$--$5^\circ$, indicating that single-view video suffices for trunk and knee angles when the side is unoccluded.
- The far side of the runner remains the main error source, and late fusion of multiple trackers recovers only about $1^\circ$ there, so occlusion handling, not tracker choice, is the limiting factor.
- For applications that require about $1.9^\circ$ accuracy, such as injury-risk assessment, the current pipeline is not sufficient and would need additional views or complementary sensing.
Reading between the lines
- Because the post-processing correction is rule-based, a natural test is to compare it against a learned denoiser or a Kalman smoother on the same trajectories; the paper does not report such a comparison.
- The reported left-side improvements may partly reflect the ground truth's own anatomical-symmetry assumption for occluded joints; a multi-camera 3D capture could separate tracker error from label bias.
- Averaging joint positions from three trackers treats all trackers as equally reliable, so weighting each tracker by per-joint confidence, such as heatmap scores, could yield further gains on occluded joints.
- The $1.9^\circ$ clinical threshold comes from one hamstring-injury study, and whether the achieved $4^\circ$--$7^\circ$ error matters depends on the specific clinical decision, which the paper does not test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper benchmarks six visual trackers—CoTracker2, CoTracker3, PoseNet, MediaPipe, MoveNet, and RTMPose—for estimating trunk, hip, and knee joint angles during sprinting, using the VideoRun2D dataset (40 sprints, 513 strides, 5,870 frames). The authors introduce a post-processing module based on SVR smoothing, a 3-sigma outlier rule, left/right swapping, and a late-fusion strategy that averages the three best joint trackers. The main results are RMSE values against manual expert annotations, with the best fused configurations reaching 3.88–6.99 degrees RMSE for the six angles considered.
Significance. The paper addresses a practical problem—markerless biomechanical analysis of sprinting—and provides a side-by-side comparison that is useful because the VideoRun2D dataset is public. The proposed post-processing and fusion ideas are simple and could be of practical value. If the ground-truth labels are reliable and the hyperparameter choices generalize, the reported RMSE values would be a useful reference point for the markerless-sprinting community. However, the manuscript does not yet establish label reliability or statistical significance, so its quantitative conclusions should be read as preliminary.
major comments (4)
- [Section II-A.1] The manual ground truth is not sufficiently characterized to support the absolute RMSE claims in Tables I–III. The paper states that three biomechanical experts labeled the joints, but it reports no inter-annotator agreement, repeated-labeling variance, or adjudication protocol, and occluded joints were filled using anatomical symmetry. Because the sagittal camera view makes one body side systematically occluded, the symmetry-based labels are not independent measurements, and their errors are likely correlated with the tracker errors on the same occluded joints. Please report inter-annotator or repeated-labeling variability and analyze symmetry-inferred joints separately from directly visible joints, or reframe the results as a relative comparison among trackers rather than an absolute accuracy benchmark.
- [Section III, Tables I-III] No uncertainty quantification accompanies the reported RMSE values. The evaluation aggregates 40 sprints and 513 strides, but there are no per-sprint standard deviations, confidence intervals, or paired statistical tests; several key differences are sub-degree (e.g., Table III right trunk 4.04 vs. 4.38, left hip 6.99 vs. 7.48). The phrase 'substantial reduction' in Section III.C is therefore not supported. Please report per-sprint error distributions and appropriate paired comparisons, or explicitly downgrade the strength of the comparative claims.
- [Section II-C.3] The post-processing module has free hyperparameters—SVR kernel, C, epsilon, the 3-sigma outlier threshold, and the left/right swapping rule—and the paper does not state how they were selected or whether the reported improvements in Table II were obtained on training data. Without a validation split, cross-validation, or sensitivity analysis, the improvements may not generalize to new sprints or camera setups. Please document the parameter-selection procedure and evaluate the module on held-out data.
- [Section IV] The Discussion claims the best systems are 'competitive compared to more expensive marker-based solutions,' but no marker-based or IMU ground truth was collected in this study; the comparison rests on references to other works. This claim goes beyond the evidence. Please either add a direct validation against marker-based or IMU measurements, or soften the conclusion to state that the systems agree with manual expert labeling within the reported RMSE.
minor comments (5)
- [Abstract] The numerical summary is internally inconsistent with the tables; Table I shows a minimum joint-tracker RMSE of 3.88 degrees (RTMPose right trunk), so the stated range '11.41 degrees to 4.37 degrees' is incorrect, and the 6.99-degree value in the abstract comes from late fusion rather than from post-processing alone. Please reconcile the abstract with Tables I–III.
- [Section II.A and Section IV] The Abstract says 'five professional runners' while Section II.A says 'five healthy amateur soccer players' and Section IV says 'amateur football players.' Please correct the participant description for consistency.
- [Section III.A, Table I] Section III.A reports CoTracker3's mean RMSE as 24.97 degrees, but Table I lists 24.83 degrees; the mean of the six angle values is 24.97 degrees, so the table entry appears to be a typo.
- [Tables I-III] Tables I–III cite RTMPose as [37], but Section II-C.2 attributes RTMPose to Jiang et al. [22]; [37] is the HRNet paper and is not the correct reference for RTMPose. Please update the citations.
- [Table III and Section II-C.3] Table III's first row reads 'MediaPipe+MoveNet+RTMPose4.14', missing a space before the numerical value, and Section II-C.3's definition of e_{m,j} should state explicitly whether the norm is computed on one coordinate at a time before thresholding, to make the outlier rule reproducible.
Circularity Check
No significant circularity: the evaluation is an empirical benchmark against manual expert labels, with only a minor self-citation of the authors' own VideoRun2D dataset as data source.
full rationale
The derivation chain is an empirical benchmark, not a derivation of predictions from fitted parameters. The reported RMSEs compare each tracker's angle output to manual expert labels from VideoRun2D [18]. The post-processing module (Section II-C.3) smooths tracker coordinates with SVR and flags outliers by 3 sigma on the residual between raw and smoothed tracker trajectories; it does not use the ground-truth angles to fit the SVR or choose the threshold, so the post-processed errors are not forced to a fitted target. The late-fusion combinations are selected using the same test data, which is a statistical selection issue rather than a circular identity. The only notable self-citation is the authors' own VideoRun2D dataset and system [18], used as the benchmark; however the labels are independent manual annotations by three experts, and the prior work is a data source rather than an assumption equivalent to the present conclusions. The paper also states in Section II-A.1 that occluded joints were 'estimated based on anatomical symmetry' and gives no inter-annotator agreement, which calls into question label accuracy, but this is a ground-truth validity concern, not a self-referential derivation: the tracker errors are still measured against externally produced labels rather than reduced to the paper's own outputs. No equation or fitted parameter is defined in terms of the target RMSEs, and no uniqueness theorem or prior result by the authors is invoked to forbid alternatives. Accordingly, no circular step can be exhibited under the required standard.
Assumptions & free parameters
free parameters (3)
- Outlier threshold multiplier (3 sigma) =
3
- SVR hyperparameters (kernel, C, epsilon) =
not reported
- Late fusion weights =
equal (1/3 each)
assumptions (3)
- domain assumption Expert manual annotations from Kinovea, with occluded joints estimated by anatomical symmetry, are accurate ground truth.
- domain assumption The keypoints produced by all six trackers correspond semantically to the shoulder, hip, knee, and ankle points annotated by the experts.
- domain assumption A 2D lateral camera view is sufficient to estimate trunk, hip, and knee angles as proxies for 3D kinematics.
Cite this review
Pith. "Pith review of Comparison of Visual Trackers for Biomechanical Analysis of Running." pith.science (2026). https://pith.science/paper/SZ23BYUG
@misc{pith2026250504713,
author = {Pith},
title = {Pith review of: Comparison of Visual Trackers for Biomechanical Analysis of Running},
year = {2026},
howpublished = {\url{https://pith.science/paper/SZ23BYUG}},
note = {Machine review of arXiv:2505.04713}
}
read the original abstract
Human pose estimation has witnessed significant advancements in recent years, mainly due to the integration of deep learning models, the availability of a vast amount of data, and large computational resources. These developments have led to highly accurate body tracking systems, which have direct applications in sports analysis and performance evaluation. This work analyzes the performance of six trackers: two point trackers and four joint trackers for biomechanical analysis in sprints. The proposed framework compares the results obtained from these pose trackers with the manual annotations of biomechanical experts for more than 5870 frames. The experimental framework employs forty sprints from five professional runners, focusing on three key angles in sprint biomechanics: trunk inclination, hip flex extension, and knee flex extension. We propose a post-processing module for outlier detection and fusion prediction in the joint angles. The experimental results demonstrate that using joint-based models yields root mean squared errors ranging from 11.41{\deg} to 4.37{\deg}. When integrated with the post-processing modules, these errors can be reduced to 6.99{\deg} and 3.88{\deg}, respectively. The experimental findings suggest that human pose tracking approaches can be valuable resources for the biomechanical analysis of running. However, there is still room for improvement in applications where high accuracy is required.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
- [22]
-
[37]
K. Sun, B. Xiao, D. Liu, and J. Wang. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5693–5703, 2019
work page 2019
- [1]
-
[2]
F. Alonso-Fernandez et al. Super-resolution for selfie biometrics: Intro- duction and application to face and iris. In A. Rattani, R. Derakhshani, and A. Ross, editors, Selfie Biometrics, pages 105–128. Springer, 2019
work page 2019
-
[3]
F. Alonso-Fernandez, J. Fierrez, and J. Ortega-Garcia. Quality mea- sures in biometric systems. IEEE Security & Privacy , 10(9):52–62, December 2012
work page 2012
-
[4]
B. J. Bastiaansen, E. Wilmes, M. S. Brink, C. J. de Ruiter, G. J. Savelsbergh, A. Steijlen, K. M. Jansen, F. C. van der Helm, E. A. Goedhart, D. van der Laan, et al. An inertial measurement unit based method to estimate hip and knee joint kinematics in team sport athletes on the field. Journal of Visualized Experiments (JoVE), 2020(159):1–8, 2020
work page 2020
- [5]
-
[6]
Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, and Y . Sheikh. OpenPose: Realtime multi-person 2D pose estimation using part affinity fields. IEEE Transactions on Pattern Analysis and Machine Intelligence , 43(1):172–186, 2019
work page 2019
Show all 42 references
-
[7]
S. M. L. Cava, R. Casula, S. Concas, G. Orr `u, R. Tolosana, M. Dra- hansky, J. Fierrez, and G. L. Marcialis. Exploiting multiple represen- tations: 3d face biometrics fusion with application to surveillance. In arXiv:2504.18886, 2025
2025 arXiv
-
[8]
S. M. L. Cava, S. Concas, R. Tolosana, R. Casula, G. Orr `u, M. Drahan- sky, J. Fierrez, and G. L. Marcialis. Exploring 3D face reconstruction and fusion methods for face verification: A case-study in video surveillance. In European Conf. on Computer Vision Workshops (ECCVw), 2024
2024
-
[9]
Chung, L.-Y
J.-L. Chung, L.-Y . Ong, and M.-C. Leow. Comparative analysis of skeleton-based human pose estimation. Future Internet , 14(12):380, 2022
2022
-
[10]
Contributors
M. Contributors. OpenMMLab pose estimation toolbox and bench- mark. https://github.com/open-mmlab/mmpose, 2020
2020
-
[11]
N. J. Cronin, J. Walker, C. B. Tucker, G. Nicholson, M. Cooke, S. Merlino, and A. Bissas. Feasibility of OpenPose markerless motion analysis in a real athletics competition. Frontiers in Sports and Active Living, 5:1298003, 2024
2024
-
[12]
Delgado-Santos, R
P. Delgado-Santos, R. Tolosana, R. Guest, R. Vera-Rodriguez, and J. Fierrez. M-gaitformer: Mobile biometric gait verification using transformers. Engineering Applications of Artificial Intelligence , 125:106682, October 2023
2023
-
[13]
Fierrez, A
J. Fierrez, A. Morales, R. Vera-Rodriguez, and D. Camacho. Multiple classifiers in biometrics. Part 2: Trends and challenges. Information Fusion, 44:103–112, November 2018
2018
-
[14]
Fierrez-Aguilar, D
J. Fierrez-Aguilar, D. Garcia-Romero, J. Ortega-Garcia, and J. Gonzalez-Rodriguez. Adapted user-dependent multimodal biometric authentication exploiting general information. Pattern Recognition Letters, 26(16):2628–2639, 2005
2005
-
[15]
Fierrez-Aguilar, D
J. Fierrez-Aguilar, D. Garcia-Romero, J. Ortega-Garcia, and J. Gonzalez-Rodriguez. Bayesian adaptation for user-dependent mul- timodal biometric authentication. Pattern Recognition, 38(8):1317– 1319, August 2005
2005
-
[16]
Fierrez-Aguilar, J
J. Fierrez-Aguilar, J. Ortega-Garcia, and J. Gonzalez-Rodriguez. Tar- get dependent score normalization techniques and their application to signature verification. IEEE Trans. on Systems, Man & Cybernetics - Part C, 35(3):418–425, August 2005
2005
-
[17]
Galasso, C
S. Galasso, C. Carissimo, G. Cerro, M. Molinara, L. Ferrigno, R. S. Calabr`o, and A. M. De Nunzio. A novel measurement procedure for error correction in single camera gait analysis. IEEE Access, 2024
2024
-
[18]
Garrido-Lopez, L
G. Garrido-Lopez, L. F. Gomez, J. Fierrez, A. Morales, R. Tolosana, J. Rueda, and E. Navarro. VideoRun2D: Cost-effective markerless mo- tion capture for sprint biomechanics. Proceedings of the International Conference on Pattern Recognition Workshops , 2024
2024
-
[19]
Hanley, S
B. Hanley, S. Merlino, and A. Bissas. Biomechanics of world-class 800m women at the 2017 IAAF world championships. Frontiers in Sports and Active Living , 4:834813, 2022
2017
-
[20]
Hernandez-Ortega, J
J. Hernandez-Ortega, J. Fierrez, A. Morales, and P. Tome. Time analysis of pulse-based face anti-spoofing in visible and NIR. In Proceddings IEEE Conference on Computer Vision and Pattern Recog- nition Workshops, June 2018
2018
-
[21]
C. S. T. Hii, K. B. Gan, N. Zainal, N. Mohamed Ibrahim, S. Azmin, S. H. Mat Desa, B. van de Warrenburg, and H. W. You. Automated gait analysis based on a marker-free pose estimation model. Sensors, 23(14):6489, 2023
2023
-
[23]
Jo and S
B. Jo and S. Kim. Comparative analysis of OpenPose, PoseNet, and MoveNet models for pose estimation in mobile devices. Traitement du Signal, 39(1):119, 2022
2022
-
[24]
Karaev, I
N. Karaev, I. Makarov, J. Wang, N. Neverova, A. Vedaldi, and C. Rupprecht. CoTracker3: Simpler and better point tracking by pseudo-labelling real videos. In arXiv:2410.11831, 2024
2024 arXiv
-
[25]
Karaev, I
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht. CoTracker: It is better to track together. In Proceedings of the European Conference on Computer Vision (ECCV) , 2024
2024
-
[26]
M. J. Lee, S. L. Reid, B. C. Elliott, and D. G. Lloyd. Running biomechanics and lower limb strength associated with prior hamstring injury. Medicine & Science in Sports & Exercise , 41(10):1942–1951, 2009
1942
-
[27]
Y .-C. Lin, K. Price, D. S. Carmichael, N. Maniar, J. T. Hickey, R. G. Timmins, B. C. Heiderscheit, S. S. Blemker, and D. A. Opar. Validity of inertial measurement units to measure lower-limb kinematics and pelvic orientation at submaximal and maximal effort running speeds. Se...
2023
-
[28]
Morin, P
J.-B. Morin, P. Edouard, and P. Samozino. Technical ability of force application as a determinant factor of sprint performance. Medicine & Science in Sports & Exercise , 43(9):1680–1688, 2011
2011
-
[29]
Nagahara, T
R. Nagahara, T. Matsubayashi, A. Matsuo, and K. Zushi. Kinematics of the thorax and pelvis during accelerated sprinting. The Journal of Sports Medicine and Physical Fitness , 58(9):1253–1263, 2017
2017
-
[30]
Nazarahari, A
M. Nazarahari, A. Khandan, A. Khan, and H. Rouhani. Foot angular kinematics measured with inertial measurement units: A reliable criterion for real-time gait event detection. Journal of Biomechanics , 130:110880, 2022
2022
-
[31]
A. Ohri, S. Agrawal, and G. S. Chaudhary. On-device realtime pose estimation & correction. International Journal of Advances in Engineering and Management (IJAEM) , 3(7), 2021
2021
-
[32]
M. Ota, H. Tateuchi, T. Hashiguchi, and N. Ichihashi. Verification of validity of gait analysis systems during treadmill walking and running using human pose tracking algorithm. Gait & Posture , 85:290–297, 2021
2021
-
[33]
Papandreou, T
G. Papandreou, T. Zhu, L.-C. Chen, S. Gidaris, J. Tompson, and K. Murphy. Personlab: Person pose estimation and instance segmen- tation with a bottom-up, part-based, geometric embedding model. In Proceedings of the European Conference on Computer Vision (ECCV), pages 269–286, 2018
2018
-
[34]
Romero-Franco, P
N. Romero-Franco, P. Jim ´enez-Reyes, A. Casta ˜no-Zambudio, F. Capelo-Ram´ırez, et al. Sprint performance and mechanical outputs computed with an iPhone app: Comparison with existing reference methods. European Journal of Sport Science , 17(4):386–392, 2017
2017
-
[35]
Schlett, C
T. Schlett, C. Rathgeb, O. Henniger, J. Galbally, J. Fierrez, and C. Busch. Face image quality assessment: A literature survey. ACM Computing Surveys, 10(54):1–49, September 2022
2022
-
[36]
Stenum, C
J. Stenum, C. Rossi, and R. T. Roemmich. Two-dimensional video- based analysis of human gait using pose estimation. PLoS Computa- tional Biology, 17(4):e1008935, 2021
2021
-
[38]
Takeichi, M
K. Takeichi, M. Ichikawa, R. Shinayama, and T. Tagawa. A mobile application for running form analysis based on pose estimation tech- nique. In Proceddings IEEE International Conference on Multimedia and Expo Workshops, 2018
2018
-
[39]
MoveNet: Ultra fast and accurate pose detection model
TensorFlow. MoveNet: Ultra fast and accurate pose detection model. https://www.tensorflow.org/hub/tutorials/movenet, 2021. Accessed: 2025-02-20
2021
-
[40]
Viswakumar, V
A. Viswakumar, V . Rajagopalan, T. Ray, P. Gottipati, and C. Parimi. Development of a robust, simple, and affordable human gait analysis system using bottom-up pose estimation with a smartphone camera. Frontiers in Physiology, 12:784865, 2022
2022
-
[41]
H. Xu, E. G. Bazavan, A. Zanfir, W. T. Freeman, R. Sukthankar, and C. Sminchisescu. Ghum & ghuml: Generative 3D human shape and articulated pose models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6184–6193, 2020
2020
-
[42]
Yang and K
J. Yang and K. Park. Improving gait analysis techniques with mark- erless pose estimation based on smartphone location. Bioengineering, 11(2):141, 2024
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.