{"id":"a6b35369-9ebc-4ddf-8862-e3dc9acd0d5d","arxiv_id":"2601.22412","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A probabilistic markerless motion capture pipeline produced well-calibrated confidence intervals for step and stride length against a GaitRite walkway, and for bias-corrected lower-limb kinematics against marker-based capture, with ECE generally below 0.1.","lead":"Researchers tested whether a camera-based, markerless gait analysis system that outputs confidence intervals can tell clinicians when its measurements are trustworthy, validating it against floor sensors and marker-based motion capture in 68 participants across two sites. In most measurements the confidence intervals were well calibrated, with step and stride length errors around 12 to 16 mm, and the system's uncertainty flagged the least reliable steps.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Kinematic calibration is evaluated only after subtracting per-participant biases fitted to the same trials; the deployment claim for calibrated joint-level uncertainty is therefore unsupported.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the kinematic calibration rests on per-participant bias correction computed from the same validation trials. The paper's own Limitations section explicitly concedes that per-participant bias removal is a limitation, which supports the conditional verdict. The spatial calibration against GaitRite is independent and the uncertainty-based filtering results in Tables II and III are internally consistent, so the work has real independent support for part of the claim. The concern is not that the bias correction is statistically invalid; it is that it is not available in the markerless-only deployment scenario the paper claims. A clinician using raw model outputs would see intervals centered at the biased predictions, and the reported calibration after subtracting per-participant offsets does not describe that scenario. The abstract's qualifier 'bias-corrected' is transparent, but the conclusion's claim about identifying unreliable outputs without ground-truth instrumentation overreaches for joint-level calibration. No change to the reader's CONDITIONAL verdict is needed; the condition should be stated as requiring per-participant bias estimation or out-of-sample validation of a fixed bias correction.","tokens_in":12488,"tokens_out":5652,"duration_ms":53170,"concrete_test":"Recompute kinematic ECE, coverage, and median absolute errors for the Shriners data using a leave-one-out or fixed population-mean bias correction instead of per-participant bias fitted on the same validation trials. If the population-mean-corrected ECE remains below 0.1 for all joints except pelvis rotation, the deployment claim holds; if ECE rises above 0.1 or median errors exceed the stated range for several joints, the abstract and conclusion must be amended to state that calibrated joint-level intervals require per-participant marker-based bias estimation, and the markerless-only deployment claim should be limited to spatial metrics and uncertainty-based filtering.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: spatial calibration against GaitRite, and kinematic calibration against marker-based motion capture. The spatial part is independent and credible. The load-bearing weakness is in the kinematic part (Section II-E3 and Table V). The systematic bias for each joint is computed as the mean marker-based minus probabilistic difference for each trial, averaged per participant, and subtracted from the probabilistic predictions before errors and ECE are computed. Because the bias is fitted on the same trials used for validation, the ECE and median errors in Table V measure only the residual, zero-mean fluctuation around a participant-specific offset; they do not measure whether the model's raw confidence intervals contain the clinically referenced joint angle in a deployment setting. At deployment, no marker-based system is present, so the per-participant bias is unknown. The reported biases are large and variable (e.g., hip flexion 23.25±6.36 degrees, pelvis tilt 19.71±4.91 degrees), so using the population mean as a fixed correction leaves per-participant residuals of several degrees and would break nominal coverage. The abstract's phrase 'bias-corrected gait kinematics' is honest, and the Limitations section acknowledges the bias removal, but the conclusion that a user can identify unreliable outputs 'without concurrent ground-truth instrumentation' is only established for spatial gait metrics and for relative uncertainty ranking, not for calibrated joint-level intervals on raw model outputs. The central claim should be restricted or an out-of-sample bias-correction procedure validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript validates a probabilistic multiview markerless motion capture (MMMC) pipeline, previously introduced by the authors, against two external references: an instrumented walkway (GaitRite) for step and stride length and marker-based motion capture for lower-extremity joint angles. Using data from 68 participants across two sites, it reports low Expected Calibration Error (ECE) values for spatial gait metrics, median step and stride errors of about 16 mm and 12 mm, and median kinematic errors of 1.5 to 3.8 degrees after per-participant bias correction. It also demonstrates that filtering steps by predicted uncertainty reduces median errors, and it argues that calibrated uncertainty allows users to identify unreliable outputs without concurrent ground-truth instrumentation.","tokens_in":12798,"tokens_out":4681,"duration_ms":43992,"significance":"The spatial validation is a genuinely independent external benchmark, and the finding that predicted uncertainty correlates with observed error in both spatial and kinematic measures is practically important for clinical gait analysis. The study addresses a real gap by testing calibration rather than only aggregate accuracy, on a diverse clinical population. However, the central kinematic calibration claim is currently undercut by the per-participant bias-correction procedure, which uses the same trials for fitting and evaluation; as written, the paper does not establish that raw markerless reconstructions provide calibrated joint-level confidence intervals in a deployment setting. The spatial filtering result is credible, but the kinematic claim needs additional analysis or substantial qualification.","major_comments":[{"comment":"The kinematic ECE and error magnitudes are computed after subtracting per-participant joint biases estimated from the same trials used for validation. This makes the reported ECE and median errors properties of residual fluctuations around a participant-specific offset, not of the raw probabilistic reconstructions. In deployment the marker-based reference is absent, so the bias is unknown; with the large between-participant standard deviations in Table IV (e.g., hip flexion 23.25 +/- 6.36 degrees, pelvis tilt 19.71 +/- 4.91 degrees), even population-mean correction would leave several degrees of systematic error and would break nominal coverage. Please report ECE and errors without any bias correction, or with an out-of-sample (e.g., leave-one-participant-out) bias estimation, so that the reader can see whether calibrated intervals survive the deployment condition.","section":"Section II-E3, Table V"},{"comment":"The conclusion that the method provides 'calibrated confidence intervals at the level of each joint and moment of time for a participant' is not supported by the analysis as presented, because the calibration was evaluated on bias-corrected residuals. The evidence supports calibrated spatial intervals and a monotone uncertainty-error relationship for kinematics, but not calibrated joint-level intervals for a video-only pipeline. Please either qualify these deployment claims or provide the additional analysis requested above.","section":"Section V (Conclusion) and Abstract"},{"comment":"The PIT/HalfNormal calibration test implicitly assumes that the errors are zero-mean with the model's predicted scale. After per-participant bias subtraction the errors are zero-mean by construction, so the test cannot detect systematic offset miscalibration. If raw errors are used, this test would be a valid check of full calibration; please also report the coverage of the nominal confidence intervals before bias removal, or explicitly state that the kinematic calibration claim is limited to residual uncertainty after a bias-correction step that requires external data.","section":"Section II-E1 (Kinematic ECE via HalfNormal CDF)"}],"minor_comments":[{"comment":"The sentence stating that kinematic patterns 'other than pelvic obliquity' are well calibrated appears to be a mistake: Table V and Figure 5 show pelvis rotation (ECE 0.18) as the poorly calibrated joint, while pelvis obliquity has ECE 0.06. Please correct this.","section":"Section IV (Discussion)"},{"comment":"There is a typo: 'Probabilty Integral Transform' should read 'Probability Integral Transform.'","section":"Section II-E1"},{"comment":"The abbreviation 'MMC' is used in a few places (e.g., 'As MMC moves from research to clinical deployment') where the text should read 'MMMC' for consistency.","section":"Abstract and Introduction"},{"comment":"The captions contain formatting inconsistencies such as 'AT AL' and 'All Participants' versus 'AL' in column headers; please proofread the table captions and headers.","section":"Tables II and III"},{"comment":"Please standardize spelling of 'GaitRite' (the reference list uses 'Gaitrite') and the use of 'pelvis obliquity' versus 'pelvic obliquity'; both variants currently appear.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The spatial validation is solid and the uncertainty-filtering result is a useful contribution. The kinematic calibration claim is the main issue: the bias-correction procedure, though acknowledged in the Limitations, is load-bearing for the abstract's and conclusion's deployment claims. This is fixable by adding raw or cross-validated calibration results, so I view it as a major revision rather than a reject, but the authors should not be allowed to keep the current framing without that analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the spatial half of this paper is genuinely good: external validation against GaitRite in 68 participants across two sites, with ECE below 0.1 across groups, and a clear demonstration that filtering by predicted uncertainty reduces median step error from ~16mm to ~12mm. Second, the kinematic half is more conditional than the abstract's tone suggests: the ECE values for joint angles are computed after subtracting per-participant bias offsets estimated from the same trials, so they describe residual fluctuations, not the coverage a clinician would see from video alone. The paper is transparent about this — it says 'bias-corrected' in the abstract and names the limitation in Section IV-C — but the conclusion that the system can identify unreliable outputs 'without concurrent ground-truth instrumentation' is only fully supported for spatial gait metrics.\n\nWhat's actually new: the model itself is the authors' prior variational inference MMMC [6]. The new contribution is external validation: two sites, diverse clinical populations (neurologic, prosthetic, pediatric), comparison to a clinical walkway and to marker-based kinematics. The uncertainty-filtering analysis — showing that the model's own confidence identifies the highest-error steps — is a useful practical result. The spatial ECE values are independently benchmarked and not circular.\n\nThe main soft spot is the kinematic bias correction. The biases are large and participant-specific (hip flexion 23.25±6.36°, pelvis tilt 19.71±4.91°). Subtracting a per-participant mean before computing errors and ECE means Table V's errors (1.5–3.8°) and low ECE values hold only after an offset is known that deployment would not have. Using the population mean would leave per-participant residuals of several degrees and break nominal coverage. This does not sink the spatial claims, but it should restrict the kinematic claim to 'bias-corrected' comparisons, which the authors do say in the contributions section. The pelvis rotation miscalibration (ECE 0.18) is an honest exception, though.\n\nMinor notes: the calibration weight λ_ece was tuned on 5 participants, but the paper reports a wide insensitivity range, so this is a minor concern. The citation pattern is appropriate — [6] is the prior method, and the external validation here is genuinely new.\n\nThis paper is for rehabilitation researchers and clinicians who want to know whether markerless gait data can be trusted per-measurement. The spatial validation and uncertainty filtering are a real step forward. The kinematic part shows tracking fidelity after model mismatch is removed, but it should not be read as out-of-the-box calibrated joint angles. Send it to review; the spatial claims are solid and the kinematic claims are clearly labeled. A good referee will ask for a deployment-oriented analysis of the bias correction.","headline":"Spatial calibration and uncertainty filtering are solid; the kinematic calibration only holds after per-participant bias removal, so the deployment claim is half-supported.","tokens_in":13316,"tokens_out":2881,"would_cite":true,"duration_ms":25356,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Probabilistic video gait analysis confirms its confidence intervals against clinical references, showing calibrated uncertainty for step, stride, and joint angles.","keywords":["probabilistic motion capture","markerless gait analysis","uncertainty calibration","expected calibration error","variational inference","clinical gait","confidence intervals"],"falsifier":"Run the probabilistic model on a new participant's video without performing per-participant bias subtraction and compare the joint-angle intervals against simultaneous marker-based kinematics; if the ECE exceeds 0.1 or median errors exceed the reported 1.5 to 3.8 degrees, the calibration claim depends on validation-time bias correction rather than the model alone.","tokens_in":12309,"feed_emoji":"🚶","tokens_out":5877,"duration_ms":48140,"temperature":0.7,"pith_summary":"For markerless motion capture to be trusted in clinical gait assessment, the system must not only measure accurately on average but also tell clinicians how accurate each individual measurement is. This paper evaluates a probabilistic multiview markerless pipeline that outputs a full distribution of possible joint angles at every time step, asking whether its confidence intervals are calibrated against external references: an instrumented walkway for step and stride length, and marker-based Vicon kinematics for lower-limb joint angles. Across 68 participants from two sites, including children, prosthesis users, and people with neurological gait impairments, the model achieves Expected Calibration Errors generally below 0.1, with median step and stride length errors near 16 mm and 12 mm and bias-corrected joint angle errors between 1.5 and 3.8 degrees. The central claim is that the model's predicted uncertainty is trustworthy enough to identify and discard unreliable steps without requiring any simultaneous ground-truth instrumentation.","feed_headline":"Video gait analysis now reports trustworthy error bars for each step","feed_subtitle":"Model uncertainty matches real error across 68 patients, so clinicians can drop bad steps without a ground-truth system.","key_machinery":"The load-bearing object is the variational posterior over joint angles, $q_{\\phi}(\\theta_t)=\\mathcal{N}(\\theta_t;\\,\\mu_{\\phi}(t),\\Sigma_{\\phi}(t))$, produced by an implicit neural function that maps time to the mean and a low-rank covariance of 40 kinematic degrees of freedom. The model is fit by optimizing an evidence lower bound augmented with geometric regularizers (site-offset and joint-limit penalties) and an internal Expected Calibration Error term that treats high-confidence detected keypoints as pseudo-ground truth. External calibration is then assessed by pushing each observed error through the predicted CDF (Probability Integral Transform) and measuring the deviation of the resulting values from uniformity via ECE.","core_discovery":"The paper claims that a variational-inference-based probabilistic MMMC model, trained only on video keypoints, produces posterior confidence intervals that are externally calibrated: ECE is 0.05 for step length and 0.04 for stride length, and for bias-corrected lower-limb joint angles ECE is below 0.1 for every joint except pelvis rotation (0.18). The magnitude of predicted uncertainty closely tracks actual error: filtering to the lowest 50% predicted uncertainty reduces median step length error from 16.2 mm to 12.0 mm, while the noisiest 10% of steps have median errors near 39 mm. The paper interprets this as evidence that the model quantifies epistemic uncertainty, enabling a workflow where the model's own confidence, not a reference instrument, gates data quality.","pith_inferences":["If the per-participant, per-joint bias is stable across sessions and camera configurations, a short marker-based calibration recording per patient could make the joint-level confidence intervals valid in routine video-only use; the paper does not test that stability.","The internal ECE regularizer, which uses high-confidence keypoints as pseudo-ground truth during training, may be the mechanism that transfers calibration to external references; an ablation study removing this term would directly test that.","The uncertainty-based filtering principle could generalize to other clinical measurement pipelines beyond gait, such as upper-limb range-of-motion or balance assessments, wherever a probabilistic reconstruction is available.","Combining predicted uncertainty with explicit occlusion or view-count features might further sharpen the trust filter, since the paper reports anecdotal occlusion effects but does not quantify them."],"forward_implications":["A clinician could use the model's predicted uncertainty to flag or exclude individual steps that are unreliable, without needing an instrumented walkway or marker system at the point of care.","Filtering by uncertainty improves data quality: retaining only the most confident half of steps reduces median step length error from roughly 16 mm to 12 mm and narrows the interquartile range.","Calibration holds across diverse populations including children, prosthetic users, and individuals with neurological gait impairments, suggesting the confidence intervals transfer beyond able-bodied adults.","Because the posterior is computed per time point and per joint, uncertainty can be propagated into derived metrics like step length, enabling per-step error bars rather than population averages.","The remaining miscalibration in pelvis rotation (ECE 0.18) indicates that aligning the biomechanical model's rotation convention with clinical models would likely restore calibration at that joint."],"supporting_citations":[{"why":"supplies the probabilistic model and variational inference approach being externally validated","marker":"[6]"},{"why":"provides the differentiable biomechanics MMMC pipeline that reconstructs pose from calibrated multiview video","marker":"[9]"},{"why":"describes how to account for local reference frame differences between markerless and marker-based analyses, motivating the bias-correction step","marker":"[3]"},{"why":"defines probabilistic calibration and the Probability Integral Transform used to compute ECE","marker":"[11]"},{"why":"supplies the ECE computation framework applied to step length, stride length, and kinematic errors","marker":"[19]"},{"why":"defines the Shriners Children's Gait Model used as the marker-based kinematic reference","marker":"[12]"},{"why":"validates the instrumented walkway used as the spatial ground truth for step and stride length","marker":"[20]"}],"fun_headline_variants":["Gait video model's own confidence flags bad step data","Uncertainty-aware gait analysis: trust the error bars","Model's uncertainty predicts real errors without ground truth","Video gait analysis: calibrated confidence filters unreliable steps","Probabilistic markerless gait: known unknowns guide clinical trust"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The kinematic calibration and reported joint-angle errors are computed only after subtracting a constant per-participant, per-joint bias measured on the same validation trials, so the results assume that the markerless-to-marker offset is stable across time and conditions.","fun_headline_variants_meta":{"raw":{"variants":["Gait video model's own confidence flags bad step data","Uncertainty-aware gait analysis: trust the error bars","Model's uncertainty predicts real errors without ground truth","Video gait analysis: calibrated confidence filters unreliable steps","Probabilistic markerless gait: known unknowns guide clinical trust"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1240,"prompt_tokens":950,"completion_tokens":290,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":212}},"tokens_in":566,"tokens_out":290,"duration_ms":3001,"temperature":1.0,"reasoning_tokens":212,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:37:47.655114+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the probabilistic model on a new participant's video without performing per-participant bias subtraction and compare the joint-angle intervals against simultaneous marker-based kinematics; if the ECE exceeds 0.1 or median errors exceed the reported 1.5 to 3.8 degrees, the calibration claim depends on validation-time bias correction rather than the model alone.","supporting_citations":[{"cited_title":"Biomechanical Reconstruction with Confidence Intervals from Multiview Markerless Motion Capture","cited_arxiv_id":"2502.06486","evidence_quote":"supplies the probabilistic model and variational inference approach being externally validated"},{"cited_title":"Differentiable biomechanics unlocks opportunities for markerless motion capture,","cited_arxiv_id":null,"evidence_quote":"provides the differentiable biomechanics MMMC pipeline that reconstructs pose from calibrated multiview video"},{"cited_title":"Comparison of markerless and marker-based motion analysis accounting for differences in local reference frame orientation,","cited_arxiv_id":null,"evidence_quote":"describes how to account for local reference frame differences between markerless and marker-based analyses, motivating the bias-correction step"},{"cited_title":"Probabilistic forecasts, calibration and sharpness,","cited_arxiv_id":null,"evidence_quote":"defines probabilistic calibration and the Probability Integral Transform used to compute ECE"},{"cited_title":"Multi- hypothesis 3D human pose estimation metrics favor miscalibrated dis- tributions,","cited_arxiv_id":null,"evidence_quote":"supplies the ECE computation framework applied to step length, stride length, and kinematic errors"},{"cited_title":"The shriners children’s gait model (scgm),","cited_arxiv_id":null,"evidence_quote":"defines the Shriners Children's Gait Model used as the marker-based kinematic reference"},{"cited_title":"Concurrent related validity of the gaitrite® walkway system for quantification of the spatial and temporal parameters of gait,","cited_arxiv_id":null,"evidence_quote":"validates the instrumented walkway used as the spatial ground truth for step and stride length"}],"review_version":2}