REVIEW 3 major objections 3 minor
Hierarchical Variable Importance with Statistical Control for Medical Data-Based Prediction
T0 review · 3 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper introduces Hierarchical-CPI, a model-agnostic variable-importance method that tests groups of variables along a tree, giving explicit family-wise error-rate control even under high correlation.
desk verdict Plausible incremental method with a strong statistical claim that the abstract alone cannot support; worth reading the full paper before trusting the FWER guarantee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a hierarchical tree over the variable set, where each node represents a group of variables. At every node, the method tests whether that group, as a whole, is conditionally associated with the outcome given the remaining variables, using a model-agnostic conditional permutation scheme. The tree-based importance allocation mechanism then distributes importance down the tree in a way that avoids vanishing importance under high correlation, while the tree structure keeps the number of tests small enough to permit explicit family-wise error-rate control.
What would settle it
Apply Hierarchical-CPI to synthetic data with a known block-correlation structure where exactly one block of variables drives the outcome and all other variables are noise, then check whether the method controls the family-wise error rate at its nominal level and recovers the predictive block; a failure in either respect would contradict the central claim.
Extended reading notes
Core claim
Hierarchical-CPI frames variable-importance inference as discovering groups of variables that are jointly predictive of an outcome. It walks a hierarchical tree over variables, testing each node as a group with a conditional permutation test, and uses a tree-based importance allocation mechanism to credit variables even when correlation would otherwise cancel their individual signals. The paper's central claim is that this procedure controls the family-wise error rate explicitly, stays computationally feasible, and outperforms current conditional-importance methods on highly correlated imaging data, as demonstrated on dementia classification from MRI and EEG analysis of the Berger effect.
Load-bearing premise
The method's validity rests on the hierarchical tree over variables matching the true dependency or grouping structure: if the tree groups variables incorrectly, the group-level tests and the error-rate control may not say anything meaningful about the actual variables of interest.
Editorial extensions
If this is right
- If the method works as claimed, it gives medical researchers a way to obtain statistically valid variable-importance rankings from black-box predictors without sacrificing power on correlated imaging features.
- The group-level tests can highlight clusters of jointly predictive variables, which may correspond to interpretable biological structures rather than isolated voxels or electrodes.
- Because the method is model-agnostic, it could be applied to any supervised predictor, broadening the use of conditional importance beyond linear or additive models.
- The computational tractability claim suggests the procedure can scale to high-dimensional datasets where pixel- or voxel-wise permutation testing would be prohibitive.
Reading between the lines
- The same hierarchical-grouping idea could transfer to other high-correlation domains such as genomics or climate modeling, where individual features are noisy but blocks of features carry signal.
- The choice of tree becomes a modeling decision with a direct effect on power and interpretability, so data-driven tree construction or stability checks over alternative trees would be a natural extension.
- The tree-based allocation mechanism might be adapted to overlapping or continuous groupings, such as spatial neighborhoods in brain images, rather than only disjoint hierarchical partitions.
- A conservative reading is that the explicit FWER control holds for the tree as tested, not necessarily for every variable discovered after allocation; the paper's claims should be read with that distinction in mind.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Hierarchical-CPI, a model-agnostic variable importance method for prediction models in medical settings. The method is described as framing variable importance as the discovery of jointly predictive variable groups along a hierarchical tree, with claims of computational tractability, explicit family-wise error rate (FWER) control, and a tree-based importance allocation mechanism that addresses vanishing conditional importance under high correlation. The authors report benchmarks against state-of-the-art methods and applications to two neuroimaging datasets, ADNI and TDBRAIN, with biologically plausible findings. The available record consists only of the abstract, so the assessment is based solely on the claims made there.
Significance. If the claims are established in the full manuscript, the paper would offer a practically valuable contribution: a model-agnostic variable importance method with explicit FWER control and a workaround for a known weakness of conditional importance measures under high correlation. These properties would be particularly relevant for interpretability in medical imaging, where correlated features are common and rigorous error control is expected. The paper clearly identifies a real problem and provides a concrete solution outline. However, the significance is conditional on the availability of proofs, algorithm details, and benchmark methodology, none of which are inspectable from the abstract alone.
major comments (3)
- [Abstract, FWER control claim] The central claim of explicit family-wise error rate control is not accompanied by any statement about whether the hierarchical tree is prespecified, estimated from data, or adapted during the search; a FWER guarantee requires the set of hypotheses and the order of exploration to be fixed independently of the data, or a valid selective-inference correction. If the tree is built from or modified using the same data, nominal FWER control can fail even when the per-node tests are individually valid, so this missing specification is load-bearing for the headline contribution.
- [Abstract, tree-based importance allocation] The proposed allocation mechanism is introduced as a remedy for vanishing conditional importance under high correlation, but the abstract does not state how the allocation affects the null distribution of the group-level test statistics or whether the per-node tests and FWER bound are adjusted for arbitrary correlation among variables. Without such an adjustment, the explicit FWER control claim is not established in the high-correlation setting that the method is specifically designed to address.
- [Abstract, benchmark evaluation] The benchmark claims against state-of-the-art methods are not verifiable from the abstract because the number of methods, their tuning, the multiple-testing procedures, and the metric used to compare selected variable sets are unspecified; these details are necessary to assess whether the reported effectiveness on ADNI and TDBRAIN is due to the method's statistical properties or to favorable implementation choices. This concern is secondary to the FWER issue but still necessary for evaluating the practical claims.
minor comments (3)
- [Abstract, terminology] The acronym CPI is not expanded in the abstract; it should be defined as Conditional Predictive Impact (or the intended phrase) at first occurrence.
- [Abstract, FWER specification] The abstract does not state whether the claimed FWER control is finite-sample or asymptotic, nor the level at which control is asserted; specifying this would help readers interpret the guarantee.
- [Abstract, biological plausibility] The statement that the identified variables are 'biologically plausible' lacks a description of the comparison basis or criterion used to determine plausibility; a brief operationalization would make this claim more meaningful.
Circularity Check
No circularity identifiable in the abstract-only record; unverified statistical guarantees are validity concerns, not circularity.
full rationale
The available material is limited to the abstract; no equations, derivations, algorithm definitions, or self-citations are present. The central claims introduce Hierarchical-CPI and state that it has explicit family-wise error rate control, computational tractability, and a tree-based importance allocation mechanism, but the abstract gives no construction that could reduce a prediction to a fitted input or define one quantity in terms of another. The benchmark evaluations on ADNI and TDBRAIN are external empirical checks, so the paper is not merely renaming its input. Concerns that the tree might be data-dependent or misspecified are substantive statistical validity risks, but without a quoted mechanism showing that the control relies on data-derived choices, they cannot be classified as circularity under the hard rules. The absence of proof details in an abstract-only record is not evidence of a circular derivation. Therefore no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Variables can be organized in a hierarchical tree that captures relevant grouping.
- domain assumption The sampling distribution and multiple testing conditions allow explicit FWER control.
Cite this review
Pith. "Pith review of Hierarchical Variable Importance with Statistical Control for Medical Data-Based Prediction." pith.science (2026). https://pith.science/paper/QTJCOJC6
@misc{pith2026250808724,
author = {Pith},
title = {Pith review of: Hierarchical Variable Importance with Statistical Control for Medical Data-Based Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/QTJCOJC6}},
note = {Machine review of arXiv:2508.08724}
}
read the original abstract
Recent advances in machine learning have greatly expanded the repertoire of predictive methods for medical imaging. However, the interpretability of complex models remains a challenge, which limits their utility in medical applications. Recently, model-agnostic methods have been proposed to measure conditional variable importance and accommodate complex non-linear models. However, they often lack power when dealing with highly correlated data, a common problem in medical imaging. We introduce Hierarchical-CPI, a model-agnostic variable importance measure that frames the inference problem as the discovery of groups of variables that are jointly predictive of the outcome. By exploring subgroups along a hierarchical tree, it remains computationally tractable, yet also enjoys explicit family-wise error rate control. Moreover, we address the issue of vanishing conditional importance under high correlation with a tree-based importance allocation mechanism. We benchmarked Hierarchical-CPI against state-of-the-art variable importance methods. Its effectiveness is demonstrated in two neuroimaging datasets: classifying dementia diagnoses from MRI data (ADNI dataset) and analyzing the Berger effect on EEG data (TDBRAIN dataset), identifying biologically plausible variables.
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.