{"id":"542fb505-8580-48e1-94ef-567b172a7d45","arxiv_id":"2501.00566","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Under compositionality, the unique nontrivial Markov boundary of a response equals the set of covariates for which every bivariate conditional independence with another covariate fails, enabling valid partial-conjunction tests.","lead":"This paper develops statistical tests for which covariates matter in a regression when the covariates are compositional, meaning they always sum to one, such as microbe abundances. Existing methods either have no power or flag every covariate; the authors test pairs of covariates jointly and control error rates for the resulting selections.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Simes/BH BCP methods advertised for power rest on an unproven PRDS assumption for base p-values; if it fails, the main FWER/FDR guarantees and all main-text power simulations lose their stated validity.","rationale":"The reader's weakest-assumption analysis and my stress-test converge on the same load-bearing point: the Simes-based BCP procedures, which are the ones whose power the paper advertises, require a PRDS property that is neither proved nor supplied by any concrete conditional independence test. This is the single most consequential gap because the main-text simulations and the recommended methods (BCP(s)-BH, BCP(s)-Holm, BCP(s)) all use Simes p-values; if PRDS fails, their FWER/FDR guarantees are not established and the empirical error control in Figures 1-3 is anecdotal rather than a consequence of the theory. The paper is honest about this gap, and the Bonferroni-based variants are valid under arbitrary dependence, so the central conceptual contribution---defining S via bivariate conditional independence and testing it with partial conjunction---is not undermined. My proposed check targets the exact simulation setting and asks whether the PRDS assumption actually holds for the dCRT p-values; a negative result would force the paper to restrict its recommendations to Bonferroni-based procedures or to prove PRDS for a specific test. Because the reader already assigned CONDITIONAL and identified this assumption, my read does not change the verdict; it reinforces it.","tokens_in":51386,"tokens_out":11394,"duration_ms":110656,"concrete_test":"Under the paper's own Dirichlet null simulation protocol (p=100, Dirichlet(2,...,2), Y independent of X), compute the dCRT base p-values exactly as in Appendix E.1 with K=1500 resamples, and directly test the PRDS assumption: estimate E[1{P in D} | P_{i,j}=x] for several increasing sets D and null indices (i,j), across a grid of x, using binned Monte Carlo; if any estimate decreases with x beyond Monte Carlo error, PRDS fails in the exact setting where the Simes-based methods are showcased. As an operational companion, run BCP(p/2)-BH on many independent null replications and test whether the empirical FDR significantly exceeds the nominal level, which would falsify the claim that the recommended procedures control FDR under the paper's simulation protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inferential claim is that, under Assumption 1, membership in S can be tested via partial conjunction and that the BCP procedures control error rates. This is rigorously established only for the Bonferroni-based procedures (Corollary 3.1, Theorem 3.1). The power-advantaged methods, which are the ones used in all main-text simulations and recommended in Section 4, are BCP(s), BCP(s)-Holm, and BCP(s)-BH built on Simes p-values P^S_j. Their validity requires the base p-values to be PRDS: Corollary 3.2 and Theorem 3.2 assume PRDS of {P_{i,j}: i in S^c\\{j}}, and Theorem 3.3 assumes PRDS of all p(p-1) base p-values. The paper explicitly says positive dependence is \"plausible (if not easily provable)\" (Section 3.1) and offers no conditional independence test with a proven PRDS property in the compositional setting. The dCRT p-values used in simulations are generated by resampling and share the same response and overlapping covariate sets, so their dependence is complex and not obviously positive regression dependent. Section 3.5's data-dependent screening speedups further alter the dependence structure, and the paper concedes these \"could in principle impact the validity\" of the positive-dependence-based procedures. Thus the advertised power gains and the corresponding error-control guarantees are conditional on an unverified distributional conjecture; if PRDS fails, the Simes-based BCP p-values and BH/Holm procedures may be liberal. The existence of valid Bonferroni-based methods means the core idea survives, but the paper's recommended, simulation-validated methods are not fully supported as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses covariate importance testing when covariates are compositional. Because each compositional covariate is a deterministic function of the others, standard conditional independence hypotheses are degenerate, and unconditional hypotheses are all false in typical sparse models. The authors define the set of important covariates via the Markov boundary, prove under conditions that it coincides with the set S of indices j for which every bivariate conditional independence hypothesis H_{i,j}: Y ⊥⊥ X_{i,j} | X_{i,j}^c is false, and then test membership in S through partial conjunction hypotheses built from bivariate p-values. They propose BCP procedures: Bonferroni-based tests valid under arbitrary dependence, and Simes-based tests and BH/Holm multiple testing procedures whose validity requires a PRDS assumption on the base p-values. Numerical experiments compare these methods with leave-one-out and univariate benchmarks in compositional, non-compositional, and sparse compositional settings.","tokens_in":51749,"tokens_out":7436,"duration_ms":75422,"significance":"If the results hold, this is a valuable contribution: it gives a principled, falsifiable definition of relevant compositional covariates and connects it to a practical testing framework. The Markov-boundary characterization in Theorem 2.1 and the Bonferroni-based error control results are rigorous under the stated assumptions, and the paper includes reproducible code and extensive simulations. The Janus-faced nature of conditional and unconditional hypotheses under compositionality is clearly explained. However, the power-oriented methods advertised and used throughout the main text rest on an unverified PRDS assumption, and two appendix proofs that support the main simplifying corollaries contain gaps. The significance is therefore high conditional on repair of these issues.","major_comments":[{"comment":"The Simes-based BCP procedures, which are the methods used in all main-text simulations and recommended for their power, are valid only if the base p-values are positively regression dependent on each null subset (PRDS). The manuscript itself says this is \"plausible (if not easily provable)\" and provides no conditional independence test, including the dCRT used in the simulations, for which PRDS is actually established. The dCRT p-values share the same response and overlapping conditioning sets, so their joint dependence is complex and not obviously PRDS. Moreover, Section 3.5's data-dependent screening changes the dependence structure, and the paper concedes it \"could in principle impact the validity\" of the positive-dependence-based procedures. Since the advertised power advantages and the FWER/FDR guarantees for BCP(s)-Holm and BCP(s)-BH rest on this unverified assumption, I request either a proven PRDS result for a concrete test class (e.g., the dCRT under the Gaussian or Dirichlet simulation models) or a restructuring that presents the Bonferroni-based procedures as the formally guaranteed methods and the Simes-based procedures as empirically supported heuristics.","section":"Section 3.1, Corollary 3.2, Theorems 3.2 and 3.3"},{"comment":"The constructed point w† in Step 3 can fail to lie in the simplex, so the claimed equivalence proof is incomplete. In the first case, the coordinate of w† corresponding to A∩B is 1 - sum_{j in A\\B} w_j - w*_{B\\A} - sum c_j, and the displayed inequality does not prevent this coordinate from being negative. The subsequent claim f_X(w†) > 0 follows only if w† is in the relevant ball within the simplex slice, which is not established. This gap affects Corollary 2.1, the main simplification of Theorem 2.1 for continuous compositional distributions, and needs to be repaired before the corollary can be considered proven.","section":"Appendix B.3, proof of Corollary 2.1"},{"comment":"The proof of Corollary 2.2 for factor covariates has a serious gap. In Step 4, after applying Lemma 2.2 to A = S_F^c ∩ F_k and B = M^c ∩ F_k^c, the displayed conditioning set and the claimed blanket are not coherent: the set (S_F ∩ F_k) ∪ (M ∪ F_k^c) is not generally a subset of M, so it cannot contradict the minimality of the Markov boundary M. The derivation of Y ⊥⊥ X_{S_F^c ∩ F_k} | X_{(S_F ∩ F_k) ∪ F_k^c} also needs justification. As stated, Corollary 2.2 is not proven and should either be supplied with a corrected argument or explicitly deferred.","section":"Appendix B.4, proof of Corollary 2.2"}],"minor_comments":[{"comment":"The word \"selelction\" should be \"selection\".","section":"Figure 1 caption"},{"comment":"The condition \"for all i, j in S^c\" should require i ≠ j, since H_{i,i} is not defined. The same comment applies to Corollary 2.2 and Corollary C.1.","section":"Corollaries 2.1 and 2.2, assumption (i)"},{"comment":"The text \"0 ≥ c < 1\" appears to be a typo and should read \"0 ≤ c < 1\".","section":"Appendix B.3, Step 2"},{"comment":"The notation P_{(i),j} is used for order statistics of the p-values in column j, but the indexing in the displayed equations is not defined explicitly; a sentence defining P_{(i),j} as the ith smallest of {P_{i',j} : i' ≠ j} would improve readability.","section":"Section 3.1, Eq. (2) and (3)"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the PRDS assumption: the Simes-based procedures carry the paper's advertised power gains, but no proof or concrete test is provided. The authors are transparent about this being unproven, which is to their credit, but it means the central practical claims are conditional on a conjecture. The two appendix proof gaps, especially in Corollary 2.1's proof, are also load-bearing for the simplified conditions. I do not see grounds for reject, because the core Theorem 2.1 argument and the Bonferroni-based procedures appear sound and the paper could be revised to make the scope of the formal guarantees precise."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Rough take: the paper earns a serious referee. The central characterization—that under compositionality the unique nontrivial Markov boundary is exactly the set S defined by bivariate conditional independence hypotheses—is new and non-obvious, and the Bonferroni-based BCP procedures built on it are rigorously valid under arbitrary dependence. That alone is a real contribution to a problem where standard conditional testing is degenerate and unconditional testing is unreliable.\n\nThe paper is also honest about its main weakness. The Simes-based BCP(s), BCP(s)-BH, and BCP(s)-Holm—the methods used in all main-text simulations and recommended for power—require the base p-values to be PRDS. The authors explicitly say this is \"plausible (if not easily provable)\" and offer no conditional independence test that provably satisfies it. The dCRT p-values share response and overlapping covariates, so their dependence is complex; the Section 3.5 screening speedups further alter it. If PRDS fails, the Simes-based p-values and the BH/Holm procedures can be liberal. The Bonferroni versions remain valid, so the core idea survives, but the advertised power gains are conditional on an unverified distributional conjecture.\n\nTwo appendix proofs also have gaps. In Corollary 2.1's proof, the constructed intermediate point w† needs to lie in the simplex; the inequality shown only gives ≤1, not ≤ 1 − sum of conditioned coordinates, so the point can have a negative coordinate. Corollary 2.2's Step 4 appears to condition on a larger set than the one implied by the preceding weak-union and intersection steps—likely a typo, but as written the proof doesn't go through. These are fixable but mean the simplified conditions for continuous and factor covariates are not fully established as stated.\n\nThe simulations are extensive, the code is shipped, and the comparison with LOO and AdaFilter is informative. I'd send this to peer review with a request for major revision: either prove PRDS for a concrete test or relegate the Simes methods to a clearly labeled conjecture, fix the two proof details, and re-run the power claims if needed. The paper is for statisticians working on compositional or constrained covariate inference; they will want to read it.","headline":"Novel and valuable theory for compositional covariate importance, but the power-advantaged methods rely on an unproven PRDS assumption and two appendix proofs need fixing.","tokens_in":52276,"tokens_out":4814,"would_cite":true,"duration_ms":44717,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H15","62G10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Compositional covariates defeat standard importance tests; this paper defines importance as the unique nontrivial Markov boundary and tests it via partial conjunction of bivariate conditional hypotheses.","keywords":["compositional data","Markov boundary","bivariate conditional independence","partial conjunction hypothesis","multiple testing","false discovery rate","familywise error rate","conditional randomization test"],"falsifier":"Generate compositional covariates with a known sparse Markov boundary (e.g., $Y = X_1 + \\varepsilon$, $X$ Dirichlet) but force the bivariate p-values to be negatively dependent — for instance by using antithetic resampling in the conditional randomization test — and apply BCP($\\bar{s}$)-BH at FDR 10%. If the empirical FDR exceeds 10%, Theorem 3.3's Simes-based guarantee is refuted. Alternatively, a distribution like Example 2 where the path-connectivity condition of Corollary 2.1 fails should produce multiple nontrivial Markov boundaries, checking the uniqueness claim.","tokens_in":51188,"feed_emoji":"📊","tokens_out":7475,"duration_ms":68203,"temperature":0.7,"pith_summary":"Compositional covariates arise whenever a total is split into parts, such as microbiome relative abundances, time-use budgets, or voting shares, and they silently break standard regression tools: each covariate is a deterministic function of the others, so every conditional importance test is trivially true and every marginal test is trivially false. This paper establishes that a workable notion of importance still exists: under mild conditions, the unique nontrivial Markov boundary of the response exists and equals the set $S$ of covariates $j$ for which every bivariate conditional hypothesis $Y \\perp\\!\\!\\perp X_{i,j} \\mid X_{i,j}^c$ (for all $i \\neq j$) is false. It then shows that testing membership in $S$ reduces to a partial conjunction hypothesis test, and builds single-test and multiple-testing procedures called BCP with proven FWER and FDR control. Simulations indicate the methods are valid and powerful across Dirichlet, logistic-normal, sparse-compositional, and even non-compositional covariate distributions.","feed_headline":"New tests reveal important covariates despite compositional sums","feed_subtitle":"Standard conditional tests are trivially true here; bivariate partial-conjunction tests restore error-controlled selection.","key_machinery":"The load-bearing object is the bivariate conditional independence hypothesis $H_{i,j}: Y \\perp\\!\\!\\perp X_{i,j} \\mid X_{i,j}^c$, which is non-degenerate under compositionality because conditioning on all but two coordinates leaves a random pair. The set $S$ aggregates these hypotheses, and Theorem 2.1 connects $S$ to the Markov boundary via an iterated set operation $(\\Delta \\circ)^k I$ built from an intersection-property lemma for supports with a single equivalence class. On the testing side, the paper's methods are partial conjunction hypothesis (PCH) tests — hypotheses stating that fewer than $r$ of a collection of base hypotheses are false — applied to the bivariate p-values, combined with Bonferroni or Simes global tests, and then wrapped in Holm-style or Benjamini-Hochberg multiple testing procedures to form the BCP family. This machinery converts an untestable definition (a Markov boundary on a measure-zero support) into a testable composite of ordinary conditional independence tests.","core_discovery":"The paper's central theoretical result is Theorem 2.1: for compositional $X$, if $S=[p]$ then no nontrivial Markov boundary exists, and otherwise, provided $S^c \\in (\\Delta \\circ)^k I$ for some $k$, $S$ is the unique nontrivial Markov boundary, where $H_{i,j}: Y \\perp\\!\\!\\perp X_{i,j} \\mid X_{i,j}^c$, $I=\\{\\{i,j\\}: H_{i,j} \\text{ true}\\}$, and $\\Delta$ collects pairs of sets for which an intersection-property lemma applies. Under Assumption 1, each hypothesis $H_{0j}: j\\notin S$ equals a partial conjunction hypothesis: fewer than $r$ of the bivariate nulls for that $j$ are false, and any strict upper bound $\\bar{s}>|S|$ yields a valid test. The paper proves validity of Bonferroni-based BCP tests under arbitrary dependence, and Simes-based BCP tests under a PRDS condition on the base p-values; it also proves FWER control for Holm-style BCP algorithms and FDR control for a Benjamini-Hochberg BCP algorithm.","pith_inferences":["The PRDS assumption could be verified for concrete conditional-independence tests: for instance, the distilled conditional randomization test with Gaussian designs may satisfy PRDS under exchangeable resampling, which would upgrade Corollary 3.2 and Theorem 3.3 from plausible to proven.","The same PCH-of-bivariate-hypotheses construction might extend to multiple deterministic constraints (e.g., $k$ constraints) by testing $k$-variate conditional independence, but at $O(p^{k+1})$ hypotheses; the paper's screening speedups suggest a possible path to tractability.","For experimental design, Remark 2 shows the Markov boundary can depend on the support of $X$, not just on $Y\\mid X$; this implies that the choice of design partially determines the scientifically meaningful target of selection.","The paper's methods could be used as a diagnostic for whether a compositional regression problem has any parsimonious structure: if $S=[p]$ (all bivariate tests false), no nontrivial Markov boundary exists and all covariates are effectively important."],"forward_implications":["Standard conditional-independence-based variable selection and testing, including parametric coefficient tests, knockoffs, and conditional randomization tests, has provably trivial power on compositional covariates; BCP methods restore nontrivial, error-controlled inference.","When an upper bound $\\bar{s}$ on the number of important covariates is known (for example $\\bar{s}=p/2$), BCP tests are substantially more powerful than the always-valid default $\\bar{s}=p-1$.","The scope extends beyond compositional vectors: the same theory and procedures apply to covariates satisfying any single deterministic constraint, including linear subspaces or the unit sphere.","Conditioning on covariates that are a priori sparse (Theorem 3.4) recovers power without sacrificing validity, as long as $|S^c \\cap D| \\neq 1$.","In non-compositional regression settings, BCP methods retain most of the power of state-of-the-art univariate conditional independence tests, so the same toolbox transfers without loss."],"supporting_citations":[{"why":"Defines the Markov boundary and supplies the weak-union and intersection properties used throughout Theorem 2.1.","marker":"Pearl, 1988"},{"why":"Provides the equivalence-class formulation of the intersection property that the paper adapts to compositional support.","marker":"Peters, 2015"},{"why":"Supplies the Markov-boundary-as-target definition and the conditional randomization test used to obtain base p-values.","marker":"Candès et al., 2018"},{"why":"Introduces partial conjunction hypothesis testing; Theorems 1 and 2 justify Bonferroni and Simes PCH p-values.","marker":"Benjamini and Heller, 2008"},{"why":"Defines PRDS and gives FDR control under dependency, assumptions used by Simes-based BCP methods.","marker":"Benjamini and Yekutieli, 2001"},{"why":"Provides the distilled conditional randomization test (dCRT) used in simulations for bivariate p-values and the screening idea for speedups.","marker":"Liu et al., 2022"},{"why":"Supplies the step-down FWER control procedure that the BCP Holm variants extend.","marker":"Holm, 1979"},{"why":"Provides the FDR-control theorem for partial conjunction hypotheses under dependency, used in Theorem 3.3's proof.","marker":"Bogomolov, 2021"},{"why":"Provides the BH procedure applied to Simes BCP p-values for FDR control.","marker":"Benjamini and Hochberg, 1995"}],"fun_headline_variants":["Covariate tests work despite compositional sums","Bivariate tests break compositional deadlock","Partial conjunction untangles compositional regression","Valid variable selection for compositional data","New tests reveal key covariates in fixed-sum data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Simes-based BCP methods, which carry the main power advantages and are used in the primary simulations, are valid only if the bivariate base p-values are positively regression dependent on each null subset (PRDS); the paper calls this plausible but not easily provable, and supplies no concrete conditional-independence test with a proven PRDS guarantee.","fun_headline_variants_meta":{"raw":{"variants":["Covariate tests work despite compositional sums","Bivariate tests break compositional deadlock","Partial conjunction untangles compositional regression","Valid variable selection for compositional data","New tests reveal key covariates in fixed-sum data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000166,"raw_usage":{"total_tokens":1281,"prompt_tokens":1001,"completion_tokens":280,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":617,"completion_tokens_details":{"reasoning_tokens":218}},"tokens_in":617,"tokens_out":280,"duration_ms":3580,"temperature":1.0,"reasoning_tokens":218,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:50:46.907766+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate compositional covariates with a known sparse Markov boundary (e.g., $Y = X_1 + \\varepsilon$, $X$ Dirichlet) but force the bivariate p-values to be negatively dependent — for instance by using antithetic resampling in the conditional randomization test — and apply BCP($\\bar{s}$)-BH at FDR 10%. If the empirical FDR exceeds 10%, Theorem 3.3's Simes-based guarantee is refuted. Alternatively, a distribution like Example 2 where the path-connectivity condition of Corollary 2.1 fails should produce multiple nontrivial Markov boundaries, checking the uniqueness claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Markov boundary and supplies the weak-union and intersection properties used throughout Theorem 2.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the equivalence-class formulation of the intersection property that the paper adapts to compositional support."},{"cited_title":"Testing partial conjunction hypotheses under dependency, with applications to meta-analysis","cited_arxiv_id":"2105.09032","evidence_quote":"Provides the FDR-control theorem for partial conjunction hypotheses under dependency, used in Theorem 3.3's proof."}],"review_version":1}