Pith. sign in

REVIEW 3 major objections 4 minor 40 references

A Modified Dependence Measure Related to Chatterjee's Rank Correlation: Theoretical Properties and Asymptotic Analysis

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper proposes a nonnegative modification of the DSS dependence measure and derives an asymptotic chi-square-mixture null limit for its estimator.

desk verdict A well-intentioned reformulation of Chatterjee's correlation whose central null distribution theorem is false. read the letter →

arxiv 2608.07844 v1 pith:U6BBU7RN submitted 2026-08-08 math.ST stat.TH

classification math.STstat.TH MSC 62G2062H2062G10
keywords dependencemeasureDSSChatterjeerankcorrelationmean-varianceindexstatisticsasymptoticdistributionstrongconsistencyindependencetest
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes a modification of the Dette–Siburg–Stoimenov dependence measure, replacing the left-continuous CDF with a right-continuous one, so the population quantity is well-defined for discrete and mixed distributions. The new coefficient still characterizes independence (value 0) and functional dependence (value 1) exactly, and the paper's plug-in estimator is designed to stay nonnegative, fixing the sign problem of Chatterjee's rank correlation, whose null distribution puts mass on negative values. The paper claims strong consistency for the estimator, an n-rate null limit that is an infinite chi-square mixture, and sqrt-n asymptotic normality under fixed alternatives, plus a family of generalizations. If those claims hold, the estimator gives an independence test with better finite-sample Type I error control than Chatterjee's coefficient.

What carries the argument

The load-bearing object is the ratio $\xi_+(X,Y) = \int \operatorname{Var}(E(1\{Y>t\}|X))\,dF_Y(t) \,/\, \int \operatorname{Var}(1\{Y>t\})\,dF_Y(t)$, written through right-continuous conditional CDFs. Its squared-distance representation, $\xi_+ = \int E[(F_{Y|X}(t)-F_Y(t))^2]\,dF_Y(t) \,/\, \int [F_Y(t)(1-F_Y(t))]\,dF_Y(t)$, connects the measure to the mean-variance index and lets the empirical version inherit U-statistic asymptotics: the categorical plug-in estimator's identity with six times that index is what transfers the chi-square-mixture null limit, and the same representation provides the influence function behind the non-null normal limit.

What would settle it

Compute equation (3.5) when $X$ has a single category ($R=1$) and $Y$ is continuous: the proposed estimator equals $3/n + 1/n^2$, while the mean-variance index is identically $0$; $n$ times the estimator therefore tends to $3$, not to the claimed null limit $0$. This one calculation is enough to test whether the premise of Theorem 3.2 is valid.

Watch

Extended reading notes

Core claim

The central discovery, stated on the paper's own terms, is that replacing the left-continuous CDF in the DSS measure by the right-continuous CDF yields a measure $\xi_+$ that coincides with the original at the extremes—zero exactly under independence and one exactly when $Y$ is a measurable function of $X$—while satisfying cleaner analytic conventions. For a categorical $X$ with $R$ classes, the paper identifies $\xi_+$ with six times the mean-variance index, and uses the known null distribution of that index to claim that under $H_0$, $n\xi_{+,n}$ converges in distribution to $6\sum_{j=1}^\infty \chi^2_j(R-1)/(\pi^2 j^2)$. It further claims strong consistency of the plug-in estimator and asymptotic normality with variance $36\sigma_\tau^2$ under fixed alternatives, and presents a parametrized family $\xi_{k,l,r,s}$ together with a conditional-dependence analogue.

Load-bearing premise

The null-distribution theorem depends on the proposed estimator being exactly six times an earlier mean-variance index; in the simplest one-category case the two are not equal, so the theorem's premise is the equality itself.

Editorial extensions

If this is right

  • Under $H_0$, for fixed $R$, the test statistic $n\xi_{+,n}$ has a non-normal limit that is an infinite mixture of chi-square variables, so critical values can be computed without bootstrapping.
  • For alternatives, $\xi_{+,n}$ is consistent and asymptotically normal at $\sqrt{n}$ rate, so confidence intervals for dependence strength follow from a plug-in variance estimator.
  • Because $\xi_+$ is nonnegative and vanishes only under independence, the proposed test repairs the finite-sample sign problem of Chatterjee's coefficient, which is negative about half the time under the null.
  • The extension family $\xi_{k,l,r,s}$ preserves the 0/1 characterization and can be tuned through $k,l,r,s$ for different distributional sensitivity.
  • The conditional version $\xi_+(Y,Z|X)$ avoids left limits and is estimated by plugging out-of-sample conditional CDFs, giving a cleaner asymptotic theory for conditional independence.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the stated equality with the mean-variance index fails for degenerate category counts, a centered or debiased version of the estimator would be needed to preserve the chi-square-mixture calibration.
  • The null limit's dependence on $R$ suggests that permutation or bootstrap critical values could be compared against the mixture quantiles, giving a robustness check that does not rely on the exact equality.
  • Because $\xi_+$ is an integrated squared distance between conditional and marginal CDFs, the same construction should extend to multivariate $Y$ or vector-valued $X$ with kernel estimators, where the rank-based baseline does not directly apply.
  • A direct simulation with $R=2$ and continuous $Y$, comparing empirical quantiles of $n\xi_{+,n}$ under $H_0$ to the claimed mixture, would reveal the theorem's practical scope at moderate sample sizes.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a modified DSS-type dependence measure ξ+(X,Y), gives equivalent representations, constructs sample estimators for categorical and continuous X, and claims strong consistency, an asymptotic null distribution as a scaled infinite chi-square mixture for categorical X, asymptotic normality under fixed alternatives, and an extension to conditional dependence. The central theoretical result is Theorem 3.2, which states that under H0: X⊥Y, nξ+,n converges to 6Σ_j χ²_j(R-1)/(π²j²) for the estimator in (3.5). The paper also reports simulations suggesting improved Type I error control relative to Chatterjee's coefficient.

Significance. If Theorem 3.2 were correct, the paper would offer a nonnegative estimator of a dependence measure with a known null distribution for categorical X, thereby addressing a known drawback of Chatterjee's coefficient. The paper properly credits earlier work (Gamboa, Klein and Lagnoux; Cui and Zhong) rather than creating a citation loop. However, the main asymptotic claim is false: the identity on which the proof rests is algebraically incorrect, so the proposed test is not asymptotically calibrated. The claimed theoretical contribution is therefore not established, and the favorable Type I error statements in the abstract and conclusions are not supported.

major comments (3)
  1. [Section 3.2, Theorem 3.1(ii)] The proof of strong consistency for continuous X is not a valid proof. It asserts 'by the strong law of large numbers' that R_i/n → F(Y_i) and R_{i,j}/n_j → F(Y_i|X_j), but these are not sums of i.i.d. random variables for fixed indices i,j; the observations Y_i and X_j depend on n, the kernel weights involve the bandwidth h_n, and the denominator n_j can be zero or small. The regularity conditions (C1)–(C4) are not used anywhere in the proof. Consequently, the almost-sure limit for the kernel estimator is unproven, which also undermines Theorem 3.3 and the simulation-based claims for the kernel version of the estimator.
  2. [Section 3.5, Asymptotic Power] The local alternative claim is unsupported. The paper simply states that under H1n: ξ+ = δ/√n, √nξ+,n → N(δ, σ²), without proving contiguity, deriving the limiting variance under the local sequence, or verifying that the fixed-alternative variance σ² applies. Moreover, the proposed test statistic uses the standardization √nξ+,n/σ̂_n and normal quantiles, whereas under H0 the correct scaling is n (as claimed in Theorem 3.2). Because Theorem 3.2 is false, the description of the level-α test and the power function β(δ) is not justified.
  3. [Section 6, Table 2 and the Abstract] The simulation study uses the kernel estimator (3.2) with continuous X, but the only null distribution provided in the paper, Theorem 3.2, applies to the categorical estimator (3.5). No null limit is derived for the kernel estimator, so the claimed 'better Type I error control' (Abstract and Section 7) is not supported by any theoretical result. The r=0 rows of Table 2 report means of ξ+,n, yet these are not compared with the claimed mixture distribution or with the correct null limit, so the simulation does not validate the asymptotic calibration of the test.
minor comments (4)
  1. [Section 2, Proposition 2.2 proof] In the proof of necessity of (iii), the displayed expression '∫ E[(F^2_{Y|X}(t) − F_Y(t)]dF_Y(t) = 0' is a typo; it should be '∫ (E[F_{Y|X}(t)^2] − F_Y(t)^2) dF_Y(t) = 0' (or the equivalent form derived from (2.5)). The subsequent reasoning is understandable but the formula as written is not.
  2. [Throughout] There are numerous typos and grammatical errors, including 'remak' (Introduction), 'As showed by Chatterjee' (Introduction), 'as follow' (Introduction), and 'the sample estimator of ξ+(X,Y) defined in (2.2)' (Section 3.1), where (2.2) is a population expression. A careful proofreading pass is needed.
  3. [Sections 3.1–3.5] The same symbol ξ+,n is used for the categorical estimator (3.1)/(3.5), the kernel estimator (3.2), and the conditional-dependence estimator (5.1), without explicit distinction. This creates ambiguity in Theorems 3.2–3.4 and in the simulation section, where the reader must infer which estimator is being discussed.
  4. [Section 3.4, Theorem 3.4] The statement 'Since ξ+,n = 6τ̂n + o_p(n^{-1/2}) by construction' is not a consequence of the construction; the exact relation is ξ+,n = 6τ̂n + 3/n + 1/n², which is indeed o_p(n^{-1/2}) under H1. The proof should be corrected to state and use the exact relation, rather than an unsupported identity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Theorem 3.2 issue is a non-circular algebraic mismatch with an external theorem, not a derivation that reduces to its own inputs.

full rationale

The paper's derivation chain does not exhibit circularity. The proposed measure ξ+ is defined directly in Definition 2.1, and the population identity ξ+ = 6τ in equation (3.4) is an algebraic expansion of the mean-variance index of Cui and Zhong, an external result cited from [13]. The estimator in (3.5) is a plug-in estimator, and the proof of Theorem 3.2 attempts to apply Cui-Zhong's Theorem 1 to it. Whether the plug-in estimator actually equals 6τ̂ is a factual algebraic question; as the reader notes, the equality fails because ∫F̂²dF̂ = 1/3 + O(1/n) rather than exactly 1/3, so nξ_{+,n} need not converge to the stated limit for the estimator in (3.5). This is a non-circular mathematical error, or a mismatch between estimators, not a case of a fitted parameter renamed as a prediction, a self-citation chain, or a definition that smuggles in the conclusion. The paper's key external inputs—Cui and Zhong's limit theorem, Chatterjee's consistency result, Gamboa et al.'s representation—are independent of the present paper's own claims, and no load-bearing step reduces to the paper's own output. The paper also contains no relevant self-citations. Therefore the appropriate circularity score is 0; the correctness defect in Theorem 3.2 is a soundness issue, not a circularity issue.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central asymptotic claims rest entirely on external theorems (Cui-Zhong) and on an unproved uniform-convergence step for kernel conditional CDF estimators. The proposed measure itself is a reformulation of existing quantities, and the main theorem is misstated for the estimator it names.

free parameters (1)
  • Kernel bandwidth h = 1.06 sigma_hat_X n^{-1/5}
    Bandwidth for kernel conditional CDF estimation in the continuous-X estimator. Chosen by a rule of thumb, not fitted to data, but the estimator's finite-sample behavior and the proof of consistency depend on it.
assumptions (3)
  • standard math Variance decomposition Var(A) = Var(E(A|X)) + E(Var(A|X)) for indicator variables.
    Used throughout Proposition 2.1 to derive equivalent representations of xi_plus.
  • domain assumption Cui and Zhong's Theorems 1 and 2 give the null and alternative asymptotic distributions of the mean-variance index tau_hat.
    The paper's central asymptotic results are direct citations of these theorems; if they are misstated or inapplicable to the actual estimator, the asymptotic claims fail.
  • domain assumption Under conditions C1-C4, kernel conditional CDF estimators converge uniformly, justifying the SLLN step for the double sum in Theorem 3.1(ii).
    The proof of Theorem 3.1(ii) asserts termwise convergence and does not prove the uniform convergence needed to pass from pointwise limits to the integral.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Modified Dependence Measure Related to Chatterjee's Rank Correlation: Theoretical Properties and Asymptotic Analysis." pith.science (2026). https://pith.science/paper/U6BBU7RN

@misc{pith2026260807844,
  author       = {Pith},
  title        = {Pith review of: A Modified Dependence Measure Related to Chatterjee's Rank Correlation: Theoretical Properties and Asymptotic Analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U6BBU7RN}},
  note         = {Machine review of arXiv:2608.07844}
}
abstract

In his recent breakthrough work [JASA, 2021], Chatterjee proposed a rank-based correlation coefficient $\xi(X,Y)$ to measure the dependence of a random variable $Y$ on $X$. Unlike classical measures such as Pearson, Spearman, or Kendall, $\xi$ satisfies $\xi=0$ if and only if $X$ and $Y$ are independent, and $\xi=1$ if and only if $Y$ is a measurable function of $X$, without requiring monotonicity or linearity. This paper proposes a refined measure of $\xi(X,Y)$ and investigates its theoretical properties. We derive several equivalent representations, construct an estimator, and establish its strong consistency as well as its asymptotic distribution. A natural extension of the proposed framework is also presented. The Monte Carlo simulations show that the proposed estimator outperforms Chatterjee's rank correlation in terms of finite-sample performance, particularly in controlling Type I error under the null.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 32 canonical work pages

  1. [1]

    Auddy, A., Deb, N., & Nandy, S. (2024). Exact detection thresholds and minimax optimality of Chatterjee’s correlation. Bernoulli 30(2), 1056-1079

  2. [2]

    Ansari, J., Langthaler, Patrick B., Fuchs, S., & Trutschnig, W. (2026). Quantifying and estimating dependence via sensitivity of conditional distributions. Bernoulli 32 (1): 179-204

  3. [3]

    Ansari, J., & Rockel, M. (2026). The exact region and an inequality between Chatter- jee’s and Spearman’s rank correlations. Journal of Multivariate Analysis 214, 105630

  4. [4]

    Azadkia, M., & Chatterjee, S. (2021). A simple measure of conditional dependence. Annals of Statistics 49(6), 3070-3102

  5. [5]

    Bahari, F. (2026). Improvement of Chatterjee’s correlation with missing at random data in theYvariable. Statistical Papers (2026) 67: 56

  6. [6]

    Bickel, P. J. (2022). Measures of independence and functional dependence. Available at arXiv:2206.13663

  7. [7]

    Cao, S., & Bickel, P. J. (2020). Correlations with tailored extremal properties. Avail- able at arXiv:2008.10177v2

  8. [8]

    & Liang, W.Q

    Chao, C.C., Bai, Z. & Liang, W.Q. (1993). Asymptotic normality for oscillation of permutation. Probability in the Engineering and Informational Sciences 7, 227-235

Show all 40 references
  1. [9]

    Chatterjee, S. (2021). A new coefficient of correlation. Journal of the American Sta- tistical Association 116(536), 2009-2022

  2. [10]

    Chatterjee, S. (2022). A survey of some recent developments in measures of associa- tion. Available at arXiv:2211.04702

  3. [11]

    Chierichetti, F., Giacchini, M., & Kumar, R. (2026). On the metricity of the Chat- terjee correlation coefficient. The American Statistician 80(2), 310-317. 27

  4. [12]

    Cui, H.J., Li, R.Z., & Zhong, W. (2015). Model-free feature screening for ultrahigh dimensional discriminant analysis. Journal of the American Statistical Association 110, 630-641

  5. [13]

    Cui, H.J., & Zhong, W. 2019. A distribution-free test of independence based on mean variance index. Computational Statistics and Data Analysis 139, 117-133

  6. [14]

    Dalitz, C., Arning, G., & Goebbels, S. (2024). A simple bias reduction for Chatterjee’s correlation. Journal of Statistical Theory and Practice 18, 51. https://doi.org/10.1007/s42519-024-00399-y

  7. [15]

    Deb, N., Ghosal, P., & Sen, B. (2020). Measuring association on topological spaces using kernels and geometric graphs. Available at arXiv:2010.01768v2

  8. [16]

    Dette, H., & Kroll, M. (2025). A simple bootstrap for Chatterjee’s rank correlation. Biometrika 112(1), 2025, asae045, https://doi.org/10.1093/biomet/asae045

  9. [17]

    Dette, H., Siburg, K. F. & Stoimenov, P. A. (2013). A copula-based non-parametric measure of regression dependence. Scandinavian Journal of Statistics 40(1), 21-41

  10. [18]

    Gamboa, F., Gremaud, P., Klein, T., & Lagnoux, A. (2022). Global sensitivity analy- sis: A novel generation of mighty estimators based on rank statistics. Bernoulli 28(4), 2345-2374

  11. [19]

    Gamboa, F., Klein, T., & Lagnoux, A. (2018). Sensitivity analysis based on Cram´ er- von Mises distance. SIAM/ASA J. Uncertain. Quantif., 6(2), 522-548

  12. [20]

    R., & Trutschnig, W

    Griessenberger, F., Junker, R. R., & Trutschnig, W. (2022). On a multivariate copula- based dependence measure and its estimation. Electron. J. Stat., 16(1), 2206-2251

  13. [21]

    Han, F. (2021). On extensions of rank correlation coefficients to multivariate spaces. Bernoulli 28, 7-11

  14. [22]

    & Huang, Z

    Han, F. & Huang, Z. (2024). Azadkia-Chatterjee’s correlation coefficient adapts to manifold data. The Annals of Applied Probability 34(6), 5172-5210

  15. [23]

    He, S., Ma, S., & Xu, W. (2019). A modified mean-variance feature-screening proce- dure for ultrahigh-dimensional discriminant analysis. Computational Statistics and Data Analysis 137, 155-169

  16. [24]

    H¨ ormann, S., & Strenger, D. (2026). Azadkia-Chatterjee’s dependence coefficient for infinite dimensional data. Bernoulli 32(1), 467-492. 28

  17. [25]

    Huang, Z., Deb, N., & Sen, B. (2022). Kernel partial correlation coefficient-a measure of conditional dependence. Journal of Machine Learning Research 23, 1-58

  18. [26]

    S., & Borovskich, Y

    Korolyuk, V. S., & Borovskich, Y. V. (1994). Theory ofU-Statistics (Mathematics and Its Applications). Dordrecht: Kluwer Academic Publishers

  19. [27]

    Kroll, M. (2025). Debiased Chatterjee’s correlation: Asymptotic normality without continuity. arXiv:2408.11547

  20. [29]

    Lin, Z., and Han, F. (2023). On Boosting the Power of Chatterjee’s Rank Correlation, Biometrika 110(2), 283-299

  21. [30]

    Lin Z, & Han F. (2025). Limit theorems of Chatterjee’s rank correlation. arXiv:2204.08031v4

  22. [31]

    Mai, Q., & Zou, H. (2015). The fused kolmogorov filter: A nonparametric model-free screening method. Ann. Statist. 43, 1471-1497

  23. [32]

    Shi, H., Drton, M., & Han, F. (2022). On the power of Chatterjee’s rank correlation. Biometrika 109(2), 317-333

  24. [33]

    Shi, H., Drton, M., & Han, F. (2023). On the asymptotic normality of the conditional Chatterjee correlation. Bernoulli(to appear)

  25. [34]

    Strothmann, C., Dette, H., & iburg, K. F. (2024). Rearranged dependence measures. Bernoulli 30(2), 1055-1078

  26. [35]

    Wang, L., Li, X.,Wang, X., & Lai, P. (2022). Unified mean-variance feature screening for ultrahigh-dimensional regression. Computational Statistics, 37, 1887-1918

  27. [36]

    Xia, L, Cao, R, Du, J, & Chen, X. (2025). The improved correlation coefficient of Chatterjee. Journal of Nonparametric Statistics 37(2), 265-281

  28. [37]

    Yan, X., Tang, N., Xie, J., Ding, X., & Wang, Z. (2018). Fused mean-variance filter for feature screening. Computational Statistics and Data Analysis 122, 18-32

  29. [38]

    Yang, S. S. (1977). General distribution theory of the concomitants of order statistics. The Annals of Statistics 5 (5), 996-1002. 29

  30. [39]

    Zhang, Q. (2023). On the asymptotic distribution of the symmetrized Chatterjee’s correlation coefficient. Stat. Probabil. Lett., 194, 109759

  31. [40]

    Zhang, Q. (2025). On the properties of distance covariance for categorical data: Robustness, sure screening, and approximate null distributions. Scandinavian Journal of Statistics 52(7), 777-804

  32. [41]

    Zhang, Q. (2026). On the extensions of the Chatterjee-Spearman test. Journal of Nonparametric Statistics 38(2), 347-376. 30

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.