REVIEW 4 major objections 6 minor 2 cited by
On a rank-based Azadkia-Chatterjee correlation coefficient
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Building the nearest-neighbor graph on coordinate-wise ranks makes the Azadkia–Chatterjee correlation coefficient invariant to rescaling while preserving its consistency and normal limit.
desk verdict A useful rank-based variant of the AC coefficient with a genuine d=1 finding, but the d≥3 CLT currently rests on a false lemma and the proof needs repair. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the rank-based nearest-neighbor graph: each observation i is connected to the j that minimizes the Euclidean distance ‖F_n(X_i) − F_n(X_j)‖ between the vectors of coordinate-wise marginal empirical CDFs. Since each coordinate of F_n is monotone in that coordinate of X, the graph is invariant under component-wise strictly increasing transformations of X, which is exactly the scale-invariance property the original graph lacks. The proof then works by showing that for d ≠ 2 the empirical-rank graph is asymptotically the same as the population-rank graph, with neighbor identities matching with probability tending to one; this reduces the variance calculation to two limits of the population-rank graph, q_d and o_d, and the central limit theorem follows from a Hájek projection together with a dependency-graph Berry–Esseen bound.
What would settle it
Simulate independent X and Y in $R^{3}$ with a continuous copula, compute ξ_n for n up to several thousand, and record whether the empirical-rank nearest neighbor matches the population-rank nearest neighbor; if that match probability fails to approach 1, or if the empirical variance of √n ξ_n does not approach 2/5 + (2/5)q_3 + (4/5)o_3, then the central CLT claim fails.
Extended reading notes
Core claim
The paper's central claim is that replacing the raw Euclidean nearest-neighbor graph with a graph computed from coordinate-wise empirical ranks does not change what the Azadkia–Chatterjee statistic estimates. Concretely, ξ_n defined in (2.1) converges in probability to the Dette–Siburg–Stoimenov measure ξ whenever Y is not a measurable function of X, with no further assumptions, and it can be computed in O(n log n) time. When X and Y are independent and the copula density is continuous, √n ξ_n converges in distribution to a centered normal with variance $σ_1^{2}$ = 1 in one dimension and $σ_d^{2}$ = 2/5 + (2/5)q_d + (4/5)o_d in dimensions d ≥ 3, where q_d is the limiting proportion of mutual nearest-neighbor pairs and o_d is the limiting proportion of two points sharing a nearest neighbor in the rank-space graph. The proof does not cover d = 2, where the estimated-rank error and the nearest-neighbor distance converge at the same speed; the authors conjecture, with simulation support, that the same variance formula holds there.
Load-bearing premise
The load-bearing premise is that, in large samples, the nearest neighbor found using estimated marginal ranks is the same neighbor the true distribution would pick; the proof of this equivalence is exactly what breaks down in dimension two.
Editorial extensions
If this is right
- Monotone rescaling or coordinate-wise transformation of X leaves ξ_n unchanged, so the statistic reports the same dependence regardless of the measurement scale of the covariates.
- Under independence, √n ξ_n has a known normal limit with explicit variance in dimensions one and three or more, giving an O(n log n) independence test whose null distribution is tractable.
- The population target is the Dette–Siburg–Stoimenov measure, which is 0 exactly under independence and 1 exactly when Y is a measurable function of X; the rank-based estimator is consistent for this same target.
- The formal CLT excludes d = 2; the paper's simulations and a related rank-based matching result suggest the variance formula carries over, but that case remains a conjecture in this paper.
Reading between the lines
- A closed-form evaluation of the integrals defining q_d and o_d would turn the variance formula into a plug-in standard error for any d ≥ 3; the paper leaves them as numerical integrals.
- The same rank-space construction transfers conceptually to other graph-based statistics, such as two-sample adjacency tests, mutual-information estimators, or conditional-dependence measures, and would plausibly give them the same scale invariance, though the paper does not pursue this.
- If the d = 2 conjecture is eventually proven, the independence test becomes uniform across dimensions; until then, finite-sample inference in dimension two should be calibrated against the conjectured variance rather than taken from the theorem.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a rank-based version of the Azadkia–Chatterjee graph-based correlation coefficient, replacing Euclidean nearest neighbours of the raw covariates X_i by nearest neighbours in the transformed variables F_n(X_i), where F_n is the vector of marginal empirical CDFs. The main claims are: (Theorem 3.1) the resulting estimator ξ_n is consistent for the Dette–Siburg–Stoimenov measure ξ whenever Y is not a measurable function of X; and (Theorem 3.2) under independence, a Lebesgue density for (X,Y), continuity of the copula density, and d ≠ 2, √n ξ_n converges in distribution to N(0, σ²_d), with σ²_1 = 1 and σ²_d = 2/5 + (2/5)q_d + (4/5)o_d for d ≥ 3, where q_d and o_d are constants from Shi et al. (2024). The paper further reports simulations comparing the rank-based and original estimators across dimensions and scaling factors.
Significance. If the results hold, the paper delivers a scale-invariant alternative to the Azadkia–Chatterjee coefficient that preserves the same population target and enjoys asymptotic normality under independence. The consistency theorem is attractive for its minimal assumptions, and the explicit variance formula is a concrete contribution. The simulation study gives a useful demonstration of the practical advantage of rank-based construction when covariates are on different scales. However, the proof of the central limit theorem is not currently rigorous: one key lemma is false as stated, and several other load-bearing steps are either omitted or contain invalid equalities. Because the main probabilistic claims are plausible and the identified gaps appear local and repairable, the paper is a candidate for major revision rather than rejection.
major comments (4)
- [§5.3, Lemmas 5.18 and 5.19] Lemma 5.12 is false as stated. It asserts that for any sequence δ_n satisfying ‖F_n−F‖∞ = o_p(δ_n), P(K̂_n < D̂_n + δ_n) → 0. But if δ_n is much larger than the nearest-neighbour scale n^{-1/d}, e.g. d = 3 and δ_n = n^{-0.1}, then D̂_n = O_p(n^{-1/3}) and K̂_n = O_p(n^{-1/3}), so the inequality K̂_n < D̂_n + δ_n holds with probability tending to 1, not 0. The proof also contains an invalid equality: P(K̂_n < D̂_n + δ_n) = P(F_n(X_2) ∈ B(F_n(X_1), D̂_n + δ_n)) drops the conditioning on M(1) = i and the summation over i = 2,...,n; the left-hand side is a union over n−2 points, not a single-point probability. Since Lemma 5.13 relies directly on this lemma, the proof of P(N(i) = N̂(i)) → 1 is incomplete, and with it the transfer of the constants q_d, o_d from the population-rank graph to the empirical rank graph in Theorem 3.2. The intended statement appears to be true for the specific δ_n = n^{-1/2+ε} with ε < 1/2 − 1/d used in Lemma 5.13, and it can likely be proved via a Poisson-process gap estimate, but the lemma and its proof must be corrected.
- [§5.3, Lemmas 5.18 and 5.19] The limiting variance formula, which is the main quantitative result of Theorem 3.2, is imported as a black box. Lemma 5.18 is declared “a special case of Lemma 7.3 of Shi et al. (2024)” and Lemma 5.19 cites “the calculations as in Han and Huang (2024)” for the expression nVar(ξ_n) = 2/5 + (2/5)E[1/n ∑ T_{i,j}] + (4/5)E[1/n ∑ C_{i,j,k}] + o(1). Since these constants are the crux of the CLT, the paper should either reproduce the derivation or state precisely which theorem in each cited paper applies, and verify that its conditions, especially those involving the dependent graph statistics T_{i,j} and C_{i,j,k}, are met. As written, the reader cannot verify the claimed variance without consulting separate papers.
- [§5.1, Lemma 5.2] Lemma 5.2, which states E[Q_n] → Q, is essential for the consistency theorem (Theorem 3.1), yet its proof is omitted with only the note “due to the similarity to Azadkia and Chatterjee (2021).” For a main theorem, an omitted proof of a central lemma is not adequate. The authors should provide the argument or give a precise correspondence to Lemma 11.8 of Azadkia and Chatterjee (2021), explaining why the replacement of the Euclidean NNG by the rank-based NNG does not change the limiting expectation.
- [§5.1, Lemma 5.4] Lemma 5.4 contains an unjustified exact identity. The proof claims P((X_{N(i)}, Y_{N(i)}) ∈ B) = P((X_1, Y_1) ∈ B) from the identical distribution of the sample, but the event i → j biases the distribution of X_j; the displayed calculation yields P((X_1, Y_1) ∈ B | i → 1) in general, not the unconditional probability. The statement may hold asymptotically in view of Lemma 5.1, but as written it is false, and the conditional-independence claims in Lemmas 5.5–5.7 depend on it. A corrected asymptotic argument should be supplied.
minor comments (6)
- [Tables 1–3] Several rows in Tables 2 and 3 appear to have misaligned or missing entries, e.g., the rows for α = 10 and α = 500 in Table 2 do not have entries in all eight columns defined by the header. Please reformat the tables so that all columns are populated and clearly labelled.
- [Abstract and Introduction] The reference “Azadkia and Chatterjee (Azadkia and Chatterjee, 2021)” is redundantly phrased; the standard parenthetical citation style should be used.
- [§5, first paragraph] The notation “Let B(x, r) denote the ball in R^n centered at x ∈ R^n” uses the sample size n; the ball should be in the rank space R^d (or the appropriate fixed dimension), not in R^n.
- [Lemma 5.7] The symbol X'_i is used before it is defined; it should be introduced as an independent copy of X, with Y'_i drawn from the conditional law of Y given X = X'_i.
- [Proof of Theorem 3.1] The statement “By the continuous mapping theorem, ξ_n − ξ̃_n a.s.→ 0” is not a standard application of the continuous mapping theorem; the argument should be spelled out, noting that Q_n is bounded and P_n → P almost surely.
- [Corollary 5.1] The final sentence says “n D̂_n^d and n D̂_n^d both converge” but the second should be n D_n^d; the typo obscures which quantities are being compared.
Circularity Check
No significant circularity: the rank-based estimator is a new construction, and the cited prior results are independent external theorems, not the paper's own fitted inputs.
full rationale
The paper's central claims are (i) consistency of the rank-based NNG estimator ξn for the DSS measure ξ (Theorem 3.1) and (ii) asymptotic normality under independence (Theorem 3.2). Theorem 3.1 is derived self-contained in Section 5.1: E[Qn] → Q (Lemma 5.2, modeled on Azadkia–Chatterjee), Var(Qn) → 0 (Lemmas 5.4–5.8), and Pn → P (Lemma 5.9). No fitted parameter is renamed as a prediction, and no target quantity is inserted into the definition of the estimator; ξn is defined through Rosenbaum rank graphs and the DSS measure is a separate population functional. Theorem 3.2's variance constants qd and od come from Shi et al. (2024) and the variance decomposition is imported from Han and Huang (2024), both peer-reviewed works with overlapping authors (Fang Han). This is self-citation, but it is not circular in the prohibited sense: those results are external, parameter-free theorems with assumptions that do not include the present paper's conclusions, and the paper independently establishes the graph-equivalence Lemmas 5.11–5.16 that justify transferring those constants to the rank-based graph. Lemma 5.18 is explicitly a special case of Shi et al. Lemma 7.3; Lemma 5.19 cites 'calculations as in Han and Huang (2024)' for algebra, with the graph limits proved here (Lemma 5.10, Corollary 5.2). Thus the central derivation does not reduce to its own inputs. Two non-circular caveats: Lemma 5.12 contains an invalid equality (dropping the n−1 summation factor), so the proof of Theorem 3.2 is currently incomplete as written; and Theorem 3.2 explicitly excludes d=2. Both are correctness or completeness issues, not definitional circularity.
Assumptions & free parameters
assumptions (8)
- domain assumption E[Q_n] → Q follows by the same proof as Azadkia-Chatterjee (2021, Lemma 11.8).
- domain assumption The conditional law in Lemma 5.4: given X_i, X_j, X_N(i), the triple (Y_i, Y_j, Y_N(i)) equals (Y_i, Y_j, Y_k) given (X_i, X_j, X_k).
- domain assumption Lemma 7.3 of Shi et al. (2024) applies verbatim to the rank-based estimator.
- domain assumption The variance decomposition nVar(ξ_n) = 2/5 + (2/5)E[ΣT/n] + (4/5)E[ΣC/n] + o(1) from Han and Huang (2024) holds.
- domain assumption Henze (1987) Lemma 2.1 / Theorem 1.4: weak convergence of (X_1, nD_n, U_n) to the Poisson-process limit.
- ad hoc to paper X_1 lies in its support with every ball around F(X_1) carrying positive F(X)-mass.
- domain assumption Continuity: (X, Y) has a Lebesgue density and F(X) has a continuous copula density on its support (Theorem 3.2).
- domain assumption Ties in the rank graph are broken uniformly by independent draws U_i ~ Uniform[0,1].
Cite this review
Pith. "Pith review of On a rank-based Azadkia-Chatterjee correlation coefficient." pith.science (2026). https://pith.science/paper/BKNV63SL
@misc{pith2026241202668,
author = {Pith},
title = {Pith review of: On a rank-based Azadkia-Chatterjee correlation coefficient},
year = {2026},
howpublished = {\url{https://pith.science/paper/BKNV63SL}},
note = {Machine review of arXiv:2412.02668}
}
read the original abstract
Azadkia and Chatterjee (Azadkia and Chatterjee, 2021) recently introduced a graph-based correlation coefficient that has garnered significant attention. The method relies on a nearest neighbor graph (NNG) constructed from the data. While appealing in many respects, NNGs typically lack the desirable property of scale invariance; that is, changing the scales of certain covariates can alter the structure of the graph. This paper addresses this limitation by employing a rank-based NNG proposed by Rosenbaum (2005) and gives necessary theoretical guarantees for the corresponding rank-based Azadkia-Chatterjee correlation coefficient.
Forward citations
Cited by 2 Pith papers
-
Spectral analysis of large dimensional Chatterjee's rank correlation matrix
A symmetrized Chatterjee rank correlation matrix has a semicircle spectral limit, plus a central limit theorem and independence tests that detect zero-linear-correlation dependence.
-
Conditional Independence Testing Using Exchangeable Pairs
A model-X conditional independence test using a Gaussian-kernel energy distance between observed data and conditionally independent variants, calibrated by random coordinate swaps.
Reference graph
Works this paper leans on
-
[1]
Ansari, J. and Fuchs, S. (2022). A simple extension of Azadki a and Chatterjee’s rank correlation to a vector of endogenous variables. arXiv preprint arXiv:2212.01621
arXiv 2022
-
[2]
Auddy, A., Deb, N., and Nandy, S. (2024). Exact detection thr esholds for Chatterjee’s correlation. Bernoulli, 30(2):1640–1668. 22
work page 2024
-
[3]
Azadkia, M. and Chatterjee, S. (2021). A simple measure of co nditional dependence. The Annals of Statistics , 49(6):3070–3102
work page 2021
-
[4]
Azadkia, M., Taeb, A., and Bühlmann, P. (2021). A fast non-par ametric approach for causal structure learning in polytrees. arXiv preprint arXiv:2111.14969
arXiv 2021
-
[5]
Bickel, P. J. (2022). Measures of independence and functiona l dependence. arXiv preprint arXiv:2206.13663. Bücher, A. and Dette, H. (2024). On the lack of weak continuity of Chatterjee’s correlation coeffi- cient. arXiv preprint arXiv:2410.11418
work page Pith review arXiv 2022
-
[6]
Cao, S. and Bickel, P. J. (2020). Correlations with tailored e xtremal properties. arXiv preprint arXiv:2008.10177
arXiv 2020
-
[7]
Cattaneo, M. D., Han, F., and Lin, Z. (2025). On Rosenbaum’s r ank-based matching estimator. Biometrika (in press)
work page 2025
-
[8]
Chatterjee, S. (2021). A new coefficient of correlation. Journal of the American Statistical Associ- ation, 116(536):2009–2022
work page 2021
Show all 32 references
-
[9]
Chatterjee, S. (2023). A survey of some recent developments in measures of association. Probability and Stochastic Processes - A Volume in Honour of Rajeeva L. Ka randikar
2023
-
[10]
Chen, L. H. and Shao, Q.-M. (2004). Normal approximation und er local dependence. The Annals of Probability, 32(3A):1985–2028
2004
-
[11]
Deb, N., Ghosal, P., and Sen, B. (2020). Measuring associatio n on topological spaces using kernels and geometric graphs. arXiv preprint arXiv:2010.01768
2020 arXiv
-
[12]
and Kroll, M
Dette, H. and Kroll, M. (2024). A simple bootstrap for Chatte rjee’s rank correlation. Biometrika (in press) , page asae045
2024
-
[13]
F., and Stoimenov, P
Dette, H., Siburg, K. F., and Stoimenov, P. A. (2013). A copul a-based non-parametric measure of regression dependence. Scandinavian Journal of Statistics , 40(1):21–41
2013
-
[14]
Dvoretzky, A., Kiefer, J., and Wolfowitz, J. (1956). Asympt otic minimax character of the sample distribution function and of the classical multinomial est imator. The Annals of Mathematical Statistics, 27(3):642–669
1956
-
[15]
Fuchs, S. (2024). Quantifying directed dependence via dime nsion reduction. Journal of Multivariate Analysis (in press) , 201:105266
2024
-
[16]
Gamboa, F., Gremaud, P., Klein, T., and Lagnoux, A. (2022). G lobal sensitivity analysis: a new generation of mighty estimators based on rank statistics. Bernoulli, 28(4):2345–2374. 23
2022
-
[17]
R., and Trutschnig, W
Griessenberger, F., Junker, R. R., and Trutschnig, W. (2022 ). On a multivariate copula-based dependence measure and its estimation. Electronic Journal of Statistics , 16(1):2206–2251
2022
-
[18]
and Huang, Z
Han, F. and Huang, Z. (2024+). Azadkia-Chatterjee’s correl ation coefficient adapts to manifold data. The Annals of Applied Probability (in press)
2024
-
[19]
Henze, N. (1987). On the fraction of random points by specifie d nearest-neighbour interrelations and degree of attraction. Advances in Applied Probability , 19(4):873–895
1987
-
[20]
Hodges, Jr., J. L. and Lehmann, E. L. (1956). The efficiency of s ome nonparametric competitors of the t-test. The Annals of Mathematical Statistics , 27(2):324–335
1956
-
[21]
Huang, Z., Deb, N., and Sen, B. (2022). Kernel partial correla tion coefficient—a measure of condi- tional dependence. The Journal of Machine Learning Research , 23(1):9699–9756
2022
-
[22]
Kroll, M. (2024). Asymptotic normality of Chatterjee’s ran k correlation. arXiv preprint arXiv:2408.11547
2024 arXiv
-
[23]
and Han, F
Lin, Z. and Han, F. (2022). Limit theorems of Chatterjee’s ra nk correlation. arXiv preprint arXiv:2204.08031
2022 arXiv
-
[24]
and Han, F
Lin, Z. and Han, F. (2023). On boosting the power of Chatterje e’s rank correlation. Biometrika, 110(2):283–299
2023
-
[25]
and Han, F
Lin, Z. and Han, F. (2025). On the failure of the bootstrap for Chatterjee’s rank correlation. Biometrika (in press) . Rényi, A. (1959). On measures of dependence. Acta Mathematica Hungarica , 10(3-4):441–451
2025
-
[26]
Rosenbaum, P. R. (2005). An exact distribution-free test co mparing two multivariate distributions based on adjacency. Journal of the Royal Statistical Society Series B: Statisti cal Methodology, 67(4):515–530
2005
-
[27]
Rosenbaum, P. R. (2010). Design of Observational Studies . Springer
2010
-
[28]
and Wolff, E
Schweizer, B. and Wolff, E. F. (1981). On nonparametric measur es of dependence for random variables. The Annals of Statistics , 9(4):879–885
1981
-
[29]
Shi, H., Drton, M., and Han, F. (2022). On the power of Chatter jee’s rank correlation. Biometrika, 109(2):317–333
2022
-
[30]
Shi, H., Drton, M., and Han, F. (2024). On Azadkia–Chatterje e’s conditional dependence coefficient. Bernoulli, 30(2):851–877
2024
-
[31]
Strothmann, C., Dette, H., and Siburg, K. F. (2024). Rearran ged dependence measures. Bernoulli, 30(2):1055–1078. 24
2024
-
[32]
Zhang, Q. (2023a). On relationships between Chatterjee’s a nd Spearman’s correlation coefficients. arXiv preprint arXiv:2302.10131
2023 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.