Pith. sign in

REVIEW 4 major objections 6 minor 2 cited by

On a rank-based Azadkia-Chatterjee correlation coefficient

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Building the nearest-neighbor graph on coordinate-wise ranks makes the Azadkia–Chatterjee correlation coefficient invariant to rescaling while preserving its consistency and normal limit.

desk verdict A useful rank-based variant of the AC coefficient with a genuine d=1 finding, but the d≥3 CLT currently rests on a false lemma and the proof needs repair. read the letter →

arxiv 2412.02668 v1 pith:BKNV63SL submitted 2024-12-03 math.ST stat.TH

classification math.STstat.TH MSC 62G2062H2062G10
keywords measureofdependencenearestneighborgraphranktransformationAzadkia-Chatterjeecorrelationscaleinvarianceasymptoticnormalitycopuladensity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper builds a dependence coefficient that first maps each coordinate of X to its marginal rank and then constructs the nearest-neighbor graph in that rank space. It claims this rank-based Azadkia–Chatterjee coefficient is consistent for the same population dependence measure as the original graph-based coefficient, reaching zero exactly under independence and one exactly when Y is a measurable function of X. It also derives the null limit: under independence, √n times the coefficient is asymptotically normal with variance 1 in dimension one and variance 2/5 + (2/5)q_d + (4/5)o_d in dimensions at least three, where q_d and o_d are geometric constants of the rank graph. The payoff is a statistic that no longer changes when covariates are rescaled or monotonically transformed, yet still supports normal inference; the main caveat is that the proof of normality excludes dimension two.

What carries the argument

The carrying object is the rank-based nearest-neighbor graph: each observation i is connected to the j that minimizes the Euclidean distance ‖F_n(X_i) − F_n(X_j)‖ between the vectors of coordinate-wise marginal empirical CDFs. Since each coordinate of F_n is monotone in that coordinate of X, the graph is invariant under component-wise strictly increasing transformations of X, which is exactly the scale-invariance property the original graph lacks. The proof then works by showing that for d ≠ 2 the empirical-rank graph is asymptotically the same as the population-rank graph, with neighbor identities matching with probability tending to one; this reduces the variance calculation to two limits of the population-rank graph, q_d and o_d, and the central limit theorem follows from a Hájek projection together with a dependency-graph Berry–Esseen bound.

What would settle it

Simulate independent X and Y in $R^{3}$ with a continuous copula, compute ξ_n for n up to several thousand, and record whether the empirical-rank nearest neighbor matches the population-rank nearest neighbor; if that match probability fails to approach 1, or if the empirical variance of √n ξ_n does not approach 2/5 + (2/5)q_3 + (4/5)o_3, then the central CLT claim fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that replacing the raw Euclidean nearest-neighbor graph with a graph computed from coordinate-wise empirical ranks does not change what the Azadkia–Chatterjee statistic estimates. Concretely, ξ_n defined in (2.1) converges in probability to the Dette–Siburg–Stoimenov measure ξ whenever Y is not a measurable function of X, with no further assumptions, and it can be computed in O(n log n) time. When X and Y are independent and the copula density is continuous, √n ξ_n converges in distribution to a centered normal with variance $σ_1^{2}$ = 1 in one dimension and $σ_d^{2}$ = 2/5 + (2/5)q_d + (4/5)o_d in dimensions d ≥ 3, where q_d is the limiting proportion of mutual nearest-neighbor pairs and o_d is the limiting proportion of two points sharing a nearest neighbor in the rank-space graph. The proof does not cover d = 2, where the estimated-rank error and the nearest-neighbor distance converge at the same speed; the authors conjecture, with simulation support, that the same variance formula holds there.

Load-bearing premise

The load-bearing premise is that, in large samples, the nearest neighbor found using estimated marginal ranks is the same neighbor the true distribution would pick; the proof of this equivalence is exactly what breaks down in dimension two.

Editorial extensions

If this is right

  • Monotone rescaling or coordinate-wise transformation of X leaves ξ_n unchanged, so the statistic reports the same dependence regardless of the measurement scale of the covariates.
  • Under independence, √n ξ_n has a known normal limit with explicit variance in dimensions one and three or more, giving an O(n log n) independence test whose null distribution is tractable.
  • The population target is the Dette–Siburg–Stoimenov measure, which is 0 exactly under independence and 1 exactly when Y is a measurable function of X; the rank-based estimator is consistent for this same target.
  • The formal CLT excludes d = 2; the paper's simulations and a related rank-based matching result suggest the variance formula carries over, but that case remains a conjecture in this paper.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A closed-form evaluation of the integrals defining q_d and o_d would turn the variance formula into a plug-in standard error for any d ≥ 3; the paper leaves them as numerical integrals.
  • The same rank-space construction transfers conceptually to other graph-based statistics, such as two-sample adjacency tests, mutual-information estimators, or conditional-dependence measures, and would plausibly give them the same scale invariance, though the paper does not pursue this.
  • If the d = 2 conjecture is eventually proven, the independence test becomes uniform across dimensions; until then, finite-sample inference in dimension two should be calibrated against the conjectured variance rather than taken from the theorem.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes a rank-based version of the Azadkia–Chatterjee graph-based correlation coefficient, replacing Euclidean nearest neighbours of the raw covariates X_i by nearest neighbours in the transformed variables F_n(X_i), where F_n is the vector of marginal empirical CDFs. The main claims are: (Theorem 3.1) the resulting estimator ξ_n is consistent for the Dette–Siburg–Stoimenov measure ξ whenever Y is not a measurable function of X; and (Theorem 3.2) under independence, a Lebesgue density for (X,Y), continuity of the copula density, and d ≠ 2, √n ξ_n converges in distribution to N(0, σ²_d), with σ²_1 = 1 and σ²_d = 2/5 + (2/5)q_d + (4/5)o_d for d ≥ 3, where q_d and o_d are constants from Shi et al. (2024). The paper further reports simulations comparing the rank-based and original estimators across dimensions and scaling factors.

Significance. If the results hold, the paper delivers a scale-invariant alternative to the Azadkia–Chatterjee coefficient that preserves the same population target and enjoys asymptotic normality under independence. The consistency theorem is attractive for its minimal assumptions, and the explicit variance formula is a concrete contribution. The simulation study gives a useful demonstration of the practical advantage of rank-based construction when covariates are on different scales. However, the proof of the central limit theorem is not currently rigorous: one key lemma is false as stated, and several other load-bearing steps are either omitted or contain invalid equalities. Because the main probabilistic claims are plausible and the identified gaps appear local and repairable, the paper is a candidate for major revision rather than rejection.

major comments (4)
  1. [§5.3, Lemmas 5.18 and 5.19] Lemma 5.12 is false as stated. It asserts that for any sequence δ_n satisfying ‖F_n−F‖∞ = o_p(δ_n), P(K̂_n < D̂_n + δ_n) → 0. But if δ_n is much larger than the nearest-neighbour scale n^{-1/d}, e.g. d = 3 and δ_n = n^{-0.1}, then D̂_n = O_p(n^{-1/3}) and K̂_n = O_p(n^{-1/3}), so the inequality K̂_n < D̂_n + δ_n holds with probability tending to 1, not 0. The proof also contains an invalid equality: P(K̂_n < D̂_n + δ_n) = P(F_n(X_2) ∈ B(F_n(X_1), D̂_n + δ_n)) drops the conditioning on M(1) = i and the summation over i = 2,...,n; the left-hand side is a union over n−2 points, not a single-point probability. Since Lemma 5.13 relies directly on this lemma, the proof of P(N(i) = N̂(i)) → 1 is incomplete, and with it the transfer of the constants q_d, o_d from the population-rank graph to the empirical rank graph in Theorem 3.2. The intended statement appears to be true for the specific δ_n = n^{-1/2+ε} with ε < 1/2 − 1/d used in Lemma 5.13, and it can likely be proved via a Poisson-process gap estimate, but the lemma and its proof must be corrected.
  2. [§5.3, Lemmas 5.18 and 5.19] The limiting variance formula, which is the main quantitative result of Theorem 3.2, is imported as a black box. Lemma 5.18 is declared “a special case of Lemma 7.3 of Shi et al. (2024)” and Lemma 5.19 cites “the calculations as in Han and Huang (2024)” for the expression nVar(ξ_n) = 2/5 + (2/5)E[1/n ∑ T_{i,j}] + (4/5)E[1/n ∑ C_{i,j,k}] + o(1). Since these constants are the crux of the CLT, the paper should either reproduce the derivation or state precisely which theorem in each cited paper applies, and verify that its conditions, especially those involving the dependent graph statistics T_{i,j} and C_{i,j,k}, are met. As written, the reader cannot verify the claimed variance without consulting separate papers.
  3. [§5.1, Lemma 5.2] Lemma 5.2, which states E[Q_n] → Q, is essential for the consistency theorem (Theorem 3.1), yet its proof is omitted with only the note “due to the similarity to Azadkia and Chatterjee (2021).” For a main theorem, an omitted proof of a central lemma is not adequate. The authors should provide the argument or give a precise correspondence to Lemma 11.8 of Azadkia and Chatterjee (2021), explaining why the replacement of the Euclidean NNG by the rank-based NNG does not change the limiting expectation.
  4. [§5.1, Lemma 5.4] Lemma 5.4 contains an unjustified exact identity. The proof claims P((X_{N(i)}, Y_{N(i)}) ∈ B) = P((X_1, Y_1) ∈ B) from the identical distribution of the sample, but the event i → j biases the distribution of X_j; the displayed calculation yields P((X_1, Y_1) ∈ B | i → 1) in general, not the unconditional probability. The statement may hold asymptotically in view of Lemma 5.1, but as written it is false, and the conditional-independence claims in Lemmas 5.5–5.7 depend on it. A corrected asymptotic argument should be supplied.
minor comments (6)
  1. [Tables 1–3] Several rows in Tables 2 and 3 appear to have misaligned or missing entries, e.g., the rows for α = 10 and α = 500 in Table 2 do not have entries in all eight columns defined by the header. Please reformat the tables so that all columns are populated and clearly labelled.
  2. [Abstract and Introduction] The reference “Azadkia and Chatterjee (Azadkia and Chatterjee, 2021)” is redundantly phrased; the standard parenthetical citation style should be used.
  3. [§5, first paragraph] The notation “Let B(x, r) denote the ball in R^n centered at x ∈ R^n” uses the sample size n; the ball should be in the rank space R^d (or the appropriate fixed dimension), not in R^n.
  4. [Lemma 5.7] The symbol X'_i is used before it is defined; it should be introduced as an independent copy of X, with Y'_i drawn from the conditional law of Y given X = X'_i.
  5. [Proof of Theorem 3.1] The statement “By the continuous mapping theorem, ξ_n − ξ̃_n a.s.→ 0” is not a standard application of the continuous mapping theorem; the argument should be spelled out, noting that Q_n is bounded and P_n → P almost surely.
  6. [Corollary 5.1] The final sentence says “n D̂_n^d and n D̂_n^d both converge” but the second should be n D_n^d; the typo obscures which quantities are being compared.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the rank-based estimator is a new construction, and the cited prior results are independent external theorems, not the paper's own fitted inputs.

full rationale

The paper's central claims are (i) consistency of the rank-based NNG estimator ξn for the DSS measure ξ (Theorem 3.1) and (ii) asymptotic normality under independence (Theorem 3.2). Theorem 3.1 is derived self-contained in Section 5.1: E[Qn] → Q (Lemma 5.2, modeled on Azadkia–Chatterjee), Var(Qn) → 0 (Lemmas 5.4–5.8), and Pn → P (Lemma 5.9). No fitted parameter is renamed as a prediction, and no target quantity is inserted into the definition of the estimator; ξn is defined through Rosenbaum rank graphs and the DSS measure is a separate population functional. Theorem 3.2's variance constants qd and od come from Shi et al. (2024) and the variance decomposition is imported from Han and Huang (2024), both peer-reviewed works with overlapping authors (Fang Han). This is self-citation, but it is not circular in the prohibited sense: those results are external, parameter-free theorems with assumptions that do not include the present paper's conclusions, and the paper independently establishes the graph-equivalence Lemmas 5.11–5.16 that justify transferring those constants to the rank-based graph. Lemma 5.18 is explicitly a special case of Shi et al. Lemma 7.3; Lemma 5.19 cites 'calculations as in Han and Huang (2024)' for algebra, with the graph limits proved here (Lemma 5.10, Corollary 5.2). Thus the central derivation does not reduce to its own inputs. Two non-circular caveats: Lemma 5.12 contains an invalid equality (dropping the n−1 summation factor), so the proof of Theorem 3.2 is currently incomplete as written; and Theorem 3.2 explicitly excludes d=2. Both are correctness or completeness issues, not definitional circularity.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The central claims rest on eight background inputs: four imported proof steps from prior literature (two from the authors' own group), one imported geometric weak-convergence result, an unproved support condition for a.s. consistency, the CLT regularity conditions, and the uniform tie-breaking convention. No free parameters are fitted; the CLT constants q_d and o_d are derivable constants from the nearest-neighbor literature, not fitted values. The paper's own derivation is concentrated in the graph-equivalence lemmas (5.11-5.15) and the d = 1 combinatorics (Lemma 5.10).

assumptions (8)
  • domain assumption E[Q_n] → Q follows by the same proof as Azadkia-Chatterjee (2021, Lemma 11.8).
    Lemma 5.2's proof is omitted 'due to the similarity'; the transfer to the rank-based graph is asserted, not shown.
  • domain assumption The conditional law in Lemma 5.4: given X_i, X_j, X_N(i), the triple (Y_i, Y_j, Y_N(i)) equals (Y_i, Y_j, Y_k) given (X_i, X_j, X_k).
    Load-bearing for the variance computation; the proof relies on the asserted independence of Y_N(i) from (X_i, Y_i) given X_N(i), which is only sketched.
  • domain assumption Lemma 7.3 of Shi et al. (2024) applies verbatim to the rank-based estimator.
    Used as black box in Lemma 5.18 to show n(Var(Q̃_n|F_n) − Var(Q̃_n)) → 0.
  • domain assumption The variance decomposition nVar(ξ_n) = 2/5 + (2/5)E[ΣT/n] + (4/5)E[ΣC/n] + o(1) from Han and Huang (2024) holds.
    Lemma 5.19 states 'After algebraic manipulation (Shi et al., 2024; Han and Huang, 2024, for instance)', without reproducing the calculation.
  • domain assumption Henze (1987) Lemma 2.1 / Theorem 1.4: weak convergence of (X_1, nD_n, U_n) to the Poisson-process limit.
    Imported in Lemmas 5.15-5.16 to obtain the joint limit used for q_d and o_d.
  • ad hoc to paper X_1 lies in its support with every ball around F(X_1) carrying positive F(X)-mass.
    Needed in Lemma 5.1 for the a.s. NN convergence, stated without proof; Remark 3.1(iii) claims consistency holds for discrete X with ties, which is stronger than what the proof shows.
  • domain assumption Continuity: (X, Y) has a Lebesgue density and F(X) has a continuous copula density on its support (Theorem 3.2).
    Stated as assumptions in Theorem 3.2; they exclude discrete X from the CLT.
  • domain assumption Ties in the rank graph are broken uniformly by independent draws U_i ~ Uniform[0,1].
    Convention inherited from the estimator definition (Section 2); the d = 1 variance value 1 depends on this tie-breaking.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On a rank-based Azadkia-Chatterjee correlation coefficient." pith.science (2026). https://pith.science/paper/BKNV63SL

@misc{pith2026241202668,
  author       = {Pith},
  title        = {Pith review of: On a rank-based Azadkia-Chatterjee correlation coefficient},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BKNV63SL}},
  note         = {Machine review of arXiv:2412.02668}
}
read the original abstract

Azadkia and Chatterjee (Azadkia and Chatterjee, 2021) recently introduced a graph-based correlation coefficient that has garnered significant attention. The method relies on a nearest neighbor graph (NNG) constructed from the data. While appealing in many respects, NNGs typically lack the desirable property of scale invariance; that is, changing the scales of certain covariates can alter the structure of the graph. This paper addresses this limitation by employing a rank-based NNG proposed by Rosenbaum (2005) and gives necessary theoretical guarantees for the corresponding rank-based Azadkia-Chatterjee correlation coefficient.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectral analysis of large dimensional Chatterjee's rank correlation matrix

    math.ST 2025-10 conditional novelty 8.0 of 10

    A symmetrized Chatterjee rank correlation matrix has a semicircle spectral limit, plus a central limit theorem and independence tests that detect zero-linear-correlation dependence.

  2. Conditional Independence Testing Using Exchangeable Pairs

    math.ST 2025-09 conditional novelty 6.0 of 10

    A model-X conditional independence test using a Gaussian-kernel energy distance between observed data and conditionally independent variants, calibrated by random coordinate swaps.

Reference graph

Works this paper leans on

32 extracted references · 26 canonical work pages · cited by 2 Pith papers

  1. [1]

    and Fuchs, S

    Ansari, J. and Fuchs, S. (2022). A simple extension of Azadki a and Chatterjee’s rank correlation to a vector of endogenous variables. arXiv preprint arXiv:2212.01621

  2. [2]

    Auddy, A., Deb, N., and Nandy, S. (2024). Exact detection thr esholds for Chatterjee’s correlation. Bernoulli, 30(2):1640–1668. 22

  3. [3]

    and Chatterjee, S

    Azadkia, M. and Chatterjee, S. (2021). A simple measure of co nditional dependence. The Annals of Statistics , 49(6):3070–3102

  4. [4]

    Azadkia, M., Taeb, A., and Bühlmann, P. (2021). A fast non-par ametric approach for causal structure learning in polytrees. arXiv preprint arXiv:2111.14969

  5. [5]

    Bickel, P. J. (2022). Measures of independence and functiona l dependence. arXiv preprint arXiv:2206.13663. Bücher, A. and Dette, H. (2024). On the lack of weak continuity of Chatterjee’s correlation coeffi- cient. arXiv preprint arXiv:2410.11418

  6. [6]

    and Bickel, P

    Cao, S. and Bickel, P. J. (2020). Correlations with tailored e xtremal properties. arXiv preprint arXiv:2008.10177

  7. [7]

    D., Han, F., and Lin, Z

    Cattaneo, M. D., Han, F., and Lin, Z. (2025). On Rosenbaum’s r ank-based matching estimator. Biometrika (in press)

  8. [8]

    Chatterjee, S. (2021). A new coefficient of correlation. Journal of the American Statistical Associ- ation, 116(536):2009–2022

Show all 32 references
  1. [9]

    Chatterjee, S. (2023). A survey of some recent developments in measures of association. Probability and Stochastic Processes - A Volume in Honour of Rajeeva L. Ka randikar

  2. [10]

    Chen, L. H. and Shao, Q.-M. (2004). Normal approximation und er local dependence. The Annals of Probability, 32(3A):1985–2028

  3. [11]

    Deb, N., Ghosal, P., and Sen, B. (2020). Measuring associatio n on topological spaces using kernels and geometric graphs. arXiv preprint arXiv:2010.01768

  4. [12]

    and Kroll, M

    Dette, H. and Kroll, M. (2024). A simple bootstrap for Chatte rjee’s rank correlation. Biometrika (in press) , page asae045

  5. [13]

    F., and Stoimenov, P

    Dette, H., Siburg, K. F., and Stoimenov, P. A. (2013). A copul a-based non-parametric measure of regression dependence. Scandinavian Journal of Statistics , 40(1):21–41

  6. [14]

    Dvoretzky, A., Kiefer, J., and Wolfowitz, J. (1956). Asympt otic minimax character of the sample distribution function and of the classical multinomial est imator. The Annals of Mathematical Statistics, 27(3):642–669

  7. [15]

    Fuchs, S. (2024). Quantifying directed dependence via dime nsion reduction. Journal of Multivariate Analysis (in press) , 201:105266

  8. [16]

    Gamboa, F., Gremaud, P., Klein, T., and Lagnoux, A. (2022). G lobal sensitivity analysis: a new generation of mighty estimators based on rank statistics. Bernoulli, 28(4):2345–2374. 23

  9. [17]

    R., and Trutschnig, W

    Griessenberger, F., Junker, R. R., and Trutschnig, W. (2022 ). On a multivariate copula-based dependence measure and its estimation. Electronic Journal of Statistics , 16(1):2206–2251

  10. [18]

    and Huang, Z

    Han, F. and Huang, Z. (2024+). Azadkia-Chatterjee’s correl ation coefficient adapts to manifold data. The Annals of Applied Probability (in press)

  11. [19]

    Henze, N. (1987). On the fraction of random points by specifie d nearest-neighbour interrelations and degree of attraction. Advances in Applied Probability , 19(4):873–895

  12. [20]

    Hodges, Jr., J. L. and Lehmann, E. L. (1956). The efficiency of s ome nonparametric competitors of the t-test. The Annals of Mathematical Statistics , 27(2):324–335

  13. [21]

    Huang, Z., Deb, N., and Sen, B. (2022). Kernel partial correla tion coefficient—a measure of condi- tional dependence. The Journal of Machine Learning Research , 23(1):9699–9756

  14. [22]

    Kroll, M. (2024). Asymptotic normality of Chatterjee’s ran k correlation. arXiv preprint arXiv:2408.11547

  15. [23]

    and Han, F

    Lin, Z. and Han, F. (2022). Limit theorems of Chatterjee’s ra nk correlation. arXiv preprint arXiv:2204.08031

  16. [24]

    and Han, F

    Lin, Z. and Han, F. (2023). On boosting the power of Chatterje e’s rank correlation. Biometrika, 110(2):283–299

  17. [25]

    and Han, F

    Lin, Z. and Han, F. (2025). On the failure of the bootstrap for Chatterjee’s rank correlation. Biometrika (in press) . Rényi, A. (1959). On measures of dependence. Acta Mathematica Hungarica , 10(3-4):441–451

  18. [26]

    Rosenbaum, P. R. (2005). An exact distribution-free test co mparing two multivariate distributions based on adjacency. Journal of the Royal Statistical Society Series B: Statisti cal Methodology, 67(4):515–530

  19. [27]

    Rosenbaum, P. R. (2010). Design of Observational Studies . Springer

  20. [28]

    and Wolff, E

    Schweizer, B. and Wolff, E. F. (1981). On nonparametric measur es of dependence for random variables. The Annals of Statistics , 9(4):879–885

  21. [29]

    Shi, H., Drton, M., and Han, F. (2022). On the power of Chatter jee’s rank correlation. Biometrika, 109(2):317–333

  22. [30]

    Shi, H., Drton, M., and Han, F. (2024). On Azadkia–Chatterje e’s conditional dependence coefficient. Bernoulli, 30(2):851–877

  23. [31]

    Strothmann, C., Dette, H., and Siburg, K. F. (2024). Rearran ged dependence measures. Bernoulli, 30(2):1055–1078. 24

  24. [32]

    Zhang, Q. (2023a). On relationships between Chatterjee’s a nd Spearman’s correlation coefficients. arXiv preprint arXiv:2302.10131

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.