{"id":"a32db789-63c4-41fb-a916-7633f28d28a5","arxiv_id":"2412.02668","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A rank-based nearest-neighbor graph yields a scale-invariant Azadkia-Chatterjee correlation coefficient that is consistent and, for d ≠ 2, asymptotically normal under independence.","lead":"This paper proposes a scale-invariant version of the Azadkia-Chatterjee dependence coefficient, built on a nearest-neighbor graph of coordinate-wise ranks instead of raw distances. It proves the new estimator is consistent for the same population dependence measure and derives its null distribution for most dimensions, leaving dimension two open.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 5.12 is false as stated (δ_n ≫ n^{-1/d} makes the displayed probability tend to 1, not 0), and its proof contains an invalid equality; Theorem 3.2's d≥3 CLT depends on this lemma, so the proof is currently incomplete.","rationale":"After reading the proof chain, the weakest point is not the consistency theorem, whose proof has gaps but follows the Azadkia-Chatterjee template, but the CLT's graph-equivalence step. The manuscript itself flags d=2 as the singular case, which is a hint that d≥3 relies on the empirical-rank NN being asymptotically the same as the population-rank NN. That is Lemma 5.13, and it depends on Lemma 5.12. The reader's report spotted the incorrect equality in Lemma 5.12; I agree this is the load-bearing spot, and I add that the lemma's quantifier is actually false, not merely its proof. This is not a disagreement with the consensus; it is an internal correctness risk. The variance formula 2/5+(2/5)q_d+(4/5)o_d is inherited by counting the same graph motifs, so if the graph-equivalence lemma fails for any d≥3, the limiting variance could differ. The false statement is easily demonstrated with a slow δ_n, but the intended use has δ_n much smaller than the NN spacing; a standard Poisson heuristic gives the required rate n^{1/d−1/2}→0. Thus the theorem is likely true but not proven as written. The appropriate disposition is conditional acceptance pending a corrected Lemma 5.12; this does not change the reader's CONDITIONAL verdict. I have no reason to doubt the authors' good faith; the issue is a technical proof gap in a central lemma.","tokens_in":20293,"tokens_out":32032,"duration_ms":321437,"concrete_test":"Analytic check: (i) Refute the stated lemma by computing P(K̂_n<D̂_n+δ_n) for d=3 and δ_n=n^{-0.1}; the probability tends to 1 because K̂_n,D̂_n=O_p(n^{-1/3})≪δ_n. (ii) For the δ_n actually used in Lemma 5.13, δ_n=n^{-1/2+ε} with 0<ε<1/2−1/d, derive the sharp bound P(K̂_n<D̂_n+δ_n) ≤ C n^{1/d−1/2} by conditioning on X_1 and using the Poisson-process approximation to the rank point process. If (ii) holds, the corrected Lemma 5.12 rescues Lemma 5.13 and the variance constants stand; if it does not, the d≥3 CLT does not follow.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 5.12 is the key step in proving Lemma 5.13, P(N(i)=N̂(i))→1, which is what transfers the Shi et al. variance constants q_d,o_d from the population-rank NNG to the empirical rank NNG in Theorem 3.2. The lemma is stated for 'any sequence δ_n satisfying ||F_n−F||∞=o_p(δ_n)'. Take d=3 and δ_n=n^{-0.1}. Then ||F_n−F||∞=O_p(n^{-1/2})=o_p(n^{-0.1}), so the hypothesis holds, but D̂_n and K̂_n are O_p(n^{-1/3}), hence K̂_n < D̂_n + δ_n with probability tending to 1. The lemma is false as stated. Its proof also asserts P(K̂_n<D̂_n+δ_n)=P(F_n(X_2)∈B(F_n(X_1),D̂_n+δ_n)), which drops the M(1)=2 conditioning and the summation over i; the left side is a union over n−2 points, not a single-point probability. For the specific δ_n=n^{-1/2+ε} used in Lemma 5.13, with 0<ε<1/2−1/d, the intended statement is plausible and can likely be repaired by a Poisson-process gap estimate, but as written the lemma and its proof do not support Theorem 3.2, so the central CLT is currently unproven.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a rank-based version of the Azadkia–Chatterjee graph-based correlation coefficient, replacing Euclidean nearest neighbours of the raw covariates X_i by nearest neighbours in the transformed variables F_n(X_i), where F_n is the vector of marginal empirical CDFs. The main claims are: (Theorem 3.1) the resulting estimator ξ_n is consistent for the Dette–Siburg–Stoimenov measure ξ whenever Y is not a measurable function of X; and (Theorem 3.2) under independence, a Lebesgue density for (X,Y), continuity of the copula density, and d ≠ 2, √n ξ_n converges in distribution to N(0, σ²_d), with σ²_1 = 1 and σ²_d = 2/5 + (2/5)q_d + (4/5)o_d for d ≥ 3, where q_d and o_d are constants from Shi et al. (2024). The paper further reports simulations comparing the rank-based and original estimators across dimensions and scaling factors.","tokens_in":20468,"tokens_out":9984,"duration_ms":91890,"significance":"If the results hold, the paper delivers a scale-invariant alternative to the Azadkia–Chatterjee coefficient that preserves the same population target and enjoys asymptotic normality under independence. The consistency theorem is attractive for its minimal assumptions, and the explicit variance formula is a concrete contribution. The simulation study gives a useful demonstration of the practical advantage of rank-based construction when covariates are on different scales. However, the proof of the central limit theorem is not currently rigorous: one key lemma is false as stated, and several other load-bearing steps are either omitted or contain invalid equalities. Because the main probabilistic claims are plausible and the identified gaps appear local and repairable, the paper is a candidate for major revision rather than rejection.","major_comments":[{"comment":"Lemma 5.12 is false as stated. It asserts that for any sequence δ_n satisfying ‖F_n−F‖∞ = o_p(δ_n), P(K̂_n < D̂_n + δ_n) → 0. But if δ_n is much larger than the nearest-neighbour scale n^{-1/d}, e.g. d = 3 and δ_n = n^{-0.1}, then D̂_n = O_p(n^{-1/3}) and K̂_n = O_p(n^{-1/3}), so the inequality K̂_n < D̂_n + δ_n holds with probability tending to 1, not 0. The proof also contains an invalid equality: P(K̂_n < D̂_n + δ_n) = P(F_n(X_2) ∈ B(F_n(X_1), D̂_n + δ_n)) drops the conditioning on M(1) = i and the summation over i = 2,...,n; the left-hand side is a union over n−2 points, not a single-point probability. Since Lemma 5.13 relies directly on this lemma, the proof of P(N(i) = N̂(i)) → 1 is incomplete, and with it the transfer of the constants q_d, o_d from the population-rank graph to the empirical rank graph in Theorem 3.2. The intended statement appears to be true for the specific δ_n = n^{-1/2+ε} with ε < 1/2 − 1/d used in Lemma 5.13, and it can likely be proved via a Poisson-process gap estimate, but the lemma and its proof must be corrected.","section":"§5.3, Lemmas 5.18 and 5.19"},{"comment":"The limiting variance formula, which is the main quantitative result of Theorem 3.2, is imported as a black box. Lemma 5.18 is declared “a special case of Lemma 7.3 of Shi et al. (2024)” and Lemma 5.19 cites “the calculations as in Han and Huang (2024)” for the expression nVar(ξ_n) = 2/5 + (2/5)E[1/n ∑ T_{i,j}] + (4/5)E[1/n ∑ C_{i,j,k}] + o(1). Since these constants are the crux of the CLT, the paper should either reproduce the derivation or state precisely which theorem in each cited paper applies, and verify that its conditions, especially those involving the dependent graph statistics T_{i,j} and C_{i,j,k}, are met. As written, the reader cannot verify the claimed variance without consulting separate papers.","section":"§5.3, Lemmas 5.18 and 5.19"},{"comment":"Lemma 5.2, which states E[Q_n] → Q, is essential for the consistency theorem (Theorem 3.1), yet its proof is omitted with only the note “due to the similarity to Azadkia and Chatterjee (2021).” For a main theorem, an omitted proof of a central lemma is not adequate. The authors should provide the argument or give a precise correspondence to Lemma 11.8 of Azadkia and Chatterjee (2021), explaining why the replacement of the Euclidean NNG by the rank-based NNG does not change the limiting expectation.","section":"§5.1, Lemma 5.2"},{"comment":"Lemma 5.4 contains an unjustified exact identity. The proof claims P((X_{N(i)}, Y_{N(i)}) ∈ B) = P((X_1, Y_1) ∈ B) from the identical distribution of the sample, but the event i → j biases the distribution of X_j; the displayed calculation yields P((X_1, Y_1) ∈ B | i → 1) in general, not the unconditional probability. The statement may hold asymptotically in view of Lemma 5.1, but as written it is false, and the conditional-independence claims in Lemmas 5.5–5.7 depend on it. A corrected asymptotic argument should be supplied.","section":"§5.1, Lemma 5.4"}],"minor_comments":[{"comment":"Several rows in Tables 2 and 3 appear to have misaligned or missing entries, e.g., the rows for α = 10 and α = 500 in Table 2 do not have entries in all eight columns defined by the header. Please reformat the tables so that all columns are populated and clearly labelled.","section":"Tables 1–3"},{"comment":"The reference “Azadkia and Chatterjee (Azadkia and Chatterjee, 2021)” is redundantly phrased; the standard parenthetical citation style should be used.","section":"Abstract and Introduction"},{"comment":"The notation “Let B(x, r) denote the ball in R^n centered at x ∈ R^n” uses the sample size n; the ball should be in the rank space R^d (or the appropriate fixed dimension), not in R^n.","section":"§5, first paragraph"},{"comment":"The symbol X'_i is used before it is defined; it should be introduced as an independent copy of X, with Y'_i drawn from the conditional law of Y given X = X'_i.","section":"Lemma 5.7"},{"comment":"The statement “By the continuous mapping theorem, ξ_n − ξ̃_n a.s.→ 0” is not a standard application of the continuous mapping theorem; the argument should be spelled out, noting that Q_n is bounded and P_n → P almost surely.","section":"Proof of Theorem 3.1"},{"comment":"The final sentence says “n D̂_n^d and n D̂_n^d both converge” but the second should be n D_n^d; the typo obscures which quantities are being compared.","section":"Corollary 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily, and in places non-transparently, on the authors' own prior work (Shi et al. 2024; Han and Huang 2024) for the variance computation. While this is not circular, the editors may wish to consider whether the level of self-citation and the deferred proofs are appropriate for the journal. The d = 2 case is explicitly left open; that is an honest limitation, though it limits the breadth of the main theorem. The false lemma and invalid equality in the proof of Lemma 5.12 are serious but appear repairable within the paper's scope, so a major revision rather than rejection seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead this one carefully before citing the CLT. The idea is good: replace the Euclidean NNG in Azadkia-Chatterjee's coefficient with Rosenbaum's rank-based NNG, making the estimator invariant to component-wise monotone transformations. The consistency theorem (3.1) is plausible, and the d=1 limiting variance of 1, versus 2/5 for the original estimator, is a genuinely new and interesting result. Section 5.10's combinatorial computation for d=1 is clean.\n\nThe soft spots are in the d≥3 CLT. Lemma 5.12 is false as stated. It claims that for any δ_n with ||F_n − F||_∞ = o_p(δ_n), P(K̂_n < D̂_n + δ_n) → 0. Take d=3 and δ_n = n^{-0.1}. Then the hypothesis holds, but D̂_n and K̂_n are O_p(n^{-1/3}), so the event occurs with probability tending to 1. The proof also drops the M(1)=2 conditioning and the summation over n−2 points, replacing the probability with a single-point probability. This lemma is the load-bearing step for Lemma 5.13, which transfers the Shi et al. variance constants to the empirical rank graph. So Theorem 3.2 for d≥3 is unproven as written. It might be repairable—the intended δ_n = n^{-1/2+ε} route could work with a Poisson process gap estimate—but the current proof does not support the conclusion.\n\nOther issues: Lemma 5.2 (E[Q_n]→Q) is omitted entirely, Lemma 5.4 hand-waves a conditional independence step, and the d≥3 variance formula is imported as a black box from the authors' own Shi et al. and Han-Huang papers. The simulations use covariance matrices that are not positive semidefinite (d=3, ρ=0.9; d=5, ρ=0.9, and singular for d=5, ρ=0.5), and several table rows are incomplete.\n\nThe core direction is sound and the d=1 result is worth publishing. The d=2 case remains open, which the authors acknowledge. If the CLT proof is repaired and the simulations corrected, this is a solid contribution to the dependence-measures literature.\n\nI'd send it to a serious referee, with a clear request to focus on Lemma 5.12 and the transfer of variance constants. I wouldn't cite it as-is.\n\nBest,","headline":"A useful rank-based variant of the AC coefficient with a genuine d=1 finding, but the d≥3 CLT currently rests on a false lemma and the proof needs repair.","tokens_in":21220,"tokens_out":6029,"would_cite":false,"duration_ms":53840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G20","62H20","62G10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Building the nearest-neighbor graph on coordinate-wise ranks makes the Azadkia–Chatterjee correlation coefficient invariant to rescaling while preserving its consistency and normal limit.","keywords":["measure of dependence","nearest neighbor graph","rank transformation","Azadkia-Chatterjee correlation","scale invariance","asymptotic normality","copula density"],"falsifier":"Simulate independent X and Y in $R^{3}$ with a continuous copula, compute ξ_n for n up to several thousand, and record whether the empirical-rank nearest neighbor matches the population-rank nearest neighbor; if that match probability fails to approach 1, or if the empirical variance of √n ξ_n does not approach 2/5 + (2/5)q_3 + (4/5)o_3, then the central CLT claim fails.","tokens_in":19862,"feed_emoji":"📊","tokens_out":10968,"duration_ms":104644,"temperature":0.7,"pith_summary":"The paper builds a dependence coefficient that first maps each coordinate of X to its marginal rank and then constructs the nearest-neighbor graph in that rank space. It claims this rank-based Azadkia–Chatterjee coefficient is consistent for the same population dependence measure as the original graph-based coefficient, reaching zero exactly under independence and one exactly when Y is a measurable function of X. It also derives the null limit: under independence, √n times the coefficient is asymptotically normal with variance 1 in dimension one and variance 2/5 + (2/5)q_d + (4/5)o_d in dimensions at least three, where q_d and o_d are geometric constants of the rank graph. The payoff is a statistic that no longer changes when covariates are rescaled or monotonically transformed, yet still supports normal inference; the main caveat is that the proof of normality excludes dimension two.","feed_headline":"Rank-based graph yields a scale-invariant dependence coefficient","feed_subtitle":"Same dependence measure, immune to rescaling, with a normal limit under independence in all dimensions except two.","key_machinery":"The carrying object is the rank-based nearest-neighbor graph: each observation i is connected to the j that minimizes the Euclidean distance ‖F_n(X_i) − F_n(X_j)‖ between the vectors of coordinate-wise marginal empirical CDFs. Since each coordinate of F_n is monotone in that coordinate of X, the graph is invariant under component-wise strictly increasing transformations of X, which is exactly the scale-invariance property the original graph lacks. The proof then works by showing that for d ≠ 2 the empirical-rank graph is asymptotically the same as the population-rank graph, with neighbor identities matching with probability tending to one; this reduces the variance calculation to two limits of the population-rank graph, q_d and o_d, and the central limit theorem follows from a Hájek projection together with a dependency-graph Berry–Esseen bound.","core_discovery":"The paper's central claim is that replacing the raw Euclidean nearest-neighbor graph with a graph computed from coordinate-wise empirical ranks does not change what the Azadkia–Chatterjee statistic estimates. Concretely, ξ_n defined in (2.1) converges in probability to the Dette–Siburg–Stoimenov measure ξ whenever Y is not a measurable function of X, with no further assumptions, and it can be computed in O(n log n) time. When X and Y are independent and the copula density is continuous, √n ξ_n converges in distribution to a centered normal with variance $σ_1^{2}$ = 1 in one dimension and $σ_d^{2}$ = 2/5 + (2/5)q_d + (4/5)o_d in dimensions d ≥ 3, where q_d is the limiting proportion of mutual nearest-neighbor pairs and o_d is the limiting proportion of two points sharing a nearest neighbor in the rank-space graph. The proof does not cover d = 2, where the estimated-rank error and the nearest-neighbor distance converge at the same speed; the authors conjecture, with simulation support, that the same variance formula holds there.","pith_inferences":["A closed-form evaluation of the integrals defining q_d and o_d would turn the variance formula into a plug-in standard error for any d ≥ 3; the paper leaves them as numerical integrals.","The same rank-space construction transfers conceptually to other graph-based statistics, such as two-sample adjacency tests, mutual-information estimators, or conditional-dependence measures, and would plausibly give them the same scale invariance, though the paper does not pursue this.","If the d = 2 conjecture is eventually proven, the independence test becomes uniform across dimensions; until then, finite-sample inference in dimension two should be calibrated against the conjectured variance rather than taken from the theorem."],"forward_implications":["Monotone rescaling or coordinate-wise transformation of X leaves ξ_n unchanged, so the statistic reports the same dependence regardless of the measurement scale of the covariates.","Under independence, √n ξ_n has a known normal limit with explicit variance in dimensions one and three or more, giving an O(n log n) independence test whose null distribution is tractable.","The population target is the Dette–Siburg–Stoimenov measure, which is 0 exactly under independence and 1 exactly when Y is a measurable function of X; the rank-based estimator is consistent for this same target.","The formal CLT excludes d = 2; the paper's simulations and a related rank-based matching result suggest the variance formula carries over, but that case remains a conjecture in this paper."],"supporting_citations":[{"why":"Defines the original graph-based coefficient whose consistency framework and key lemmas are adapted here.","marker":"(Azadkia and Chatterjee, 2021)"},{"why":"Introduces the rank-based nearest-neighbor graph that the paper uses to achieve scale invariance.","marker":"(Rosenbaum, 2005)"},{"why":"Defines the population dependence measure ξ that the rank-based coefficient is shown to estimate.","marker":"(Dette et al., 2013)"},{"why":"Supplies the asymptotic-normality approach, the constants q_d and o_d, and Lemma 7.3 comparing conditional variances.","marker":"(Shi et al., 2024)"},{"why":"Provides Lemma D.1 used to reduce the estimator to its Hájek projection in the CLT proof.","marker":"(Deb et al., 2020)"},{"why":"Establishes the weak limit of nearest-neighbor distances and directions used to identify q_d and o_d.","marker":"(Henze, 1987)"},{"why":"Gives the uniform empirical-process bound that controls the error of estimated ranks throughout the proof.","marker":"(Dvoretzky et al., 1956)"},{"why":"Supplies the dependency-graph Berry–Esseen bound used to prove conditional asymptotic normality.","marker":"(Chen and Shao, 2004)"},{"why":"A related rank-based matching result cited as evidence that dimension two is not fundamentally different.","marker":"(Cattaneo et al., 2025)"}],"fun_headline_variants":["Rank-based twist makes dependence measure scale-free","Scale-invariant dependence: ranks fix the nearest-neighbor flaw","Ranks give Azadkia-Chatterjee coefficient a scale-free graph","New rank-based graph preserves dependence limit, resists rescaling","Dependence measure now immune to coordinate scaling"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that, in large samples, the nearest neighbor found using estimated marginal ranks is the same neighbor the true distribution would pick; the proof of this equivalence is exactly what breaks down in dimension two.","fun_headline_variants_meta":{"raw":{"variants":["Rank-based twist makes dependence measure scale-free","Scale-invariant dependence: ranks fix the nearest-neighbor flaw","Ranks give Azadkia-Chatterjee coefficient a scale-free graph","New rank-based graph preserves dependence limit, resists rescaling","Dependence measure now immune to coordinate scaling"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1332,"prompt_tokens":876,"completion_tokens":456,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":375}},"tokens_in":492,"tokens_out":456,"duration_ms":4233,"temperature":1.0,"reasoning_tokens":375,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:17:19.091477+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate independent X and Y in $R^{3}$ with a continuous copula, compute ξ_n for n up to several thousand, and record whether the empirical-rank nearest neighbor matches the population-rank nearest neighbor; if that match probability fails to approach 1, or if the empirical variance of √n ξ_n does not approach 2/5 + (2/5)q_3 + (4/5)o_3, then the central CLT claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the rank-based nearest-neighbor graph that the paper uses to achieve scale invariance."},{"cited_title":"F., and Stoimenov, P","cited_arxiv_id":null,"evidence_quote":"Defines the population dependence measure ξ that the rank-based coefficient is shown to estimate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the asymptotic-normality approach, the constants q_d and o_d, and Lemma 7.3 comparing conditional variances."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the weak limit of nearest-neighbor distances and directions used to identify q_d and o_d."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the uniform empirical-process bound that controls the error of estimated ranks throughout the proof."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the dependency-graph Berry–Esseen bound used to prove conditional asymptotic normality."},{"cited_title":"D., Han, F., and Lin, Z","cited_arxiv_id":null,"evidence_quote":"A related rank-based matching result cited as evidence that dimension two is not fundamentally different."}],"review_version":1}