{"id":"2dd3dc40-98df-4183-ac54-4ebf59c86ab0","arxiv_id":"2607.05273","paper_version":1,"verdict":"CONDITIONAL","confidence":"UNKNOWN","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"A cycle-counting-ratio estimator for the β-model achieves minimax-optimal MSE and consistency under the weak conditions θ_max→0 and θ_t‖θ‖₁→∞, even at network densities near log n/n.","lead":"The paper introduces a closed-form estimator for node-level parameters in the β-model of networks that works even when networks are extremely sparse. A smart generalist might read it because it solves a long-standing estimation problem in network statistics without requiring strong assumptions about parameter structure.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The minimax rate claim requires a stronger condition than advertised: Theorem 2's upper bound needs θ_t‖θ‖₁/(log n)²→∞, while the abstract and Theorem 4's parameter space Θ(ε_n) quietly inherit this restriction.","rationale":"The reader correctly identified the most load-bearing concern: the gap between the upper bound condition (15) requiring θ_t‖θ‖₁/(log n)² → ∞ and the lower bound/consistency condition (7) requiring only θ_t‖θ‖₁ → ∞. This is the central tension in the paper's claims. The abstract advertises minimax optimality under the weak conditions, but the proof of the upper bound (Theorem 2) requires the stronger condition, and the minimax parameter space Θ(ε_n) inherits this restriction through Θ_0. The reader also noted the threshold truncation at log n, which is a secondary concern—the truncation is accounted for in the theorems (they analyze β̂*_t, not β̂_t), and the paper argues |β_t| < log n is reasonable for real networks. The primary concern about the condition gap is well-placed and is the single most important issue for the paper's central claim. The CONDITIONAL verdict is appropriate: the contribution is real (explicit estimator, minimax framework, consistency near Erdős-Rényi lower bound), but the gap between advertised and proven conditions needs to be either closed in the proof or acknowledged more transparently. Proofs are in unavailable supplementary material, which adds to the uncertainty. The verdict should remain CONDITIONAL pending verification of the supplementary proofs, particularly whether the (log n)² factor in condition (15) is essential or can be relaxed.","tokens_in":17050,"tokens_out":2155,"duration_ms":62643,"concrete_test":"Run simulations in the regime where θ_t‖θ‖₁ → ∞ but θ_t‖θ‖₁/(log n)² → 0 (e.g., θ_t‖θ‖₁ ~ (log n)^{1.5}), and compute the empirical MSE of β̂*_t across 1000+ replications for n = 1000, 2000, 4000. If the MSE scales as 1/(θ_t‖θ‖₁) (matching r_n(β)), the stronger condition (15) is likely an artifact of the proof technique and the weak-condition claim may hold. If the MSE deviates substantially from r_n(β) in this regime, the gap is genuine and the abstract's claim is overstated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states: 'Under the very weak conditions that max_t θ_t → 0 and θ_t‖θ‖₁ → ∞, we show that the CCR estimator is consistent and achieves the minimax rate.' However, the MSE upper bound (Theorem 2) explicitly requires the stronger condition (15): θ_t‖θ‖₁/(log n)² → ∞, which is (log n)² more restrictive than condition (7). The minimax optimality (Theorem 4) is established over Θ(ε_n) ⊂ Θ_0, where Θ_0 includes the requirement θ_t‖θ‖₁/(log n)² ≥ (1/2)log(log n). So the parameter space over which minimax optimality is proven already embeds the stronger condition. The lower bound (Theorem 3) only needs condition (7), creating a genuine gap: the lower bound holds under weaker conditions than the upper bound. The CLT (Theorem 5) does hold under condition (7) alone, giving convergence in distribution with variance ≈ 1/(θ_t‖θ‖₁), but convergence in distribution does not directly imply the MSE bound without a uniform integrability argument, which is not provided. Thus the claim that the minimax rate is achieved 'under the very weak conditions' is not fully established: the MSE rate is proven only under the stronger condition (15). The gap between (7) and (15) is not merely technical—it is a (log n)² factor that determines whether the central claim of minimax optimality under weak conditions actually holds.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper proposes a Cycle Counting Ratio (CCR) estimator for node-specific parameters in the β-model of undirected networks, targeting moderate to extremely sparse regimes. The estimator is based on the log-ratio of two 3-cycle counting statistics and has a closed-form expression, avoiding the iterative algorithms and non-existence issues of the MLE in sparse settings. The authors establish a signal-to-noise ratio (SNR) of θ_t‖θ‖_1, derive an MSE upper bound (Theorem 2), a matching minimax lower bound via the two-point χ² method (Theorems 3–4), and a CLT (Theorem 5). The central claim is that the CCR estimator achieves the minimax MSE rate 1/(θ_t‖θ‖_1) under the conditions θ_max → 0 and θ_t‖θ‖_1 → ∞, which are weaker than those in prior work.","tokens_in":17856,"tokens_out":1360,"duration_ms":95684,"significance":"The paper addresses an important open problem: estimation in the β-model under weak conditions in sparse networks, without requiring the structured parameter assumptions of Chen et al. (2021) or Shao et al. (2023). The closed-form, computationally scalable estimator is a practical strength. The minimax lower bound via the two-point method and the matching upper bound constitute a genuine theoretical contribution. The CLT and variance estimator (Theorems 5–6) enable inference. Simulations and a real-data application illustrate the method's viability. The gap between the upper and lower bound conditions (see major comments) does not negate the contribution but does affect how the central claim should be stated.","major_comments":[{"comment":"The abstract states: 'Under the very weak conditions that max_t θ_t → 0 and θ_t‖θ‖_1 → ∞, we show that the CCR estimator is consistent and achieves the minimax rate.' However, the MSE upper bound (Theorem 2, §3.2) explicitly requires the stronger condition (15): θ_t‖θ‖_1/(log n)² → ∞. The minimax optimality (Theorem 4) is established over Θ(ε_n) ⊂ Θ_0, where Θ_0 includes the requirement θ_t‖θ‖_1/(log n)² ≥ (1/2)log(log n). Thus the parameter space over which minimax optimality is proven already embeds the stronger condition. The lower bound (Theorem 3) only needs condition (7). This creates a genuine gap: the lower bound holds under weaker conditions than the upper bound. The CLT (Theorem 5) does hold under condition (7) alone, but convergence in distribution does not directly imply the MSE bound without a uniform integrability argument, which is not provided. The claim that the minimax率","section":null},{"comment":"The relationship between Theorem 5 (CLT under condition (7)) and Theorem 2 (MSE bound under condition (15)) should be clarified. If the authors can provide a uniform integrability argument to bridge the CLT to the MSE bound under condition (7) alone, the gap would be closed. If not, the abstract and Theorem 4 statement should be revised to accurately reflect that the MSE rate is proven under condition (15), not condition (7). The current phrasing 'Under a slight stronger condition' in the abstract for asymptotic normality suggests the authors are aware of the distinction, but the minimax claim does not make the same distinction clear. The parameter space Θ_0 in §3.2 should be discussed more transparently in relation to condition (7), not only in Section 3.2.","section":null}],"minor_comments":[{"comment":"In the definition of Θ_0 (§3.2), the term log(log n) is said to be chosen 'only for convenience and can be replaced by other diverging sequences.' A brief remark on how sensitive the results are to this choice would help the reader.","section":null},{"comment":"The threshold estimator (3) truncates at log n. The sensitivity of the theoretical results to this specific threshold is not rigorously bounded. A brief discussion of whether other thresholds (e.g., C log n for some constant C) would yield the same asymptotic properties would improve clarity.","section":null},{"comment":"Table 1 caption mentions 'based on 100 generated networks in each simulation,' but the text in §4.1 states 'Each simulation is repeated 500 times.' Please reconcile.","section":null},{"comment":"In Section 4.2, the phrase 'the leading estimator is meaningfulness' appears to be a grammatical error; consider revising to 'the leading estimator is not meaningful' or similar.","section":null},{"comment":"The reference to 'Feng et al. (2026)' in the introduction and reference list appears to be a future-dated preprint (arXiv:2601.01325). Confirming the correct year and providing complete bibliographic details would be helpful.","section":null},{"comment":"Remark 2 states that the asymptotic variance of the CCR estimator matches that of the MLE. This is an interesting point; a brief discussion of whether this holds only under θ_max → 0 or under broader conditions would strengthen the remark.","section":null},{"comment":"The condition θ_t‖θ‖_1 → ∞ in (7) is described as 'necessary' in §3.1. A one-sentence justification (e.g., SNR → 0 otherwise) is given, but making this more explicit would help readers less familiar with the signal-to-noise framework.","section":null}],"recommendation":"major_revision","confidential_remarks":"The core methodological contribution (closed-form estimator for sparse β-models without structured parameter assumptions) is novel and valuable. The main issue is a presentation gap between the advertised conditions and the proven conditions for the minimax MSE rate. This is fixable either by (a) providing a uniform integrability argument to bridge the CLT to the MSE bound under condition (7), or (b) revising the abstract and theorem statements to accurately state that the MSE rate requires condition (15). Option (b) is straightforward and would likely suffice for acceptance. The paper is a good fit for the journal once this is addressed."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The referee correctly identifies a gap between the conditions under which our lower and upper bounds hold, and between the CLT and the MSE bound. We address both major comments below and commit to revising the manuscript accordingly.","responses":[{"response":"The referee is correct on all counts. There is a genuine gap between the conditions for the lower bound (condition (7)) and the upper bound (condition (15)). We have carefully re-examined whether a uniform integrability argument can bridge the CLT (Theorem 5, under condition (7)) to the MSE bound under condition (7) alone. While the CLT gives convergence in distribution at the correct rate, we have not been able to establish the uniform integrability of the squared studentized statistic under condition (7) alone; the technical difficulty is that the tail behavior of the log-ratio statistic depends on the probability that the cycle counts T_{n,t}(a) or T_{n,t}(b) are very small, and controlling this requires the stronger condition (15). We will therefore revise the abstract, the introduction, and the statement of Theorem 4 to accurately reflect that the MSE rate 1/(θ_t‖θ‖_1) is proven under condition (15), not condition (7). The lower bound (Theorem 3) holds under condition (7), and the CLT (Theorem 5) holds under condition (7), but the matching upper bound requires the stronger condition. We will state this distinction clearly throughout the paper.","revision_made":"yes","referee_comment":"The abstract states the minimax rate is achieved under conditions (7) alone, but Theorem 2 (upper bound) requires the stronger condition (15), and the parameter space Θ_0 for Theorem 4 embeds condition (15). The lower bound (Theorem 3) only needs condition (7). The CLT (Theorem 5) holds under (7) but does not directly imply the MSE bound without a uniform integrability argument. The minimax claim should be revised."},{"response":"We agree that the relationship between these results and the role of Θ_0 need to be stated more transparently. In the revision, we will: (1) add an explicit remark explaining that Theorem 5 (CLT) holds under condition (7) alone, while Theorem 2 (MSE upper bound) requires the stronger condition (15), and that the gap reflects the absence of a uniform integrability argument; (2) clarify in Section 3.2 that Θ_0 embeds condition (15) via the requirement θ_t‖θ‖_1/(log n)^2 ≥ (1/2)log(log n), and that this is stronger than condition (7); (3) add a sentence noting that if one neglects the log n factors, Θ_0 essentially corresponds to condition (7), but that the log n factors are essential for the upper bound proof. We will also adjust the abstract's phrasing to make the distinction between the consistency/minimax result (under condition (15)) and the CLT (under condition (7)) unambiguous.","revision_made":"yes","referee_comment":"The relationship between Theorem 5 (CLT under condition (7)) and Theorem 2 (MSE bound under condition (15)) should be clarified. The parameter space Θ_0 should be discussed more transparently in relation to condition (7)."}],"tokens_in":16870,"tokens_out":727,"duration_ms":54582,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper introduces a closed-form estimator for β-model degree parameters that works in sparse regimes where the MLE doesn't exist. That is a genuine contribution. The estimator is a log-ratio of two 3-cycle counts (one involving edges, one involving non-edges), and it has O(n·d_max) complexity — practical for large sparse networks. The CLT (Theorem 5) holds under the weak condition θ_t‖θ‖₁ → ∞, and the asymptotic variance matches the MLE's variance in the dense regime. The simulations show the estimator is comparable to the MLE when the MLE exists and still works when it doesn't. This is real and useful work. The cycle-counting technique itself isn't new — the same authors used it for community detection and reciprocity estimation — but the application to β-model parameter estimation is new, and Proposition 1 (the structural identity that makes the ratio work) is clean. The two-point lower bound (Theorem 3) is standard but correctly executed. Now the soft spots. The stress-test concern lands. The abstract says the CCR estimator 'achieves the minimax rate' under the weak conditions max θ_t → 0 and θ_t‖θ‖₁ → ∞. But the MSE upper bound (Theorem 2) requires the stronger condition θ_t‖θ‖₁/(log n)² → ∞. The minimax optimality (Theorem 4) is stated over a parameter space Θ(ε_n) that embeds this stronger condition via Θ₀. So the lower bound holds under weaker conditions than the upper bound, and the gap is a (log n)² factor — not a rounding error. The CLT does hold under the weak condition alone, but convergence in distribution doesn't give you the MSE bound without a uniform integrability argument, which isn't provided. The abstract oversells what's proven. The threshold truncation at log n (equation 3) is justified heuristically, not rigorously. The sensitivity of the results to this choice isn't bounded. In the extremely sparse simulation regime (density ≈ 1/n), the errors are notably larger, which is consistent with the theory being on thinner ice there. Proofs are in supplementary material that I can't verify. These are fixable issues — the authors could either prove the upper bound under the weaker condition or simply state the minimax result under condition (15) and be honest about the gap. The paper is for network statisticians working on sparse graph models. It deserves a serious referee who can check the supplementary proofs and push the authors to align the abstract with what the theorems actually deliver. I'd recommend revise-and-resubmit with attention to the gap between conditions (7) and (15).","headline":"Cycle-counting estimator for β-model parameters in sparse networks: novel construction, real contribution, but the minimax claim is overstated relative to what is proven.","tokens_in":18076,"tokens_out":649,"would_cite":true,"duration_ms":41036,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Triangle counts crack sparse-network parameter estimation","keywords":["beta-model","sparse networks","subgraph counting","cycle counting ratio estimator","minimax estimation","degree heterogeneity","asymptotic normality","network density"],"falsifier":"If networks with density near log n / n and heterogeneous degree parameters were generated and the CCR estimator systematically failed to concentrate (e.g., MSE not scaling as 1/(theta_t ||theta||_1)), the main theoretical claims would be refuted. Additionally, if the threshold choice at log n were shown to introduce non-vanishing bias for specific parameter configurations, the consistency results would not hold as stated.","tokens_in":17294,"feed_emoji":"三角形","tokens_out":1117,"duration_ms":51615,"temperature":0.7,"pith_summary":"The paper proposes the Cycle Counting Ratio (CCR) estimator for the beta-model of undirected networks. The beta-model assigns each node a degree-heterogeneity parameter that controls its propensity to form edges, and the link probability between two nodes is a logistic function of the sum of their parameters. The CCR estimator works by counting two specific types of 3-node subgraphs (triangles) involving the target node -- one where the node has two edges and one non-edge to its neighbors, and one where it has two non-edges and one edge -- and taking half the log of their ratio. This ratio isolates the target node's parameter because the contributions from all other nodes cancel exactly. The paper proves that under the conditions that the maximum link parameter goes to zero (ensuring sparsity) and the product of a node's parameter with the sum of all parameters diverges, the CCR estimator is consistent and achieves the minimax mean-squared-error rate of 1/(theta_t times the L1 norm of the parameter vector). This rate persists even when overall network density approaches the Erdos-Renyi connectivity threshold of log n / n. The authors also establish asymptotic normality with variance matching that of the maximum likelihood estimator, and uniform consistency across all nodes under a slightly stronger condition.","feed_headline":"Triangle counts crack sparse-network parameter estimation","feed_subtitle":"A log-ratio of two triangle types recovers node-level degree parameters at minimax rate even near the Erdos-Renyi threshold.","key_machinery":"The estimator is half the log of the ratio of T_{n,t}(a) = sum of A_{ti} B_{ij} A_{jt} to T_{n,t}(b) = sum of B_{ti} A_{ij} B_{jt}, where A is the adjacency matrix and B records non-edges. By Lemma 1, these sums equal diagonal entries of ABA and BAB respectively, enabling matrix-based computation. A threshold at log n prevents infinite values when counts are zero. The signal-to-noise ratio is shown to be sqrt(theta_t ||theta||_1), making the squared SNR the inverse of the minimax rate.","core_discovery":"The central mechanism is that the ratio of expected counts of two complementary triangle types around a node equals exp(2*beta_t), so beta_t can be recovered by a simple log-ratio of observable subgraph counts. This converts a high-dimensional likelihood problem into n independent counting problems, each solvable in time proportional to the maximum degree. The minimax rate 1/(theta_t ||theta||_1) is established by matching an upper bound on the CCR estimator's MSE against an information-theoretic lower bound obtained via a two-point testing argument, where the chi-squared distance between distributions differing only in one node's parameter goes to zero precisely when the perturbation is on,","pith_inferences":["The cycle-counting ratio approach could extend to directed networks or models with reciprocity parameters, since the cancellation mechanism relies only on the additive structure of parameters in the link probability.","The gap between the condition needed for the upper bound (theta_t ||theta||_1 / (log n)^2 -> infinity) and the lower bound (theta_t ||theta||_1 -> infinity) suggests that a sharper concentration argument for the subgraph counts could close this log-squared gap.","If the asymptotic independence of estimators extends to growing dimensions, simultaneous confidence bands for the full parameter vector could be constructed, enabling global goodness-of-fit testing for the beta-model.","The choice of 3-cycles over longer cycles trades statistical efficiency (negligible improvement per simulations) for computational simplicity, but in regimes where 3-cycle counts are frequently zero, longer cycles might provide a fallback with weaker but non-degenerate estimates."],"forward_implications":["Networks near the Erdos-Renyi connectivity threshold log n / n -- where the MLE frequently fails to exist -- can still yield per-node parameter estimates with provable optimality guarantees.","The explicit closed-form estimator avoids iterative optimization, making it scalable to networks with millions of nodes provided the maximum degree is not too large.","The asymptotic variance matching the MLE suggests no statistical efficiency is lost by abandoning likelihood-based methods in sparse regimes.","The testing framework based on pairwise parameter differences enables hypothesis tests about degree heterogeneity structure even in very sparse networks."],"fun_headline_variants":["Triangle ratios estimate node parameters at minimax rate","Cycle log-ratios recover beta-model parameters in sparse networks","Complementary triangle counts decouple high-dimensional estimation","Subgraph ratios achieve consistency near Erdos-Renyi threshold","Triangle log-ratios turn likelihood into per-node counting problems"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The upper bound on the estimator's error requires a condition that is log-squared-n times more restrictive than the condition needed for the matching lower bound, so the minimax optimality is established over a parameter space that requires both conditions to hold simultaneously. Whether the minimax rate holds under the weaker condition alone remains open.","fun_headline_variants_meta":{"raw":{"variants":["Triangle ratios estimate node parameters at minimax rate","Cycle log-ratios recover beta-model parameters in sparse networks","Complementary triangle counts decouple high-dimensional estimation","Subgraph ratios achieve consistency near Erdos-Renyi threshold","Triangle log-ratios turn likelihood into per-node counting problems"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1138,"prompt_tokens":1075,"completion_tokens":63,"prompt_tokens_details":null},"tokens_in":1075,"tokens_out":63,"duration_ms":128493,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T20:09:28.425748+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If networks with density near log n / n and heterogeneous degree parameters were generated and the CCR estimator systematically failed to concentrate (e.g., MSE not scaling as 1/(theta_t ||theta||_1)), the main theoretical claims would be refuted. Additionally, if the threshold choice at log n were shown to introduce non-vanishing bias for specific parameter configurations, the consistency results would not hold as stated.","supporting_citations":[],"review_version":1}