{"id":"dcba02b2-1b68-4d4c-ba0c-05229c93f0aa","arxiv_id":"2608.02507","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The worst-case (1−δ) quantile of the logistic log-likelihood ratio is d log(en/d)+log(1/δ) for n≥d≥3, with d=2 at log log log n and d=1 at log(1/δ).","lead":"This paper gives worst-case finite-sample bounds for the log-likelihood ratio in logistic regression, uniform over all designs and parameters: order d log(en/d)+log(1/δ) for d≥3, with a surprising log-log-log n scale in dimension 2. Gaussian designs recover the classical chi-square scale once n≳d+log(1/δ).","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The d=2 upper bound (Theorem A.1) rests on an AI-generated proof sketch with two unverified key steps—the scalar reduction (58) and the pathwise planar Helly aggregation; until these are fully proved, the advertised d=2 log log log n scale is not established.","rationale":"The stress-test confirms the reader's identification of the d=2 upper bound as the load-bearing weakness. I examined the d≥3 proof and the d=2 lower bound and found no internal inconsistency. The Shtarkov upper bound is a standard counting argument; the Vandermonde lower bound is detailed and the constants check out. The d=2 lower bound construction (Lemma 4.1) is explicit and plausible. The remaining risk is entirely in Theorem A.1, whose proof sketch is AI-generated and leaves the two hardest steps asserted. This is not a criticism of AI assistance per se—the paper discloses it—but a correctness risk: a headline result that depends on an incomplete proof. The reader's CONDITIONAL verdict is exactly right. My proposed check—verifying (58) in the simplest nontrivial block and then the Helly aggregation—would settle whether the concern lands. If (58) is verified and a formal version of the Helly step is supplied, the d=2 result should be accepted; otherwise the conditional remains.","tokens_in":54057,"tokens_out":22248,"duration_ms":187709,"concrete_test":"Independently re-derive the scalar reduction (58) for a dyadic block of size |I_s|=2 with arbitrary q_i∈(0,1/2] and t_i, and verify by enumeration over b∈{0,1}^2 that G_s^(0)(b) ≤ G_m(b)+C+2C_m(b) for a universal C. If the inequality fails for any configuration, the key reduction behind Theorem A.1 is invalid; if it holds, the main missing link is the Helly aggregation, which should then be checked by explicitly applying the contrapositive of the Bárány–Katchalski–Pach theorem to the sets C_s.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Reader's weakest assumption is the right one. The d≥3 theorem (2.1) is supported by detailed arguments: Shtarkov upper bound (Prop 2.1, Lemma 2.2, Thm 4.1(i)) and Vandermonde lower bound (Prop 2.4) appear sound, as does the d=1 bound via Corollary 3.1. The d=2 lower bound (Prop 4.1/Cor A.1) is a rigorous construction giving log log log n. The missing piece is the matching upper bound Theorem A.1, placed in an appendix and explicitly credited to GPT-5.6-Sol with hints. The proof sketch's critical steps are (i) the scalar reduction (58) asserting a 2D dyadic-block LLR is bounded by a constant plus two 1D LLRs (intercept-only G_m and slope-only C_m) after profiling and conditioning; and (ii) the 'pathwise' application of planar quantitative Helly to data-dependent sublevel sets to aggregate H~log log n blocks with only a log H overhead. Neither is proved in full; the appendix says 'simplified proof sketch.' If either has a hidden gap—for instance, if (58) fails for some q_i,t_i, or if the Helly step requires extra regularity not present for all designs—the headline d=2 scale collapses. This is load-bearing for the d=2 claim, which is prominently advertised, though not for the d≥3 main theorem.","agreement_with_reader":"agree"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real contribution is Theorem 2.1: for n≥d≥3, the worst-case (1−δ) quantile of the logistic log-likelihood ratio, over every deterministic design and every target parameter, is d log(en/d)+log(1/δ) up to universal constants. The upper bound goes through Shtarkov sums and hyperplane arrangements; the lower bound uses Vandermonde subspaces. The argument is detailed, self-contained, and does not require MLE existence or design regularity. That is a genuinely new nonasymptotic Wilks analogue, and it yields honest finite-sample confidence sets. The d=1 bound log(1/δ), the Gaussian-design d+log(1/δ) bound, and the d^{3/2}/n boundary for Wilks approximation are also new and appear well supported. Credit where due: the core d≥3 theorem is serious mathematics.\n\nThe soft spot is exactly what the reader flagged. The d=2 upper bound, Theorem A.1, is in an appendix, explicitly credited to GPT-5.6-Sol, and presented as a 'simplified proof sketch.' Two steps carry the whole argument: the scalar reduction (58), which bounds a 2D dyadic-block LLR by intercept-only and slope-only 1D LLRs after profiling and conditioning, and the 'pathwise' application of planar quantitative Helly to data-dependent sublevel sets that aggregates H ~ log log n blocks at only log H cost. Neither is fully proved. If (58) fails for some margin configurations, or the Helly step needs regularity that worst-case designs do not provide, the log log log n scale collapses. The d=2 lower bound is rigorous, but a lower bound alone doesn't establish the headline scale.\n\nThe citation pattern is fine. The paper imports Gaussian-design quadratic bounds from a co-authored paper [14] and mentions transductive regret [45], but the main fixed-design theorem doesn't depend on them. Self-citation is not the issue here.\n\nWho should read this: anyone working on finite-sample likelihood inference, logistic MLE existence, or minimax regret. The d≥3 theorem is citable now. The d=2 claim should be labeled a conjecture until a complete proof appears. My recommendation: send to a serious referee. A good referee can verify the d≥3 proofs and demand that Theorem A.1 either be completed or removed from the abstract. The paper is a conditional reject/major revision as it stands.","headline":"The d≥3 worst-case LLR characterization is a genuine finite-sample Wilks analogue and deserves a serious referee; the d=2 log log log n upper bound is not yet proven — it sits in an appendix as an AI-generated proof sketch with load-bearing gaps.","tokens_in":54932,"tokens_out":2602,"would_cite":true,"duration_ms":27003,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62E20","62F25","62J12"],"pacs":[],"model":"deepseek-v4-flash","headline":"In binary logistic regression, the worst-case (1−δ) quantile of the log-likelihood ratio is d log(en/d) + log(1/δ) for n ≥ d ≥ 3, up to universal constants.","keywords":["logistic regression","log-likelihood ratio","Wilks phenomenon","Shtarkov sum","worst-case design","Vandermonde subspaces","confidence sets","Gaussian random design"],"falsifier":"For n = exp(exp(exp(20))), simulate the two-dimensional block design from Proposition 4.1 with k=20 and estimate the 90% quantile of the log-likelihood ratio; the paper's d=2 claim predicts it to be about 20 + log(1/δ) up to a universal constant, so a value outside a constant factor of that would refute the d=2 scale.","tokens_in":54003,"feed_emoji":"📊","tokens_out":9056,"duration_ms":90826,"temperature":0.7,"pith_summary":"The paper proves a finite-sample, worst-case characterization of the log-likelihood-ratio statistic in logistic regression. For n≥d≥3, the largest possible (1−δ) quantile over every fixed design matrix and every target parameter is, up to universal constants, d log(en/d) + log(1/δ); the d=2 worst case is only log log log n + log(1/δ), and d=1 is log(1/δ). This replaces the asymptotic Wilks chi-square calibration with a nonasymptotic guarantee that holds uniformly, with no regularity conditions on the design and no requirement that the maximum likelihood estimator exist. The payoff is a confidence set for the target parameter that covers with probability at least 1−δ for every design and every parameter. The proof turns the likelihood ratio into a supremum over a linear subspace, bounds its exponential moment by a Shtarkov sum, and shows the bound is tight via Vandermonde designs whose polynomial sign patterns force the logarithmic inflation; the d=2 upper bound is presented as a proof sketch in the appendix.","feed_headline":"Logistic likelihood-ratio worst case equals d log(en/d)","feed_subtitle":"Finite-sample, design-free confidence sets for logistic regression—no MLE required.","key_machinery":"The argument rests on three mechanisms. First, the log-likelihood ratio depends on the design only through the subspace W = ran(X) and on the target only through v = Xθ*, reducing the statistic to Γ_W(ε;v) = sup_{w∈W} log(p_w(ε)/p_v(ε)). Second, the exponential moment of Γ_W evaluated at λ=1 is exactly the Shtarkov sum S_log(W) = Σ_ε sup_{w∈W} p_w(ε), which is independent of v and bounded by the number of regions in an arrangement of at most n hyperplanes, at most Σ_{ℓ=0}^k C(n,ℓ) ≤ (en/k)^k; Markov's inequality then yields the upper tail. Third, matching lower bounds come from Vandermonde subspaces—ranges of polynomial-evaluation maps—whose sign patterns realize all patterns with few negati","core_discovery":"The paper's central discovery is a nonasymptotic Wilks-type law for logistic regression that is uniform in the worst possible way. For n≥d≥3, no matter which fixed design vectors x_1,...,x_n in R^d and no matter which target parameter θ* in R^d, the (1−δ) quantile of the log-likelihood ratio Λ satisfies Q_{1−δ} ≈ d log(en/d) + log(1/δ). The same scale holds for random design when the supremum is taken over all distributions. In dimension 2 the worst case drops to log log log n + log(1/δ), and in dimension 1 to log(1/δ). Under i.i.d. Gaussian design the logarithmic factor disappears: with n ≳ d + log(1/δ), the bound is d + log(1/δ), uniformly over θ*. The upper bounds require no MLE existence","pith_inferences":["The same subspace-and-Shtarkov machinery likely gives sharp finite-sample regret bounds for transductive sequential prediction of logistic labels; the paper's upper bound is effectively a worst-case regret of order d log(en/d) for the logistic class, and the lower bounds indicate this is minimax in the worst case.","The d=2 anomaly suggests that the plane has special geometry for this problem; one can ask whether other low-dimensional substructures inside higher-dimensional designs could create intermediate scales, a question the paper leaves open.","The Gaussian-design guarantee is proved only for Gaussian covariates, but the route through Hessian concentration suggests a testable extension to other random designs with well-conditioned covariance and sub-Gaussian tails; this is an editorial conjecture, not a paper claim.","The d=2 upper bound rests on an appendix proof sketch whose key reductions are asserted rather than fully fleshed out; until that sketch is completed, the log log log n scale should be treated as conditional, while the d≥3 statement does not inherit this caveat."],"forward_implications":["The confidence set {θ : Λ(θ) ≤ d log(en/d) + log(1/δ)} has coverage at least 1−δ for every fixed design and every θ*, with no regularity and no need for the MLE to exist.","The logarithmic factor log(en/d) is an unavoidable price of uniformity: some fixed designs and target parameters (Vandermonde designs with θ* along a coordinate axis) have likelihood-ratio quantiles of this order, so the classical chi-square scale cannot hold in worst-case fixed design.","Under i.i.d. Gaussian design, the classical chi-square-type scale d + log(1/δ) is restored uniformly over θ* as soon as n ≳ d + log(1/δ).","The validity of the classical Wilks approximation at the origin is governed by d^{3/2}/n → 0, not by the aspect ratio d/n; the distribution can be far from chi-square even when d/n is small.","The worst-case scale is sharply dimension-sensitive: d=1 costs only log(1/δ), d=2 costs log log log n, and d≥3 costs d log(en/d)."],"fun_headline_variants":["Worst-case LLR in logistic: d log(en/d) + log(1/δ)","Finite-sample Wilks: logistic worst case = d log(en/d)","Dimension 2 surprise: LLR worst case log log log n","Uniform over design and θ: LLR worst case d log(en/d)","No MLE needed: worst-case LLR in logistic is d log(en/d)"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything in the d≥3 theorem rests on standard inequalities, but the d=2 upper bound rests on the appendix's proof sketch—after fixing the common offset and conditioning on the number of mislabeled points, the remaining uncertainty is handled by a one-dimensional bound—and if that reduction or the planar Helly step has a gap, the log log log n scale for d=2 is not established.","fun_headline_variants_meta":{"raw":{"variants":["Worst-case LLR in logistic: d log(en/d) + log(1/δ)","Finite-sample Wilks: logistic worst case = d log(en/d)","Dimension 2 surprise: LLR worst case log log log n","Uniform over design and θ: LLR worst case d log(en/d)","No MLE needed: worst-case LLR in logistic is d log(en/d)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001071,"raw_usage":{"total_tokens":4365,"prompt_tokens":827,"completion_tokens":3538,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":3436}},"tokens_in":571,"tokens_out":3538,"duration_ms":26335,"temperature":1.0,"reasoning_tokens":3436,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T05:51:07.731269+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For n = exp(exp(exp(20))), simulate the two-dimensional block design from Proposition 4.1 with k=20 and estimate the 90% quantile of the log-likelihood ratio; the paper's d=2 claim predicts it to be about 20 + log(1/δ) up to a universal constant, so a value outside a constant factor of that would refute the d=2 scale.","supporting_citations":[],"review_version":1}