{"id":"60e0a4af-d0cb-4330-962c-ec321197f521","arxiv_id":"2506.00330","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Neural MI estimators can become reliable in low-latent-dimension settings with a protocol of max-test early stopping, subsampling extrapolation, and probabilistic critics.","lead":"This paper proposes a practical protocol for estimating mutual information (MI) from high-dimensional data using neural networks, adding early-stopping rules, bias checks, and confidence intervals. If valid, it could help scientists trust MI estimates in regimes where data are scarce relative to dimensionality, such as neuroscience and imaging.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The protocol's most load-bearing assumption is its max-test stopping rule: reporting training MI at the checkpoint with peak test MI assumes the trained critic is near-optimal at that point, a premise only loosely justified in Appx. A.3.","rationale":"I read the paper as a practical protocol claim: reliable MI estimates with error bars in undersampled high-dimensional settings, provided a low-dimensional latent structure exists. The synthetic teacher-network experiments, the sample-size scaling in Fig. 5, the noisy MNIST result, and the comparisons to KSG, SMI, LMI, and the Czyz benchmark suite are genuine evidence and are presented with enough detail to reproduce. The latent-dimensionality claim is plausible and is the paper's central contribution. The weakest point is not the existence of latent structure but the stopping rule that turns finite-sample training curves into a bias-corrected estimate. The reader's Appx. A.3 concern is exactly the load-bearing one, and I agree with it. The appendix itself labels the proof \"loose\" and relies on \"empirical observations,\" and the tested families do not stress the assumption that the trained critic is close to the DV-optimal critic. A distribution with sharp, discrete latent structure would provide such a stress test. Separately, the abstract promises CIFAR-10/100 validation with a ResNet-20 backbone, but the full text contains no such experiments; this is a completeness gap that further supports keeping the verdict conditional, but I do not treat it as the main correctness risk. Because my concern reinforces rather than overturns the reader's CONDITIONAL verdict, I recommend no change to the verdict.","tokens_in":27119,"tokens_out":9046,"duration_ms":93055,"concrete_test":"Run the full protocol on a synthetic family with known MI where the optimal critic is deliberately outside the MLP family—for example, X and Y are deterministic functions of a discrete latent class with sharp cluster boundaries plus heavy-tailed noise, so the log-density-ratio is nearly a sum of delta-like modes. For sample sizes N=256 and N=1024 and 50 random seeds, compare (i) the reported training MI at the max-test checkpoint, (ii) the test MI at that checkpoint, and (iii) the true MI. If the training value misses the true MI by more than the reported prediction interval, or if it is farther from the truth than the test value, the stopping rule's key premise fails. A second check: re-derive Eq. (33) without the T_train*≈T* assumption to see whether any two-sided bound on |I_train−I_true| follows; if not, the gap is formal as well as empirical.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every headline estimate in Figs. 4-7 and Tables 1-2 is produced by the max-test stopping heuristic of Sec. 4.2: train a critic, pick the epoch with maximal held-out test MI, and report the training MI at that epoch. The formal justification in Appx. A.3 is an inequality (Eq. 33) showing only that the expected test MI is no larger than the expected training MI under the assumption that T_train* is close to the global optimum. That inequality does not establish that the training value at the selected checkpoint is close to the true MI; it is compatible with the training value being systematically high or low, and the argmax over epochs selects an extremum of a noisy test curve rather than the average case. The appendix explicitly calls the proof \"loose\" and then falls back on empirical observation. The empirical support is confined to Gaussian/teacher-network synthetic data and noisy MNIST; no distribution is tested where the true optimal critic is far from the MLP family, and no calibration study of the reported confidence intervals is presented. Because this stopping rule is the component that converts neural training curves into a point estimate with error bars, a failure for distributions outside the tested families would invalidate the central claim that the protocol provides reliable estimates and trustworthy confidence intervals even in the presence of low-dimensional latent structure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a practical protocol for neural mutual information (MI) estimation in high-dimensional, finite-sample settings. The protocol combines a max-test early-stopping rule, subsampling-based bias extrapolation, explicit unreliability flags, and confidence intervals, and it introduces probabilistic VSIB critics claimed to reduce bias and variance at high MI values. The central claim is that reliable MI estimation is governed by the true latent dimensionality K_Z rather than the ambient dimension K, so that MI can be estimated from hundreds to thousands of samples in problems where K is 500-784, provided the dependence is low-dimensional. The evidence includes synthetic teacher-network experiments with K=500, a subset of the Czyz et al. (2023) benchmark suite, and a noisy MNIST experiment with K=784. The abstract also claims validation on CIFAR-10/100 with a ResNet-20 backbone, but no such results appear in the body.","tokens_in":27460,"tokens_out":4353,"duration_ms":47280,"significance":"If the central claims hold, this would be a practically valuable contribution: it would turn neural MI estimators into more trustworthy scientific instruments by adding consistency checks, confidence intervals, and a clear statement of when estimates should not be trusted. The synthetic experiments are well designed for the latent-dimensionality hypothesis, and the MNIST result is a plausible demonstration of the proposed regime shift. The comparison to the Czyz et al. (2023) benchmarks is useful and generally favorable. However, the paper's theoretical grounding is a heuristic random-matrix calculation, and the load-bearing max-test stopping rule is justified only by a loose inequality plus empirical observation. The abstract's CIFAR-10/100 claim is unsupported in the body. These issues need to be resolved before the stronger claims of the paper can be accepted.","major_comments":[{"comment":"The max-test stopping rule is the component that converts neural training curves into point estimates with error bars, and it is used for every headline estimate in Figs. 4-7 and Tables 1-2. However, the formal justification in Appx. A.3 proves only that the expected test MI is bounded above by the expected training MI, and this only under the assumption that the trained critic is close to the globally optimal critic. That inequality does not establish that the training MI evaluated at the checkpoint selected by the argmax of a noisy test curve is close to the true MI; it is compatible with selection bias in either direction. The appendix itself labels the proof 'loose' and then falls back on empirical observation. Given the central role of this rule, the authors should either provide a sharper bound that accounts for selection over epochs or an explicit calibration/ablation study showing that the reported training value at the peak-test checkpoint is unbiased (or at least consistently closer to the truth than the test value) across a broader family of distributions.","section":"Sec. 4.2, Appx. A.3, Eq. (33)"},{"comment":"The abstract states that the protocol is validated on CIFAR-10/100 with a ResNet-20 backbone, but the full text contains no CIFAR experiments, no ResNet-20 results, and no tables or figures reporting CIFAR numbers. This is a load-bearing discrepancy because the abstract's claim of reliable MI detection 'well below the ambient pixel dimension on real images' rests on that validation. The authors should either add the missing experiments or remove the CIFAR-10/100 claim from the abstract and any summary statements.","section":"Abstract, Sec. 4.3"},{"comment":"The claim that sample complexity is 'grounded theoretically via random matrix theory' is stronger than what the appendix delivers. Equation (38) and the quadratic scaling in Eq. (39) are derived from a spiked-covariance detection threshold for linear Gaussian latent-variable models, and the derivation explicitly requires conditions ('if K_Z >> 1 and rank v << 2K_Z') that the authors state are 'neither strictly true in our model.' The extension to nonlinear teacher networks and to neural critics is asserted rather than derived, and the appendix acknowledges that the bound is optimistic because it ignores the cost of learning the nonlinear embedding. As written, this is a heuristic analogy, not a proof. The authors should either present the RMT analysis as a heuristic that motivates the empirical scaling, or prove a formal sample-complexity result for the nonlinear latent-variable setting.","section":"Appx. A.5, Eqs. (36)-(39)"},{"comment":"The paper's claim to be 'the only approach to report confidence intervals and flag unreliable estimates' is not backed by any calibration study. The reported intervals in Figs. 6-7 and Tables 1-2 are prediction intervals from weighted least squares fits and subset standard deviations, but there is no experiment measuring coverage, i.e., the fraction of trials in which the reported interval contains the true MI. Similarly, the unreliability thresholds (delta > 0.1 and gamma_max <= 5) are introduced in Appx. A.4 without sensitivity analysis. Because the practical value of the protocol depends on these intervals and flags being meaningful, the authors should add a coverage analysis on synthetic data where the true MI is known, and should test how the estimates and flags vary with the choice of delta and the gamma cutoff.","section":"Sec. 4.3, Appx. A.4, Figs. 6-7, Tables 1-2"}],"minor_comments":[{"comment":"There are typographical errors, e.g., 'estimators do not provide' should be 'estimators do not provide' with subject-verb agreement; a careful proofread is needed.","section":"Abstract and Sec. 1"},{"comment":"The vertical lines N*_Z and N* are defined only in Appx. A.5; the caption should give a one-sentence definition so the figure is self-contained.","section":"Fig. 5 caption"},{"comment":"The phrase 'we do not show the negative values' is ambiguous: it should state whether the displayed curves were clipped at zero or whether negative values simply fall outside the plotted range.","section":"Fig. 3 caption"},{"comment":"The symbol k_Z is used for the critic's embedding dimension, while K_Z is the true latent dimension; the distinction is important and should be stated consistently in every figure caption where both appear.","section":"Notation throughout"},{"comment":"The grey rows that mark unreliable fits may not be visually distinguishable in all rendering environments; consider adding an explicit symbol or column so the flag is readable regardless of color/greyscale reproduction.","section":"Tables 2-3"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision. The most serious issue is the unsupported CIFAR-10/100 claim in the abstract, which must either be backed by the missing experiments or removed. The max-test stopping rule is the load-bearing methodological choice, and its current justification in Appx. A.3 is explicitly loose; the authors should be asked either for a principled argument that accounts for checkpoint selection or for a calibration study. The RMT-based sample-complexity argument in Appx. A.5 is presented as a theory when it is a heuristic; this should be relabeled and, ideally, replaced by a formal statement or at least a transparent caveat in the main text. The self-citation to Abdelaleem et al. (2025) for the VSIB objective is appropriate, but the novelty boundary between that prior work and the present protocol should be clarified. The synthetic benchmarks and the MNIST result are valuable and, if the above points are fixed, the paper could become a strong contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth engaging. The paper's central claim—that neural MI estimators become reliable when the shared information lives in a low-dimensional latent space, and the sample complexity is set by that latent dimension—is well supported by the synthetic teacher-network experiments and the noisy MNIST result. I read the RMT analogy as a heuristic, not a derivation, and the authors mostly present it that way. The larger contribution is the assembly: max-test stopping, subsample extrapolation, critic-dimension plateau, and error bars. That protocol is new as a package, and the benchmark tables on the Czyz et al. suite show it consistently matches or beats a simple InfoNCE baseline while also saying when it does not trust itself. That is real value for practitioners.\n\nThe VSIB probabilistic critics are imported from the authors' own prior work, but they do real work here: SMILE-VSIB is the only estimator that tracks the 8-bit teacher model without saturating or exploding. That is a concrete, non-obvious result.\n\nSoft spots. First, the abstract promises CIFAR-10/100 with a ResNet-20 backbone; the body contains no such experiment, not even in the appendix. That is a direct abstract-body mismatch and needs fixing before publication. Second, the max-test stopping heuristic—report the training MI at the checkpoint with peak test MI—is load-bearing. The Appx. A.3 justification is an inequality showing the test value is expected to be no larger than the train value, which is compatible with the train value being anywhere above the true MI. The authors call it loose and lean on empirics. The empirics cover teacher networks and noisy MNIST, which is decent but not enough to establish that the reported confidence intervals have the right coverage outside those families. A calibration study on the CIs would go a long way. Third, the RMT sample-complexity argument is a spike-detection heuristic applied to nonlinear embeddings; it is suggestive and useful for intuition, but not a proof. The paper would be stronger if that were said in the main text, not just in the appendix.\n\nThere are minor issues: the workflow has several free parameters (smoothing windows, patience, beta, tau, delta) and the paper does not systematically study sensitivity to them; a user following the protocol may not know how robust their answer is to those choices. Also, the claim that the protocol never significantly overshoots is based on the tested families; it is not a guarantee.\n\nWho is this for? Anyone estimating MI from finite neural or imaging data would get practical guidance. It deserves a serious referee, with the expectation that the CIFAR gap be filled or the abstract claim removed, and the stopping-rule justification either tightened or tested with a CI coverage study. I would engage with it.","headline":"A practical protocol for neural MI estimation with error bars and honest diagnostics, but the promised CIFAR results are missing and the stopping rule needs stronger justification.","tokens_in":27945,"tokens_out":2238,"would_cite":true,"duration_ms":22098,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62B10","94A17"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims neural mutual-information estimators become reliable in high-dimensional data when the dependence between the variables lives on a low-dimensional latent space, and builds a protocol with error bars and built-in failure…","keywords":["mutual information estimation","neural estimators","high-dimensional data","latent dimensionality","sample complexity","early stopping heuristic","confidence intervals","random matrix theory"],"falsifier":"On a teacher-network generator with known latent dimension ($K_Z = 10$) embedded into $K = 500$ dimensions but with latent variables drawn from a distribution the paper did not test—e.g., heavy-tailed, discrete, or concentrated on a low-dimensional curved manifold—run the full protocol at $N = 256$ and $N = 1024$ and check whether the reported estimate plus prediction interval covers the known true MI. If the coverage fails, or if the bias changes sign when the embedding map is varied with $K_Z$ held fixed, the max-test stopping rule's core premise is falsified. A sharper test removes the paper's own escape hatch: restrict the critic family so that the optimal density-ratio critic is provably outside it, since the Appx. A.3 justification for preferring the training value over the test value requires the trained critic to be near-optimal.","tokens_in":26905,"feed_emoji":"📊","tokens_out":13801,"duration_ms":119508,"temperature":0.7,"pith_summary":"Mutual information (MI) is the standard measure of statistical dependence, but in high-dimensional data with few samples ($N \\lesssim K$) all known estimators either break down or give the user no way to know they have broken down. This paper's central claim is that neural MI estimators become reliable precisely when the dependence between the two variables lives on a low-dimensional latent space: the sample size needed is set by the latent dimension $K_Z$, not the ambient dimension $K$. The authors support the claim with a random-matrix-theory argument tying the onset of a nonzero estimate to the detection threshold of a spiked covariance signal in the latent space, and they assemble a practical protocol around it—a max-test early-stopping rule, a bias-removing subsampling extrapolation, and prediction intervals. They also introduce probabilistic (VSIB) critics that stabilize estimation at high MI values where standard estimators saturate or overfit. Across synthetic $K=500$ problems, a standard 40-dataset benchmark suite, and noisy MNIST ($K=784$), the pipeline matches or beats existing methods while being the only one that reports confidence intervals and declares when it cannot be trusted.","feed_headline":"Sample needs track latent dimension, not data dimension","feed_subtitle":"A protocol with error bars recovers near-true mutual information on 784-pixel images from 16,384 samples.","key_machinery":"Three mechanisms carry the argument. First, the generalized critic $T(x,y) = f(g(x), h(y))$: separable embeddings force the estimator through a latent bottleneck of dimension $k_Z$, and the paper shows a critic needs $k_Z \\ge K_Z$ to capture all dependence, while modestly exceeding $K_Z$ is harmless—turning the embedding dimension into a dial that reveals the structure of the data. Second, the max-test stopping rule: because MI is a nonlinear functional of the distribution, unbiased density estimates do not yield unbiased MI, and the held-out MI curve rises then collapses as the critic overfits; the paper selects the checkpoint with peak test MI and reports the corresponding training MI, arguing in Appx. A.3 that the test value is systematically biased downward while the training value at the best-generalizing checkpoint tracks the truth. Third, the random-matrix-theory detection bound: in a spiked-covariance model of the latent dependence, a signal spike separates from sampling noise only when $N$ exceeds $N^*_Z \\approx 2K_Z/\\theta^2$, which scales as $K_Z^2$ in the weak-signal limit—this is what makes 'sample the latent space, not the data space' quantitative. The VSIB probabilistic critics add a fourth piece: stochastic encoders with the loss $I_E(X;Z_X) + I_E(Y;Z_Y) - \\beta I_D(Z_X;Z_Y)$ regularize the critic and control variance in the high-MI regime.","core_discovery":"The paper establishes a regime-shift principle: the difficulty of estimating MI is governed by the dimensionality of the statistical dependence itself, not of the observed variables. For data generated by a latent model with $K_Z \\le 10$ hidden variables embedded into $K = 500$ observed dimensions, trained critics recover the ground-truth MI once the number of samples satisfies $N \\gg K_Z$, even when $N$ is far below $K$; the estimate only begins to form once $N$ passes a latent-space detection threshold $N^*_Z$ that scales roughly as $K_Z^2/I$, matching the spiked-covariance phase transition of random matrix theory. The paper also establishes a protocol that makes this usable: stop the critic at the epoch where held-out MI peaks and report the training value at that checkpoint; grow the critic embedding until the estimate plateaus at $k^*_Z$; subsample into $\\gamma$ equal parts and extrapolate a weighted linear fit to $\\gamma \\to 0$; report the intercept as the MI estimate with a prediction interval, and refuse to report anything if the fit is nonlinear. A new family of probabilistic critics, VSIB, wraps InfoNCE or SMILE in stochastic encoders regularized by an information-bottleneck objective, and the paper shows this suppresses SMILE's severe overestimation at high MI. Empirically, the pipeline stays within error bars of the true MI on the benchmark suite and recovers $3.13 \\pm 0.12$ bits against a true $\\log_2 10 \\approx 3.32$ bits on 784-dimensional noisy MNIST from 16,384 samples, without ever significantly overshooting.","pith_inferences":["If the latent-dimension principle holds generally, practitioners can pre-register a data budget: estimate $K_Z$ from the critic-plateau curve and collect $N \\gtrsim K_Z^2$ samples, rather than treating ambient dimension as the driver of sample complexity.","The random-matrix-theory connection suggests a diagnostic the paper does not itself build: comparing the observed onset of nonzero MI against the $N^*_Z$ prediction for candidate $K_Z$ values could estimate the effective latent dimension of a real dataset directly from the data.","The validation set is concentrated on smooth, continuous dependence structures (teacher networks, Gaussian links, image labels); an untested stress case is dependence carried by discrete or non-smooth structure, where the critic's smoothness assumptions and the linear $\\gamma$-extrapolation could fail even with small $K_Z$.","The paper's inversion of standard practice—trust the training value at the peak-test checkpoint rather than the test value—rests on MI being a nonlinear functional; the same logic may apply to other nonlinear population functionals such as entropy or divergences, but the paper does not test that generalization."],"forward_implications":["Mutual information can be estimated with honest error bars from a few hundred samples in problems with ambient dimension $K \\approx 500$, provided the true dependence is low-dimensional ($K_Z \\approx 10$).","The data requirement grows roughly quadratically with the latent dimension, so the plateau in the critic-embedding curve both diagnoses the latent dimension and sets the sample budget needed for a trustworthy estimate.","Compressive embeddings are not an optional convenience in high dimensions; an estimator that does not project into a low-dimensional space cannot exploit the latent structure that makes estimation possible.","In high-MI regimes, the VSIB probabilistic critics keep estimates stable where InfoNCE saturates near $\\log(\\text{batch size}) \\approx 7$ bits and plain SMILE overfits upward.","The protocol returns a falsifiable output: if the $\\gamma$-extrapolation is nonlinear or the fit range shrinks below $\\gamma = 5$, the pipeline refuses to report a number, giving scientists a built-in failure signal that existing neural estimators lack."],"supporting_citations":[{"why":"Supplies the 40-dataset benchmark suite and the baseline InfoNCE numbers the protocol must match or beat.","marker":"Czyz et al. (2023)"},{"why":"Defines the InfoNCE estimator and its saturation near log(batch size), the gap that motivates the VSIB variant.","marker":"van den Oord et al. (2018)"},{"why":"Defines the SMILE estimator whose high-MI overfitting the VSIB wrapper is designed to control.","marker":"Song & Ermon (2019)"},{"why":"Introduces MINE, the Donsker-Varadhan neural estimator that all critics in this paper build upon.","marker":"Belghazi et al. (2018)"},{"why":"Catalogues variational MI bounds and shows Monte Carlo-normalized DV estimates lose the strict lower-bound property.","marker":"Poole et al. (2019)"},{"why":"Provides the KSG k-nearest-neighbor estimator, the classical baseline that fails in the $K=500$ latent-teacher regime.","marker":"Kraskov et al. (2004)"},{"why":"Establishes the eigenvalue phase transition used to derive the latent-space detection threshold $N^*_Z$.","marker":"Baik et al. (2005)"},{"why":"Supplies the spiked-covariance random matrix theory formalism behind the sample-complexity bound.","marker":"Potters & Bouchaud (2020)"},{"why":"Supplies the subsampling, extrapolation, and error-bar best practices that the protocol extends to neural estimators.","marker":"Holmes & Nemenman (2019)"},{"why":"Provides the VSIB variational objective that defines the new probabilistic critics.","marker":"Abdelaleem et al. (2025)"}],"fun_headline_variants":["Latent dim, not data dim, sets sample complexity for MI","Reliable MI estimates with error bars from few samples","Sample size scales with latent complexity, not ambient","Protocol provides confidence intervals for mutual information","Neural MI estimators made reliable via latent structure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the training MI at the checkpoint where held-out MI peaks is a better estimate of true MI than the held-out value itself; the paper's justification of this rule indirectly assumes the trained critic is already close to the ideal critic, so distributions for which that closeness fails could make the reported values and their error bars wrong even when low-dimensional latent structure exists.","fun_headline_variants_meta":{"raw":{"variants":["Latent dim, not data dim, sets sample complexity for MI","Reliable MI estimates with error bars from few samples","Sample size scales with latent complexity, not ambient","Protocol provides confidence intervals for mutual information","Neural MI estimators made reliable via latent structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000697,"raw_usage":{"total_tokens":3261,"prompt_tokens":1167,"completion_tokens":2094,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":783,"completion_tokens_details":{"reasoning_tokens":2020}},"tokens_in":783,"tokens_out":2094,"duration_ms":15436,"temperature":1.0,"reasoning_tokens":2020,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:07:37.010906+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a teacher-network generator with known latent dimension ($K_Z = 10$) embedded into $K = 500$ dimensions but with latent variables drawn from a distribution the paper did not test—e.g., heavy-tailed, discrete, or concentrated on a low-dimensional curved manifold—run the full protocol at $N = 256$ and $N = 1024$ and check whether the reported estimate plus prediction interval covers the known true MI. If the coverage fails, or if the bias changes sign when the embedding map is varied with $K_Z$ held fixed, the max-test stopping rule's core premise is falsified. A sharper test removes the paper's own escape hatch: restrict the critic family so that the optimal density-ratio critic is provably outside it, since the Appx. A.3 justification for preferring the training value over the test value requires the trained critic to be near-optimal.","supporting_citations":[{"cited_title":"Beyond normal: On the evaluation of mutual information estimators","cited_arxiv_id":null,"evidence_quote":"Supplies the 40-dataset benchmark suite and the baseline InfoNCE numbers the protocol must match or beat."},{"cited_title":"Mutual information neural estimation","cited_arxiv_id":null,"evidence_quote":"Introduces MINE, the Donsker-Varadhan neural estimator that all critics in this paper build upon."},{"cited_title":"On variational bounds of mutual information","cited_arxiv_id":null,"evidence_quote":"Catalogues variational MI bounds and shows Monte Carlo-normalized DV estimates lose the strict lower-bound property."},{"cited_title":"Estimating mutual information","cited_arxiv_id":null,"evidence_quote":"Provides the KSG k-nearest-neighbor estimator, the classical baseline that fails in the $K=500$ latent-teacher regime."},{"cited_title":"Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices","cited_arxiv_id":null,"evidence_quote":"Establishes the eigenvalue phase transition used to derive the latent-space detection threshold $N^*_Z$."},{"cited_title":"A first course in random matrix theory: for physicists, engineers and data scientists","cited_arxiv_id":null,"evidence_quote":"Supplies the spiked-covariance random matrix theory formalism behind the sample-complexity bound."},{"cited_title":"Estimation of mutual information for real-valued data with error bars and controlled bias","cited_arxiv_id":null,"evidence_quote":"Supplies the subsampling, extrapolation, and error-bar best practices that the protocol extends to neural estimators."},{"cited_title":"Deep variational multivariate information bottleneck-a framework for variational losses","cited_arxiv_id":null,"evidence_quote":"Provides the VSIB variational objective that defines the new probabilistic critics."}],"review_version":1}