Pith. sign in

REVIEW 2 cited by

Estimating Information Flow in Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.05728 v4 pith:JMGYU6CC submitted 2018-10-12 cs.LG stat.ML

classification cs.LGstat.ML
keywords compressioninformationrepresentationsnoisyclusteringhiddenmutualpast
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We study the flow of information and the evolution of internal representations during deep neural network (DNN) training, aiming to demystify the compression aspect of the information bottleneck theory. The theory suggests that DNN training comprises a rapid fitting phase followed by a slower compression phase, in which the mutual information $I(X;T)$ between the input $X$ and internal representations $T$ decreases. Several papers observe compression of estimated mutual information on different DNN models, but the true $I(X;T)$ over these networks is provably either constant (discrete $X$) or infinite (continuous $X$). This work explains the discrepancy between theory and experiments, and clarifies what was actually measured by these past works. To this end, we introduce an auxiliary (noisy) DNN framework for which $I(X;T)$ is a meaningful quantity that depends on the network's parameters. This noisy framework is shown to be a good proxy for the original (deterministic) DNN both in terms of performance and the learned representations. We then develop a rigorous estimator for $I(X;T)$ in noisy DNNs and observe compression in various models. By relating $I(X;T)$ in the noisy DNN to an information-theoretic communication problem, we show that compression is driven by the progressive clustering of hidden representations of inputs from the same class. Several methods to directly monitor clustering of hidden representations, both in noisy and deterministic DNNs, are used to show that meaningful clusters form in the $T$ space. Finally, we return to the estimator of $I(X;T)$ employed in past works, and demonstrate that while it fails to capture the true (vacuous) mutual information, it does serve as a measure for clustering. This clarifies the past observations of compression and isolates the geometric clustering of hidden representations as the true phenomenon of interest.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation

    cs.LG 2026-07 accept novelty 6.5 of 10

    A copula-theoretic benchmark suite reveals that non-parametric, discriminative and generative MI estimators each dominate only in specific regimes, with no universal winner.

  2. Entropy-Based Non-Invasive Reliability Monitoring of Convolutional Neural Networks

    cs.CV 2025-08 reject novelty 3.0 of 10

    A CNN's activation entropy separates clean from FGSM-attacked image batches in a small VGG-16 test, but fitted binning, tiny samples, and contradictory numbers weaken the claim.

Pith tools