REVIEW 1 major objections 5 minor 12 references
Separable Computation of Information Measures
T0 review · 1 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper proves that any sufficient statistics s(X), t(Y) preserve a broad class of information measures, and it pins down Gács–Körner common information as the entropy of the singular functions with unit singular values.
desk verdict Sound and honestly scoped: exact sufficiency gives clean invariance theorems, with the Gács–Körner spectral characterization the real new money; the omitted proofs and approximation gap are fixable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the canonical dependence kernel (CDK), $i_{X;Y}(x,y)=\frac{P_{X,Y}(x,y)}{P_X(x)P_Y(y)}-1$, together with its modal decomposition: a singular value decomposition $i_{X;Y}=\sum_i \sigma_i f_i^*(x)g_i^*(y)$, with $\sigma_1\ge\sigma_2\ge\cdots>0$ and orthonormal singular functions. Two facts carry the argument: first, $S$ and $T$ are sufficient exactly when the density ratio of $(X,Y)$ factors through the density ratio of $(S,T)$, so the CDK is unchanged by reducing to sufficient features; second, the singular functions $f^*(X)$, $g^*(Y)$ are minimal sufficient statistics. The CDK also pinpoints the unit singular values: functions $f(X)$ and $g(Y)$ that coincide with probability 1 are exactly linear transforms of the singular functions with $\sigma_i=1$, which yields the Gács–Körner entropy formula and the other invariance proofs.
What would settle it
Take a small finite joint distribution satisfying sufficiency for a chosen $s(X)$, $t(Y)$, compute $C_{GK}$ both by brute-force enumeration of all functions $f,g$ with $P\{f(X)=g(Y)\}=1$ and by the formula $H(f_0^*(X),\ldots,f_k^*(X))$ from the modal decomposition; a single distribution where they differ would refute Theorem 2. Separately, choose an intentionally non-sufficient feature map, such as one that merges two symbols with different conditional distributions of $Y$, and compute $I(X;Y)$ versus $I(s(X);Y)$; seeing the equality fail confirms the condition is doing the work.
Extended reading notes
Core claim
The central claim is the invariance theorem: for finite-alphabet $X,Y$, if $S=s(X)$ and $T=t(Y)$ are sufficient statistics in the sense $X-S-Y$ and $X-T-Y$ (equivalently $X-S-T-Y$), then $I(X;Y)=I(S;T)$, $I_f(X;Y)=I_f(S;T)$ for every $f$-information, $C(X,Y)=C(S,T)$ for Wyner's common information, $C_{GK}(X,Y)=C_{GK}(S,T)$ for Gács–Körner common information, $L^*_{IB}(X,Y;\beta)=L^*_{IB}(S,T;\beta)$ for every $\beta>0$, and $\vartheta_{X,Y}(R)=\vartheta_{S,T}(R)$ for every $R\ge 0$ for the information bottleneck. Along the way the paper characterizes Gács–Körner common information as $H(f_0^*(X),\ldots,f_k^*(X))$, where $f_i^*$ are the left singular functions of the canonical dependence kernel and $k$ is the largest index with singular value $\sigma_k=1$. The optimal auxiliary variables in Wyner's problem and in the bottleneck are shown to be functions of the sufficient features, which is what makes separable computation possible.
Load-bearing premise
The entire set of equalities rests on the learned features being exactly sufficient statistics, so that $X-s(X)-t(Y)-Y$; if that Markov condition fails, the paper provides no guarantee and the equalities can break.
Editorial extensions
If this is right
- Feature-based mutual information estimation is exact whenever the learned features are sufficient statistics; the estimator no longer needs access to raw $X$ and $Y$.
- The optimal Wyner common-information channel $W$ can be restricted to depend on $(S,T)$, so common information can be computed from sufficient features without loss.
- Gács–Körner common information is a spectral quantity: it is the entropy of the top segment of canonical features whose singular values equal 1.
- For any $\beta>0$ and any $R\ge 0$, the information bottleneck Lagrangian optimum and the $\vartheta(R)$ curve are invariant under sufficient-statistic reduction, so bottleneck computations can be carried out on features.
- All $f$-information measures, including mutual information and divergences based on other convex functions with $f(1)=0$, are invariant under sufficient-statistic reduction.
Reading between the lines
- An extension the authors leave open is quantification: when features are only approximately sufficient, the equalities degrade, and the modal decomposition suggests a perturbation bound based on discarded singular modes with $\sigma_i<1$. The paper itself stops at the exact case.
- The Gács–Körner formula suggests a practical null hypothesis for learning common information: after estimating the CDK singular spectrum, the number of singular values at 1, together with $H$ of the corresponding $f_i^*$, is the quantity to test; finite-sample versions would require deciding how close to 1 counts as 1.
- Because Proposition 1 phrases sufficiency purely in terms of the CDK, any information measure that is a function of the CDK may inherit the same separable-computation property; the paper demonstrates this for the listed measures but does not give a general criterion.
- In continuous or weak-dependence settings, universal features have known analytic forms, so the same invariance might be derivable outside finite alphabets; the paper's proofs are restricted to finite alphabets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies when an information measure theta(X,Y) can be computed from learned feature representations s(X) and t(Y) without loss. Under the standing assumption that s(X) and t(Y) are sufficient statistics for X and Y, equivalently X-s(X)-t(Y)-Y, it proves four invariance results: mutual information and f-information satisfy I_f(X;Y)=I_f(S;T); Wyner's common information satisfies C(X,Y)=C(S,T); Gács-Körner common information is characterized by H(f_0^*(X),...,f_k^*(X)) with k the largest index such that sigma_k=1 in the modal decomposition; and the information bottleneck curve satisfies theta_{X,Y}(R)=theta_{S,T}(R) for all R>=0, with L*_IB(X,Y;beta)=L*_IB(S,T;beta) for all beta>0. The proofs use the canonical dependence kernel and its modal decomposition, together with several Markov-chain lemmas.
Significance. If the results hold, they give a clean conditional guarantee for modular, representation-based estimation of several information measures. The characterization of Gács-Körner common information through the singular modes with unit singular values is a nice connection. The proofs are mostly complete and self-contained, and the Markov-chain structure is made explicit. The main practical caveat, which the paper itself acknowledges, is that the guarantee is conditional on exact sufficiency; no approximation bounds or finite-sample statements are given for approximately sufficient learned features, so the advertised practical scope should be phrased carefully.
major comments (1)
- [Section V, Lemma 3] The lemma is stated with 'Proof: Omitted' and is subsequently used in the proof of Lemma 5 and in the proof of Theorem 3, so the information bottleneck invariance arguments depend on an unproved statement. The lemma is true, and a short proof can be supplied via the chain rule, but it should appear in the manuscript for the proof of Theorem 3 to be complete.
minor comments (5)
- [Section II-A, Eq. (1)] The denominator of the canonical dependence kernel is written as P_X(y)P_Y(y); it should be P_X(x)P_Y(y).
- [Section III-A, Corollary 1] The proof of Corollary 1 is omitted with 'We omit the proof.' Since it follows immediately from Proposition 1, please add the two-line derivation so that the f-information invariance is fully supported.
- [Section V-C, Eq. (25)] In the expansion of E[(f(X)-g(Y))^2], the term f_i^*(Y) should be f_i^*(X); the displayed formula currently misstates the argument.
- [Section IV] The text reads 'invariance to the choices of sufficient statics'; 'statics' should be 'statistics'.
- [Abstract and Introduction] The phrase 'mild assumptions' overstates the condition of exact sufficiency; consider wording that reflects that the theorems are conditional on s(X) and t(Y) being sufficient statistics, since approximate or learned features are outside the proven scope.
Circularity Check
No significant circularity: the sufficiency-based invariance theorems are proved from the Markov assumptions and modal decomposition, not from their conclusions.
full rationale
The derivation chain is self-contained conditional on the stated sufficiency assumption. Proposition 1 is proven from the definition of sufficiency and elementary density-factorization, and Corollary 1 follows immediately from that factorization. Theorem 1 constructs W' via Lemma 4 and uses the identity I(W';X,Y)=I(W;X,Y)-I(W;X,Y|S,T) to force the optimal W to satisfy W-(X,Y)-(S,T); this is a direct optimality argument, not an assumption of the conclusion. Theorem 2's Gacs-Korner proof is an independent spectral proof: Eq. (25) decomposes E[(f(X)-g(Y))^2] into a sum of nonnegative terms using the bound sigma_i <= 1, forcing f=A f-hat and g=A g-hat, and hence H(f(X)) <= H(f*_0,...,f*_k). The use of Proposition 2, cited as [4, Proposition 2], is a citation to a standalone minimal-sufficiency theorem whose assumptions do not include Gacs-Korner invariance or the separable-computation conclusion; it is therefore independent support rather than a self-referential premise. Theorem 3 follows from Lemma 5, which is proved from the Markov assumptions, and the omitted Lemma 3 is a standard graphoid composition that is readily supplied. No fitted parameter is renamed as a prediction, and no equality is assumed in the form of its conclusion. The only caveats are scope limitations the paper itself states: exact sufficiency is required, and no approximation bounds or finite-sample guarantees are given for learned approximately-sufficient features. The omitted proof of Lemma 3 is a completeness gap, not a circular step.
Assumptions & free parameters
assumptions (5)
- standard math Data processing inequality (Lemma 1)
- domain assumption Modal decomposition of the canonical dependence kernel via SVD (Section II-A), including Proposition 2 that f*(X), g*(Y) are minimal sufficient statistics
- standard math Lemma 3 (U-X-Y, U-(X,Y)-Z, X-Z-Y implies U-X-Z-Y), stated without proof
- domain assumption Finite alphabets with positive marginals for X and Y
- domain assumption S=s(X) and T=t(Y) are exact sufficient statistics, i.e., X-S-T-Y
Cite this review
Pith. "Pith review of Separable Computation of Information Measures." pith.science (2026). https://pith.science/paper/JC5OTKME
@misc{pith2026250115301,
author = {Pith},
title = {Pith review of: Separable Computation of Information Measures},
year = {2026},
howpublished = {\url{https://pith.science/paper/JC5OTKME}},
note = {Machine review of arXiv:2501.15301}
}
abstract
We study a separable design for computing information measures, where the information measure is computed from learned feature representations instead of raw data. Under mild assumptions on the feature representations, we demonstrate that a class of information measures admit such separable computation, including mutual information, $f$-information, Wyner's common information, G{\'a}cs--K{\"o}rner common information, and Tishby's information bottleneck. Our development establishes several new connections between information measures and the statistical dependence structure. The characterizations also provide theoretical guarantees of practical designs for estimating information measures through representation learning.
Figures
Reference graph
Works this paper leans on
-
[3]
Approximating mutual information of high-dimensional variables using le arned representations,
G. Gowri, X. Lun, A. M. Klein, and P . Yin, “Approximating mutual information of high-dimensional variables using le arned representations,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024. [Online]. Available: https://openreview.net/forum?id=HN05DQxyLl
work page 2024
-
[1]
Prediction and entropy of printed englis h,
C. E. Shannon, “Prediction and entropy of printed englis h,” Bell system technical journal , vol. 30, no. 1, pp. 50–64, 1951
work page 1951
-
[2]
Mutual information neural esti mation,
M. I. Belghazi, A. Baratin, S. Rajeshwar, S. Ozair, Y . Ben gio, A. Courville, and D. Hjelm, “Mutual information neural esti mation,” in International conference on machine learning . PMLR, 2018, pp. 531–540
work page 2018
-
[4]
Dependence induced representations ,
X. Xu and L. Zheng, “Dependence induced representations ,” in 2024 60th Annual Allerton Conference on Communication, Control , and Computing. IEEE, 2024, pp. 1–8
work page 2024
-
[5]
The common information of two dependent rando m vari- ables,
A. Wyner, “The common information of two dependent rando m vari- ables,” IEEE Transactions on Information Theory , vol. 21, no. 2, pp. 163–179, 1975
work page 1975
-
[6]
Common information is far less t han mutual information
P . Gács, , and J. Körner, “Common information is far less t han mutual information.” Problems of Control and Information Theory , vol. 2, pp. 149–162, 1973
work page 1973
-
[7]
The information bottleneck method,
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” arXiv preprint physics/0004057 , 2000
arXiv 2000
-
[8]
Universal features for high-dimensional learning and inf erence,
S.-L. Huang, A. Makur, G. W. Wornell, and L. Zheng, “Universal features for high-dimensional learning and inf erence,” F oundations and Trends® in Communications and Information Theory, vol. 21, no. 1-2, pp. 1–299, 2024. [Online]. Available: http://dx.doi.org/10.1561/0100000107
Show all 12 references
-
[9]
Neural feature learning in function s pace,
X. Xu and L. Zheng, “Neural feature learning in function s pace,” Journal of Machine Learning Research , vol. 25, no. 142, pp. 1–76, 2024
2024
-
[10]
On sequences of pairs of dependent random variables,
H. S. Witsenhausen, “On sequences of pairs of dependent random variables,” SIAM Journal on Applied Mathematics , vol. 28, no. 1, pp. 100–113, 1975
1975
-
[11]
Hypothesis testing with co mmunication constraints,
R. Ahlswede and I. Csiszár, “Hypothesis testing with co mmunication constraints,” IEEE transactions on information theory , vol. 32, no. 4, pp. 533–542, 1986
1986
-
[12]
T. M. Cover and J. A. Thomas, Elements of infor- mation theory (2. ed.) . Wiley, 2006. [Online]. Available: http://www.elementsofinformationtheory.com/
2006
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.