Pith. sign in

REVIEW 2 cited by

DiME: Maximizing Mutual Information by a Difference of Matrix-Based Entropies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.08164 v3 pith:AOI42RXB submitted 2023-01-19 cs.LG cs.ITmath.IT

classification cs.LGcs.ITmath.IT
keywords dimeinformationmutualmatrix-baseddifferenceeigenvaluesentropiesquantity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce an information-theoretic quantity with similar properties to mutual information that can be estimated from data without making explicit assumptions on the underlying distribution. This quantity is based on a recently proposed matrix-based entropy that uses the eigenvalues of a normalized Gram matrix to compute an estimate of the eigenvalues of an uncentered covariance operator in a reproducing kernel Hilbert space. We show that a difference of matrix-based entropies (DiME) is well suited for problems involving the maximization of mutual information between random variables. While many methods for such tasks can lead to trivial solutions, DiME naturally penalizes such outcomes. We compare DiME to several baseline estimators of mutual information on a toy Gaussian dataset. We provide examples of use cases for DiME, such as latent factor disentanglement and a multiview representation learning problem where DiME is used to learn a shared representation among views with high mutual information.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

    cs.SD 2026-07 conditional novelty 5.0 of 10

    InsideSSL analyzes self-supervised speech models layer-by-layer using entropy, curvature, robustness metrics, and a cross-layer Generative Compatibility Matrix, finding that training objectives induce distinct compres...

  2. Enhancing Cross-task Transfer of Large Language Models via Activation Steering

    cs.CL 2025-07 conditional novelty 5.0 of 10

    CAST transfers knowledge across tasks by adding the average few-shot minus zero-shot activation difference from a high-resource task to a low-resource task's hidden state, improving accuracy without training or longer...

Pith tools