Pith. sign in

REVIEW 2 cited by

Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09366 v1 pith:JFVFAZ62 submitted 2024-06-13 cs.LG cs.CVq-bio.NC

classification cs.LGcs.CVq-bio.NC
keywords mmcrmvssldataperspectiveunderstandingbettercapacityembeddings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Maximum Manifold Capacity Representations (MMCR) is a recent multi-view self-supervised learning (MVSSL) method that matches or surpasses other leading MVSSL methods. MMCR is intriguing because it does not fit neatly into any of the commonplace MVSSL lineages, instead originating from a statistical mechanical perspective on the linear separability of data manifolds. In this paper, we seek to improve our understanding and our utilization of MMCR. To better understand MMCR, we leverage tools from high dimensional probability to demonstrate that MMCR incentivizes alignment and uniformity of learned embeddings. We then leverage tools from information theory to show that such embeddings maximize a well-known lower bound on mutual information between views, thereby connecting the geometric perspective of MMCR to the information-theoretic perspective commonly discussed in MVSSL. To better utilize MMCR, we mathematically predict and experimentally confirm non-monotonic changes in the pretraining loss akin to double descent but with respect to atypical hyperparameters. We also discover compute scaling laws that enable predicting the pretraining loss as a function of gradients steps, batch size, embedding dimension and number of views. We then show that MMCR, originally applied to image data, is performant on multimodal image-text data. By more deeply understanding the theoretical and empirical behavior of MMCR, our work reveals insights on improving MVSSL methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simplifying DINO via Coding Rate Regularization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Replacing DINO's complex anti-collapse machinery with an explicit coding rate regularizer yields simpler, more stable, and higher-performing self-supervised models.

  2. Generalized Category Discovery via Token Manifold Capacity Learning

    cs.LG 2025-05 conditional novelty 4.0 of 10

    MTMC adds a nuclear-norm loss on class tokens to GCD objectives and reports small accuracy gains on several image benchmarks.

Pith tools