Pith. sign in

hub

Language Modeling Is Compression

24 Pith papers cite this work, alongside 27 external citations. Polarity classification is still indexing.

24 Pith papers citing it
27 external citations · Pith
abstract

It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model.

hub tools

citation-role summary

background 3 other 1

citation-polarity summary

polarities

background 3 unclear 1

representative citing papers

Are Flat Minima an Illusion?

cs.LG · 2026-03-24 · conditional · novelty 5.0

The paper argues flat minima are an illusion and that a reparameterization-invariant 'weakness' score predicts generalization where raw sharpness fails.

Efficient compression of neural networks and datasets

cs.LG · 2025-05-23 · unverdicted · novelty 5.0

Refined probabilistic and smooth l0 pruning techniques approximate minimum description length for neural networks, achieving high compression with minimal accuracy loss and empirically verifying better sample efficiency and generalization on image and text tasks.

Interestingness as an Inductive Heuristic for Future Compression Progress

cs.AI · 2026-05-14 · unverdicted · novelty 4.0

Interestingness is defined as an inductive signal for future compression progress, with proofs that expected progress decays exponentially with time since last breakthrough and that the Algorithmic Prior yields quadratic gains over the Length Prior.

The Rhetoric of Machine Learning

cs.LG · 2026-04-08 · unverdicted · novelty 4.0

Machine learning is inherently rhetorical and is often deployed as 'manipulation as a service' in business models.

citing papers explorer

Showing 24 of 24 citing papers.