Pith. sign in

REVIEW 2 cited by

Training Foundation Models as Data Compression: On Information, Model Weights and Copyright Law

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13493 v4 pith:7L2QCRDT submitted 2024-07-18 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords trainingcopyrightfoundationmodelsweightsdatalegalmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The training process of foundation models as for other classes of deep learning systems is based on minimizing the reconstruction error over a training set. For this reason, they are susceptible to the memorization and subsequent reproduction of training samples. In this paper, we introduce a training-as-compressing perspective, wherein the model's weights embody a compressed representation of the training data. From a copyright standpoint, this point of view implies that the weights can be considered a reproduction or, more likely, a derivative work of a potentially protected set of works. We investigate the technical and legal challenges that emerge from this framing of the copyright of outputs generated by foundation models, including their implications for practitioners and researchers. We demonstrate that adopting an information-centric approach to the problem presents a promising pathway for tackling these emerging complex legal issues.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Probabilistic "Copies" in Generative AI Models

    cs.CY 2026-07 conditional novelty 6.0 of 10

    An LLM is an infringing copy of a work only when the work can be extracted from it with relatively little effort, so some models are copies of some works and no model is a copy of everything it trained on.

  2. LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A frozen LLM can process very long contexts by recurrently compressing them with a Perceiver and injecting the compressed memory through gated cross-attention, with query-dependent compression boosting QA performance.

Pith tools