Pith. sign in

REVIEW 7 cited by

Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14985 v5 pith:PYUHV3XS submitted 2024-07-20 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords memorizationpretrainingtasksdatamodelsansweringfactualgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The impressive capabilities of large language models (LLMs) have sparked debate over whether these models genuinely generalize to unseen tasks or predominantly rely on memorizing vast amounts of pretraining data. To explore this issue, we introduce an extended concept of memorization, distributional memorization, which measures the correlation between the LLM output probabilities and the pretraining data frequency. To effectively capture task-specific pretraining data frequency, we propose a novel task-gram language model, which is built by counting the co-occurrence of semantically related $n$-gram pairs from task inputs and outputs in the pretraining corpus. Using the Pythia models trained on the Pile dataset, we evaluate four distinct tasks: machine translation, factual question answering, world knowledge understanding, and math reasoning. Our findings reveal varying levels of memorization, with the strongest effect observed in factual question answering. Furthermore, while model performance improves across all tasks as LLM size increases, only factual question answering shows an increase in memorization, whereas machine translation and reasoning tasks exhibit greater generalization, producing more novel outputs. This study demonstrates that memorization plays a larger role in simpler, knowledge-intensive tasks, while generalization is the key for harder, reasoning-based tasks, providing a scalable method for analyzing large pretraining corpora in greater depth.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    Cognitive biases in LLMs are largely set during pretraining, while finetuning data and seed randomness only modulate them.

  2. Rethinking Memorization Measures and their Implications in Large Language Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Contextual memorization, defined by comparing a string's training loss against the best loss without training on that string, is stricter than counterfactual memorization and suggests that zero-memorization optimal le...

  3. Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

    cs.AI 2025-05 conditional novelty 6.0 of 10

    RL-trained VLMs generalize compositionally far better than SFT-trained ones on synthetic geometry and spatial tasks, but cross-modal combination remains weak, and a caption-before-thinking plus progress-reward recipe ...

  4. An Annotated Reading of 'The Singer of Tales' in the LLM Era

    cs.CY 2025-02 conditional novelty 6.0 of 10

    LLM generation resembles oral-formulaic composition: single-pass, pattern-based, and non-authorial, so AI output should be treated as a new post-literate medium.

  5. Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Properly spaced multi-epoch reuse of high-quality data, guided by a measured memorization window, continues to improve LLM performance far beyond the common four-epoch heuristic.

  6. Adaptive Multi-Agent Reasoning via Automated Workflow Generation

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Automated workflow generation and iterative prompt refinement let a standard GPT-4.1 model outperform state-of-the-art reasoning models on a revised riddle benchmark.

  7. Counterfactual Influence as a Distributional Quantity

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Near-duplicate training samples lower a model's self-influence on a record while raising its extractability, so self-influence alone underestimates memorization risk.

Pith tools