Pith. sign in

REVIEW 4 cited by

An Empirical Study of Scaling Laws for Transfer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.16947 v1 pith:OOWOVAOJ submitted 2024-08-30 cs.LG

classification cs.LG
keywords transferscalingdatadownstreamperformancewhendistributionempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a limited empirical study of scaling laws for transfer learning in transformer models. More specifically, we examine a scaling law that incorporates a "transfer gap" term, indicating the effectiveness of pre-training on one distribution when optimizing for downstream performance on another distribution. When the transfer gap is low, pre-training is a cost-effective strategy for improving downstream performance. Conversely, when the gap is high, collecting high-quality fine-tuning data becomes relatively more cost effective. Fitting the scaling law to experiments from diverse datasets reveals significant variations in the transfer gap across distributions. In theory, the scaling law can inform optimal data allocation strategies and highlights how the scarcity of downstream data can bottleneck performance. Our findings contribute to a principled way to measure transfer learning efficiency and understand how data availability affects capabilities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distillation Scaling Laws

    cs.LG 2025-02 conditional novelty 7.0 of 10

    A distillation scaling law predicts student cross-entropy from teacher loss, student size, and data, and gives compute-optimal teacher-student allocations.

  2. Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning

    cs.LG 2025-06 reject novelty 6.0 of 10

    From Open LLM Leaderboard data grouped by base model, the authors recover a three-factor ordering of LLM capabilities and claim instruction-following causally supports math reasoning.

  3. Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Finetuning forgetting follows a multiplicative scaling law in model size, finetuning tokens, and injected pretraining fraction, with 1% injection nearly eliminating forgetting.

  4. GUST: Quantifying Free-Form Geometric Uncertainty of Metamaterials Using Small Data

    cs.LG 2025-05 conditional novelty 5.0 of 10

    GUST combines synthetic-data pretraining with transfer learning on a conditional diffusion model to quantify free-form geometric uncertainty in manufactured metamaterials from small real-world datasets.

Pith tools