Pith. sign in

REVIEW 3 cited by

An Empirical Investigation of the Role of Pre-training in Lifelong Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.09153 v2 pith:55KSCABL submitted 2021-12-16 cs.LG cs.AIcs.CLcs.CV

classification cs.LGcs.AIcs.CLcs.CV
keywords learningforgettingpre-trainingtaskscatastrophiclifelonglossmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The lifelong learning paradigm in machine learning is an attractive alternative to the more prominent isolated learning scheme not only due to its resemblance to biological learning but also its potential to reduce energy waste by obviating excessive model re-training. A key challenge to this paradigm is the phenomenon of catastrophic forgetting. With the increasing popularity and success of pre-trained models in machine learning, we pose the question: What role does pre-training play in lifelong learning, specifically with respect to catastrophic forgetting? We investigate existing methods in the context of large, pre-trained models and evaluate their performance on a variety of text and image classification tasks, including a large-scale study using a novel data set of 15 diverse NLP tasks. Across all settings, we observe that generic pre-training implicitly alleviates the effects of catastrophic forgetting when learning multiple tasks sequentially compared to randomly initialized models. We then further investigate why pre-training alleviates forgetting in this setting. We study this phenomenon by analyzing the loss landscape, finding that pre-trained weights appear to ease forgetting by leading to wider minima. Based on this insight, we propose jointly optimizing for current task loss and loss basin sharpness to explicitly encourage wider basins during sequential fine-tuning. We show that this optimization approach outperforms several state-of-the-art task-sequential continual learning algorithms across multiple settings, occasionally even without retaining a memory that scales in size with the number of tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MG-CLIP preserves CLIP's modality gap by adaptively limiting fine-tuning epochs and compensates for its limits with a visual-space classifier, improving class-incremental learning without replay.

  2. Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A questions-only replay plus attention distillation method outperforms image-replay baselines for continual visual question answering while storing no past images.

  3. Transfer of Structural Knowledge from Synthetic Languages

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A new synthetic language, flat_shuffle, transfers more structure to English fine-tuning than earlier synthetic bracket languages, though still far short of training on English from scratch.

Pith tools