Pith. sign in

REVIEW 2 cited by

Understanding Forgetting in Continual Learning with Linear Regression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.17583 v1 pith:IMNAQHKA submitted 2024-05-27 cs.LG

classification cs.LG
keywords forgettingtheoreticallearninglinearregressiontasksanalysiscontinual
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Continual learning, focused on sequentially learning multiple tasks, has gained significant attention recently. Despite the tremendous progress made in the past, the theoretical understanding, especially factors contributing to catastrophic forgetting, remains relatively unexplored. In this paper, we provide a general theoretical analysis of forgetting in the linear regression model via Stochastic Gradient Descent (SGD) applicable to both underparameterized and overparameterized regimes. Our theoretical framework reveals some interesting insights into the intricate relationship between task sequence and algorithmic parameters, an aspect not fully captured in previous studies due to their restrictive assumptions. Specifically, we demonstrate that, given a sufficiently large data size, the arrangement of tasks in a sequence, where tasks with larger eigenvalues in their population data covariance matrices are trained later, tends to result in increased forgetting. Additionally, our findings highlight that an appropriate choice of step size will help mitigate forgetting in both underparameterized and overparameterized settings. To validate our theoretical analysis, we conducted simulation experiments on both linear regression models and Deep Neural Networks (DNNs). Results from these simulations substantiate our theoretical findings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Drift and Dependence: Layer-wise Information-Theoretic Bounds for Replay-Based Continual Learning

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A layer-wise information-theoretic decomposition bounds replay-based continual learning's generalization gap, predicting memory scaling, an interior stabilization layer, and gradient-alignment signals that track forgetting.

  2. Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Representation discrepancy, a new metric with theoretical bounds, shows continual learning forgets features faster in deeper layers and slower in wider networks.

Pith tools