REVIEW 2 cited by
On the Diminishing Returns of Width for Continual Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
While deep neural networks have demonstrated groundbreaking performance in various settings, these models often suffer from \emph{catastrophic forgetting} when trained on new tasks in sequence. Several works have empirically demonstrated that increasing the width of a neural network leads to a decrease in catastrophic forgetting but have yet to characterize the exact relationship between width and continual learning. We design one of the first frameworks to analyze Continual Learning Theory and prove that width is directly related to forgetting in Feed-Forward Networks (FFN). Specifically, we demonstrate that increasing network widths to reduce forgetting yields diminishing returns. We empirically verify our claims at widths hitherto unexplored in prior studies where the diminishing returns are clearly observed as predicted by our theory.
Forward citations
Cited by 2 Pith papers
-
Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective
Representation discrepancy, a new metric with theoretical bounds, shows continual learning forgets features faster in deeper layers and slower in wider networks.
-
Rethinking Traffic Flow Forecasting: From Transition to Generatation
EMBSFormer splits traffic forecasting into a multi-period similarity lookup branch and a transition branch, and reports better accuracy with roughly 82% fewer parameters than GMAN on the PEMS benchmarks.
Discussion (0). Continue with ORCID to comment.