Pith. sign in

REVIEW 7 cited by

Gradient Episodic Memory for Continual Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1706.08840 v6 pith:5NIRUFXP submitted 2017-06-26 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningcontinualknowledgemodelstasksabilityepisodicforgetting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One major obstacle towards AI is the poor ability of models to solve new problems quicker, and without forgetting previously acquired knowledge. To better understand this issue, we study the problem of continual learning, where the model observes, once and one by one, examples concerning a sequence of tasks. First, we propose a set of metrics to evaluate models learning over a continuum of data. These metrics characterize models not only by their test accuracy, but also in terms of their ability to transfer knowledge across tasks. Second, we propose a model for continual learning, called Gradient Episodic Memory (GEM) that alleviates forgetting, while allowing beneficial transfer of knowledge to previous tasks. Our experiments on variants of the MNIST and CIFAR-100 datasets demonstrate the strong performance of GEM when compared to the state-of-the-art.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 497 citations worldwide. Full citation record

  1. Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SFT lessons — reason-based training, on-model replay, and wash-out robustness — transfer across toy models, model organisms, and alignment SFT, improving the capability–safety tradeoff.

  2. Fine-Tuning Regimes Define Distinct Continual Learning Problems

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    The relative rankings of continual learning methods are not preserved across different fine-tuning regimes defined by trainable parameter depth.

  3. Class Incremental Continual Learning with Self-Organizing Maps and Variational Autoencoders Using Synthetic Replay

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A SOM-VAE generative replay method stores per-unit Gaussian statistics instead of raw data and reports competitive class-incremental accuracy on standard benchmarks.

  4. When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models

    cs.LG 2025-12 conditional novelty 5.0 of 10

    Quantized (INT8/INT4) LLMs can outperform FP16 in later-task forward accuracy and retention during continual learning, though single-seed runs leave the effect unquantified.

  5. SWE-Bench-CL: Continual Learning for Coding Agents

    cs.LG 2025-06 conditional novelty 5.0 of 10

    SWE-Bench-CL reorganizes SWE-Bench Verified into 8 time-ordered sequences of 273 total tasks to measure continual learning in coding agents, adding CL-specific metrics and a semantic memory agent design.

  6. The Future of Continual Learning in the Era of Foundation Models: Three Key Directions

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Continual learning should pivot from weight-update-based methods to continual compositionality and orchestration of foundation models and agents.

  7. Class Incremental Learning for Algorithm Selection

    cs.LG 2025-06 conditional novelty 4.0 of 10

    In an algorithm-selection stream where new solver classes appear, rehearsal-based class-incremental learning retains 82.3% accuracy across four solver classes, about 7% below an oracle that sees all data.

Pith tools