Pith. sign in

REVIEW 10 cited by

TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06762 v1 pith:JQRVUTKU submitted 2023-10-10 cs.CL

classification cs.CL
keywords llmscontinuallearningtasksalignedcapabilitiestracedatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Aligned large language models (LLMs) demonstrate exceptional capabilities in task-solving, following instructions, and ensuring safety. However, the continual learning aspect of these aligned LLMs has been largely overlooked. Existing continual learning benchmarks lack sufficient challenge for leading aligned LLMs, owing to both their simplicity and the models' potential exposure during instruction tuning. In this paper, we introduce TRACE, a novel benchmark designed to evaluate continual learning in LLMs. TRACE consists of 8 distinct datasets spanning challenging tasks including domain-specific tasks, multilingual capabilities, code generation, and mathematical reasoning. All datasets are standardized into a unified format, allowing for effortless automatic evaluation of LLMs. Our experiments show that after training on TRACE, aligned LLMs exhibit significant declines in both general ability and instruction-following capabilities. For example, the accuracy of llama2-chat 13B on gsm8k dataset declined precipitously from 28.8\% to 2\% after training on our datasets. This highlights the challenge of finding a suitable tradeoff between achieving performance on specific tasks while preserving the original prowess of LLMs. Empirical findings suggest that tasks inherently equipped with reasoning paths contribute significantly to preserving certain capabilities of LLMs against potential declines. Motivated by this, we introduce the Reasoning-augmented Continual Learning (RCL) approach. RCL integrates task-specific cues with meta-rationales, effectively reducing catastrophic forgetting in LLMs while expediting convergence on novel tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethinking Transfer in Continual Learning: A Replay-Based Realisation

    cs.LG 2026-07 conditional novelty 7.0 of 10

    In continual learning, forward transfer requires target headroom, a persistent carrier, and a compatible source; routing replay by gradient signatures improves accuracy and stability over uniform replay.

  2. CEO-Bench: Can Agents Play the Long Game?

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Only two of ten advanced AI agents finish a 500-day simulated CEO challenge above the starting cash, and none surpass a hand-tuned rule-based baseline.

  3. Omega-S: A Functional Resilience Index for LLM Fine-Tuning

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Omega-S, a penalty on node-degree variance in the weight matrix, improves code retention during LoRA fine-tuning of Llama-3-8B, while its advertised clustering/topological channel is inert.

  4. The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Automatically grouping and sequencing tasks into multiple QLoRA adapters improves continual fine-tuning performance over a single shared adapter at matched trainable capacity.

  5. Bisecle: Binding and Separation in Continual Learning for Video Language Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Bisecle reduces catastrophic forgetting and improves accuracy in sequential VideoQA learning using multi-directional auxiliary losses and contrastive prompt regularization.

  6. TreeLoRA: Efficient Continual Learning via Layer-Wise LoRAs Guided by a Hierarchical Gradient-Similarity Tree

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A gradient-similarity tree of LoRA adapters, searched by an LCB bandit algorithm, delivers state-of-the-art continual learning accuracy with up to 3.2x faster training on ViTs and 2.4x on LLMs.

  7. T2I-ConBench: Text-to-Image Benchmark for Continual Post-training

    cs.CV 2025-05 conditional novelty 6.0 of 10

    T2I-ConBench provides a unified multi-metric benchmark for continual post-training of text-to-image models and shows that all tested methods have notable weaknesses.

  8. Continual Learning in Transition

    cs.LG 2026-08 accept novelty 5.0 of 10

    A tri-axial framework of When, Where, and How organizes the ongoing transition of continual learning from parameter-centric updates to system-level capability evolution.

  9. Continual Gradient Low-Rank Projection Fine-Tuning for LLMs

    cs.LG 2025-07 conditional novelty 5.0 of 10

    GORP jointly trains LoRA and full-rank parameters inside a low-rank gradient subspace built from Adam first moments, reporting higher average accuracy and lower forgetting than O-LoRA and N-LoRA on LLM continual learn...

  10. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

Pith tools