Pith. sign in

REVIEW 8 cited by

Continual learning with hypernetworks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1906.00695 v4 pith:G36A5QQG submitted 2019-06-03 cs.LG cs.AIstat.ML

Continual learning with hypernetworks

classification cs.LG cs.AIstat.ML
keywords hypernetworkstask-conditionedlearningtaskcontinualhypernetworklongmemory
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Artificial neural networks suffer from catastrophic forgetting when they are sequentially trained on multiple tasks. To overcome this problem, we present a novel approach based on task-conditioned hypernetworks, i.e., networks that generate the weights of a target model based on task identity. Continual learning (CL) is less difficult for this class of models thanks to a simple key feature: instead of recalling the input-output relations of all previously seen data, task-conditioned hypernetworks only require rehearsing task-specific weight realizations, which can be maintained in memory using a simple regularizer. Besides achieving state-of-the-art performance on standard CL benchmarks, additional experiments on long task sequences reveal that task-conditioned hypernetworks display a very large capacity to retain previous memories. Notably, such long memory lifetimes are achieved in a compressive regime, when the number of trainable hypernetwork weights is comparable or smaller than target network size. We provide insight into the structure of low-dimensional task embedding spaces (the input space of the hypernetwork) and show that task-conditioned hypernetworks demonstrate transfer learning. Finally, forward information transfer is further supported by empirical results on a challenging CL benchmark based on the CIFAR-10/100 image datasets.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Instance-Adaptive Parametrization for Amortized Variational Inference

    cs.LG 2026-04 unverdicted novelty 7.0

    IA-VAE augments amortized variational inference with hypernetwork-generated instance-adaptive modulations, strictly containing the standard variational family and improving held-out ELBO on synthetic and image data.

  2. Architecture Generalization with MetaNCA

    cs.LG 2026-07 conditional novelty 6.0

    A learned local rule (Weight Transformer) iteratively self-organizes task-network weights from local graph neighborhoods and generalizes across unseen MLP, CNN, and ResNet architectures up to ~2M parameters.

  3. Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction

    cs.LG 2026-06 unverdicted novelty 6.0

    Proposes feature splitting and a closed-form bound on extrapolation range to enable zero-shot topological out-of-domain generalization in dynamical systems reconstruction across tipping points.

  4. RareCP: Regime-Aware Retrieval for Efficient Conformal Prediction

    cs.LG 2026-05 unverdicted novelty 6.0

    RareCP improves interval efficiency for time series conformal prediction by retrieving and weighting regime-specific calibration examples while adapting to drift and maintaining coverage.

  5. Context-Aware Force Estimation for Deformable Tool Manipulation in Robotic Environmental Swabbing via Few-Shot Continual Adaptation

    cs.RO 2026-07 conditional novelty 5.0

    A frozen LSTM backbone with FiLM-based few-shot context adaptation estimates tip-level contact forces for deformable swabbing tools across nine surface-tool regimes using only wrist-mounted proprioception.

  6. CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning

    cs.LG 2026-05 unverdicted novelty 5.0

    CP-MoE uses a transient expert, consistency-preserving routing bias, and guided regularization to reduce catastrophic forgetting in MoE-based LLMs and VLMs while preserving cross-task transfer, reporting SOTA on Super...

  7. Adaptive Reorganization of Neural Pathways for Continual Learning with Spiking Neural Networks

    cs.NE 2023-09 unverdicted novelty 4.0

    SOR-SNN employs Self-Organizing Regulation networks to reorganize a single SNN into sparse pathways, achieving better performance, energy efficiency, memory use, backward transfer, and self-repair on continual learnin...

  8. An Optimal Transport-driven Approach for Cultivating Latent Space in Online Incremental Learning

    cs.LG 2022-11 unverdicted novelty 3.0

    Proposes MMOT, an optimal transport-based online mixture model with dynamic preservation strategy, for improved handling of multimodal data in online incremental learning.