Pith. sign in

REVIEW 1 cited by

Supermasks in Superposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.14769 v3 pith:5MUTCUY3 submitted 2020-06-26 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords supermaskstasksuperpositionsupsupcorrectfindidentitynetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the Supermasks in Superposition (SupSup) model, capable of sequentially learning thousands of tasks without catastrophic forgetting. Our approach uses a randomly initialized, fixed base network and for each task finds a subnetwork (supermask) that achieves good performance. If task identity is given at test time, the correct subnetwork can be retrieved with minimal memory usage. If not provided, SupSup can infer the task using gradient-based optimization to find a linear superposition of learned supermasks which minimizes the output entropy. In practice we find that a single gradient step is often sufficient to identify the correct mask, even among 2500 tasks. We also showcase two promising extensions. First, SupSup models can be trained entirely without task identity information, as they may detect when they are uncertain about new data and allocate an additional supermask for the new training distribution. Finally the entire, growing set of supermasks can be stored in a constant-sized reservoir by implicitly storing them as attractors in a fixed-sized Hopfield network.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The impact of allocation strategies in subset learning on the expressive power of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    In a teacher-student setup, maximal expressive power for a fixed learnable-weight budget is characterized by even row or column distribution in linear RNNs and feedforward networks.

Pith tools