Pith. sign in

REVIEW 2 cited by

The Expressive Power of Tuning Only the Normalization Layers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.07937 v2 pith:CLCLBMVM submitted 2023-02-15 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords networksnormalizationlayerstuningexpressivefine-tuningonlypower
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Feature normalization transforms such as Batch and Layer-Normalization have become indispensable ingredients of state-of-the-art deep neural networks. Recent studies on fine-tuning large pretrained models indicate that just tuning the parameters of these affine transforms can achieve high accuracy for downstream tasks. These findings open the questions about the expressive power of tuning the normalization layers of frozen networks. In this work, we take the first step towards this question and show that for random ReLU networks, fine-tuning only its normalization layers can reconstruct any target network that is $O(\sqrt{\text{width}})$ times smaller. We show that this holds even for randomly sparsified networks, under sufficient overparameterization, in agreement with prior empirical work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The impact of allocation strategies in subset learning on the expressive power of neural networks

    cs.LG 2025-02 conditional novelty 7.0 of 10

    In a teacher-student setup, maximal expressive power for a fixed learnable-weight budget is characterized by even row or column distribution in linear RNNs and feedforward networks.

  2. Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    LayerNorm parameter shifts after fine-tuning encode domain-transition information; rescaling them via an FSR-dependent scalar lambda plus a cyclic step improves ViT classification under data scarcity and domain shift.

Pith tools