Pith. sign in

REVIEW 5 cited by

Exploring Low Rank Training of Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.13569 v1 pith:LGEPWDFE submitted 2022-09-27 cs.LG stat.ML

classification cs.LGstat.ML
keywords trainingranknetworksdeepneuralpracticeworkablations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training deep neural networks in low rank, i.e. with factorised layers, is of particular interest to the community: it offers efficiency over unfactorised training in terms of both memory consumption and training time. Prior work has focused on low rank approximations of pre-trained networks and training in low rank space with additional objectives, offering various ad hoc explanations for chosen practice. We analyse techniques that work well in practice, and through extensive ablations on models such as GPT2 we provide evidence falsifying common beliefs in the field, hinting in the process at exciting research opportunities that still need answering.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SLORR: Simple and Efficient In-Training Low-Rank Regularization

    cs.LG 2026-07 accept novelty 6.0 of 10

    A stateless, SVD-free regularizer approximates polar factors to induce low-rank weight structure during training, enabling better post-training compression of vision models and LLMs at under 8% overhead.

  2. Taming LLMs by Scaling Learning Rates with Gradient Grouping

    cs.LG 2025-06 conditional novelty 5.0 of 10

    An optimizer wrapper that clusters per-layer momentum and scales learning rates by cluster-wise median deviations improves perplexity and accuracy across LLM and MLLM training, and lets LoRA pretraining approach full-...

  3. ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers

    cs.LG 2026-07 conditional novelty 4.0 of 10

    ELSAA fuses sparse exact-block attention with a low-rank bucket-sketch branch using a denominator-aware scalar, achieving linear-time attention with competitive long-context accuracy.

  4. Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A benchmark and two low-cost tricks (weight refactorization and momentum reset) that make low-rank LLM pre-training competitive with GaLore and Fira at about 25% lower memory.

  5. Accelerating Attention with Basis Decomposition

    cs.LG 2025-10 reject novelty 3.0 of 10

    A low-rank factorization of attention projection matrices (basis plus coefficients) gives modest FLOP savings in exact arithmetic, but the claimed losslessness and novelty are not supported.

Pith tools