LMD, a log-normal multiplicative-weight optimizer, trains ViT and GPT-2 from scratch and keeps accuracy under MXFP6 forward-pass quantization.
Perceptron learning with sign-constrained weights
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
LMD, a log-normal multiplicative-weight optimizer, trains ViT and GPT-2 from scratch and keeps accuracy under MXFP6 forward-pass quantization.