REVIEW 4 major objections 6 minor 1 cited by
Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A biologically inspired log-normal optimizer trains transformers from scratch with multiplicative updates.
desk verdict LMD is a genuinely new optimizer with a striking ViT result, but the GPT-2 low-precision claim rests on a confounded baseline comparison; worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the log-normal posterior over weights and the multiplicative-group update on its median. LMD maintains $m^+$ and $m^-$ (positive and negative copies of each weight), samples log-normal noise $\varepsilon$, forms $\theta = m \odot \varepsilon$, computes the gradient scaled by $\theta$, and updates $m$ multiplicatively as $m \leftarrow m \odot \exp(-\eta(\operatorname{sign}(\nu_{\mathrm{temp}}) + r))$. The regularizer $r = \tau(\log \theta - \log m_r)/\sigma^2$ is chosen from a log-normal prior, so in expectation the update performs weight decay in log space, pulling $m$ toward $m_r$ instead of toward zero; this is what prevents the exponential weight growth that broke earlier multiplicative methods like Madam. Multiplicative noise injection serves as a regularizer and, because perturbations scale with weight magnitude, they are not wiped out when weights are quantized to MXFP6 or MXFP4, giving stability in low-precision forward passes.
What would settle it
Train a large model such as BERT or a ResNet from scratch with LMD under the paper's default $\sigma=0.125$ and $m_r=0.01\exp(\sigma^2/2)$ and an MXFP6 forward pass; if the run diverges, the weight norm explodes, or accuracy collapses relative to bfloat16, the claimed drop-in stability does not generalize. A second check: disable weight sampling by using only the mean $\theta=m$ under MXFP6; the paper predicts the weight-norm dynamics become inconsistent and regularization weakens, so if mean training matches sampled training exactly, multiplicative noise is not the active stabilizing mechanism.
Extended reading notes
Core claim
The discovery is that a multiplicative update rule can be made stable at scale by making both the noise and the regularization multiplicative, in direct analogy to the noisy multiplicative dynamics of biological synapses. Specifically, the paper derives LMD from the Lie-group Bayesian learning rule over log-normal posteriors: weights are sampled as $\theta = m \odot \varepsilon$ with $\varepsilon \sim \mathrm{LogN}(0, \sigma^2 I)$, and the median $m$ is updated as $m \leftarrow m \odot \exp(-\eta(\operatorname{sign}(\nu_{\mathrm{temp}}) + r))$, where $r = \tau(\log \theta - \log m_r)/\sigma^2$ acts as weight decay in logarithmic space. In this scheme the sign of a weight is fixed by keeping separate positive and negative copies (the EG$\pm$ trick), gradient scaling by the weight magnitude replaces additive updates, and multiplicative noise injection survives low-precision quantization because perturbations scale with weight size. On ViT/ImageNet and GPT-2/OpenWebText, LMD reportedly trains from scratch with no degradation under MXFP6 forward passes, reaching 77.06\% test accuracy on ViT compared with 68.11\% for AdamW, and reaching a 2.925 validation loss on GPT-2 at sequence length 4096.
Load-bearing premise
The load-bearing premise is that the hand-set defaults for the noise width, $\sigma=0.125$, and the target weight scale, $m_r=0.01\exp(\sigma^2/2)$, transfer across architectures without per-model tuning; the paper itself calls the scale-parameter initialization heuristic and reports no sensitivity study.
Editorial extensions
If this is right
- If LMD is correct, multiplicative weight updates are no longer confined to small networks: ViT and GPT-2 can be trained from scratch with them, so exponentiated-gradient optimizers re-enter the practical deep-learning toolbox.
- MXFP6 forward passes without performance loss would let training run on hardware built around microscaling formats, reducing the memory and energy cost of forward matrix multiplications.
- The log-space weight decay pulls weights toward $m_r$ instead of zero, so LMD changes how regularization and pruning interact: weights near $m_r$ are implicitly treated as emulated activation perturbations rather than dead parameters.
- Because LMD keeps the weight norm close to its initial value, it may make training dynamics more predictable and remove the need for gradient-norm clipping (the paper uses none for ViT).
- The method costs one extra state vector over AdamW ($4P$ parameters versus $3P$), which is small enough for a drop-in optimizer replacement in existing code.
Reading between the lines
- The fixed hyperparameters across ViT and GPT-2 hint at a scale-free property of the log-space dynamics; a natural test is whether the same $\sigma=0.125$ and $m_r=0.01\exp(\sigma^2/2)$ transfer to convolutional nets, encoder-only LLMs, or fine-tuning, none of which the paper examines.
- If multiplicative noise's main role is preserving perturbations through quantization, LMD's sampling could replace or complement stochastic rounding in low-precision forward passes; a direct comparison would isolate which mechanism stabilizes MXFP6 training.
- The log-space sign update resembles signSGD on log-weights, so convergence and generalization analyses from online learning and exponentiated-gradient theory may carry over to LMD, giving a theoretical handle the paper does not develop.
- Extending LMD to low-precision backward passes, which the paper keeps in bfloat16, would determine whether multiplicative dynamics unlock fully low-precision training or only low-precision inference-side computation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Log-Normal Multiplicative Dynamics (LMD), an optimizer that trains neural networks with multiplicative updates. LMD is derived from a Lie-group Bayesian learning rule with a fixed-variance log-normal posterior; the update multiplies positive/negative median parameters by exp(-eta(sign(nu_temp)+r)), where the regularizer r pulls each EG+- component toward a preset median m_r. The paper reports experiments on ViT/ImageNet and GPT-2/OpenWebText, including forward passes emulated in MXFP6 (and MXFP4 for ViT ablations), and claims stable from-scratch training with no degradation under low-precision forward passes, plus suppressed weight growth relative to Madam. It also includes ablations of multiplicative versus additive weight decay and of sampled versus mean training.
Significance. If the central claim holds, this would be the first demonstration that multiplicative optimizers can train large transformers from scratch and tolerate low-precision forward passes, which is relevant to energy-efficient hardware. The derivation of the multiplicative weight-decay step (Eqs. 8-10) is internally consistent, the code is public, the main table reports n=3 repeated runs, and the ablation structure cleanly separates regularization from noise. However, the GPT-2 low-precision claim currently rests on a sequence-length/batch-size confound, no Lion baseline isolates the effect of the sign-momentum component, and the new hyperparameters are not studied for sensitivity. The paper's own Limitations section correctly notes that backward passes remain bf16 and no real speedups were measured, so the abstract should be qualified accordingly.
major comments (4)
- [§4.1, Table 1 and Figure 1 caption] The GPT-2 comparison confounds optimizer with sequence length and batch size: AdamW is evaluated at sequence length 1024 with batch size 64, while LMD and Madam use sequence length 4096 and batch size 16. The figure caption's statement that the token count per step is the same does not control for sequence-length-dependent optimization dynamics, and the table shows that AdamW at sequence length 4096 in bf16 is already unstable (4.790±2.017). Hence the claim that "AdamW cannot stably learn using low-precision forward passes" is not established: the degradation from 2.937 to 3.015 could be a sequence-length effect rather than an MXFP6 effect. This weakens the GPT-2 part of Contribution 1. Notably, at the same sequence length 1024 in bf16, AdamW (2.937±0.001) is slightly better than LMD (2.961±0.002), so aligning the configuration also changes the bf16 comparison. Please add an AdamW run at sequence length 4096 with MXFP6 forward passes, or an LMD run at sequence length 1024, before claiming low-precision stability for GPT-2.
- [§3.2 and §4 (Experimental Settings)] The method's ease-of-use claim depends on the default hyperparameters sigma=0.125 and m_r=0.01*exp(sigma^2/2), and Section 3.2 explicitly calls the scale-parameter initialization "heuristic." No sensitivity study is reported for either parameter, and the paper's "drop-in replacement" claim implies these values should transfer without per-model tuning. The fact that the same values work for ViT and GPT-2 is encouraging, but a small grid over sigma and m_r, or a statement of the working range, is needed before the transferability claim is supported.
- [§3, Algorithm 1; §4.1] LMD's update rule is built from Lion's signed momentum and interpolation coefficients, yet no Lion baseline is run in Table 1. Because Lion is a strong and widely used transformer optimizer, the comparison with AdamW and Madam alone does not identify whether the reported gains come from the multiplicative dynamics or from the Lion-style sign momentum with decoupled multiplicative regularization. A Lion baseline, with and without the log-normal noise, would make the contribution of the multiplicative mechanism explicit.
- [§4.1, Table 1 (ViT rows)] The ViT comparison shows a 9-point accuracy gap between LMD (77.06±0.08) and AdamW (68.11±0.38). This is much larger than typical optimizer effects for ViT training, so the reader cannot tell whether the gap reflects the optimizer or an under-tuned AdamW baseline. Please provide a learning-rate sweep or reference values for the ViT setup to show that 0.001 is well chosen for AdamW in this configuration, or explicitly discuss the comparison as a fixed-configuration comparison rather than a tuned one.
minor comments (6)
- [§4.1] Please state how test-time predictions are produced for LMD: with the median weights m only, or with Monte Carlo sampling. If sampling is used at test time, the accuracy comparison would reflect ensemble averaging rather than training dynamics alone.
- [§3.1, Eqs. (8)-(9)] The variance notation is inconsistent: Eq. (8) uses sigma_p^2 for the prior variance while the noise variance in Eq. (2) is sigma^2, and alpha=eta*gamma/sigma^2 introduces gamma without prior definition (it appears to be tau). Please align the notation.
- [§4.3, Figure 4] The "mean training" ablation in Figure 4 is not defined in Algorithm 1. Please specify whether it sets epsilon=1, uses m directly, or replaces the sampled weights by their mean, and confirm that all other hyperparameters are identical to the sampled runs.
- [§4, ViT and GPT-2 settings] There are several typos and formatting errors: "begingJ= 8" in the ViT settings, "m_r = 0.01×exp(sigma^2/2)1" with a stray superscript, and "2.1×10 6 tokens per step" should be typeset consistently.
- [§3.1] The log-normal prior is imposed on the positive EG+- components theta+ and theta-, not on the effective network weights theta+-theta-; the statement that the penalty "does not force weights to zero" should be restricted to the components, because the effective weights are already centered at zero by construction.
- [§5 (Limitations) and Abstract] The abstract says "low-precision inference and learning on future energy-efficient hardware," but the paper only evaluates low-precision forward passes, with backward passes in bf16 and no hardware speedups. Please align the abstract and contribution claims with Section 5's limitations.
Circularity Check
No significant circularity: LMD's central results are external benchmark measurements; the self-cited Lie-Group BLR derivation supplies independent, non-self-targeting support.
full rationale
The paper's main claims are empirical: LMD trains ViT on ImageNet and GPT-2 on OpenWebText from scratch, including with MXFP6 forward-pass emulation. These results are measured against external benchmarks and are not derived from the optimizer's defining equations. The update rule is imported from the authors' prior Lie-Group BLR work (Kiral et al., 2023, App. A.4), and the paper explicitly says 'We use Algorithm 2 in Kiral et al. (2023, App. A.4)' and 'We specialize the updates in Kiral et al. (2023, App. A.4)'. This is a self-citation, but the cited work is a peer-reviewed, published derivation with stated assumptions that do not include the present target result (stable large-scale low-precision training); under the review rules this counts as independent evidence rather than circularity. The multiplicative regularizer in Eq. 8 is deliberately constructed to pull the median m toward m_r, so the observed suppression of weight growth is a designed property of the update rather than a predicted consequence; the paper does not disguise this as a fitted prediction, and the stability/accuracy conclusions rest on external evaluations and ablations. The initialization in Eq. 12 is defined so that the mean of Aθ equals θ0, again a construction rather than a fitted input renamed as a prediction. The hyperparameters σ=0.125 and m_r=0.01 exp(σ^2/2) are stated as fixed settings rather than fitted to the reported results. The GPT-2 low-precision comparison uses different sequence lengths for AdamW versus LMD/Madam, but that is an experimental-validity concern, not a circularity of derivation. No load-bearing step in the derivation chain reduces to its own inputs by construction.
Assumptions & free parameters
free parameters (4)
- sigma (log-normal noise std) =
0.125
- m_r (log-normal prior median) =
0.01 * exp(sigma^2/2)
- learning rate eta =
0.005 (LMD for ViT and GPT-2)
- momentum coefficients beta1, beta2 =
0.95, 0.99
assumptions (4)
- standard math The Lie-Group Bayesian Learning Rule update in Eq. (5) is a correct optimization rule for the variational objective (Eq. 4).
- domain assumption Weights can be restricted to positive values and represented as theta = theta+ - theta- via the EG± trick without losing expressivity for large transformers.
- ad hoc to paper A log-normal posterior with fixed variance sigma^2 I and median m is an adequate variational family for large-scale network training.
- domain assumption Quantization error in MX low-precision formats is effectively multiplicative, so multiplicative noise is preserved after quantization.
Cite this review
Pith. "Pith review of Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks." pith.science (2026). https://pith.science/paper/WIEMSXC2
@misc{pith2026250617768,
author = {Pith},
title = {Pith review of: Log-Normal Multiplicative Dynamics for Stable Low-Precision Training of Large Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/WIEMSXC2}},
note = {Machine review of arXiv:2506.17768}
}
read the original abstract
Studies in neuroscience have shown that biological synapses follow a log-normal distribution whose transitioning can be explained by noisy multiplicative dynamics. Biological networks can function stably even under dynamically fluctuating conditions arising due to unreliable synaptic transmissions. Here we ask: Is it possible to design similar multiplicative training in artificial neural networks? To answer this question, we derive a Bayesian learning rule that assumes log-normal posterior distributions over weights which gives rise to a new Log-Normal Multiplicative Dynamics (LMD) algorithm. The algorithm uses multiplicative updates with both noise and regularization applied multiplicatively. The method is as easy to implement as Adam and only requires one additional vector to store. Our results show that LMD achieves stable and accurate training-from-scratch under low-precision forward operations for Vision Transformer and GPT-2. These results suggest that multiplicative dynamics, a biological feature, may enable stable low-precision inference and learning on future energy-efficient hardware.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
M+Adam: Low-Precision Training via Additive-Multiplicative Optimization
M+Adam combines additive and multiplicative update branches, avoiding low-precision rounding stalls and improving LLaMA-style pretraining perplexity compared with AdamW at BF16, FP8, and FP4 master-weight storage.
Reference graph
Works this paper leans on
-
[1]
Laurence Aitchison, Jannes Jegminat, Jorge Aurelio Menendez, Jean-Pascal Pfister, Alexandre Pouget, and Peter E. Latham. Synaptic plasticity as B ayesian inference. Nature Neuroscience, 24 0 (4): 0 565--571, 2021
work page 2021
-
[2]
Perceptron learning with sign-constrained weights
D J Amit, K Y M Wong, and C Campbell. Perceptron learning with sign-constrained weights. Journal of Physics A: Mathematical and General, 22 0 (12): 0 2039, 1989
work page 1989
-
[3]
The Effects of Adding Noise During Backpropagation Training on a Generalization Performance
Guozhong An. The Effects of Adding Noise During Backpropagation Training on a Generalization Performance . Neural Computation, 8 0 (3): 0 643--674, 1996
work page 1996
-
[4]
The Multiplicative Weights Update Method: a Meta-Algorithm and Applications
Sanjeev Arora, Elad Hazan, and Satyen Kale. The Multiplicative Weights Update Method: a Meta-Algorithm and Applications . Theory of Computing, 8 0 (6): 0 121--164, 2012
work page 2012
-
[5]
Nanoconnectomic upper bound on the variability of synaptic plasticity
Jr Bartol, Thomas M, Cailey Bromer, Justin Kinney, Michael A Chirillo, Jennifer N Bourne, Kristen M Harris, and Terrence J Sejnowski. Nanoconnectomic upper bound on the variability of synaptic plasticity. eLife, 4: 0 e10778, 2015
work page 2015
-
[6]
sign SGD: compressed optimisation for non-convex problems
Jeremy Bernstein, Yu - Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar. sign SGD: compressed optimisation for non-convex problems. In International Conference on Machine Learning, 2018
work page 2018
-
[7]
Learning compositional functions via multiplicative weight updates
Jeremy Bernstein, Jiawei Zhao, Markus Meister, Ming-Yu Liu, Anima Anandkumar, and Yisong Yue. Learning compositional functions via multiplicative weight updates. In Advances in Neural Information Processing Systems, 2020
work page 2020
-
[8]
Devansh Bisla, Jing Wang, and Anna Choromanska. Low- P ass F iltering SGD for R ecovering F lat O ptima in the D eep L earning O ptimization L andscape. In International Conference on Artificial Intelligence and Statistics, 2022
work page 2022
Show all 66 references
-
[9]
Understanding D ecoupled and E arly W eight D ecay
Johan Bjorck, Kilian Weinberger, and Carla Gomes. Understanding D ecoupled and E arly W eight D ecay. In AAAI Conference on Artifial Intelligence, 2021
2021
-
[10]
Weight U ncertainty in N eural N etworks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. Weight U ncertainty in N eural N etworks. In International Conference on Machine Learning, 2015
2015
-
[11]
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V. Le. Symbolic D iscovery of O ptimization A lgorithms. In Advances in Neural Information Processing Systems, 2023
2023
-
[12]
Exploring Q uantization for E fficient P re- T raining of T ransformer L anguage M odels
Kamran Chitsaz, Quentin Fournier, Goncalo Mordido, and Sarath Chandar. Exploring Q uantization for E fficient P re- T raining of T ransformer L anguage M odels. In Findings of the Association for Computational Linguistics: EMNLP, 2024
2024
-
[13]
The D ynamic S ynapse
Daniel Choquet and Antoine Triller. The D ynamic S ynapse. Neuron, 80 0 (3): 0 691--703, 2013
2013
-
[14]
Brain-like learning with exponentiated gradients
Jonathan Cornford, Roman Pogodin, Arna Ghosh, Kaiwen Sheng, Brendan A Bicknell, Olivier Codol, Beverley A Clark, Guillaume Lajoie, and Blake A Richards. Brain-like learning with exponentiated gradients. bioRxiv, 2024
2024
-
[15]
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. Training deep neural networks with low precision multiplications. In Workshop Track Proceedings of the International Conference on Learning Representations, 2015
2015
-
[16]
Higham, Theo Mary, and Mantas Mikaitis
Matteo Croci, Massimiliano Fasi, Nicholas J. Higham, Theo Mary, and Mantas Mikaitis. Stochastic rounding: implementation, error analysis and applications. Royal Society Open Science, 9: 0 211631, 2022
2022
-
[17]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009
2009
-
[18]
An I mage is W orth 16x16 W ords: T ransformers for I mage R ecognition at S cale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An I mage is W orth 16x16 W ords: T ransformers for I mage R ecognit...
2021
-
[19]
A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting
Yoav Freund and Robert E Schapire. A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting . Journal of Computer and System Sciences, 55 0 (1): 0 119--139, 1997
1997
-
[20]
Exponentiated G radient M eets G radient D escent
Udaya Ghai, Elad Hazan, and Yoram Singer. Exponentiated G radient M eets G radient D escent. In International Conference on Algorithmic Learning Theory, 2020
2020
-
[21]
Openwebtext corpus
Aaron Gokaslan and Vanya Cohen. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019
2019
-
[22]
Practical Variational Inference for Neural Networks
Alex Graves. Practical Variational Inference for Neural Networks . In Advances in Neural Information Processing Systems, volume 24, 2011
2011
-
[23]
Grigoriadis and Leonid G
Michael D. Grigoriadis and Leonid G. Khachiyan. A sublinear-time randomized approximation algorithm for matrix games. Oper. Res. Lett., 18 0 (2): 0 53–58, 1995
1995
-
[24]
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. Deep learning with limited numerical precision. In International Conference on Machine Learning, 2015
2015
-
[25]
Bridging the G ap B etween LLM s and LNS with Dynamic Data Format and Architecture Codesign
Pouya Haghi, Chunshu Wu, Zahra Azad, Yanfei Li, Andrew Gui, Yuchen Hao, Ang Li, and Tony Tong Geng. Bridging the G ap B etween LLM s and LNS with Dynamic Data Format and Architecture Codesign . In IEEE/ACM International Symposium on Microarchitecture, 2024
2024
-
[26]
A study of bfloat16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, et al. A study of bfloat16 for deep learning training. arXiv preprint arXiv:1905.12322, 2019
1905 arXiv
-
[27]
Network P lasticity as B ayesian I nference
David Kappel, Stefan Habenschuss, Robert Legenstein, and Wolfgang Maass. Network P lasticity as B ayesian I nference. PLOS Computational Biology, 11 0 (11): 0 e1004485, 2015
2015
-
[28]
Andrej Karpathy. NanoGPT . https://github.com/karpathy/nanoGPT, 2022
2022
-
[29]
Ziv, Hitoshi Okazaki, Sho Yagishita, and Taro Toyoizumi
Haruo Kasai, Noam E. Ziv, Hitoshi Okazaki, Sho Yagishita, and Taro Toyoizumi. Spine dynamics in the brain, mental disorders and artificial neural networks. Nature Reviews Neuroscience, 22: 0 407--422, 2021
2021
-
[30]
The B ayesian L earning R ule
Mohammad Emtiyaz Khan and Håvard Rue. The B ayesian L earning R ule. Journal of Machine Learning Research, 24 0 (281): 0 1--46, 2023
2023
-
[31]
Fast and S calable B ayesian D eep L earning by W eight- P erturbation in A dam
Mohammad Emtiyaz Khan, Didrik Nielsen, Voot Tangkaratt, Wu Lin, Yarin Gal, and Akash Srivastava. Fast and S calable B ayesian D eep L earning by W eight- P erturbation in A dam. In International Conference on Machine Learning, 2018
2018
-
[32]
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling. Variational dropout and the local reparameterization trick. In Advances in Neural Information Processing Systems, 2015
2015
-
[33]
The L ie- G roup B ayesian L earning R ule
Eren Mehmet Kiral, Thomas Moellenhoff, and Mohammad Emtiyaz Khan. The L ie- G roup B ayesian L earning R ule. In International Conference on Artificial Intelligence and Statistics, 2023
2023
-
[34]
Jyrki Kivinen and Manfred K. Warmuth. Exponentiated Gradient versus Gradient Descent for Linear Predictors . Information and Computation, 132 0 (1): 0 1--63, 1997
1997
-
[35]
Lee, Daisuke Miyashita, Elaina Chai, Boris Murmann, and S
Edward H. Lee, Daisuke Miyashita, Elaina Chai, Boris Murmann, and S. Simon Wong. Lognet: Energy-efficient neural networks using logarithmic computation. In International Conference on Acoustics, Speech and Signal Processing, 2017
2017
-
[36]
Learning Quickly When Irrelevant Attributes Abound: A New Linear-Threshold Algorithm
Nick Littlestone. Learning Quickly When Irrelevant Attributes Abound: A New Linear-Threshold Algorithm . Mach. Learn., 2 0 (4): 0 285–318, 1988
1988
-
[37]
Multiplicative Dynamics Underlie the Emergence of the Log-Normal Distribution of Spine sizes in the Neocortex In Vivo
Yonatan Loewenstein, Annerose Kuras, and Simon Rumpel. Multiplicative Dynamics Underlie the Emergence of the Log-Normal Distribution of Spine sizes in the Neocortex In Vivo . Journal of Neuroscience, 31 0 (26): 0 9481--9488, 2011
2011
-
[38]
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization . In International Conference on Learning Representations, 2019
2019
-
[39]
Diamos, Erich Elsen, David Garc \' a, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David Garc \' a, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. Mixed Precision Training . In International Conference on Learning Representations, 2018
2018
-
[40]
FP8 Formats for Deep Learning
Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, et al. FP8 Formats for Deep Learning . arXiv preprint arXiv:2209.05433, 2022
2022 arXiv
-
[41]
Lee, and Boris Murmann
Daisuke Miyashita, Edward H. Lee, and Boris Murmann. Convolutional Neural Networks using Logarithmic Data Representation . arXiv preprint arXiv:1603.01025, 2016
2016 arXiv
-
[42]
Up or Down? Adaptive Rounding for Post-Training Quantization
Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort. Up or Down? Adaptive Rounding for Post-Training Quantization . In International Conference on Machine Learning, 2020
2020
-
[43]
Ex Uno Pluria: Insights on Ensembling in Low Precision Number Systems
Giung Nam and Juho Lee. Ex Uno Pluria: Insights on Ensembling in Low Precision Number Systems . In Advances in Neural Information Processing Systems, 2024
2024
-
[44]
Anticorrelated Noise Injection for Improved Generalization
Antonio Orvieto, Hans Kersting, Frank Proske, Francis Bach, and Aurelien Lucchi. Anticorrelated Noise Injection for Improved Generalization . In International Conference on Machine Learning, 2022
2022
-
[45]
FP8-LM: Training FP8 Large Language Models
Houwen Peng, Kan Wu, Yixuan Wei, Guoshuai Zhao, Yuxiang Yang, Ze Liu, Yifan Xiong, Ziyue Yang, Bolin Ni, Jingcheng Hu, Ruihang Li, Miaosen Zhang, Chen Li, Jia Ning, Ruizhe Wang, Zheng Zhang, Shuguang Liu, Joe Chau, Han Hu, and Peng Cheng. FP8-LM: Training FP8 Large Language Mo...
-
[46]
MX Pytorch Emulation Library
Project. MX Pytorch Emulation Library . https://github.com/microsoft/microxcaling, 2023
2023
-
[47]
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language Models are Unsupervised Multitask Learners . OpenAI blog, 2019
2019
-
[48]
OpenAI Triton on NVIDIA Blackwell Boosts AI Performance and Programmability , 2025
Pradeep Ramani, Jason Knight, Philippe Tillet, Thomas Raoux, Pawe Szczerbuk, and Peter Bell. OpenAI Triton on NVIDIA Blackwell Boosts AI Performance and Programmability , 2025. URL https://developer.nvidia.com/blog/openai-triton-on-nvidia-blackwell-boosts-ai-performance-and-pr...
2025
-
[49]
OCP Microscaling (MX) Specification
Bita Rouhani, Nitin Garegrat, Tom Savell, Ankit More, Kyung-Nam Han, Mathew Zhao, Ritchie amd Hall, Jasmine Klar, Eric Chung, Yuan Yu, Michael Schulte, Ralph Wittig, Ian Bratt, Nigel Stephens, Jelena Milanovic, John Brothers, Pradeep Dubey, Marius Cornea, Alexander Heinecke, A...
2023
-
[50]
With Shared Microexponents, A Little Shifting Goes a Long Way
Bita Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, Levi Melnick, Maximilian Golub, Girish Varatkar, Lai Shao, Gaurav Kolhe, Dimitry Melts, Jasmine Klar, Renee L'Heureux, Matt Perry, Doug Burger, Eric Chung, Zhaoxia (...
2023
-
[51]
Microscaling Data Formats for Deep Learning
Bita Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Lan...
-
[52]
Learning in Spiking Neural Networks by Reinforcement of Stochastic Synaptic Transmission
H.Sebastian Seung. Learning in Spiking Neural Networks by Reinforcement of Stochastic Synaptic Transmission . Neuron, 40 0 (6): 0 1063--1073, 2003
2003
-
[53]
Variational Learning is Effective for Large Deep Networks
Yuesong Shen, Nico Daheim, Bai Cong, Peter Nickl, Gian Maria Marconi, Clement Bazan, Rio Yokota, Iryna Gurevych, Daniel Cremers, Mohammad Emtiyaz Khan, and Thomas M\" o llenhoff. Variational Learning is Effective for Large Deep Networks . In International Conference on Machine...
2024
-
[54]
Improving robustness to corruptions with multiplicative weight perturbations
Trung Trinh, Markus Heinonen, Luigi Acerbi, and Samuel Kaski. Improving robustness to corruptions with multiplicative weight perturbations. In Advances in Neural Information Processing Systems, 2024
2024
-
[55]
Training LLMs with MXFP4
Albert Tseng, Tao Yu, and Youngsuk Park. Training LLMs with MXFP4 . In Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, 2025
2025
-
[56]
QUEST: A 7.49TOPS multi-purpose log-quantized DNN inference engine stacked on 96MB 3D SRAM using inductive-coupling technology in 40nm CMOS
Kodai Ueyoshi, Kota Ando, Kazutoshi Hirose, Shinya Takamaeda-Yamazaki, Junichiro Kadomoto, Tomoki Miyata, Mototsugu Hamada, Tadahiro Kuroda, and Masato Motomura. QUEST: A 7.49TOPS multi-purpose log-quantized DNN inference engine stacked on 96MB 3D SRAM using inductive-coupling...
2018
-
[57]
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is All you Need . In Advances in Neural Information Processing Systems, volume 30, 2017
2017
-
[58]
Qualcomm Cloud AI 100 Accelerates Large Language Model Inference by 2x Using Microscaling (Mx) Formats
Colin Verrilli. Qualcomm Cloud AI 100 Accelerates Large Language Model Inference by 2x Using Microscaling (Mx) Formats . Qualcomm Developer Blog, 2024. URL https://www.qualcomm.com/developer/blog/2024/01/qualcomm-cloud-ai-100-accelerates-large-language-model-inference-2x-using...
2024
-
[59]
Training Deep Neural Networks with 8-bit Floating Point Numbers
Naigang Wang, Jungwook Choi, Daniel Brand, Chia-Yu Chen, and Kailash Gopalakrishnan. Training Deep Neural Networks with 8-bit Floating Point Numbers . In Advances in Neural Information Processing Systems, 2018
2018
-
[60]
Optimizing Large Language Model Training Using FP4 Quantization
Ruizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao, Ziyue Yang, Baining Guo, Zhengjun Zha, and Peng Cheng. Optimizing Large Language Model Training Using FP4 Quantization . arXiv preprint arXiv:2501.17116, 2025
2025 arXiv
-
[61]
Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, and Roger Grosse. Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches . In International Conference on Machine Learning, 2018
2018
-
[62]
PyTorch Image Models
Ross Wightman. PyTorch Image Models . https://github.com/rwightman/pytorch-image-models, 2019
2019
-
[63]
Collage: Light-Weight Low-Precision Strategy for LLM Training
Tao Yu, Gaurav Gupta, Karthick Gopalswamy, Amith Mamidala, Hao Zhou, Jeffrey Huynh, Youngsuk Park, Ron Diamant, Anoop Deoras, and Luke Huan. Collage: Light-Weight Low-Precision Strategy for LLM Training . In International Conference on Machine Learning, 2024
2024
-
[64]
Optimal Information Processing and Bayes's Theorem
Arnold Zellner. Optimal Information Processing and Bayes's Theorem . The American Statistician, 42 0 (4): 0 278--280, 1988
1988
-
[65]
Low-Precision Stochastic Gradient L angevin Dynamics
Ruqi Zhang, Andrew Gordon Wilson, and Christopher De Sa. Low-Precision Stochastic Gradient L angevin Dynamics . In International Conference on Machine Learning, 2022
2022
-
[66]
Dally, and Anima Anandkumar
Jiawei Zhao, Steve Dai, Rangharajan Venkatesan, Brian Zimmer, Mustafa Ali, Ming-Yu Liu, Brucek Khailany, William J. Dally, and Anima Anandkumar. LNS-Madam: Low-Precision Training in Logarithmic Number System Using Multiplicative Weight Update . IEEE Transactions on Computers, ...
2022
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.