Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:41:00.476703Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 9 inbound Pith citation observations for arXiv:2412.11768.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:41:00.476703Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:28:29.191879Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T13:24:53.244110Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5a23ca41-e8d1-43a6-904f-ab03cf556daa · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e7eee76-3c22-4850-865b-ab454744731a · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need S., Mehrotra, A., Dudziak, ., and Lane, N
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70034392-ec84-4bd3-8ccb-3e8ed37e489c · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need and Lavie, A
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c9fc9a5-4336-495e-b260-8f8f443e1401 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need signSGD: Compressed Optimisation for Non-Convex Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd9e5b54-53a7-4f03-aa07-f83b16207700 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Better plain ViT baselines for ImageNet-1k
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64e9856-5b81-4d10-8f54-e95aff095684 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Big vision
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46a82b51-02cb-4b4f-94f8-923ba24ef230 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need How does topology influence gradient propagation and model performance of deep networks with DenseNet-type skip connections?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b4bca43-ca98-4031-8e9a-b268f7c34380 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need A downsampled variant of imagenet as an alternative to the cifar datasets, 2017
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fd50a9df-52b0-47bd-b89e-a5a480944992 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Imagenet: A large-scale hierarchical image database
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfdb851c-d78f-4b8c-b253-2d17c3a35a72 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need and Zettlemoyer, L
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4dd42e93-584e-4f38-a86c-57b19a161d3b · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need 8-bit Optimizers via Block-wise Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe0f9c9-9af5-445b-b222-b30bf7f31d8b · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26231635-8a70-4e14-8997-9362671b1f4c · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7a57760-50ab-4763-b067-9cd6a06d30fe · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need NATS-Bench : Benchmarking nas algorithms for architecture topology and size
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c369c1-e0a9-45c9-9339-b5872222103c · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9018f376-683a-4227-b46a-15b243e5e34e · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Incorporating Nesterov Momentum into Adam
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation de17ac81-8889-4062-84f1-8627e2608bdf · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Adaptive subgradient methods for online learning and stochastic optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c41f2f4-da72-4268-9b6a-6e812d49d4ac · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eccf62c8-7213-4717-bd66-77cc462288cc · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Pruning Neural Networks at Initialization: Why are We Missing the Mark?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e9d33e-5ad5-4198-bebe-5de0ba7413cb · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Improving Robustness with Adaptive Weight Decay
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4402df0-377b-47a1-aaea-96a0e24f1c46 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Adaptive Gradient Methods at the Edge of Stability
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7519f124-418b-4358-80de-dcdf73fcf5ca · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need and Cohen, V
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1cd2074-91dc-42fb-9869-baca4a7608e8 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Generating Sequences With Recurrent Neural Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8da474ab-e724-4b5a-b9ce-b87174212ef6 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Z., Shi, Y., Chen, Y., Fan, Z., Xiao, W., Zhao, R., Chang, S., Wu, W., et al
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7928531-dfb2-4d62-af9a-38473cacb35e · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Deep Residual Learning for Image Recognition
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b65eb7d-ecf1-4c4d-85d0-3aaf66469f48 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7099f1fe-2fcc-4188-9bb6-dbb3eb9291fc · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 582356df-2de9-4254-b0a9-47c3509c02c2 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Denoising diffusion probabilistic models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb65f2f-e8df-4321-b96f-36838c1a07da · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need LoRA: Low-Rank Adaptation of Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52193b27-5753-4ac8-9988-8df84e39c961 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Scaling Laws for Neural Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5a57a9-42c7-4f34-b52e-b159f74340db · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Adam: A Method for Stochastic Optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe80ac8-24c4-4ee0-b533-ef67b9ab2557 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need and Hinton, G
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6688607b-0369-4907-83ae-57d2fd1e6aad · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need On weight initialization in deep neural networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99af92b6-4ce7-4725-b528-a1be92ec9af6 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60a2f80e-634c-4856-b321-07b3d94b4a82 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need SNIP: Single-shot Network Pruning based on Connection Sensitivity
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 742582f8-c51e-454e-b0e5-54f8e95529ed · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Balance is Essence: Accelerating Sparse Training via Adaptive Gradient Correction
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef551001-d278-4bbf-8acd-24eeb31204c7 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Memory Efficient Optimizers with 4-bit States
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eadaa16-86d9-41ed-887a-d64a5ac464fb · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Zico: Zero-shot nas via inverse coefficient of variation on gradients
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ce442a3-b804-4bdb-9985-12fff3f9e427 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Convergence of adam under relaxed assumptions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91102306-bbff-4c09-98b3-8b96bb675b87 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Rouge: A package for automatic evaluation of summaries
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddff11e7-46f3-4f80-83da-3fa88889e118 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need On the Variance of the Adaptive Learning Rate and Beyond
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eedcda6-fbec-4afd-aeb7-1378be580899 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need On the variance of the adaptive learning rate and beyond, 2021
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d9d4379-4e18-4018-8f1d-45a53ff98e04 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need and Hutter, F
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 167d5afc-e193-428f-bcfc-d8346e176657 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Prodigy: An Expeditiously Adaptive Parameter-Free Learner
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e89c90d-157a-4caf-8912-92490e98f85d · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a40c3e4-987b-423d-8251-bc7346a0ce27 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need The E2E Dataset: New Challenges For End-to-End Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373f4ce3-1a37-40c1-aa71-cf20b221934e · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Bleu: a method for automatic evaluation of machine translation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15c4f053-3d79-4c1e-ad3b-4e58998ed5df · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Language models are unsupervised multitask learners
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24231697-645e-4201-9634-2c76dcbe9813 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need High-Resolution Image Synthesis with Latent Diffusion Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1dc02a-325c-44dc-be1c-b84fb5dc8ceb · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need and Stern, M
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7401dcd3-6bd2-4212-9c77-edd1e297e6a1 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545de7b4-971e-4bbd-8667-61221d5c32ea · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need L., and Ganguli, S
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6180ba-acf1-4b70-b51d-5c1ccab5160a · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Gemini: A Family of Highly Capable Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34e7ab6-9b21-4e17-ad2e-01fa55cb08a1 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Attention Is All You Need
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a9fb59-75dd-44f1-a4ce-8613026788c2 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Cider: Consensus-based image description evaluation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dcb21ad-ea67-4cbb-8539-8088058cb09a · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4f03eb27-e772-4db4-8146-8b5a2f54ccbd · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Exploiting network compressibility and topology in zero-cost nas
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f7776041-aaae-4300-ae83-0542ecef541a · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Early Convolutions Help Transformers See Better
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d978a1-bd6f-440f-ae8c-6cf9dc4cd46e · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need ADADELTA: An Adaptive Learning Rate Method
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ac07514-e2bd-4385-b6dd-19edac1129d8 · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bbed0d-d504-4638-8602-38806799ab8d · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Why Transformers Need Adam: A Hessian Perspective
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7856eda-43d6-47fa-8ab5-68f7db99463a · outbound
No More Adam: Learning Rate Scaling at Initialization is All You Need Adam-mini: Use Fewer Learning Rates To Gain More
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42800392-4b12-4e39-aad4-2377e9ba907e · inbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e15bbc-102e-42d2-ac99-0f888910f114 · inbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4028448-fec2-4f51-9294-9d575a3c015a · inbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938f875f-cad8-42ce-916b-ecc6ec5e8f30 · inbound
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c732579d-e615-4f62-8939-4f4a3b6dc7df · inbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46d1497f-dc9c-4ad4-9a10-66c3666877b6 · inbound
Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fffd8c75-6370-485e-bd02-c6bfaf3be839 · inbound
Layerwise LQR for Geometry-Aware Optimization of Deep Networks No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90d18fcd-5287-4f14-b605-6fccffadd7e3 · inbound
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2da8cb7e-00cc-4a91-b77c-518ee26566e6 · inbound
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.