Pith. sign in

Paper Citation Record · LEDGER

No More Adam: Learning Rate Scaling at Initialization is All You Need

As of 12 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 9 inbound Pith citation observations for arXiv:2412.11768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11768 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:41:00.476703Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:28:29.191879Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T13:24:53.244110Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a23ca41-e8d1-43a6-904f-ab03cf556daa · outbound

This paper cites write newline.

No More Adam: Learning Rate Scaling at Initialization is All You Need write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.154159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.154159Z digest=sha256:3ad627e26a2ce90c62a12549abefbff22b5d36e23d9a5fe153c74f9a14d34789

Observation 0e7eee76-3c22-4850-865b-ab454744731a · outbound

This paper cites S., Mehrotra, A., Dudziak, ., and Lane, N.

No More Adam: Learning Rate Scaling at Initialization is All You Need S., Mehrotra, A., Dudziak, ., and Lane, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.603289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.160319Z digest=sha256:db41c6789af20373f0005128e3fed7dae1d238159eea4d8fdb22a7506e0368f7

Observation 70034392-ec84-4bd3-8ccb-3e8ed37e489c · outbound

This paper cites and Lavie, A.

No More Adam: Learning Rate Scaling at Initialization is All You Need and Lavie, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.165526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.165526Z digest=sha256:688b2e4feebdc61b0d6b2a21c825d29a7572244017ebf5b949260ff08eb3a5b5

Observation 1c9fc9a5-4336-495e-b260-8f8f443e1401 · outbound

This paper cites signSGD: Compressed Optimisation for Non-Convex Problems.

No More Adam: Learning Rate Scaling at Initialization is All You Need signSGD: Compressed Optimisation for Non-Convex Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.170163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.170163Z digest=sha256:e31c022c1916495bfe5b6cbb51e1bf6a352196cb7c64d2422b259a014ccb20de

Observation cd9e5b54-53a7-4f03-aa07-f83b16207700 · outbound

This paper cites Better plain ViT baselines for ImageNet-1k.

No More Adam: Learning Rate Scaling at Initialization is All You Need Better plain ViT baselines for ImageNet-1k

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.175888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.175888Z digest=sha256:0bd0959c54e56bdd488166e6f414ccbea3efb388ac133a4ca734e22d69c41bb3

Observation a64e9856-5b81-4d10-8f54-e95aff095684 · outbound

This paper cites Big vision.

No More Adam: Learning Rate Scaling at Initialization is All You Need Big vision

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.583062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.181788Z digest=sha256:625870deeb96803ad14b3d9d0bb00d70dc688e5052c6e77168b6193352f5b18f

Observation 46a82b51-02cb-4b4f-94f8-923ba24ef230 · outbound

This paper cites How does topology influence gradient propagation and model performance of deep networks with DenseNet-type skip connections?.

No More Adam: Learning Rate Scaling at Initialization is All You Need How does topology influence gradient propagation and model performance of deep networks with DenseNet-type skip connections?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:41:01.179531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.188312Z digest=sha256:a91c601c0c311684dab646e9e36608e1f0c4452a79ba5f3e19e49d3ed3d19968

Observation 4b4bca43-ca98-4031-8e9a-b268f7c34380 · outbound

This paper cites A downsampled variant of imagenet as an alternative to the cifar datasets, 2017.

No More Adam: Learning Rate Scaling at Initialization is All You Need A downsampled variant of imagenet as an alternative to the cifar datasets, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.565581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.195445Z digest=sha256:2d1c6fe8f6b98ad454f5d399f09632048a3d132a8ed5dd6ad2e3fa6806fd606d

Observation fd50a9df-52b0-47bd-b89e-a5a480944992 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

No More Adam: Learning Rate Scaling at Initialization is All You Need Imagenet: A large-scale hierarchical image database

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.201114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.201114Z digest=sha256:6a78b8028d81b806f5b9baf40c64593d254ade865abafc99eb9274d4c5384b03

Observation cfdb851c-d78f-4b8c-b253-2d17c3a35a72 · outbound

This paper cites and Zettlemoyer, L.

No More Adam: Learning Rate Scaling at Initialization is All You Need and Zettlemoyer, L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.529366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.205808Z digest=sha256:a535ccf94b6d1665d842db0fed1694fae5ec55cdd14a7958e312f165d11621ae

Observation 4dd42e93-584e-4f38-a86c-57b19a161d3b · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

No More Adam: Learning Rate Scaling at Initialization is All You Need 8-bit Optimizers via Block-wise Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.210604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.210604Z digest=sha256:418241c669d76bc0872a9703fdea18ec204c182d4a5067a9ffd48140be279b79

Observation ffe0f9c9-9af5-445b-b222-b30bf7f31d8b · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

No More Adam: Learning Rate Scaling at Initialization is All You Need LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.215908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.215908Z digest=sha256:65f7032c59cab835e2d4a174ba0034f80942515e46e15121ff3fbc52a3988ebc

Observation 26231635-8a70-4e14-8997-9362671b1f4c · outbound

This paper cites Automatic evaluation of machine translation quality using n-gram co-occurrence statistics.

No More Adam: Learning Rate Scaling at Initialization is All You Need Automatic evaluation of machine translation quality using n-gram co-occurrence statistics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.513161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.221681Z digest=sha256:a4c7cbe5087e0030903e3b7de468dd556d4f2bb342b77455f0dd145fd1fe4da5

Observation a7a57760-50ab-4763-b067-9cd6a06d30fe · outbound

This paper cites NATS-Bench : Benchmarking nas algorithms for architecture topology and size.

No More Adam: Learning Rate Scaling at Initialization is All You Need NATS-Bench : Benchmarking nas algorithms for architecture topology and size

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.225895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.225895Z digest=sha256:e70392a6c0ef8feeb92027df1e30cc3c0909ea01fbdd40638724fd74da504a24

Observation a1c369c1-e0a9-45c9-9339-b5872222103c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

No More Adam: Learning Rate Scaling at Initialization is All You Need An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.231895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.231895Z digest=sha256:11d7c5db7b12419c1ca8a5285bb27147f882c8f3ee1f7192ce005d971451ea2f

Observation 9018f376-683a-4227-b46a-15b243e5e34e · outbound

This paper cites Incorporating Nesterov Momentum into Adam.

No More Adam: Learning Rate Scaling at Initialization is All You Need Incorporating Nesterov Momentum into Adam

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.497487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.236651Z digest=sha256:5a41fc4d37b1cc563f61a9d3085855060898760874bcd3f135986e76f0b231e6

Observation de17ac81-8889-4062-84f1-8627e2608bdf · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

No More Adam: Learning Rate Scaling at Initialization is All You Need Adaptive subgradient methods for online learning and stochastic optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.240925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.240925Z digest=sha256:0aaab22567cdfdac7c32c918e89f2cd177e9daa81c81a2af85f4d99ac453985b

Observation 3c41f2f4-da72-4268-9b6a-6e812d49d4ac · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

No More Adam: Learning Rate Scaling at Initialization is All You Need The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.246691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.246691Z digest=sha256:791393a65dfe9d48ae3dec01a640ec195bded718f8f1498a0974e041b93f0e5c

Observation eccf62c8-7213-4717-bd66-77cc462288cc · outbound

This paper cites Pruning Neural Networks at Initialization: Why are We Missing the Mark?.

No More Adam: Learning Rate Scaling at Initialization is All You Need Pruning Neural Networks at Initialization: Why are We Missing the Mark?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.252287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.252287Z digest=sha256:a579f96b89c37eaacb5c58de06e3e104caee6c08b9bfd749c03100bcee2ed599

Observation a2e9d33e-5ad5-4198-bebe-5de0ba7413cb · outbound

This paper cites Improving Robustness with Adaptive Weight Decay.

No More Adam: Learning Rate Scaling at Initialization is All You Need Improving Robustness with Adaptive Weight Decay

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:41:01.006008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.256707Z digest=sha256:500c41ce1c67515d521895efb3bfa3a2df6f68dd28cfd0774d45374f25944244

Observation c4402df0-377b-47a1-aaea-96a0e24f1c46 · outbound

This paper cites Adaptive Gradient Methods at the Edge of Stability.

No More Adam: Learning Rate Scaling at Initialization is All You Need Adaptive Gradient Methods at the Edge of Stability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.261965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.261965Z digest=sha256:b985a9047cfc62208b4943072baea19bd7568e4496a0484b2949aaff0da39561

Observation 7519f124-418b-4358-80de-dcdf73fcf5ca · outbound

This paper cites and Cohen, V.

No More Adam: Learning Rate Scaling at Initialization is All You Need and Cohen, V

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.269787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.269787Z digest=sha256:dabd784353c87010af4b18a9aae8cd5489591d155110e62adf7d6f51ea8ef5c8

Observation c1cd2074-91dc-42fb-9869-baca4a7608e8 · outbound

This paper cites Generating Sequences With Recurrent Neural Networks.

No More Adam: Learning Rate Scaling at Initialization is All You Need Generating Sequences With Recurrent Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.275533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.275533Z digest=sha256:a8c9269080bba75a7548548c7a1bbc72e470f3f5cd00686a83d7924d6a7602e1

Observation 8da474ab-e724-4b5a-b9ce-b87174212ef6 · outbound

This paper cites Z., Shi, Y., Chen, Y., Fan, Z., Xiao, W., Zhao, R., Chang, S., Wu, W., et al.

No More Adam: Learning Rate Scaling at Initialization is All You Need Z., Shi, Y., Chen, Y., Fan, Z., Xiao, W., Zhao, R., Chang, S., Wu, W., et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.461417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.281134Z digest=sha256:7b9a06884636225a5500215e4afc4c562c0a0885bc655efaf878aed524a443d0

Observation a7928531-dfb2-4d62-af9a-38473cacb35e · outbound

This paper cites Deep Residual Learning for Image Recognition.

No More Adam: Learning Rate Scaling at Initialization is All You Need Deep Residual Learning for Image Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.285857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.285857Z digest=sha256:8d3506bde1c38f07636a9272999ff1ee942aea4732eee88c932f3a67e5e316a1

Observation 4b65eb7d-ecf1-4c4d-85d0-3aaf66469f48 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

No More Adam: Learning Rate Scaling at Initialization is All You Need Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.444567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.292206Z digest=sha256:8e5c597d40712d38183c736bd0bc98666b836a9984c045f8b31666338e42319c

Observation 7099f1fe-2fcc-4188-9bb6-dbb3eb9291fc · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.

No More Adam: Learning Rate Scaling at Initialization is All You Need Neural networks for machine learning lecture 6a overview of mini-batch gradient descent

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.296488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.296488Z digest=sha256:d1f50a20041a379ea50f83d57bd411fc5d3d5ede8d936718df4ffc35bac30615

Observation 582356df-2de9-4254-b0a9-47c3509c02c2 · outbound

This paper cites Denoising diffusion probabilistic models.

No More Adam: Learning Rate Scaling at Initialization is All You Need Denoising diffusion probabilistic models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.301262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.301262Z digest=sha256:7ec364cef65e88a851b09b65c04c049aa1b00f5edd903615a1eab4bfa50a91a6

Observation 9cb65f2f-e8df-4321-b96f-36838c1a07da · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

No More Adam: Learning Rate Scaling at Initialization is All You Need LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.306792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.306792Z digest=sha256:70bf037d110c350e2f5ad31acb8376855627dd9830e346d8ff1d7c4250b83a93

Observation 52193b27-5753-4ac8-9988-8df84e39c961 · outbound

This paper cites Scaling Laws for Neural Language Models.

No More Adam: Learning Rate Scaling at Initialization is All You Need Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.311610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.311610Z digest=sha256:5c603a4f86151297da52a9abca02e533db58c85ae084308a981893d280f9e0c0

Observation 1a5a57a9-42c7-4f34-b52e-b159f74340db · outbound

This paper cites Adam: A Method for Stochastic Optimization.

No More Adam: Learning Rate Scaling at Initialization is All You Need Adam: A Method for Stochastic Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.316291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.316291Z digest=sha256:4b4d9fcb5d3fb138efd0d46211ba392a280660e287ae3747b0cb3dae98f46e84

Observation 5fe80ac8-24c4-4ee0-b533-ef67b9ab2557 · outbound

This paper cites and Hinton, G.

No More Adam: Learning Rate Scaling at Initialization is All You Need and Hinton, G

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.320903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.320903Z digest=sha256:a73c9d6544bbea6b37b3139f89183b2be7ae9801849439c930dfdf0fdcce7467

Observation 6688607b-0369-4907-83ae-57d2fd1e6aad · outbound

This paper cites On weight initialization in deep neural networks.

No More Adam: Learning Rate Scaling at Initialization is All You Need On weight initialization in deep neural networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.325872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.325872Z digest=sha256:b7a2e8af12bbb9d69d9d98e400a3f6da6090bd2c6190499190a836bb6e0a032e

Observation 99af92b6-4ce7-4725-b528-a1be92ec9af6 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

No More Adam: Learning Rate Scaling at Initialization is All You Need Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.331376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.331376Z digest=sha256:b7c550f3ff4f5673734fe317dfc1a91dd5315ad235980a72f9061098f552b35d

Observation 60a2f80e-634c-4856-b321-07b3d94b4a82 · outbound

This paper cites SNIP: Single-shot Network Pruning based on Connection Sensitivity.

No More Adam: Learning Rate Scaling at Initialization is All You Need SNIP: Single-shot Network Pruning based on Connection Sensitivity

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.336215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.336215Z digest=sha256:e3b958bdf403e373c326b6b0ac5b2f47eb2e7155ed3849ddf5355f11cf6de655

Observation 742582f8-c51e-454e-b0e5-54f8e95529ed · outbound

This paper cites Balance is Essence: Accelerating Sparse Training via Adaptive Gradient Correction.

No More Adam: Learning Rate Scaling at Initialization is All You Need Balance is Essence: Accelerating Sparse Training via Adaptive Gradient Correction

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:41:00.840623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.342292Z digest=sha256:fb6bd456632d5d4bdc8144e72ed1df5ebda5220e907ff38ae284bb40ddc42707

Observation ef551001-d278-4bbf-8acd-24eeb31204c7 · outbound

This paper cites Memory Efficient Optimizers with 4-bit States.

No More Adam: Learning Rate Scaling at Initialization is All You Need Memory Efficient Optimizers with 4-bit States

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.347140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.347140Z digest=sha256:b1b6e3cd025d6b91f96bcfda8f953e13a47e46a0785810563297d5ab3f14b805

Observation 6eadaa16-86d9-41ed-887a-d64a5ac464fb · outbound

This paper cites Zico: Zero-shot nas via inverse coefficient of variation on gradients.

No More Adam: Learning Rate Scaling at Initialization is All You Need Zico: Zero-shot nas via inverse coefficient of variation on gradients

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.397293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.352525Z digest=sha256:5c70d7864f33cbdf1ef2c9cf54890e6ea4576b8737ac35a4f865d96ffbaf4e80

Observation 1ce442a3-b804-4bdb-9985-12fff3f9e427 · outbound

This paper cites Convergence of adam under relaxed assumptions.

No More Adam: Learning Rate Scaling at Initialization is All You Need Convergence of adam under relaxed assumptions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.372884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.356987Z digest=sha256:5a596445831eff6c3b38af6f1e4dc6af9890cac5dd3db3c274fa945b0d31820e

Observation 91102306-bbff-4c09-98b3-8b96bb675b87 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

No More Adam: Learning Rate Scaling at Initialization is All You Need Rouge: A package for automatic evaluation of summaries

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.362073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.362073Z digest=sha256:46dd9164a653eb7cf3b2214cc03d76a652a9e8230dab6417f918700f4258649b

Observation ddff11e7-46f3-4f80-83da-3fa88889e118 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

No More Adam: Learning Rate Scaling at Initialization is All You Need On the Variance of the Adaptive Learning Rate and Beyond

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.366975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.366975Z digest=sha256:ef90f711198b2950566b80916d798e5f21f45fc3cc8edea661f8f9feb622f380

Observation 0eedcda6-fbec-4afd-aeb7-1378be580899 · outbound

This paper cites On the variance of the adaptive learning rate and beyond, 2021.

No More Adam: Learning Rate Scaling at Initialization is All You Need On the variance of the adaptive learning rate and beyond, 2021

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.342186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.376468Z digest=sha256:9a113ff4aa2d958aa50c48ddb83fb011300f68cd218977c0c5bd818be5bc5489

Observation 5d9d4379-4e18-4018-8f1d-45a53ff98e04 · outbound

This paper cites and Hutter, F.

No More Adam: Learning Rate Scaling at Initialization is All You Need and Hutter, F

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.381707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.381707Z digest=sha256:2e20383acc30c240c6f62497f85ad5ccce2efcfefb3294f2f3108ae4614276fb

Observation 167d5afc-e193-428f-bcfc-d8346e176657 · outbound

This paper cites Prodigy: An Expeditiously Adaptive Parameter-Free Learner.

No More Adam: Learning Rate Scaling at Initialization is All You Need Prodigy: An Expeditiously Adaptive Parameter-Free Learner

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.387287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.387287Z digest=sha256:d172bd17374c93ffae424634d0bbe6ba592d01b41037fa92cb661f37c398b788

Observation 8e89c90d-157a-4caf-8912-92490e98f85d · outbound

This paper cites an unresolved cited work.

No More Adam: Learning Rate Scaling at Initialization is All You Need Unresolved cited work

Reference 45

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:41:00.761892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.393613Z digest=sha256:2fc89ee6b6937eee88328bbf86156ae7aa79a1de996d3a12f2e53c728837c04f

Observation 8a40c3e4-987b-423d-8251-bc7346a0ce27 · outbound

This paper cites The E2E Dataset: New Challenges For End-to-End Generation.

No More Adam: Learning Rate Scaling at Initialization is All You Need The E2E Dataset: New Challenges For End-to-End Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.397834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.397834Z digest=sha256:d0536d4ac17e3e504177fea77f46761ac1f2e4a4ce6193e4a5e55a8a9ac07fc9

Observation 373f4ce3-1a37-40c1-aa71-cf20b221934e · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

No More Adam: Learning Rate Scaling at Initialization is All You Need Bleu: a method for automatic evaluation of machine translation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.403703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.403703Z digest=sha256:973fc12d6fa82a5a52e05c52c1f2e573645129ce686d3d9f3c1b298a829d77da

Observation 15c4f053-3d79-4c1e-ad3b-4e58998ed5df · outbound

This paper cites Language models are unsupervised multitask learners.

No More Adam: Learning Rate Scaling at Initialization is All You Need Language models are unsupervised multitask learners

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.408353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.408353Z digest=sha256:f03752d7948c36bd658588fd6d590d379922c941d79488872eb0e0e4c3986f0b

Observation 24231697-645e-4201-9634-2c76dcbe9813 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

No More Adam: Learning Rate Scaling at Initialization is All You Need High-Resolution Image Synthesis with Latent Diffusion Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.412871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.412871Z digest=sha256:5a6b74ad98f4ca8db9ea7aebf897d36cff3e73ff6b41a2e6efda0ccda3a54e0b

Observation 7a1dc02a-325c-44dc-be1c-b84fb5dc8ceb · outbound

This paper cites and Stern, M.

No More Adam: Learning Rate Scaling at Initialization is All You Need and Stern, M

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.418014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.418014Z digest=sha256:448b076f39e0fee695f84f61c13f281b0d0090b5e9b9ed3083ddad16dd7ded63

Observation 7401dcd3-6bd2-4212-9c77-edd1e297e6a1 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

No More Adam: Learning Rate Scaling at Initialization is All You Need How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.422171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.422171Z digest=sha256:ed627e3d9fe3288c7dd5a6a480f66cf7877dbb9b6646e529729db89e240b8060

Observation 545de7b4-971e-4bbd-8667-61221d5c32ea · outbound

This paper cites L., and Ganguli, S.

No More Adam: Learning Rate Scaling at Initialization is All You Need L., and Ganguli, S

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.426486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.426486Z digest=sha256:da95f00a75bd1abd9c8245b9d9afa3ad5f7a9f9dc23b65014440ec503ca59851

Observation 2c6180ba-acf1-4b70-b51d-5c1ccab5160a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

No More Adam: Learning Rate Scaling at Initialization is All You Need Gemini: A Family of Highly Capable Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.430238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.430238Z digest=sha256:bb32447de3ba43153ddc12e1499ff701edf035248bcfd1fec0b7f0da7f4950a2

Observation d34e7ab6-9b21-4e17-ad2e-01fa55cb08a1 · outbound

This paper cites Attention Is All You Need.

No More Adam: Learning Rate Scaling at Initialization is All You Need Attention Is All You Need

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.434176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.434176Z digest=sha256:92324b9405bbe079e7e519c998d2ff4b58ad9b90978107ef118d3c7c0969b287

Observation e3a9fb59-75dd-44f1-a4ce-8613026788c2 · outbound

This paper cites Cider: Consensus-based image description evaluation.

No More Adam: Learning Rate Scaling at Initialization is All You Need Cider: Consensus-based image description evaluation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.438884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.438884Z digest=sha256:72c4b779d4f0aa0c92fdb157ac5ed9d0382a4839f3a98e36a9f6ef72ea12dd6f

Observation 5dcb21ad-ea67-4cbb-8539-8088058cb09a · outbound

This paper cites Adagrad stepsizes: Sharp convergence over nonconvex landscapes.

No More Adam: Learning Rate Scaling at Initialization is All You Need Adagrad stepsizes: Sharp convergence over nonconvex landscapes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.246624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.443039Z digest=sha256:96edb81fd46fa390b4aaede627ca7b0ff7280e2e954f9f290adac8e1db2f1e06

Observation 4f03eb27-e772-4db4-8146-8b5a2f54ccbd · outbound

This paper cites Exploiting network compressibility and topology in zero-cost nas.

No More Adam: Learning Rate Scaling at Initialization is All You Need Exploiting network compressibility and topology in zero-cost nas

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:41:01.228816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:41:00.447107Z digest=sha256:6550fe5318d1ea0226d42d9fab02e067d852e8e293fdf55f928d5ab247268a7d

Observation f7776041-aaae-4300-ae83-0542ecef541a · outbound

This paper cites Early Convolutions Help Transformers See Better.

No More Adam: Learning Rate Scaling at Initialization is All You Need Early Convolutions Help Transformers See Better

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.454756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.454756Z digest=sha256:575734361aadbf856d15d8698e978bbfe4cf92ffe5a50269763de8c8f5fd4f80

Observation 07d978a1-bd6f-440f-ae8c-6cf9dc4cd46e · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

No More Adam: Learning Rate Scaling at Initialization is All You Need ADADELTA: An Adaptive Learning Rate Method

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.459733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.459733Z digest=sha256:f91a5b6b14bd0c72bbc9395f53dd069de82874f633f9211dcc2d58c90eafee53

Observation 0ac07514-e2bd-4385-b6dd-19edac1129d8 · outbound

This paper cites Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models.

No More Adam: Learning Rate Scaling at Initialization is All You Need Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.465623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.465623Z digest=sha256:f31de4d084a5017814a7b96f36bcae002b266b9985ca75ee2568e8b6a3f2481b

Observation 96bbed0d-d504-4638-8602-38806799ab8d · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

No More Adam: Learning Rate Scaling at Initialization is All You Need Why Transformers Need Adam: A Hessian Perspective

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.472499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.472499Z digest=sha256:88a6fb72c432e398b55918aa346aba6bab50357f6729db6d34ba950474436dc2

Observation a7856eda-43d6-47fa-8ab5-68f7db99463a · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

No More Adam: Learning Rate Scaling at Initialization is All You Need Adam-mini: Use Fewer Learning Rates To Gain More

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.476703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.476703Z digest=sha256:518f4f6b6a6be3a3d5e09ce946f533998f3622778c3440f764b5f0e559b676ba

Pith citing papers

Observation 42800392-4b12-4e39-aad4-2377e9ba907e · inbound

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training cites this paper.

SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T13:28:29.191879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:28:29.191879Z digest=sha256:8a012dbbbb4ca92205d97e4a6459913c87759c4c335de86af9391f5b63665558

Observation 16e15bbc-102e-42d2-ac99-0f888910f114 · inbound

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training cites this paper.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.802232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.802232Z digest=sha256:9483dcc35c7695a7e53ad3a53f82ba666b8de3a6cbf16e134211168a6a0e8cc6

Observation f4028448-fec2-4f51-9294-9d575a3c015a · inbound

Gradient Multi-Normalization for Stateless and Scalable LLM Training cites this paper.

Gradient Multi-Normalization for Stateless and Scalable LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.786570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.786570Z digest=sha256:c36bcf8f4e4a33bf71f65c93faf988f189908bab16517e7171d12b505079dc0d

Observation 938f875f-cad8-42ce-916b-ecc6ec5e8f30 · inbound

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension cites this paper.

Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T11:46:47.782800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:46:47.782800Z digest=sha256:8ba23916bc2604c1ce299273b9e5ebac11a4b90be848020adac4cb2489cfd60c

Observation c732579d-e615-4f62-8939-4f4a3b6dc7df · inbound

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design cites this paper.

Memory-Efficient LLM Pretraining via Minimalist Optimizer Design No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:24:53.247110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T13:23:55.233840Z digest=sha256:93f6f7d46ffaa62a9652b6dd1b8345051eb3fa63a92c64aa96c07c791e9ca31b

Observation 46d1497f-dc9c-4ad4-9a10-66c3666877b6 · inbound

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations cites this paper.

Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:51:20.166965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:16:58.221358Z digest=sha256:0a42b28d6f4703b875e160a760d62e4a6bc0eed9fe16336c046f08de44c96aa9

Observation fffd8c75-6370-485e-bd02-c6bfaf3be839 · inbound

Layerwise LQR for Geometry-Aware Optimization of Deep Networks cites this paper.

Layerwise LQR for Geometry-Aware Optimization of Deep Networks No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.813669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T16:59:22.216757Z digest=sha256:aeb6a582edee8d0f7bf3cae5382fe121304d3e7fffc9ec9758912bd8d07abc05

Observation 90d18fcd-5287-4f14-b605-6fccffadd7e3 · inbound

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio cites this paper.

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:41:05.035119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T15:39:51.611115Z digest=sha256:6885ad9c0d6c6f1c4fe2ccc9099e4b9d226cd012cae52f4903952ea04b218956

Observation 2da8cb7e-00cc-4a91-b77c-518ee26566e6 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 118

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:b17d5d00521563a5e645b9745ecc1118ac9637ca08477987377dab642298aaf9