Pith. sign in

Paper Citation Record · LEDGER

Taming Transformer Without Using Learning Rate Warmup

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.21910.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21910 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:12.948856Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 257220fa-7306-4f3c-b984-ee6914197076 · outbound

This paper cites Rezero is all you need: Fast convergence at large depth.

Taming Transformer Without Using Learning Rate Warmup Rezero is all you need: Fast convergence at large depth

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.871651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:05.383126Z digest=sha256:4d0f2e6d026622fad00561ad6b695833c081a384b62e948530fab29c62cdc3c7

Observation 627bff35-0fc5-4536-92e4-04b298512953 · outbound

This paper cites Language models are few-shot learners.

Taming Transformer Without Using Learning Rate Warmup Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.480251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.480251Z digest=sha256:9097416ed15938c81f67d9c38b20cc1bb2b35e065efd5b36afe201e936c0910a

Observation 17532fd4-7a2f-4153-97f8-e259ec0ba405 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Taming Transformer Without Using Learning Rate Warmup Palm: Scaling language modeling with pathways

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.525225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.525225Z digest=sha256:56eea45ffb6b418ec4d7ac4dedc20b37eefdd1b1c7802a9466ec81d7ecf4d404

Observation 42ade93d-5437-4261-bc70-9188d76b02f4 · outbound

This paper cites The Road Less Scheduled.

Taming Transformer Without Using Learning Rate Warmup The Road Less Scheduled

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.649809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.649809Z digest=sha256:0dda0110adf506f1fa2b064c0b819077b4a93a823597784076df4607d80139c1

Observation f28c3483-64e4-4317-a86e-c5d6bc0d941b · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Taming Transformer Without Using Learning Rate Warmup Scaling vision transformers to 22 billion parameters

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.715302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.715302Z digest=sha256:19e1362abbf32cb568ad1206332fe039c6074d9bdda254a08dd707f7f5e96f15

Observation e0321042-1743-4bb4-b921-325db581498a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Taming Transformer Without Using Learning Rate Warmup Imagenet: A large-scale hierarchical image database

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.764111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.764111Z digest=sha256:232a6b0c9b0c3c94e128683ae21129465b1889096d7110d9ff07fb08f6caf60e

Observation 1e99ca12-4aa5-42fd-855b-6fc2a6055561 · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Taming Transformer Without Using Learning Rate Warmup Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.593547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:05.838849Z digest=sha256:544e007350bce9ac2b8727d902b9307b544ec8b1412dbae0859c6ead84a1b8c5

Observation 7372e20c-c47f-4841-82bb-5da079e2573d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Taming Transformer Without Using Learning Rate Warmup An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.881498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.881498Z digest=sha256:ccbfa66c447e05f036d2f90fb92fc906cfda9524fb8f5801adc43e7ab69e0e55

Observation 854b518f-d178-43ed-95c7-47b460d6c760 · outbound

This paper cites The Llama 3 Herd of Models.

Taming Transformer Without Using Learning Rate Warmup The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.929214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.929214Z digest=sha256:6b44a2ff943b7aefeaf396d70f1fc5f86f9ca55297e495a36a2b3ff5a47fbaf6

Observation 2d753c21-bfb2-4229-9cd0-22251710c090 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Taming Transformer Without Using Learning Rate Warmup Adaptive subgradient methods for online learning and stochastic optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.005636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.005636Z digest=sha256:e5a1b51296796176ebb0f3955b639612462219d4c3a486723b72d04162b62636

Observation 08543823-8ff1-4e2a-8b78-f5b61eade4d0 · outbound

This paper cites Openwebtext corpus.

Taming Transformer Without Using Learning Rate Warmup Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.065481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.065481Z digest=sha256:1e15341a2177f74ed0255a74d160de881ce263b1df8a2a55fd34ec0c4abe1bb6

Observation 4272df13-9591-4304-bd64-5ceac5594f86 · outbound

This paper cites Matrix computations.

Taming Transformer Without Using Learning Rate Warmup Matrix computations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.120967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.120967Z digest=sha256:94b6103f9f820e7ec83849119871c20a9fe1f6b54ba354aade6c7f35ceb51273

Observation 0c820b89-57cd-4d03-bd41-95f95040d735 · outbound

This paper cites Kronecker products and matrix calculus with applications.

Taming Transformer Without Using Learning Rate Warmup Kronecker products and matrix calculus with applications

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.410854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.267817Z digest=sha256:629e71a354653739a8f7da17a196ddbdea742374fe211d66b1cc0e9c447cd1ca

Observation 8f33d8bf-4ad0-4313-90a8-9e874efbb050 · outbound

This paper cites Flatten transformer: Vision transformer using focused linear attention.

Taming Transformer Without Using Learning Rate Warmup Flatten transformer: Vision transformer using focused linear attention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.242762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.354068Z digest=sha256:f190dc233d1737d2ecb4f829b5f42ea17a47f9057a84019dce0c79d6dbbc6298

Observation 0fe4f361-0ce1-4460-8fee-f92622469b58 · outbound

This paper cites Query-key normalization for transformers.

Taming Transformer Without Using Learning Rate Warmup Query-key normalization for transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.997633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.453396Z digest=sha256:8cbbaa62ff583cc039956f118edc673a6585496d430aa6d2f6498ae4e2686dca

Observation 49aee7c9-c154-4248-94e0-cb2289f9f928 · outbound

This paper cites Topics in matrix analysis, 1991.

Taming Transformer Without Using Learning Rate Warmup Topics in matrix analysis, 1991

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.735652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.516079Z digest=sha256:3fef9baa3ddb05751c5fe0af99c49c4a0fd98016d98deaffd621f578342e0eb1

Observation ffb6ea69-f6dd-4d8a-a5f6-2dc73da0c39e · outbound

This paper cites Matrix analysis.

Taming Transformer Without Using Learning Rate Warmup Matrix analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.586362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.586362Z digest=sha256:4c76e99f84de3c3112e0af9e8998cb5dd75fbf5984179f98e0dc675f490b0780

Observation f477f785-e69f-4b32-b923-bdf9488594cd · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:26:15.518844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.722088Z digest=sha256:4c10051062950a5de9c7a136ac10fd82612771ad7ca7e478486b092976ce25ef

Observation 169cfeff-2ab4-421b-a40d-5cbd51a47a97 · outbound

This paper cites The lipschitz constant of self-attention.

Taming Transformer Without Using Learning Rate Warmup The lipschitz constant of self-attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.345563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.837314Z digest=sha256:468b21994eb667520d797eb37a14d1bc770245d9135f0235ec802a7fbf8aa39c

Observation 4d968143-41b1-4785-b4cf-c4f63278e786 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Taming Transformer Without Using Learning Rate Warmup Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.941052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.941052Z digest=sha256:c93347d21fa20dba6efbbaf420609aa163085eeee285b53c7b4a89d5e0455a76

Observation cd387253-ef1e-484f-82d0-1ba700eba51c · outbound

This paper cites Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.

Taming Transformer Without Using Learning Rate Warmup Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:07.038954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:07.038954Z digest=sha256:425e66b36bc0b6588e0c5077e38fd3e5510562ec2436d799625ea636b68a5b57

Observation 6944f4a5-f7d5-4d2e-bb92-c4b0edd843f4 · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Taming Transformer Without Using Learning Rate Warmup Analyzing & reducing the need for learning rate warmup in gpt training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.212279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:07.233908Z digest=sha256:84d65c1a167071094a6621c5d8aac976218f673f5f5ccfa69d77b7978604927a

Observation 0933c51d-06cf-4d25-ad53-f083e3f78eac · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

Taming Transformer Without Using Learning Rate Warmup Backpropagation applied to handwritten zip code recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.994693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:07.415329Z digest=sha256:274142587c33f360c43807d149773189829ce8b4ced1ed16359fdc4bd1b9a0ec

Observation c57049ec-059f-4631-93b5-4436c488d358 · outbound

This paper cites Gradient-based learning applied to document recognition.

Taming Transformer Without Using Learning Rate Warmup Gradient-based learning applied to document recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:07.676301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:07.676301Z digest=sha256:fbe65fe973f001792f84c7ae733eb451195f98056159a53da1fab2b836be78ee

Observation fb3209bf-7f47-48f9-ab17-046a1f130195 · outbound

This paper cites Efficient backprop.

Taming Transformer Without Using Learning Rate Warmup Efficient backprop

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.840983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:08.240797Z digest=sha256:880f4a90045aa30a0fe76eed8da978b09db331ed22e30db651cb944b61b3d2f1

Observation 79db7c08-6028-4904-8bae-c0774fd91aa3 · outbound

This paper cites Understanding the difficulty of training transformers.

Taming Transformer Without Using Learning Rate Warmup Understanding the difficulty of training transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.661451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:09.058294Z digest=sha256:ff6c3663b1353c224343d6a97e66d11402e2b6d45c3d4f301ff7fc0d7e56fb0b

Observation f2732e95-40cc-46f9-8159-874e53d30c51 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Taming Transformer Without Using Learning Rate Warmup Swin transformer: Hierarchical vision transformer using shifted windows

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.453992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:09.982728Z digest=sha256:dc9c4dbc2d5e012c3089f5432c96a563e092e55747647c5a8705fad527d67338

Observation cd256782-b797-43e7-9712-d932a9feac97 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Taming Transformer Without Using Learning Rate Warmup SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.126111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.126111Z digest=sha256:d9b3b3b0f9a2781006bbfb2c017403fdecd3ba3f493e5824c1d466cf23f332b5

Observation 5f59261d-ae14-444f-ad00-38a8c63336dc · outbound

This paper cites Fixing weight decay regularization in adam.

Taming Transformer Without Using Learning Rate Warmup Fixing weight decay regularization in adam

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.237757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.206724Z digest=sha256:159facb02b2b2e42cce75b19d3ca26643bcd012082ee741973aadf7d655d1a45

Observation dd65d404-4628-46fa-a0c1-358733015b3b · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.

Taming Transformer Without Using Learning Rate Warmup Signal propagation in transformers: Theoretical perspectives and the role of rank collapse

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.020113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.257907Z digest=sha256:bd2c42fc1ec4bc86e505e174e876a492656866e837c4276ee31618c708dc14c4

Observation 598dda0a-3158-4bbf-b619-47ff234103bf · outbound

This paper cites Scalable diffusion models with transformers.

Taming Transformer Without Using Learning Rate Warmup Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.332842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.332842Z digest=sha256:36df4a0b9d6ac2e80801a2d07b4e6fba27fd924ea3120c3956a048e8e68059f6

Observation f31e6bdf-8d8d-4933-90cb-1cea43146c53 · outbound

This paper cites The matrix cookbook.

Taming Transformer Without Using Learning Rate Warmup The matrix cookbook

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.441366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.441366Z digest=sha256:5d92b77897533efe8d555e3bb1f2dc0a3a52e32abdc0131ae820c162374e69a7

Observation d456acfd-fcd2-4553-857e-3f24fc2b9c91 · outbound

This paper cites Lipsformer: Introducing lipschitz continuity to vision transformers.

Taming Transformer Without Using Learning Rate Warmup Lipsformer: Introducing lipschitz continuity to vision transformers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.882027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.517077Z digest=sha256:c831e17b5e3700046cbe1407a208cc7a396fd0217c36b4d7926e87d9cc5e8207

Observation b97858a9-b3d1-466c-989f-5c2b8ab1aff8 · outbound

This paper cites Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant.

Taming Transformer Without Using Learning Rate Warmup Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.591325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.591325Z digest=sha256:73292e42b8896a909cf6b9d1d51549b08d4a047634954a715b88b00dc9016c34

Observation fb65b2dd-5629-4c47-bf8a-3b19f731b856 · outbound

This paper cites Improving language understanding by generative pre-training.

Taming Transformer Without Using Learning Rate Warmup Improving language understanding by generative pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.643909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.643909Z digest=sha256:0c02cf195714473ff65715952b595c2e76cf6f8738fdeb12c06b54d979b44841

Observation fa51fde8-5b81-49ae-86f2-e89ce4479b27 · outbound

This paper cites Language models are unsupervised multitask learners.

Taming Transformer Without Using Learning Rate Warmup Language models are unsupervised multitask learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.768095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.768095Z digest=sha256:a03e9366933fbe02897bc988220029ba64643aae5908d5ed3ef6e6fc397cfb69

Observation 023b964b-3eb8-430f-b40e-14da3b647713 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Taming Transformer Without Using Learning Rate Warmup Learning transferable visual models from natural language supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.824743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.824743Z digest=sha256:6f5d6d181d02fb15cab45eed45656f53fac8a7a9a245656978f280ee9dd1a666

Observation a7c0e253-242e-48df-aecf-9a40aa387c4f · outbound

This paper cites Zero-shot text-to-image generation.

Taming Transformer Without Using Learning Rate Warmup Zero-shot text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.911598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.911598Z digest=sha256:ef48be501fe344f1cb6255a087976695f5b43e3fe7c8e53e0df39582a40c8bfd

Observation 7f679c03-3b1a-42c7-850d-ed29c4038639 · outbound

This paper cites A stochastic approximation method.

Taming Transformer Without Using Learning Rate Warmup A stochastic approximation method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.015079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.015079Z digest=sha256:9d557263bde05027557a98fe555730e5dd70bdd76d11058ceca3e0134cd3bd3d

Observation 1e3f62cb-71d6-4f6f-a084-f403cb1c1f0d · outbound

This paper cites Learning representations by back-propagating errors.

Taming Transformer Without Using Learning Rate Warmup Learning representations by back-propagating errors

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.102419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.102419Z digest=sha256:a301053a903f8d9ced87f30c8b4a34d618a67d661dd777b52283f8cf1b5da951

Observation 16719b8a-8d27-4e6f-8747-3abe37ca8fdc · outbound

This paper cites Cyclical learning rates for training neural networks.

Taming Transformer Without Using Learning Rate Warmup Cyclical learning rates for training neural networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.656492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:11.173579Z digest=sha256:d0e253cb48dd399b2647a5306ab4f70b9f68cf58bb3c3cb8ac99fbb18c1dc00f

Observation 9c2748a7-0246-48b5-81c1-211a72517b3a · outbound

This paper cites Scan and snap: Understanding training dynamics and token composition in 1-layer transformer.

Taming Transformer Without Using Learning Rate Warmup Scan and snap: Understanding training dynamics and token composition in 1-layer transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.453488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:11.280585Z digest=sha256:4b96dd35cacebaf0f66f5470cf798844d37a72634e150748a4fdfa4d29e62e1d

Observation 795361a3-6a79-46e7-ab88-0175b600b518 · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

Taming Transformer Without Using Learning Rate Warmup JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.405603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.405603Z digest=sha256:cbde3ec1d6447ea3e307b39fb1636215e6cf2b823f54ef7765f4c9f55b162a09

Observation ffe7fa62-6a19-4c8a-bdf5-0d1005c938a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Taming Transformer Without Using Learning Rate Warmup Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.518646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.518646Z digest=sha256:82ecb0429269749e22525c1f33952d8d923e7ebd7e1a65eb1f160f181be6f646

Observation 15ca0245-aad0-4545-a921-8d27e25009d2 · outbound

This paper cites Attention is all you need.

Taming Transformer Without Using Learning Rate Warmup Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.617287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.617287Z digest=sha256:db31ed76b684d46a45dd49683abafe10113c84ec057125431329d97342dec094

Observation 61e6a18f-7c0a-40fe-9936-df9204792b23 · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Taming Transformer Without Using Learning Rate Warmup High-dimensional probability: An introduction with applications in data science, volume 47

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.746254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.746254Z digest=sha256:af7f588c69d6bc0d84f760e665f0c6f40ccf5d1daba7269feff053f6188659ed

Observation c5041f34-d033-4787-b3bf-16da13f0a5b6 · outbound

This paper cites DeepNet: Scaling Transformers to 1,000 Layers.

Taming Transformer Without Using Learning Rate Warmup DeepNet: Scaling Transformers to 1,000 Layers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.857239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.857239Z digest=sha256:b4878df9f83da2b3fd66ff34cffcb09fc1b3682dfe4a4950e8d6357c76e6ce8b

Observation 81f2774c-3042-4510-900d-333208208d21 · outbound

This paper cites Learning Deep Transformer Models for Machine Translation.

Taming Transformer Without Using Learning Rate Warmup Learning Deep Transformer Models for Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.909056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.909056Z digest=sha256:7f30e47592fb5922f3f7a361dc30be74cc6d533aa929d991f5f64c26c3380d85

Observation a3150a55-86c8-429c-8061-24f545beced3 · outbound

This paper cites Pytorch image models.

Taming Transformer Without Using Learning Rate Warmup Pytorch image models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.056554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.056554Z digest=sha256:5e87c0eb39ccee3d46170e951f0cdfd274630dae73a95151c8765824c2111c5c

Observation 46f3d844-ff54-4cfb-bdc5-1da34e717569 · outbound

This paper cites High-dimensional data analysis with low-dimensional models: Principles, computation, and applications.

Taming Transformer Without Using Learning Rate Warmup High-dimensional data analysis with low-dimensional models: Principles, computation, and applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.154560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.154560Z digest=sha256:62e27491b39d47b7f7fcb6c73808865e86fa50644ad5f6a0a9658518bd9831f9

Observation 7a1d6b40-b659-4b7d-858c-64b2334feedc · outbound

This paper cites On layer normalization in the transformer architecture.

Taming Transformer Without Using Learning Rate Warmup On layer normalization in the transformer architecture

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.239775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:26:12.285606Z digest=sha256:bf9f9f5c7d1b1ca04cd0c63db607376f1b795265f80135356df43ed5feaf226a

Observation 10660810-ac57-42c8-835a-697d9671cbea · outbound

This paper cites Stabilizing transformer training by preventing attention entropy collapse.

Taming Transformer Without Using Learning Rate Warmup Stabilizing transformer training by preventing attention entropy collapse

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.409813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.409813Z digest=sha256:dfe08a312f2dbaedf5a75cd2df1850c5d17247db1e2f64449bf2db0bb7c9ecf2

Observation cc1460b6-f112-4c9c-b2ed-ea07f901c0d4 · outbound

This paper cites Root mean square layer normalization.

Taming Transformer Without Using Learning Rate Warmup Root mean square layer normalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.544561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.544561Z digest=sha256:005685498ff4f02fe1b225ece644f952e7b398e2fc2982cbd41f370de6ded305

Observation 42508aa3-13a1-450e-8e89-2b0aae900f3f · outbound

This paper cites write newline.

Taming Transformer Without Using Learning Rate Warmup write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.625867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.625867Z digest=sha256:a8797f26f60fb4529efbd992542ddaa61ba8ebdf7f23a3e988f47f41681dce6f

Observation 0a02c986-71ee-4c54-8c00-217fe0d1b11c · outbound

This paper cites @esa (Ref.

Taming Transformer Without Using Learning Rate Warmup @esa (Ref

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.700048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.700048Z digest=sha256:420cea9e9553f89c9f7c38e0d486f1dacca15ac420c01d949603e6fd8d29aaef

Observation 55fd245d-a651-4a37-866a-fd997f6cb4e7 · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.817304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.817304Z digest=sha256:e43fd9820b152a5181e7158b190d2284a761197812c253c01ca124b876a86eea

Observation 5978d486-dbf6-47f8-a283-ca2b4fba5aaf · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.948856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.948856Z digest=sha256:ec9b1168192f8437f31188d04121a5dbaa0a051a42f6138cf4e42996da2b10cc

Pith citing papers

No inbound Pith citation observations are available.