Pith. sign in

Paper Citation Record · LEDGER

Taming Transformer Without Using Learning Rate Warmup

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2505.21910.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21910 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:12.948856Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 257220fa-7306-4f3c-b984-ee6914197076 · outbound

This paper cites Rezero is all you need: Fast convergence at large depth.

Taming Transformer Without Using Learning Rate Warmup Rezero is all you need: Fast convergence at large depth

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.871651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:05.383126Z digest=sha256:f81b6750436b7b1bb76977a4b6ffee67dcf12d325fae06019f51c17042583f89

Observation 627bff35-0fc5-4536-92e4-04b298512953 · outbound

This paper cites Language models are few-shot learners.

Taming Transformer Without Using Learning Rate Warmup Language models are few-shot learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.480251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.480251Z digest=sha256:7f9997fb5d23fa40b3cc755f348999d75366124a8ff4ba66354d898d2bfbeb59

Observation 17532fd4-7a2f-4153-97f8-e259ec0ba405 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Taming Transformer Without Using Learning Rate Warmup Palm: Scaling language modeling with pathways

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.525225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.525225Z digest=sha256:a5c9d31d55af1bd711570fbabacce8e24312439ef83b79e112434ac99ce019e9

Observation 42ade93d-5437-4261-bc70-9188d76b02f4 · outbound

This paper cites The Road Less Scheduled.

Taming Transformer Without Using Learning Rate Warmup The Road Less Scheduled

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.649809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.649809Z digest=sha256:241694786bb4b928c80c975b3746b4a3d0b383e834ecc589c3f63a086c08d6c4

Observation f28c3483-64e4-4317-a86e-c5d6bc0d941b · outbound

This paper cites Scaling vision transformers to 22 billion parameters.

Taming Transformer Without Using Learning Rate Warmup Scaling vision transformers to 22 billion parameters

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.715302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.715302Z digest=sha256:a115ab4e0881c44bd876b45f6c1b7e59de80a19c16ee9e07cc22b2c94b1e01b3

Observation e0321042-1743-4bb4-b921-325db581498a · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Taming Transformer Without Using Learning Rate Warmup Imagenet: A large-scale hierarchical image database

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.764111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.764111Z digest=sha256:9b16387427f04cce143501efe59ad879dd39101af0a72cec7470da5cda0722d4

Observation 1e99ca12-4aa5-42fd-855b-6fc2a6055561 · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Taming Transformer Without Using Learning Rate Warmup Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.593547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:05.838849Z digest=sha256:a66c3e7c2ed76a312fc702aaf73028bc65dbd9567330f049930a97d89854d0f2

Observation 7372e20c-c47f-4841-82bb-5da079e2573d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Taming Transformer Without Using Learning Rate Warmup An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.881498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.881498Z digest=sha256:224c58c600bfada5ad520869f644412d7812295a591b129e4dd19b12af347066

Observation 854b518f-d178-43ed-95c7-47b460d6c760 · outbound

This paper cites The Llama 3 Herd of Models.

Taming Transformer Without Using Learning Rate Warmup The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:05.929214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:05.929214Z digest=sha256:df09ce586feee8cd8e224b9587aeb235cae1455d4cc9310b043075d626d8000b

Observation 2d753c21-bfb2-4229-9cd0-22251710c090 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Taming Transformer Without Using Learning Rate Warmup Adaptive subgradient methods for online learning and stochastic optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.005636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.005636Z digest=sha256:430f53d662961d3769966f8226c0b49b9fb5ce439cca76893c56116955cece7e

Observation 08543823-8ff1-4e2a-8b78-f5b61eade4d0 · outbound

This paper cites Openwebtext corpus.

Taming Transformer Without Using Learning Rate Warmup Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.065481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.065481Z digest=sha256:49c528842cdce3a711fbce5014bcfd1daa0659dcb4fb8846b2c25f9f2e29d88a

Observation 4272df13-9591-4304-bd64-5ceac5594f86 · outbound

This paper cites Matrix computations.

Taming Transformer Without Using Learning Rate Warmup Matrix computations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.120967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.120967Z digest=sha256:2f440b8650fe87ea737528816980fcf8125308e9c1d7da29137b2662014720f4

Observation 0c820b89-57cd-4d03-bd41-95f95040d735 · outbound

This paper cites Kronecker products and matrix calculus with applications.

Taming Transformer Without Using Learning Rate Warmup Kronecker products and matrix calculus with applications

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.410854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.267817Z digest=sha256:600cbe1030282b66b6a5c905f563cd2f83c9e0b0d3335a5b826fc47082969f26

Observation 8f33d8bf-4ad0-4313-90a8-9e874efbb050 · outbound

This paper cites Flatten transformer: Vision transformer using focused linear attention.

Taming Transformer Without Using Learning Rate Warmup Flatten transformer: Vision transformer using focused linear attention

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:16.242762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.354068Z digest=sha256:2831fcb56a170ed8e23cb915ad10b6864d58be069a054f0601a3fc0a217f6dc3

Observation 0fe4f361-0ce1-4460-8fee-f92622469b58 · outbound

This paper cites Query-key normalization for transformers.

Taming Transformer Without Using Learning Rate Warmup Query-key normalization for transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.997633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.453396Z digest=sha256:235df53450721dc55218b9052a4867911bb6328e914257973d968216eadc658b

Observation 49aee7c9-c154-4248-94e0-cb2289f9f928 · outbound

This paper cites Topics in matrix analysis, 1991.

Taming Transformer Without Using Learning Rate Warmup Topics in matrix analysis, 1991

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.735652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.516079Z digest=sha256:df70eba45fd349bc4a2c051ac448f4bf6d7a130c957baf1c33f205fb06739c86

Observation ffb6ea69-f6dd-4d8a-a5f6-2dc73da0c39e · outbound

This paper cites Matrix analysis.

Taming Transformer Without Using Learning Rate Warmup Matrix analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.586362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.586362Z digest=sha256:2fbe66189e7798e37ff4ed1acc38176845b318541c8cf8e0eb0afc78388f7f14

Observation f477f785-e69f-4b32-b923-bdf9488594cd · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:26:15.518844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.722088Z digest=sha256:4ba7769b8923a9592dc362e2cc7f8ed5998c413ec2b51a23df23f5396842972a

Observation 169cfeff-2ab4-421b-a40d-5cbd51a47a97 · outbound

This paper cites The lipschitz constant of self-attention.

Taming Transformer Without Using Learning Rate Warmup The lipschitz constant of self-attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.345563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:06.837314Z digest=sha256:03117a3992582923b8167cb681a7541e475f67384c370227c0ac11cd710fe864

Observation 4d968143-41b1-4785-b4cf-c4f63278e786 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Taming Transformer Without Using Learning Rate Warmup Adam: A Method for Stochastic Optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:06.941052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:06.941052Z digest=sha256:f6924e1395708c4e47c6be52a0d48f3f91e17cb24e1cf0870694ad53ae22f12a

Observation cd387253-ef1e-484f-82d0-1ba700eba51c · outbound

This paper cites Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks.

Taming Transformer Without Using Learning Rate Warmup Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:07.038954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:07.038954Z digest=sha256:4facd84bfc9347116eec7cb865c2ca90d3ab4d0ea12924aaaa9c78ec86cf85cd

Observation 6944f4a5-f7d5-4d2e-bb92-c4b0edd843f4 · outbound

This paper cites Analyzing & reducing the need for learning rate warmup in gpt training.

Taming Transformer Without Using Learning Rate Warmup Analyzing & reducing the need for learning rate warmup in gpt training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:15.212279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:07.233908Z digest=sha256:9bd4628163149c55cf32d7f68d86c6912956f3d7d9a3446ef856d2acb1ace35e

Observation 0933c51d-06cf-4d25-ad53-f083e3f78eac · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

Taming Transformer Without Using Learning Rate Warmup Backpropagation applied to handwritten zip code recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.994693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:07.415329Z digest=sha256:da809fbbac95a613b73d15f9ebabb7b543b2594a1924e66b4f10404b4a58ddd6

Observation c57049ec-059f-4631-93b5-4436c488d358 · outbound

This paper cites Gradient-based learning applied to document recognition.

Taming Transformer Without Using Learning Rate Warmup Gradient-based learning applied to document recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:07.676301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:07.676301Z digest=sha256:11a87b308b43977c09e1fae69a32911e0fc5b5c9f98b19aed292bf9a0a4a6129

Observation fb3209bf-7f47-48f9-ab17-046a1f130195 · outbound

This paper cites Efficient backprop.

Taming Transformer Without Using Learning Rate Warmup Efficient backprop

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.840983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:08.240797Z digest=sha256:43cae90ca153cb018284950c428c22feb2e4b5dd2850bb26456be917cf783e42

Observation 79db7c08-6028-4904-8bae-c0774fd91aa3 · outbound

This paper cites Understanding the difficulty of training transformers.

Taming Transformer Without Using Learning Rate Warmup Understanding the difficulty of training transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.661451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:09.058294Z digest=sha256:710150192b794f05f504a0b6e3234edf10322a5cb321077ce5be47531df8ccbe

Observation f2732e95-40cc-46f9-8159-874e53d30c51 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Taming Transformer Without Using Learning Rate Warmup Swin transformer: Hierarchical vision transformer using shifted windows

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.453992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:09.982728Z digest=sha256:3eb51ccc36495830a47c4a6b8c5500958dfd3d45c1d116dcc895d93559c5d845

Observation cd256782-b797-43e7-9712-d932a9feac97 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Taming Transformer Without Using Learning Rate Warmup SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.126111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.126111Z digest=sha256:5099b1b0f658cb6d7624335a87e95e649bfebec3dab16c2bb9ebea9595c26876

Observation 5f59261d-ae14-444f-ad00-38a8c63336dc · outbound

This paper cites Fixing weight decay regularization in adam.

Taming Transformer Without Using Learning Rate Warmup Fixing weight decay regularization in adam

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.237757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.206724Z digest=sha256:82b87d73843f7b1262df80c0ae5fa3e0b6320e591e221c3b66aabd12df2e51c3

Observation dd65d404-4628-46fa-a0c1-358733015b3b · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.

Taming Transformer Without Using Learning Rate Warmup Signal propagation in transformers: Theoretical perspectives and the role of rank collapse

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:14.020113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.257907Z digest=sha256:1b3f120bf6496b783c4ef268ec98b87023f03f33fd4d6fd1e10a102dd2832377

Observation 598dda0a-3158-4bbf-b619-47ff234103bf · outbound

This paper cites Scalable diffusion models with transformers.

Taming Transformer Without Using Learning Rate Warmup Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.332842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.332842Z digest=sha256:3c4dcaec8a242be619727767df62c4feea2f8769f5e92ce0605058e20e8cbfd7

Observation f31e6bdf-8d8d-4933-90cb-1cea43146c53 · outbound

This paper cites The matrix cookbook.

Taming Transformer Without Using Learning Rate Warmup The matrix cookbook

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.441366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.441366Z digest=sha256:5b28ab49557392d3bb045cf88e9ac809579ef3e126db023e039c045fc6ceefa6

Observation d456acfd-fcd2-4553-857e-3f24fc2b9c91 · outbound

This paper cites Lipsformer: Introducing lipschitz continuity to vision transformers.

Taming Transformer Without Using Learning Rate Warmup Lipsformer: Introducing lipschitz continuity to vision transformers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.882027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:10.517077Z digest=sha256:e04c30606031dc02cbb18d24c97c7b4bb788aaa1848e5c858b3f260d1129ee13

Observation b97858a9-b3d1-466c-989f-5c2b8ab1aff8 · outbound

This paper cites Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant.

Taming Transformer Without Using Learning Rate Warmup Understanding Optimization of Deep Learning via Jacobian Matrix and Lipschitz Constant

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.591325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.591325Z digest=sha256:b6b51b56c3f080fdc3de91cdd102f6f6dc3c7ebd5d29e082555aa41978c03a6f

Observation fb65b2dd-5629-4c47-bf8a-3b19f731b856 · outbound

This paper cites Improving language understanding by generative pre-training.

Taming Transformer Without Using Learning Rate Warmup Improving language understanding by generative pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.643909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.643909Z digest=sha256:5ff41031065e2ae62f1a5d2e7073020afa88d8c875e85978b8b391550ccc1442

Observation fa51fde8-5b81-49ae-86f2-e89ce4479b27 · outbound

This paper cites Language models are unsupervised multitask learners.

Taming Transformer Without Using Learning Rate Warmup Language models are unsupervised multitask learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.768095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.768095Z digest=sha256:a798aad7e1b7e2f3575622bd3e2e284006a06e348475e59c238490d6eda117e3

Observation 023b964b-3eb8-430f-b40e-14da3b647713 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Taming Transformer Without Using Learning Rate Warmup Learning transferable visual models from natural language supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.824743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.824743Z digest=sha256:cc762dbdfc0544e9431fb2591fad1c08fa547d991ad145158339b74c9f83042f

Observation a7c0e253-242e-48df-aecf-9a40aa387c4f · outbound

This paper cites Zero-shot text-to-image generation.

Taming Transformer Without Using Learning Rate Warmup Zero-shot text-to-image generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:10.911598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:10.911598Z digest=sha256:e5218f681dabf802d34dfd82a37fece9b25f4446365f95852770f9c77e5ae74d

Observation 7f679c03-3b1a-42c7-850d-ed29c4038639 · outbound

This paper cites A stochastic approximation method.

Taming Transformer Without Using Learning Rate Warmup A stochastic approximation method

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.015079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.015079Z digest=sha256:19fb909aff5e2bba9319e9dcfc609e1509d3f31ac56e241eb6f2991d7fdbe476

Observation 1e3f62cb-71d6-4f6f-a084-f403cb1c1f0d · outbound

This paper cites Learning representations by back-propagating errors.

Taming Transformer Without Using Learning Rate Warmup Learning representations by back-propagating errors

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.102419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.102419Z digest=sha256:fb23aa9e921cc1b85ff2ee4ff29bf76e4ad4703447cf5f9a15ac6fd79e0eacba

Observation 16719b8a-8d27-4e6f-8747-3abe37ca8fdc · outbound

This paper cites Cyclical learning rates for training neural networks.

Taming Transformer Without Using Learning Rate Warmup Cyclical learning rates for training neural networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.656492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:11.173579Z digest=sha256:f6e53fa290b74c398946a654909a999b631d5c66c05f0b62c4e20a6ff7fb65f2

Observation 9c2748a7-0246-48b5-81c1-211a72517b3a · outbound

This paper cites Scan and snap: Understanding training dynamics and token composition in 1-layer transformer.

Taming Transformer Without Using Learning Rate Warmup Scan and snap: Understanding training dynamics and token composition in 1-layer transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.453488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:11.280585Z digest=sha256:870b9f8033736ae391530ec7858977413e95c2f60d54e85f54a005e192eb0a2b

Observation 795361a3-6a79-46e7-ab88-0175b600b518 · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

Taming Transformer Without Using Learning Rate Warmup JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.405603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.405603Z digest=sha256:aa6651c906bc9a51cb6b6da03e245f3fc042fcf3398a0fe089422db08cd0d412

Observation ffe7fa62-6a19-4c8a-bdf5-0d1005c938a4 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Taming Transformer Without Using Learning Rate Warmup Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.518646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.518646Z digest=sha256:118642a8c918b75b8fccfb5cdb12d08e7ef43ef3379331f379c3e1cdf5c34722

Observation 15ca0245-aad0-4545-a921-8d27e25009d2 · outbound

This paper cites Attention is all you need.

Taming Transformer Without Using Learning Rate Warmup Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.617287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.617287Z digest=sha256:be9215c8ef7cf04e5f773c5aa5f325ca4211d6f314c4ceafbe6142876e418356

Observation 61e6a18f-7c0a-40fe-9936-df9204792b23 · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Taming Transformer Without Using Learning Rate Warmup High-dimensional probability: An introduction with applications in data science, volume 47

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.746254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.746254Z digest=sha256:5f2702449530fbb630ac2c2d51765886d46310eede7c4ab86975ffa9e5391754

Observation c5041f34-d033-4787-b3bf-16da13f0a5b6 · outbound

This paper cites DeepNet: Scaling Transformers to 1,000 Layers.

Taming Transformer Without Using Learning Rate Warmup DeepNet: Scaling Transformers to 1,000 Layers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.857239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.857239Z digest=sha256:13afedaf9504cf7efc64572f5f6ca8da152c4b57a8bcd925cc02477f7af0b036

Observation 81f2774c-3042-4510-900d-333208208d21 · outbound

This paper cites Learning Deep Transformer Models for Machine Translation.

Taming Transformer Without Using Learning Rate Warmup Learning Deep Transformer Models for Machine Translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.909056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.909056Z digest=sha256:dec688707a1fc12cb1c765d5f3a301c407bdaeae1ddfa326ef55963424d1cbcd

Observation a3150a55-86c8-429c-8061-24f545beced3 · outbound

This paper cites Pytorch image models.

Taming Transformer Without Using Learning Rate Warmup Pytorch image models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.056554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.056554Z digest=sha256:84ec06b1a16cd5f7c85b4fea876f3dc6163a4b6431d8e69c273ab7a2ca51f014

Observation 46f3d844-ff54-4cfb-bdc5-1da34e717569 · outbound

This paper cites High-dimensional data analysis with low-dimensional models: Principles, computation, and applications.

Taming Transformer Without Using Learning Rate Warmup High-dimensional data analysis with low-dimensional models: Principles, computation, and applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.154560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.154560Z digest=sha256:fea52214b3620028e52b91d7463634ee6ecc809a583c93b984b33b41a056c0ec

Observation 7a1d6b40-b659-4b7d-858c-64b2334feedc · outbound

This paper cites On layer normalization in the transformer architecture.

Taming Transformer Without Using Learning Rate Warmup On layer normalization in the transformer architecture

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:26:13.239775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T13:26:12.285606Z digest=sha256:b93e7d3ad4e7a328baed862505a2fca833d51ec5684b61852c14c9e9a837c18a

Observation 10660810-ac57-42c8-835a-697d9671cbea · outbound

This paper cites Stabilizing transformer training by preventing attention entropy collapse.

Taming Transformer Without Using Learning Rate Warmup Stabilizing transformer training by preventing attention entropy collapse

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.409813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.409813Z digest=sha256:bd6b7ef4ac2a45eea9830d70cac37db7481c42de549792b0ad29990f143bcf8b

Observation cc1460b6-f112-4c9c-b2ed-ea07f901c0d4 · outbound

This paper cites Root mean square layer normalization.

Taming Transformer Without Using Learning Rate Warmup Root mean square layer normalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.544561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.544561Z digest=sha256:1db7d1d3c0cd400b6cf588189132ecdf13b14f57bed1436ae246ec1981fea818

Observation 42508aa3-13a1-450e-8e89-2b0aae900f3f · outbound

This paper cites write newline.

Taming Transformer Without Using Learning Rate Warmup write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.625867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.625867Z digest=sha256:b62c2cb2d248bc6a05fdbb9a19913668c715ac1b53b8271732d07a6c9a17e4df

Observation 0a02c986-71ee-4c54-8c00-217fe0d1b11c · outbound

This paper cites @esa (Ref.

Taming Transformer Without Using Learning Rate Warmup @esa (Ref

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.700048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.700048Z digest=sha256:958bd61b52ee6fe3efce9f97572cf9ca39d274c46c536752b9e6e264c40f5f0d

Observation 55fd245d-a651-4a37-866a-fd997f6cb4e7 · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.817304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.817304Z digest=sha256:deb218ae894adce2a920af63fc0ccbd198da99d7d6f77577d0c4a8359a60ca72

Observation 5978d486-dbf6-47f8-a283-ca2b4fba5aaf · outbound

This paper cites an unresolved cited work.

Taming Transformer Without Using Learning Rate Warmup Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:12.948856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:12.948856Z digest=sha256:367f15ff2d8b1cd2fa7a0e2f477979c169857a7f4dec01963b2294a38268022b

Pith citing papers

No inbound Pith citation observations are available.