Pith. sign in

Paper Citation Record · LEDGER

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 10 inbound Pith citation observations for arXiv:2502.02732.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02732 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:23:41.055567Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:10.661789Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:30:07.740339Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f434b86a-30db-4c69-9e5d-f50cb7796c1b · outbound

This paper cites write newline.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.851628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.851628Z digest=sha256:f8672b1d69baafc5528a627c9c2e9751bbb90348827b564a4630f35f223a00a1

Observation 33036fcf-68a6-40d5-9d13-e66ae0bcde08 · outbound

This paper cites Layer Normalization.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.877867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.877867Z digest=sha256:af6e5cb02b12bd365383d5744485b2c1373e57668602df3004d24fe176ab409d

Observation 489b6201-372d-49a7-abd8-e83b6d28aec4 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Piqa: Reasoning about physical commonsense in natural language

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.892330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.892330Z digest=sha256:3ada5e7093b151b7a80133a52243b1507c5a5d5a0918435ddb92c3d75d386bb4

Observation 6fa722ab-77cf-46aa-9809-f9ffe06d2493 · outbound

This paper cites L., and Simonyan, K.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture L., and Simonyan, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.573228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.895994Z digest=sha256:7be5b19dd06c1ee4918f3e53df94beb79804aea6c7c988223fa60471ac32dc61

Observation ae05f490-bfe2-4729-8b53-be2f80d62ad1 · outbound

This paper cites L., and Simonyan, K.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture L., and Simonyan, K

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.562745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.899516Z digest=sha256:a991c862bf8eafeb308469617e7c3084519de2eeb324420e0870e6d69dec64a5

Observation 1b316606-3123-4869-a6db-25d10e720cde · outbound

This paper cites M., Thorne, J., and Yun, S.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture M., Thorne, J., and Yun, S

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.550491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.902568Z digest=sha256:524992652b1081770cd840fd3bfb1e97b1994d538d82a903054598bd15f2bbec

Observation a70a568f-fcf9-44cd-abf6-3f6c792b270e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.905850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.905850Z digest=sha256:66654865d70731deafeb14b42280cda092a53b57feecb23ca3faef5fcdb7066f

Observation 1a540c13-cf09-4887-992e-30b8a00c4bbb · outbound

This paper cites MoEUT: Mixture-of-Experts Universal Transformers.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture MoEUT: Mixture-of-Experts Universal Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.909701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.909701Z digest=sha256:fddc844adb16e2eb4eed148b69ad2cc9082a49c86cbd0ec34d5fae9f1054a26b

Observation daa6c897-9273-45f5-a1e6-dad8f48b6d34 · outbound

This paper cites and Smith, S.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture and Smith, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.540017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.914052Z digest=sha256:b4033ab403438ab2abff094a001aede8eb06e7d60281655874182f135ed2dae0

Observation 8229a91f-b51e-4e77-bebb-c2f276e0c3d5 · outbound

This paper cites P., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Ruiz, C.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture P., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Ruiz, C

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.528861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.918117Z digest=sha256:5159ad344486f5b70b1f6d218109fbfd739482e37e3b58e54f00eb82121d1793

Observation 4ff7c646-81ee-4f4f-a085-e65cd0e673bd · outbound

This paper cites Gpt3.int8(): 8-bit matrix multiplication for transformers at scale.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Gpt3.int8(): 8-bit matrix multiplication for transformers at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.518073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.922084Z digest=sha256:3484b1bf7e473145aa52bd9331ba382a3ea5bdf4aca2f2d4603b60c0b9d55bdb

Observation 90beec48-6507-4b43-85cb-26e22324facb · outbound

This paper cites The Llama 3 Herd of Models.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.925710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.925710Z digest=sha256:ac73c1a463b67d764f8f7149b917cb151dae3d6fbb53670c61715a844aecd361

Observation f804b551-985d-46b4-9dc7-9f50336ff2f0 · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Scaling FP8 training to trillion-token LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.929720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.929720Z digest=sha256:af49dd92c79e776a3dcc942ff66b403780695b78fb035def00c96f556bc3b95f

Observation 682aae9f-fafc-44c9-9a0d-e5c5c0bda128 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture A framework for few-shot language model evaluation, 07 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.933794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.933794Z digest=sha256:d52fb748c946e322ad7e631068d448dff098d6f32d35cb1446bf7d53b4cf4ddf

Observation d09e8149-6d69-44e0-bad1-c5ce28e1ffdb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.937603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.937603Z digest=sha256:15167c94a0998409b0779e7045a4fa5b30d53f2236be9079619320a64162f0a2

Observation 41f1e2a8-b304-4d18-bc08-be09ac3352ef · outbound

This paper cites Identity mappings in deep residual networks.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Identity mappings in deep residual networks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.941936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.941936Z digest=sha256:f6e76467f5e76860335c48e85175d296db8781365e7dde204a0b4bfb306ee85a

Observation 3797cbe1-dd16-44b3-ad2e-ab9d53d4116e · outbound

This paper cites Training Compute-Optimal Large Language Models.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Training Compute-Optimal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.945771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.945771Z digest=sha256:fe2aa947950d9392b3cf4ff34c78738ecddb8618b60c3b3f8e63ce14bebef366

Observation f1bb609d-33cc-42c1-8579-99c6a4319361 · outbound

This paper cites A., Khyalia, S., Jung, J., Goka, H., and Lee, H.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture A., Khyalia, S., Jung, J., Goka, H., and Lee, H

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.506469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.949975Z digest=sha256:a082c12496dd1cc0a796e1c600c4a579eb4905485632ae614e12f3054482a95d

Observation 7c0e1cdc-9a3b-4385-864c-ac3734dfee30 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture DataComp-LM: In search of the next generation of training sets for language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.953362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.953362Z digest=sha256:327fd780264f95709a860ae803f35a152b9ded28cebd09eb906abd56babf5480

Observation a31f75b4-e492-4d28-a9f1-391160482c09 · outbound

This paper cites Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.957150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.957150Z digest=sha256:098454e235d7558a11eecfc62b2a309a7af882711e3b6e96a0462a1e91fea1a7

Observation 387a7cb2-6a81-4870-b072-6f18762890be · outbound

This paper cites nGPT: Normalized Transformer with Representation Learning on the Hypersphere.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture nGPT: Normalized Transformer with Representation Learning on the Hypersphere

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.961082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.961082Z digest=sha256:8c545f6132a711dc03fac8659aa014623887b8e47e599b057716a515d3a27d2e

Observation c45c6c18-8285-42b1-a488-f82ce8b71647 · outbound

This paper cites Pointer sentinel mixture models, 2016.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Pointer sentinel mixture models, 2016

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.965013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.965013Z digest=sha256:070b3a15c45f5dc7434d1b463a12fe15086fa4ce19c1cd58b99b519fd7cc48a6

Observation 1d8a9af5-cecf-4a19-90e4-fe0ef20036a3 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.968440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.968440Z digest=sha256:1f03d4cb952d6c5c1f0cc6b208721dad108b79e332acefab04b291b2a7ba39c5

Observation 18b61c7c-d3ba-4f10-9f51-c300bec588ee · outbound

This paper cites 2 OLMo 2 Furious.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture 2 OLMo 2 Furious

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.972128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.972128Z digest=sha256:ba8210a989e6fff69382b30457ea49f7da4a9c7b9064c2c680fd3eb11564a43a

Observation 20b3394c-4b23-4505-8dbc-1851da4d4e73 · outbound

This paper cites L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.483607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:40.976010Z digest=sha256:249070b50e158ed326718f4273f12bb0d75dc1f81ff154608eefdf37c8c29e61

Observation 85ec0cc4-3662-442c-97cf-fd1d460adbfd · outbound

This paper cites an unresolved cited work.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.979249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.979249Z digest=sha256:69859170afdd33e8b2c0bcd3e4abb90fb1dddcba37f47d030f58ded8fff5322e

Observation 8b1e79a5-d6b3-4751-9ea3-df53f9195c07 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Gemma 2: Improving Open Language Models at a Practical Size

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.982032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.982032Z digest=sha256:ec84c030c18ad8333c0265b21ebadd351ebe99f9e2f1b0c08867e6a42a2f7272

Observation cf1532cd-400e-44f5-a65b-adc92efedd93 · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture L., Bhagavatula, C., and Choi, Y

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.985508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.985508Z digest=sha256:a06dfc9755666bb0115325f2974b531567ef1d91e021b7ac0981a90cc5ca5c09

Observation 5abc72c6-91bb-45f4-81cb-d482a5edef7b · outbound

This paper cites L., and Choi, Y.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture L., and Choi, Y

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.989356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.989356Z digest=sha256:63611adefb58f868e1bc7038903c9d9e5f4053bbb272215cf0c0741177649a85

Observation 94b3edb7-d041-4212-bc15-c788e160fcef · outbound

This paper cites Massive Activations in Large Language Models.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Massive Activations in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.993209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.993209Z digest=sha256:c9d8747a0df54347f9d2944e613b9b0ad606372285bbf319aec3a9ed0538932b

Observation 8a60c67d-ea8b-4711-8474-e9476ac3d168 · outbound

This paper cites Spike No More: Stabilizing the Pre-training of Large Language Models.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.997471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.997471Z digest=sha256:97c4b610564bf9d5c8943cf441d84c3ed001065c7801cbdd9095ed551f933a56

Observation d3a15d2e-cd47-41d2-9885-18a4cf15b8f9 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.001809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.001809Z digest=sha256:e11af142b4fb06dbd71ac2a3d5bd7fa37d91c8753c4b068e7481840d62842534

Observation 1a0ffefa-e06a-4d59-bcc2-249464d75863 · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture N., Kaiser, L., and Polosukhin, I

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.005504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.005504Z digest=sha256:2ff162d4c47b505c842e896e8098cc3aec85f4326a27fa59d805405c5f519563

Observation fb1782f0-9542-4ac3-97ac-ff1c35bec19b · outbound

This paper cites GLUE : A multi-task benchmark and analysis platform for natural language understanding.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture GLUE : A multi-task benchmark and analysis platform for natural language understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.009462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.009462Z digest=sha256:1d07bca2661e713c70ff699f32b464e68ab83418c5d6f2f981af3d449e15ea37

Observation 4f061457-a891-4704-a02e-6304a0928a86 · outbound

This paper cites L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.013426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.013426Z digest=sha256:a3dc771c426c1a7d324db8ebcf36828c45d80d57bb0a36ceb5c63dc2661cbae8

Observation bdf134be-3901-4f65-9e91-2c71f1564e43 · outbound

This paper cites J., Xiao, L., Everett, K.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture J., Xiao, L., Everett, K

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.454954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.017054Z digest=sha256:48760db21146c58852e839afd74110f4695535ec3a7aec925d4a365d87bd9159

Observation 235e7f01-f60e-464a-9ec5-ef9bdb045830 · outbound

This paper cites On layer normalization in the transformer architecture.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture On layer normalization in the transformer architecture

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.444567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.020942Z digest=sha256:d106ffb3c940a6ec72839db2493a1c7101480a4181f82a78d7eabfe3875ed98c

Observation 23483fe6-e21b-45e7-8bee-6191beb2f6ce · outbound

This paper cites and Hu, E.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture and Hu, E

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.433266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.024962Z digest=sha256:6bd99a03a7373e92e9bb0a386044fba4d8821f224ed5474d5a92d436f389cff2

Observation f5e1b34c-27a7-47fb-828b-3d361a3956a8 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.028667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.028667Z digest=sha256:e4a5fce036b5e3318de1c4761fd673f5849590b89f450cb340fb0411496a04bb

Observation c2c44df5-cfa7-4c95-8f64-380ab4e7127c · outbound

This paper cites Tensor programs VI: feature learning in infinite depth neural networks.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Tensor programs VI: feature learning in infinite depth neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.422449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.032836Z digest=sha256:4c4eb91a73b3c3e44b08a2031f73ca9c40eacae212e8116f5b5c85411533c3a2

Observation 72684422-acc7-4d9c-b5e4-255bd4e51eb3 · outbound

This paper cites The Super Weight in Large Language Models.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture The Super Weight in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.036654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.036654Z digest=sha256:7b26ab5de965a2d471073b3602415fc663a4a8433be0e8a4550e3f5464b0eec9

Observation daa9108c-0e48-4ae2-8efb-73377e0be55d · outbound

This paper cites H ella S wag: Can a machine really finish your sentence? In Korhonen, A., Traum, D., and M \`a rquez, L.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture H ella S wag: Can a machine really finish your sentence? In Korhonen, A., Traum, D., and M \`a rquez, L

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.043808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.043808Z digest=sha256:24787fee9eacdccac2e114aaed399eb65cc8856c147c008b5f7f739f465f99ec

Observation 7ae72bcf-7902-42ec-973e-de252bd75d5f · outbound

This paper cites an unresolved cited work.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:23:41.411931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.047603Z digest=sha256:fca084f24dc66cd8e27ec329a4958a59abfc938fe0177a4cddecad12c8b588c6

Observation 15357f38-1a24-48bf-a6b7-06ac4b227d9a · outbound

This paper cites and Sennrich, R.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture and Sennrich, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.400476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.052010Z digest=sha256:0e7c57a9c6d4a3ff254265617e17e9ccd1195c838595ba6f78ed5f14d6f4d6e3

Observation 71f6c8eb-862d-46d5-83e1-99b8114cfe6b · outbound

This paper cites LIMA: less is more for alignment.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture LIMA: less is more for alignment

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:23:41.389771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:23:41.055567Z digest=sha256:0e8895e9bc493a1195e754e37446b4a4d113ee6076b2132e16e93cf4d30d60f9

Pith citing papers

Observation 5ed40db9-0ac8-49d2-9536-daa3fb9bbafe · inbound

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling cites this paper.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.661789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.661789Z digest=sha256:386ba0352d5fc388854c0519fd3d023230e70baca458949134c226c9d3520385

Observation 06bcf884-d9c7-4e6d-947d-a84f125f271c · inbound

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers cites this paper.

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:37:52.586372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:37:52.586372Z digest=sha256:c011f258f5a2b4c7767284f80d7f3727d68042723ad9e6fb835ceefe7cb0145e

Observation b04c715f-4d2d-4480-a55f-e04f410edc69 · inbound

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication cites this paper.

LoRDO: Distributed Low-Rank Optimization with Infrequent Communication Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:44:55.564583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:44:55.564583Z digest=sha256:84af35d8587a7c34739f65c4998e49523fb5f7f07f2cd62913d78b62b6761ff0

Observation a43618bd-4db6-4598-b4fb-d7af19f0727e · inbound

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm cites this paper.

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:14:47.951414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T11:11:51.440058Z digest=sha256:87249e166d77b123d53b60cdec575608e4dc947045394a82e481f99d9eb44d36

Observation 78dcf9be-8b7a-4367-933a-18d8627f13ed · inbound

Stability and Generalization in Looped Transformers cites this paper.

Stability and Generalization in Looped Transformers Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:25:22.716182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T12:21:16.015362Z digest=sha256:a527bbd6596f5418a5c47b9d9b99316f2ad9cbd3898afb9a08690881a71b1a67

Observation 9a9dcb7a-7994-4f6a-85c5-7ebfefe6c14a · inbound

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer cites this paper.

When Does Removing LayerNorm Help? Activation Bounding as a Regime-Dependent Implicit Regularizer Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.321368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T08:29:58.495518Z digest=sha256:bc9cf6b1bb8cd347d9dfd684b66b46daf77707d9653ee1910312800ee373f396

Observation c0dc39b1-a0ee-4ea3-b7d2-21ab322a31f0 · inbound

Dead Directions: Geometric Singular Learning cites this paper.

Dead Directions: Geometric Singular Learning Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:36:55.173661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T03:20:09.365073Z digest=sha256:323964dc4ea38ce3a6a7a6589eb5ef95ef934ab454bb50296b3a927515b7a4cb

Observation e0605a6f-b5f6-43ad-b366-a132f84ce771 · inbound

Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale cites this paper.

Algebraic Dead Directions in LayerNorm Transformers: A Forward-Pass-Only Diagnostic at LLM Scale Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.314353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T21:14:11.815522Z digest=sha256:f13cb7bb400a7ac2227b443eb38939f1d5b9c38191913d1b6f8c174fbd11e825

Observation 04607fde-e876-48eb-a945-022b466d8c50 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 98

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.742267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:e0d3b72267ff80bac74c7d142b3da13e54e96b94e092b8786cc259f103f0d91c

Observation 688d377c-aa05-4460-9e65-c8e76ea8e7f5 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:07.666821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:07.666821Z digest=sha256:b97aeab8228837d342bb54e2cf97330eeda4fcf0acabf8f1416cea05d92eda7d