Pith. sign in

Paper Citation Record · LEDGER

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD

As of 14 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.17501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17501 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:54:58.891090Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce36854f-edba-4a91-9615-b9841d518ab7 · outbound

This paper cites Layer Normalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.796795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.796795Z digest=sha256:db0b9b208192e342fad856b13cfacb4515c8667a86ce77dae3e6a356a9cc961b

Observation d219215c-3b81-437b-81be-c36abb8427cd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Adam: A Method for Stochastic Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.818406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.818406Z digest=sha256:252b8c2c24d179511405bd2fb2f579ccef190de602cf10791a77212e0c583e58

Observation 05b68beb-efea-4b68-b80d-19d4b1057d2d · outbound

This paper cites DeepSeek-V3 Technical Report.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.823828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.823828Z digest=sha256:f5be89371a642612262b5365f706a16bdd8045fb75914bbd17215918a3615155

Observation fc2c3735-121c-44aa-80ff-0ef1832a9784 · outbound

This paper cites A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.844812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.844812Z digest=sha256:62f66fb718f04ea339cea7086cab9dea92cd2571f5d84f241007f0a123f19fa8

Observation c3f65aae-289e-4c31-be1a-4d27a400a888 · outbound

This paper cites Large Batch Training of Convolutional Networks.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Large Batch Training of Convolutional Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.851371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.851371Z digest=sha256:e0d55780f2ce8dc5f3bcbe271cf42835bb925aa3e0931feaae433fe6ac8f24ae

Observation a4a68fd6-09ed-4b1f-b3d8-372136ec03ff · outbound

This paper cites Transformers without Normalization.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Transformers without Normalization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.862382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.862382Z digest=sha256:91e80a42648d7ce7e3464bb2bf52abf54330851ea87c02e027994d55961e3424

Observation ca7a4ae5-d782-4073-8b6c-93d8e67b6134 · outbound

This paper cites In a backpropagation, since we have obtained ∂L ∂vec(Y ), we would like to further analyze ∂L ∂vec(Wq) , ∂L ∂vec(Wk) , ∂L ∂vec(Wv).

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD In a backpropagation, since we have obtained ∂L ∂vec(Y ), we would like to further analyze ∂L ∂vec(Wq) , ∂L ∂vec(Wk) , ∂L ∂vec(Wv)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.233831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T14:54:58.868032Z digest=sha256:85a382cc95ccbf5abb5b061489d48dbefcc37f542c94420d0b0f62749b0f90e1

Observation 55947400-400a-4948-8d24-66b547135587 · outbound

This paper cites Défossez et al.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Défossez et al

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.160647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T14:54:58.891090Z digest=sha256:c84bfccd7ef36f010efcef42a878d7f0607603d176c7137ecd0788b6f5ff91e4

Observation 38da0786-31b5-44cb-a816-8b63bc9fd89e · outbound

This paper cites This method augments the gradient direction with a fraction of the update vector from the previous step, allowing faster convergence and helping escape local minima.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD This method augments the gradient direction with a fraction of the update vector from the previous step, allowing faster convergence and helping escape local minima

Reference 1983

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.180836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T14:54:58.884769Z digest=sha256:f26044b8fe263b43d0ef63c6ab4d0123c9abe294545e7838c97adeaffee6dd24

Observation e20bf3e9-e019-4bb4-8551-e98612ef886b · outbound

This paper cites Qwen Technical Report.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Qwen Technical Report

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.834431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.834431Z digest=sha256:2024d10a2fed0a9c0308ebb23b8aec979f2837c8bcf257337370c6fd5309b113

Observation a11c4c39-2f8d-4a92-a779-1caf95de60bb · outbound

This paper cites MARS: Unleashing the Power of Variance Reduction for Training Large Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.857286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.857286Z digest=sha256:16f69f7bd3d9b95579c47f144fe6c4d3e8cb80cc57897dba21eb18b6a1075bcf

Observation e4d077cf-903a-497f-9a1a-68817dae4d15 · outbound

This paper cites Language models are few-shot learners.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Language models are few-shot learners

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.802656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.802656Z digest=sha256:6c57210f175a3cd860bd0d2c8e6ffe1bd352bd2fe000b800c50016b69e5ca201

Observation bfe13c81-6ebe-4925-ab8b-32859137c785 · outbound

This paper cites All language models were trained on OpenWebText, using GPT-2 tokenizer.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD All language models were trained on OpenWebText, using GPT-2 tokenizer

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.215255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T14:54:58.873604Z digest=sha256:b42b9600c9545fbcb269b0fbea4cf275cbd828edf10189164c59ecf2c51ca6bd

Observation c3ff864a-a53e-40a7-ae5c-7b7fbb974fe1 · outbound

This paper cites The Llama 3 Herd of Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.808030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.808030Z digest=sha256:8e6028acf6f39d9a8f47eca93b6fefcda903613f315b3b2a988dae95999cb24c

Observation 271477db-e557-4ed9-9925-3ced02bd0778 · outbound

This paper cites Query-key normalization for transformers.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Query-key normalization for transformers

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:54:59.253015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T14:54:58.812922Z digest=sha256:7c85ec48c55580077a26a332de8c2af97b056c9ca77995173b2fda0345b90c16

Observation 2937fc9b-9bca-464a-9646-ea17b6a75406 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.840074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.840074Z digest=sha256:cd2afb5ceda564adcd603e7984ce77300b09bbbac7a0210856b87c01aea4e227

Observation 8ff146d4-94b4-4b36-9368-70d78bbf45b9 · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:58.829273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:58.829273Z digest=sha256:f6b5a91f923050270f635a28ec21d5d76abe26846368a138c78975cc6311fc92

Observation a612d0c9-61dc-4944-b71d-eba64cdf5b11 · outbound

This paper cites an unresolved cited work.

DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD Unresolved cited work

Reference 2242

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:54:59.197679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T14:54:58.879365Z digest=sha256:4205bc4bc8a65d0cbfceb26f567f2806db20883f1b124030b401dcd046f03d19

Pith citing papers

No inbound Pith citation observations are available.