Pith. sign in

Paper Citation Record · LEDGER

DeepNet: Scaling Transformers to 1,000 Layers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2203.00555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.00555 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:26:11.857239Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

54
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d3a80d7-d25f-4d34-b7a9-5dae2e110da1 · inbound

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness cites this paper.

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness DeepNet: Scaling Transformers to 1,000 Layers

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:22:08.887842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T16:22:08.801066Z digest=sha256:ad23959286694d55ad2189214a5bf6098280e9ee03fba0e4c471910767ad852a

Observation f30cc71f-b14a-40bb-b558-4b59c1b0fced · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:32:22.873613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:c9cd849fd1d8d012113f81f54909dd9a727ae18e4ff235ddd0ef6ceac305ad46

Observation 4809d27e-dd9a-4897-b631-e2f5347ecf70 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 282

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.725925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:ec5b6ced02b55af1bee08f75e6b85b0963fab3392e3527159756c9f47a8387cc

Observation 1298970c-a84f-4672-b77e-b7ccfaab61fa · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.233161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:1d9c7c68671dc0484d70bdedbcbb10c583721cde74071b405b9fac31f81ef9b0

Observation 1b15f8d7-71c9-4b3f-876f-084f39a99204 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.286410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:39f344c2956e6d39a5e04e94d0ef584bc312141f4b8547b77312e86ec9578c75

Observation 6f5ab099-0aec-47a4-b763-4e01a88edbb7 · inbound

Retentive Network: A Successor to Transformer for Large Language Models cites this paper.

Retentive Network: A Successor to Transformer for Large Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:29:59.737427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T20:29:59.633357Z digest=sha256:ec48f2173ef6653b53d4318f9b568079eeece2f8895d62e2bd39b3e86ac80d3c

Observation a1019b78-1466-4b3d-8e4a-39007fabba98 · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free DeepNet: Scaling Transformers to 1,000 Layers

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:04:34.938109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:27dd9740e560e7eb918d9b9a8224315023bb6273097dae0a759b713a00d0c56d

Observation c5041f34-d033-4787-b3bf-16da13f0a5b6 · inbound

Taming Transformer Without Using Learning Rate Warmup cites this paper.

Taming Transformer Without Using Learning Rate Warmup DeepNet: Scaling Transformers to 1,000 Layers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:11.857239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:26:11.857239Z digest=sha256:b4878df9f83da2b3fd66ff34cffcb09fc1b3682dfe4a4950e8d6357c76e6ce8b

Observation c3fabdea-9f73-48c5-8c82-7cf8c357bba6 · inbound

The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network cites this paper.

The Algorithm Is Not the Behavior: Learned Priors Override Look-Ahead in a Chess-Playing Neural Network DeepNet: Scaling Transformers to 1,000 Layers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T14:24:16.800803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:24:16.800803Z digest=sha256:8f02562915a87192a68653efd1dbda1b3dee66964c158c6c0a90972a5482fbc9

Observation 4979db94-aa97-440e-a0f4-16e9ca7455cf · inbound

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers cites this paper.

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers DeepNet: Scaling Transformers to 1,000 Layers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.500759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:19:46.400753Z digest=sha256:159c66dad7b63402ec30833f9fd9d00423e2e3ea615c9a2af1e2b7b588c86160

Observation 9f5b1146-dfed-49fb-8341-7c72123b16ed · inbound

Attention Residuals cites this paper.

Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:04.447997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T06:39:04.312270Z digest=sha256:367f61ebc8b4cc710894e744074cbecd560800ebecc834362f446089f36aa48d

Observation ae21a8cb-1144-436d-9275-5c41827c84ea · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs DeepNet: Scaling Transformers to 1,000 Layers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:7ebf825396e807b2010c95eef9194af0f585adbf07ddace9f41c8bf3a96e2357

Observation eb4ec1ad-925a-4957-bf1b-2ff4e44b2321 · inbound

Delta Attention Residuals cites this paper.

Delta Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:59:01.842019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T20:55:29.257097Z digest=sha256:8d7dc2adc7c1e505bbe094f43412a7ec276cd551f23a6881a6004366b30b1e2f

Observation a55527ee-9603-4700-980d-b86a5cc646a0 · inbound

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis cites this paper.

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis DeepNet: Scaling Transformers to 1,000 Layers

Reference 274

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T15:43:26.396967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T15:41:31.181025Z digest=sha256:44b7e3333db3de957b1a74681407766fcdb329467c9648b4a2ff7e737d05faa3

Observation bf866434-9986-41cb-9eb1-a100cf4345bc · inbound

HAARES Half-Split Residual Basis Routing for Deep Transformers cites this paper.

HAARES Half-Split Residual Basis Routing for Deep Transformers DeepNet: Scaling Transformers to 1,000 Layers

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:36:55.424558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T03:12:38.612580Z digest=sha256:a4eca5f8764bdd0f3a28965dc8d5a198473334da6b6df1f152c6a031d2826eec

Observation 25bf54da-2c13-4d58-83eb-668bb44d8c1b · inbound

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models cites this paper.

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models DeepNet: Scaling Transformers to 1,000 Layers

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:08:44.043563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:27:48.590008Z digest=sha256:34642ab1c0265a732da12c8647808cf06feac86430affadb04aed5b67a3e819c

Observation 29b3ba39-91ba-46e7-8dac-bc818def2ea9 · inbound

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry cites this paper.

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry DeepNet: Scaling Transformers to 1,000 Layers

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-26T05:29:00.082328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T05:22:26.818078Z digest=sha256:cc56e1c3071e993f30defd1fea74b9bea0c2ea3ae328173e1ea640d50e89ee40

Observation d1beed07-4ff1-4138-b2c3-1e679ea7d257 · inbound

Review Residuals: Update-Conditioned Residual Gating for Transformers cites this paper.

Review Residuals: Update-Conditioned Residual Gating for Transformers DeepNet: Scaling Transformers to 1,000 Layers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.860060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:38:06.546686Z digest=sha256:d54cec72fe79bb7eb4dc2becdcf738ecaa93c974f1415fb1d38db63f48f9b012

Observation a5f73626-9260-4004-bab3-cfcd8552fb0b · inbound

AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating cites this paper.

AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating DeepNet: Scaling Transformers to 1,000 Layers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T10:36:03.467256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:36:03.467256Z digest=sha256:1058a1275809b14644c676e39b19379f2363dd88c065f6a010b340ec88281841

Observation 41a76f9c-c9e8-4c07-9061-6ae57e9dfcc8 · inbound

Dynamic Parameterization Is Not Dynamic Inference cites this paper.

Dynamic Parameterization Is Not Dynamic Inference DeepNet: Scaling Transformers to 1,000 Layers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:35:16.992424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T00:35:16.992424Z digest=sha256:ba339cfd513467f73e17ff89a9ade147306531d30adcbe14a3fc92dc964ebe0f

Observation a5b95787-2b6c-41e1-8e86-3381df5de39a · inbound

Multi-Head Attention Residuals cites this paper.

Multi-Head Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:39:23.618991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:39:23.618991Z digest=sha256:019504b71f30a257226695b1f5bca2b5c79ceb21a9c43cbbba360c185bc921d4

Observation 3b4514cd-824d-47b6-a1ff-8b93512acde5 · inbound

Multi-Head Attention Residuals cites this paper.

Multi-Head Attention Residuals DeepNet: Scaling Transformers to 1,000 Layers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:48.498444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:48.498444Z digest=sha256:9b8356d1612f7d58414fdf441648ade914bcb30d13439de672b5ca21baf75c5c