Pith. sign in

Paper Citation Record · LEDGER

A multilevel approach to accelerate the training of Transformers

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2504.18590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18590 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:35.336584Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy10
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b728c857-212f-4087-bdef-ac34a6a7c28e · outbound

This paper cites Avelin and K.

A multilevel approach to accelerate the training of Transformers Avelin and K

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.864843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.189309Z digest=sha256:dfdf432c326fa89b7a54a68149baa72ccf0c15952f7278b3529a49abfb50b466

Observation 9d6f7dbd-ab6d-4e3e-b06c-73b6ac5c6935 · outbound

This paper cites N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations.

A multilevel approach to accelerate the training of Transformers N-ODE Transformer: A Depth-Adaptive Variant of the Transformer Using Neural Ordinary Differential Equations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.194923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.194923Z digest=sha256:dcd693a316d4615dfab143fa0fe828d6c35c79e4fd0eaf3077c62ba2543d5015

Observation 7a400bd7-f962-48c1-9774-c890f73bfbdb · outbound

This paper cites Brown, B.

A multilevel approach to accelerate the training of Transformers Brown, B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.849873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.200292Z digest=sha256:46413eec6203589b1dd7d1c711de286fa2ab52c6dbd91ae2f0feeed210dbe778

Observation 28bf26ff-921c-4538-9a49-1a4261f46b44 · outbound

This paper cites Chang, W.

A multilevel approach to accelerate the training of Transformers Chang, W

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.830869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.205590Z digest=sha256:8e5cd12d4f70139423267d097608e70c8596f4738f8feeb09984f269144d9aac

Observation c26b12c7-9df5-4539-b7e7-df7cd162f1fa · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

A multilevel approach to accelerate the training of Transformers bert2BERT: Towards Reusable Pretrained Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.210795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.210795Z digest=sha256:264dff8863dcb6c41b938eb1b66e82970dd5aca6cf4555d8972f894e28b505ad

Observation f457159a-da31-4971-a761-283c595066ab · outbound

This paper cites Neural Ordinary Differential Equations.

A multilevel approach to accelerate the training of Transformers Neural Ordinary Differential Equations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.216045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.216045Z digest=sha256:239ad2e66342f355ac1ed9b30d9b44521b8786c6ef4d14f1e2e7897df4e148a6

Observation 4707cc99-010e-4f24-b773-94edf14dc961 · outbound

This paper cites Net2Net: Accelerating Learning via Knowledge Transfer.

A multilevel approach to accelerate the training of Transformers Net2Net: Accelerating Learning via Knowledge Transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.221598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.221598Z digest=sha256:01633ec0aa77b5f62406d715596d56f8abe6a8b431db3ad9974c1ed21c260a83

Observation c0a3eff5-2460-4d3e-a662-09cdeec8b0ad · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.815550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.226671Z digest=sha256:f78e135776aff60ad8635ef54a0cb054d4068c18e27f28e96b0bb4794bfeb70d

Observation 505345ce-d9bd-4783-974c-c3be0b8802be · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A multilevel approach to accelerate the training of Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.231248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.231248Z digest=sha256:bc3a72a0c6eb87e54cfc99acec5097bbc45dc6f9a722a89473c5544af93f8733

Observation 132c4ec8-514f-4ddc-8667-ad41fefcd172 · outbound

This paper cites Gaedke-Merzhäuser, A.

A multilevel approach to accelerate the training of Transformers Gaedke-Merzhäuser, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.799927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.236367Z digest=sha256:fe1d21bdb19ef629b1b7f36fc9aaf7d5b2158668d4a1e95a3e745c257badd74d

Observation 75745099-7f8c-4d52-a522-0d47aaf49c69 · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.783616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.241107Z digest=sha256:602275725844cd8a13865f65349b70f9d0274c3646ec01ca36a8999445a2ca49

Observation 1f19f6bf-a763-4515-bbb0-4ecf31ca5fe5 · outbound

This paper cites Gratton, V.

A multilevel approach to accelerate the training of Transformers Gratton, V

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.767006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.245979Z digest=sha256:d8263559b3822e9172e3a0b2a570e9af38fa8cc763793f7cf18a1d11d353764d

Observation 31a7479a-bc73-4642-a71a-028bc7d50c46 · outbound

This paper cites On the Transformer Growth for Progressive BERT Training.

A multilevel approach to accelerate the training of Transformers On the Transformer Growth for Progressive BERT Training

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.546153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.251108Z digest=sha256:7b47c98be54ddf2395c43faaa9aa5ee38149e3f5a4bd3d411bd2988153390ab5

Observation a0ceb076-e790-41a4-824e-bf8b9b3125c5 · outbound

This paper cites Scaling Laws for Neural Language Models.

A multilevel approach to accelerate the training of Transformers Scaling Laws for Neural Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.256820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.256820Z digest=sha256:cae8549ec9dcab88e3a6e662976d1d96f625449ea4129197ea04a71b3ab9053f

Observation 232b848b-4c0f-4b07-9fd7-30af0633d1b8 · outbound

This paper cites Kopaniˇcáková and R.

A multilevel approach to accelerate the training of Transformers Kopaniˇcáková and R

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.750989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.261885Z digest=sha256:0ffc1732acb586fd2cc3ca8c371e3256003335ffa9024844de9dc84dbb257eed

Observation 3cb45f2f-e467-426a-8da9-a6932d1bf5cc · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.734119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.266782Z digest=sha256:2d0d9550073b8ce2e47b176dd5a1a799a742e5ddb8f5dc52f9fea0478399b956

Observation 588b1dc4-8636-4783-be1f-38e77e8d8df4 · outbound

This paper cites Lauga, E.

A multilevel approach to accelerate the training of Transformers Lauga, E

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.718722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.271991Z digest=sha256:e5088f1d6a73ab3e5d5aeb720aeb31be89824325ea7931cf63bcfb0bb662b8b5

Observation 8c2f8830-7220-476b-affd-18c20cc474e3 · outbound

This paper cites ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation.

A multilevel approach to accelerate the training of Transformers ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.506056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.277017Z digest=sha256:35809e151a313b0841d49a2a109cceab46bbc94d3311f8f4e5873aea7231cc38

Observation 7a61d24e-f828-4e9b-b5b7-3a64fe57e723 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

A multilevel approach to accelerate the training of Transformers Understanding the Difficulty of Training Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.282331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.282331Z digest=sha256:12a646aa028d9e82fe3c835dbcf490663508963686402622ff1c54c3e54787eb

Observation de837293-0650-49db-ba50-e0525d447baf · outbound

This paper cites an unresolved cited work.

A multilevel approach to accelerate the training of Transformers Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:47:35.703031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.287755Z digest=sha256:10709c03615bd06df581aa17345ff2a3df58a4850e77ebd8a596db149051717e

Observation c73b403a-9811-4f42-9249-20da5fcef414 · outbound

This paper cites Exploring Transformers for Large-Scale Speech Recognition.

A multilevel approach to accelerate the training of Transformers Exploring Transformers for Large-Scale Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.292408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.292408Z digest=sha256:7010b991953b9ed3495748d3b355392457f5f4199d458ad45e9d409c41551de5

Observation c1ab8717-6967-42c0-9909-0c2039fb4ba6 · outbound

This paper cites Penedo, H.

A multilevel approach to accelerate the training of Transformers Penedo, H

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.686816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.297850Z digest=sha256:584cf96c681547a0af29653bd9f794225e405c8858f67ae05e09c3ee8eb9c8a7

Observation e282aa30-7704-413c-ae16-182c669bb71a · outbound

This paper cites Quemener and M.

A multilevel approach to accelerate the training of Transformers Quemener and M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.670791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.302570Z digest=sha256:ec6501edbeffead00793e16af1cc0089ce34d1d771db1ed6081b63cc8fb544d4

Observation 7795f1e0-724b-4a88-99db-188c588e0399 · outbound

This paper cites Radford, J.

A multilevel approach to accelerate the training of Transformers Radford, J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.307185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.307185Z digest=sha256:cde6faf0f932b4c17b8aa3a46218a2242cadf3dc0f1a8745e3e41fc0b70386c3

Observation 0c27f562-a2f7-4ae1-a25a-a92ed8a62b65 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A multilevel approach to accelerate the training of Transformers LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.311821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.311821Z digest=sha256:ea0600aee88824fddeec1d346080486cab7b4b97bbafa06d079c18f27539184a

Observation 3f157f21-6be3-440b-ab52-2fbf396242b6 · outbound

This paper cites Vaswani, N.

A multilevel approach to accelerate the training of Transformers Vaswani, N

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:47:35.645250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.316770Z digest=sha256:2f8cb2325bf6381430766c8259fdd88c0021403cf1f28dcc4d7e5f40a3c11df4

Observation b4feb1a5-ce67-46aa-962c-446a901e70a7 · outbound

This paper cites Learning to Grow Pretrained Models for Efficient Transformer Training.

A multilevel approach to accelerate the training of Transformers Learning to Grow Pretrained Models for Efficient Transformer Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.321589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.321589Z digest=sha256:aabeafc7f519f0d97593b6e3250b286671d90e0dfefab8965743758d92ac856e

Observation f2488d91-bc84-4c69-9bb7-e3e961e3f049 · outbound

This paper cites Speeding up Deep Model Training by Sharing Weights and Then Unsharing.

A multilevel approach to accelerate the training of Transformers Speeding up Deep Model Training by Sharing Weights and Then Unsharing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.326873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.326873Z digest=sha256:98b43b9092b34036324369f5879ee9160701c296d2a8ef5c758986485e9b4a73

Observation d6f58139-59c4-4296-9cc3-9bc13904363c · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

A multilevel approach to accelerate the training of Transformers Deconstructing What Makes a Good Optimizer for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:35.331803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:35.331803Z digest=sha256:43e78bfde56d4916373184e4162769577ebbe5db013d39562a238302c71378fd

Observation 129c71b4-b63a-4541-b10f-b69a28dfa524 · outbound

This paper cites A Multi-Level Framework for Accelerating Training Transformer Models.

A multilevel approach to accelerate the training of Transformers A Multi-Level Framework for Accelerating Training Transformer Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-16T10:47:35.381340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T10:47:35.336584Z digest=sha256:c451e9b745c2d07564ec4c0c7758cd86d35168fd7b93a904ca36c7635402bd27

Pith citing papers

No inbound Pith citation observations are available.