Pith. sign in

Paper Citation Record · LEDGER

Spike No More: Stabilizing the Pre-training of Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.16903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16903 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:34:18.057936Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:38:43.634840Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4bb26cd6-ba68-43ab-b817-04db150cf91b · inbound

SMMF: Square-Matricized Momentum Factorization for Memory-Efficient Optimization cites this paper.

SMMF: Square-Matricized Momentum Factorization for Memory-Efficient Optimization Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:18.057936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:34:18.057936Z digest=sha256:49409c588e7ecff8a5beeec70457f0d72375958667097862a4ee9f690d533aad

Observation d434e55d-86b0-4de5-b89e-531de6c7414a · inbound

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN cites this paper.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.510977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.510977Z digest=sha256:e86c3bdc43de2d8b8939eef49310b7f9fc69f5307cc0a943ebdd1c0f19e35da3

Observation ddb73ce6-27b2-4e9e-a0ec-878948fd8ead · inbound

Technical Report: Small Language Model for Japanese Clinical and Medicine cites this paper.

Technical Report: Small Language Model for Japanese Clinical and Medicine Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T10:39:13.184328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:39:13.184328Z digest=sha256:250235823cd3367d2029af8476f77cd598d8b57f26b6c7c0b7c62883fe28dec5

Observation ebafdae7-2f71-4a6e-8c3a-b718eab9b0c1 · inbound

YuLan-Mini: An Open Data-efficient Language Model cites this paper.

YuLan-Mini: An Open Data-efficient Language Model Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T05:17:55.863452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:17:55.863452Z digest=sha256:fdcce06eca532f881b15018bb9c3fd47fc845a6a973f301052daedc4fe1df7d7

Observation 524ec8bc-f2e0-4fec-8c91-549b02490826 · inbound

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training cites this paper.

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:25.911183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:56:25.911183Z digest=sha256:852a7bd4bbdc5f6d4ca9bae72819e703099d588fa8f68d4e3d6135a9d6051ec8

Observation 8a60c67d-ea8b-4711-8474-e9476ac3d168 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.997471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.997471Z digest=sha256:7cf74f565c672b481fec5a39762fd6173a3b4cb5faf3f0f54f3820645ed56b67

Observation 17df7458-4181-46be-abcf-a8ad4fba1259 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.474681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:a12c9b5f8fad2f6ce599c0302b480995ceb55dae03113947971f71e1df47cd91

Observation 6184a1d1-ae4a-47ff-ae75-e7bbb2156631 · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:13.667569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:13.667569Z digest=sha256:3910109e7e1ba47047d8235b78e5a9306ec815b6b4aab378b57859c00dc5ab65

Observation 992f5f56-b36b-417e-bdae-16dc89ca8156 · inbound

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling cites this paper.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.325852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.325852Z digest=sha256:cf140a1f6dc83c8eae8467793814785012e9511e5eff630972508e37030c247f

Observation b17893d6-0fc5-462c-8b2f-03367d4f9bed · inbound

Foundation Models for Discovery and Exploration in Chemical Space cites this paper.

Foundation Models for Discovery and Exploration in Chemical Space Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:52:25.010557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T05:52:10.848118Z digest=sha256:ffc2c764901ee8933764c42084b2fdedb168f829afe028292ed655fa1a264da9

Observation c63a0b42-d534-471a-9761-203a00acd1b7 · inbound

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers cites this paper.

DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:30:55.527546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:30:55.527546Z digest=sha256:05c2261c0f6cca7334dbac1e6931c49933d7de6fc9a3d242485dceea2ca3f83d

Observation 092d188c-22f4-48ec-aecf-9283e1ad4ed6 · inbound

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers cites this paper.

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:37:53.414766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:37:53.414766Z digest=sha256:70ca0f558a60dbd1d154e127a4585a18537d61808cbbc4797074632d03bcf0cf

Observation ae562f90-4d5a-4e53-a27e-1f3a0d64a612 · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:1cd6761d619b062095d374edecf2a78cb5b34c4c9cf1fd15d7175761d8d07550

Observation aef623dd-8646-42fb-b7cc-058c7ff0384b · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:01.089544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:7da17177493947343bbb520dac425183333cd068c3250ed12d55cadcda90daad

Observation 48f2859d-28c4-47df-bc13-7f552cded153 · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.440694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:bc8ca5545fac7f4e4a71932282454a38eeadf4140b8ce03bd3adfcf9963682c5

Observation e9ce5ca0-8ae6-4b13-a4dc-d9c285ad2854 · inbound

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training cites this paper.

GNMR: Runtime Stability Control for Low-Precision Large Language Model Training Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:42:36.035343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T18:53:47.187437Z digest=sha256:ba26ad291f9eab99a35f5fc14f8bcc31e4fcea8364b43c325f043f22a914d49b

Observation 500ad647-176e-40c1-9fcc-d05e13a516cb · inbound

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding cites this paper.

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:38:43.636171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T17:30:39.458521Z digest=sha256:fcee577be3f5a7682477cc60f68a44800829ba36af188daea0de5912fa17ec29

Observation 39d98781-3a3c-4b92-afaa-a6f7a00e3b52 · inbound

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions cites this paper.

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T09:17:47.212034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:17:47.212034Z digest=sha256:24393e0a17a207b666e3fb7f13a99dc11ac808410369c7f9d8f9f0f944ecbab7

Observation 948613b1-a859-4158-8167-07321c08220c · inbound

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions cites this paper.

Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:08:20.137063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:08:20.137063Z digest=sha256:b2a8219b4905ce2abadce2e376592ed277ededb38610815d71c3441342d41f55

Observation 05c4c720-1162-4aa0-ad63-e86d1da21025 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:45.563563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:45.563563Z digest=sha256:b0d5a0fe2cf62b79f60abc338b54cdfd5d1e5a51d4d2eb8012244d8e94b01390