Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:53:34.006065Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 9 inbound Pith citation observations for arXiv:2412.13795.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:53:34.006065Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:23:40.957150Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
42 of 42 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 9e2913ca-7b58-473a-a07b-d225fe891b32 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1deeb8-dd1a-4640-9e3f-f5e1bb95b207 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ace1b8-c061-4a61-9d14-4e8a0ca68314 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d133d81a-690d-4f03-b290-0a9c5698b328 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f74acf5-21fb-4bfe-8626-94f4232b2402 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Measuring Massive Multitask Language Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed06f0fd-8f44-42a0-b0d5-8b117cd559df · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ac470d-ec52-4677-9a70-f0252183be11 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1f7586-7abf-4540-9b99-d99cde294344 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Mistral 7B
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14024b53-9046-4958-b7fb-8c969f243c8a · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Adam: A Method for Stochastic Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb5ebb1c-00c7-4aa0-be79-e1ce97dacfae · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Outlier-weighed Layerwise Sampling for LLM Fine-tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 124b9d85-224e-4ce7-a286-632dcd38daa0 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN ReLoRA: High-Rank Training Through Low-Rank Updates
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d682455c-7642-42e1-8fc2-ec22fd81fda7 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf80b2ee-bcfd-4997-822a-9a4c671cd851 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa6d420-7513-403b-91f7-c1a4ef59eafd · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Transformers without Tears: Improving the Normalization of Self-Attention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd28a2c-8728-4634-9af2-5ed19faba107 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN What Language Model to Train if You Have One Million GPU Hours?
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf177aae-d0d4-4d2d-9508-ae8c7fc92ea7 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN GLU Variants Improve Transformer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977e0f45-021d-4008-ab43-24c2cb03d611 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN A deeper look at depth pruning of LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e0d9f6-8fd2-4870-a4a6-b194b14a8c1c · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN LLM Pruning and Distillation in Practice: The Minitron Approach
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 120870c7-ed8e-40e2-9fd6-7bee2805dec7 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The curse of depth in large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6af3a28-7b4a-4d08-afee-3df905087401 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN B2T Connection: Serving Stability and Performance in Deep Transformers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d434e55d-86b0-4de5-b89e-531de6c7414a · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Spike No More: Stabilizing the Pre-training of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 702efa29-9c32-496a-b11e-03216cf85266 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN LLaMA: Open and Efficient Foundation Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 262e1765-1de1-4fe8-a16e-9acc0ef6d240 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Learning Deep Transformer Models for Machine Translation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d087676-a2f3-4ba3-8d1d-598c84008a3b · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74ef1976-1e15-49df-8559-fd95f273015c · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e90714-fd9c-4bae-8892-7250baca33e9 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Adam-mini: Use Fewer Learning Rates To Gain More
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd14319-3cd1-4860-8a31-7bbe07af6f5a · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f9c56b-72c8-4fb1-8ffe-3db3596a63db · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN BlockPruner: Fine-grained Pruning for Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fadcd210-41f2-4f4b-955c-e95ed7cd0822 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Table 9 shows the most hyperparameters of LLaMA models across model sizes
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e080ea2-cc63-4c9e-a09d-54772a6b0d3e · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN We observe that both Pre-LN and Mix-LN work effectively with Scaled Initialization
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d8f817e-316a-415c-b1e5-5dcff6d3bb35 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Xiong et al
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 22b549b4-613c-4f3b-a89a-9b43ed948b43 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Curse of Depth
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 926cebec-f703-4a47-9dbe-6da861ad3fb7 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The Remarkable Robustness of LLMs: Stages of Inference?
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce600bf-ace7-48d1-b9a4-e9ed27fe4726 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Adaptive Input Representations for Neural Language Modeling
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3bee34-4311-4279-adc1-3cd8527c6b0b · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Training Deeper Neural Machine Translation Models with Transparent Attention
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d419124-8153-42aa-accd-42ea74970e0d · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41bb98ca-2250-4811-88ab-50ea42387b39 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e63a6e-4406-4c1e-9e33-5a50942ae36b · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Exploiting Deep Representations for Neural Machine Translation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd8136e4-e427-483d-95cb-56f0c8e19f97 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8a39a3-aaef-4d16-aeb3-1239b78177d9 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Layer Normalization
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2373516-110d-4731-9a76-b6e8ad42e68a · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The Unreasonable Ineffectiveness of the Deeper Layers
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20078ba-f349-4cc3-a83a-8f194ff167f6 · outbound
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31f75b4-e492-4d28-a9f1-391160482c09 · inbound
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ce60a9a-7ba9-4126-bdf4-772657ddd940 · inbound
NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ead766-3b43-466f-b27b-331c1a18ecf0 · inbound
GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3e22a9-0f0c-4540-bc49-1ae4682b283b · inbound
EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2844be7e-0df4-4760-9536-d9621466bdba · inbound
On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031a2822-67d9-4a01-98b9-916d66d7168d · inbound
When Does Sparsity Mitigate the Curse of Depth in LLMs Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b954a05-1d0a-4d8f-b9ed-6508d39da1a8 · inbound
Layer Collapse in Diffusion Language Models Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 842d8124-f168-400f-93e9-c1028eab88f4 · inbound
Layer Collapse in Diffusion Language Models Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a024ae2e-df91-435a-b74e-9d398979b863 · inbound
CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.