Pith. sign in

Paper Citation Record · LEDGER

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 9 inbound Pith citation observations for arXiv:2412.13795.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13795 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:53:34.006065Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:23:40.957150Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9e2913ca-7b58-473a-a07b-d225fe891b32 · outbound

This paper cites GPT-4 Technical Report.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.634085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.634085Z digest=sha256:dc3395ec60d215c8ba81adce99390e6c5e6aea9251d71c1d2feb8cfc57495eba

Observation 2b1deeb8-dd1a-4640-9e3f-f5e1bb95b207 · outbound

This paper cites Language Models are Few-Shot Learners.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.853082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.853082Z digest=sha256:ea0cd25aa7d44dbf40df7e5e38725c9d3780f3dc2839454f958ef39ab82405de

Observation f6ace1b8-c061-4a61-9d14-4e8a0ca68314 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.864581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.864581Z digest=sha256:14f24e7aa3aba6769fa4efe14672931f0e1771fa007fe27a6530073e2a5c311c

Observation d133d81a-690d-4f03-b290-0a9c5698b328 · outbound

This paper cites The Llama 3 Herd of Models.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.897987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.897987Z digest=sha256:a86263e06c1de5929e398d4ebd1e1e8f91a16f9feb42890bdbc7e73a277d462a

Observation 7f74acf5-21fb-4bfe-8626-94f4232b2402 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Measuring Massive Multitask Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.126229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.126229Z digest=sha256:ffb674693bf30ad017e6217621855e22b8c21dee198abc29aa061bed4e758c28

Observation ed06f0fd-8f44-42a0-b0d5-8b117cd559df · outbound

This paper cites LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.131213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.131213Z digest=sha256:bc1446f6de76bcda67db1d8061a85213f7a99f2358f9762a29de343ab2612a03

Observation f4ac470d-ec52-4677-9a70-f0252183be11 · outbound

This paper cites SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.136998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.136998Z digest=sha256:ff8a26a327fc87c9b5cf9acb222114a7b6b23da385895a1819adbd7b76d070fc

Observation bc1f7586-7abf-4540-9b99-d99cde294344 · outbound

This paper cites Mistral 7B.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Mistral 7B

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.147804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.147804Z digest=sha256:3566148f9690b67889ba9928c1af51d76dec339e69fbaa1becdedc44260f15c0

Observation 14024b53-9046-4958-b7fb-8c969f243c8a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.153057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.153057Z digest=sha256:861cddd394fda31bbb5bc1b5adba56a22272b5bbf21c3c66d8109699b8c811a2

Observation fb5ebb1c-00c7-4aa0-be79-e1ce97dacfae · outbound

This paper cites Outlier-weighed Layerwise Sampling for LLM Fine-tuning.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Outlier-weighed Layerwise Sampling for LLM Fine-tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.163976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.163976Z digest=sha256:bf5d7e64136621ed2fc2486b7c000c3a759be9272fc1d230958eeeaf23450084

Observation 124b9d85-224e-4ce7-a286-632dcd38daa0 · outbound

This paper cites ReLoRA: High-Rank Training Through Low-Rank Updates.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN ReLoRA: High-Rank Training Through Low-Rank Updates

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.168970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.168970Z digest=sha256:f4eca1b07860fd3a1b3f1335e316128a2481576ec10179bb401292c784a4fbc4

Observation d682455c-7642-42e1-8fc2-ec22fd81fda7 · outbound

This paper cites More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using Sparsity

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.175070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.175070Z digest=sha256:45c8263f949ab46e9cfbafee72bac46c654aec8484ff887e4aea42d1b45c83cc

Observation bf80b2ee-bcfd-4997-822a-9a4c671cd851 · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.230566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.230566Z digest=sha256:b3d46a1bd970db4a6e1dcc6fd47171ce26ab15db194adf3ed0f485e1b8932349

Observation efa6d420-7513-403b-91f7-c1a4ef59eafd · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Transformers without Tears: Improving the Normalization of Self-Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.339703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.339703Z digest=sha256:490928e4f24f7b329abb9cd553631faa35e919dea78d0b95da4f1fb8b2f3b355

Observation cfd28a2c-8728-4634-9af2-5ed19faba107 · outbound

This paper cites What Language Model to Train if You Have One Million GPU Hours?.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN What Language Model to Train if You Have One Million GPU Hours?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.472454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.472454Z digest=sha256:24cc4c624799d71faff0f39905072e2c16ede748d4307a8cc550fa10ad18a762

Observation bf177aae-d0d4-4d2d-9508-ae8c7fc92ea7 · outbound

This paper cites GLU Variants Improve Transformer.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN GLU Variants Improve Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.478696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.478696Z digest=sha256:0ef91f483f5ef7a085abaea6e70d9b45919aa8da5aa4980c91ecfedca12eaf8d

Observation 977e0f45-021d-4008-ab43-24c2cb03d611 · outbound

This paper cites A deeper look at depth pruning of LLMs.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN A deeper look at depth pruning of LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.484113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.484113Z digest=sha256:7b73a1d9c773219eed329f13178ca2d0f86993363433a07b8fa2908e6e7559bd

Observation 54e0d9f6-8fd2-4870-a4a6-b194b14a8c1c · outbound

This paper cites LLM Pruning and Distillation in Practice: The Minitron Approach.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN LLM Pruning and Distillation in Practice: The Minitron Approach

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.490401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.490401Z digest=sha256:8812208fda085f1dd4c49eb1b3f6098f8dccf6a16a53a14328971f953fff0d0c

Observation 120870c7-ed8e-40e2-9fd6-7bee2805dec7 · outbound

This paper cites The curse of depth in large language models.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The curse of depth in large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.495804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.495804Z digest=sha256:0276a596a95596b6dcc90e1569843e3c2c6e2c51178eac9efd3f131310284984

Observation b6af3a28-7b4a-4d08-afee-3df905087401 · outbound

This paper cites B2T Connection: Serving Stability and Performance in Deep Transformers.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN B2T Connection: Serving Stability and Performance in Deep Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.504460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.504460Z digest=sha256:4371a976d48cdd7a44cab712656802f7dfef4d8fbecbc29a78ec86a50261d0f3

Observation d434e55d-86b0-4de5-b89e-531de6c7414a · outbound

This paper cites Spike No More: Stabilizing the Pre-training of Large Language Models.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.510977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.510977Z digest=sha256:527bf75ddf790b6c0676a6e6ac42638cab37d5801ee86646205d40a3cbeb15a3

Observation 702efa29-9c32-496a-b11e-03216cf85266 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.517266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.517266Z digest=sha256:0a89f2158527df9ebc5685ef9e2c6812bbce1e21a6f7d6d4ba84b9e7208f160a

Observation 262e1765-1de1-4fe8-a16e-9acc0ef6d240 · outbound

This paper cites Learning Deep Transformer Models for Machine Translation.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Learning Deep Transformer Models for Machine Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.524262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.524262Z digest=sha256:586940f111a4b7e2acc1bfc9ff8936eacd1ab0f0cc31f6bcc3f347d3c20b19fe

Observation 1d087676-a2f3-4ba3-8d1d-598c84008a3b · outbound

This paper cites Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.530446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.530446Z digest=sha256:bda512cdbf718e4eee79ded6af2f9ed0adc3856a5fa06144810476a96cfa16eb

Observation 74ef1976-1e15-49df-8559-fd95f273015c · outbound

This paper cites Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High Sparsity

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.568841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.568841Z digest=sha256:ff6cb16dbf0537267f382cc423f61b7aa07e6b3191bbcecabe8e173178c8b885

Observation 73e90714-fd9c-4bae-8892-7250baca33e9 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Adam-mini: Use Fewer Learning Rates To Gain More

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.695082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.695082Z digest=sha256:287eb6c3dc9aa132229eaf979401e93115fd4bdb5155856bb2a000fe8407e5d1

Observation 2fd14319-3cd1-4860-8a31-7bbe07af6f5a · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.818519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.818519Z digest=sha256:46ebd9de5ddccea5c86ff18e4ea96d38fc4083e38d3705e37d725993e4785c84

Observation a1f9c56b-72c8-4fb1-8ffe-3db3596a63db · outbound

This paper cites BlockPruner: Fine-grained Pruning for Large Language Models.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN BlockPruner: Fine-grained Pruning for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.916955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.916955Z digest=sha256:7171514a99e49af82d548d1591be36652bd24daba48af87a34c019bc264dcd64

Observation fadcd210-41f2-4f4b-955c-e95ed7cd0822 · outbound

This paper cites Table 9 shows the most hyperparameters of LLaMA models across model sizes.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Table 9 shows the most hyperparameters of LLaMA models across model sizes

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:53:35.128252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:53:33.979438Z digest=sha256:feea04636a0e55a036c14c38f1177a63a49728d3552d4f9f02f62cb7cefe9613

Observation 1e080ea2-cc63-4c9e-a09d-54772a6b0d3e · outbound

This paper cites We observe that both Pre-LN and Mix-LN work effectively with Scaled Initialization.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN We observe that both Pre-LN and Mix-LN work effectively with Scaled Initialization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:53:35.111567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:53:33.992747Z digest=sha256:813362ba9e0d9eaea8d7c2830b3a7b38046142c2fcc14b7399e023bd71125b12

Observation 5d8f817e-316a-415c-b1e5-5dcff6d3bb35 · outbound

This paper cites Xiong et al.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Xiong et al

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:53:35.095228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:53:33.999175Z digest=sha256:957f1d4e8b2eb4208fe4a385f43b854f4fcf1403d7a94cba03a0de494d11ee55

Observation 22b549b4-613c-4f3b-a89a-9b43ed948b43 · outbound

This paper cites Curse of Depth.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Curse of Depth

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:53:35.077225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:53:34.006065Z digest=sha256:a7332ed6813d456e00324f120ebae1d2e32c5c021ff781357d3da83b4e1e1c84

Observation 926cebec-f703-4a47-9dbe-6da861ad3fb7 · outbound

This paper cites The Remarkable Robustness of LLMs: Stages of Inference?.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The Remarkable Robustness of LLMs: Stages of Inference?

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.158887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.158887Z digest=sha256:cdd0d0d846896bfa18a3463df00a577db1699c355d109fb92dc5cadb118cc45c

Observation 2ce600bf-ace7-48d1-b9a4-e9ed27fe4726 · outbound

This paper cites Adaptive Input Representations for Neural Language Modeling.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Adaptive Input Representations for Neural Language Modeling

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.841009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.841009Z digest=sha256:1c2bcf805bbcc4d2e313a9370e7e7891624eb7d3481748674ec6fe86151e691c

Observation 3e3bee34-4311-4279-adc1-3cd8527c6b0b · outbound

This paper cites Training Deeper Neural Machine Translation Models with Transparent Attention.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Training Deeper Neural Machine Translation Models with Transparent Attention

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.847128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.847128Z digest=sha256:10b81057d8202eaa9ccbd30913abfbf1a5e79aeaa35a784a827905230fe30960

Observation 9d419124-8153-42aa-accd-42ea74970e0d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.869572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.869572Z digest=sha256:8078fad3b6db188e3534050be4a551b7d6a504a77707a159a686a27484e9bc38

Observation 41bb98ca-2250-4811-88ab-50ea42387b39 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.858885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.858885Z digest=sha256:5378bf4be3e35e919c4b0db96794eabf987b2ea17980a0544f21238456b8b70a

Observation 40e63a6e-4406-4c1e-9e33-5a50942ae36b · outbound

This paper cites Exploiting Deep Representations for Neural Machine Translation.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Exploiting Deep Representations for Neural Machine Translation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.875722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.875722Z digest=sha256:a0c78677dfc884e5de139294602be6d5d58ac689f90fdf6bcaa9a140baff293b

Observation cd8136e4-e427-483d-95cb-56f0c8e19f97 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.440522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.440522Z digest=sha256:14bed3d95f24ef0211f50a0ab13fcfbe468d449080defd8e1134be9956a35102

Observation ac8a39a3-aaef-4d16-aeb3-1239b78177d9 · outbound

This paper cites Layer Normalization.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN Layer Normalization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:32.727326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:32.727326Z digest=sha256:5dfcfcc27f50df8fe95bf166a62f49188cb8a3450405226c167d5a0788f69ea1

Observation f2373516-110d-4731-9a76-b6e8ad42e68a · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN The Unreasonable Ineffectiveness of the Deeper Layers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.024299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.024299Z digest=sha256:43d11d74ffa5014d295fa2c2ecab2ed85c854acc30596921a8e25fb73262bf43

Observation c20078ba-f349-4cc3-a83a-8f194ff167f6 · outbound

This paper cites From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.142019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.142019Z digest=sha256:2b00a998c8116d49fe85d842a84c6116713e1f74e8db26b9b1b65c4605c2a46d

Pith citing papers

Observation a31f75b4-e492-4d28-a9f1-391160482c09 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:40.957150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:40.957150Z digest=sha256:3399a2f6feadf53fa98b0b25e60e07b35da8d25f87af8c025400dcf7a0f61c61

Observation 6ce60a9a-7ba9-4126-bdf4-772657ddd940 · inbound

NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling cites this paper.

NeuroTrails: Training with Dynamic Sparse Heads as the Key to Effective Ensembling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:18.461312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:44:18.461312Z digest=sha256:e9c44971e6050543bffa0a26fc1fa2a07d0c1d4ef90099752224a54bf871404f

Observation a7ead766-3b43-466f-b27b-331c1a18ecf0 · inbound

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling cites this paper.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.135364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.135364Z digest=sha256:612fea72e3f945f138f63b649e392b375b33d4948467d9a9c1009a23308be357

Observation 0c3e22a9-0f0c-4540-bc49-1ae4682b283b · inbound

EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens cites this paper.

EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register Tokens Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:18:47.507878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:18:47.507878Z digest=sha256:999a1008c095149b79bfff73a87177d44eb601d96e2997160986ef3ab899b0f7

Observation 2844be7e-0df4-4760-9536-d9621466bdba · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:48.672256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:48.672256Z digest=sha256:53d944591542d3e1ae82878653b47f05559866ee8e517728f6e3538df7332284

Observation 031a2822-67d9-4a01-98b9-916d66d7168d · inbound

When Does Sparsity Mitigate the Curse of Depth in LLMs cites this paper.

When Does Sparsity Mitigate the Curse of Depth in LLMs Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T20:29:33.439034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:29:33.439034Z digest=sha256:6b54454706ee8de4fb5b5acb4609c70c615f62ec5b5eef2df96459792b5ffa6e

Observation 8b954a05-1d0a-4d8f-b9ed-6508d39da1a8 · inbound

Layer Collapse in Diffusion Language Models cites this paper.

Layer Collapse in Diffusion Language Models Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:18.899651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T12:52:12.104539Z digest=sha256:66603d7458589dd1b4d390eab706ae296d9b32e6ac55db4af359c96909f33528

Observation 842d8124-f168-400f-93e9-c1028eab88f4 · inbound

Layer Collapse in Diffusion Language Models cites this paper.

Layer Collapse in Diffusion Language Models Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.183050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:13:32.889304Z digest=sha256:6f8bd280851e03104ac1104709fa1a088695f6959b2c626d2474c6cd64cef95f

Observation a024ae2e-df91-435a-b74e-9d398979b863 · inbound

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry cites this paper.

CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-26T05:29:00.085809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T05:22:26.818078Z digest=sha256:752adfde58bc6935c426fc8b0ec7b0b179003dc1519de3796bd8b431ca45a8cd