Pith. sign in

Paper Citation Record · LEDGER

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling

As of 10 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.22049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22049 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:21:11.749664Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1e56cb1-b782-4045-8322-3306d5d4fc46 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.855557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:09.405154Z digest=sha256:bce5785f02b3b644df5ee8cbd3b1a26065d3233c6fcefe1fe739d9b1c5f63fe0

Observation f6ae2540-0da1-4afd-8286-7f3a109c9cc2 · outbound

This paper cites The Llama 3 Herd of Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.452244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.452244Z digest=sha256:882a65abc19b8b47e314011c3a4640fe09c607c052e30cc3d3fad4c502ae4886

Observation 0bd11535-26f4-403f-a4b6-860fed6d6cc7 · outbound

This paper cites Qwen2 Technical Report.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Qwen2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.516987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.516987Z digest=sha256:8e112e116568511e6a6249f43dd23ed116b80418c5c2f22db450341817c2d2a3

Observation b4bb5521-5c02-4c3a-b75c-b1b989834350 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.605367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.605367Z digest=sha256:aaeb68c74fef3d98712e770136ebdc53fde5a41a867562e79f27a8ca04b9fd19

Observation caca3065-fec6-45b6-a839-e549e05e583c · outbound

This paper cites Adaptive input representations for neural language modeling.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Adaptive input representations for neural language modeling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.683118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:09.660591Z digest=sha256:2e8619b18b09dcd2dbaf018927c0677348408ff6b2677472e6538c64c2b60800

Observation e77e122a-9f1d-40f1-9926-4ca44946bdcc · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.708524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.708524Z digest=sha256:355fabdfc70a7362a1dd1809df975266e46a0a43597e154e2343e0c112307f27

Observation bc0dcd1f-e364-4b0a-bc50-bd37b63d1aa5 · outbound

This paper cites Learning deep transformer models for machine translation.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Learning deep transformer models for machine translation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.545995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:09.764755Z digest=sha256:122eb454f4e2aff50579316cdf71cd2687a82d6fe9bedba27bf719e1b74e87e2

Observation b457ab1c-8055-469a-b204-a5afb9ca33d8 · outbound

This paper cites Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.387759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:09.865572Z digest=sha256:45fa6b2b593e4ab78f0fe74cb6cd60d09e2b9efac79562776380bab13f03a037

Observation b716bc9a-db6d-4c36-bcd7-c917ef91e4cc · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The Unreasonable Ineffectiveness of the Deeper Layers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:09.909236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:09.909236Z digest=sha256:6c0870df7ab0ecf094185941cd4b99a4ea2594fc793f6e264233b48c33229375

Observation bbda0a57-d6e7-4e8b-84bd-8262e9067d5a · outbound

This paper cites ShortGPT: Layers in Large Language Models are More Redundant Than You Expect.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling ShortGPT: Layers in Large Language Models are More Redundant Than You Expect

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.026345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.026345Z digest=sha256:4633fd112b658c0d027f5f664e177f5c1da66579deb2215a3a1d7e76a7ac6ad9

Observation 6d9c51fc-52b9-45eb-8887-1129ca86a587 · outbound

This paper cites Layer Normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Layer Normalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.068870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.068870Z digest=sha256:e5ae41fd88812d94f6a425cf04dd13178f83815817cfceeb27d70583d14c3f3f

Observation a7ead766-3b43-466f-b27b-331c1a18ecf0 · outbound

This paper cites Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.135364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.135364Z digest=sha256:57fb099ee5157fc64f80efc25a323f05f4936deb0a0e8809942130c96445dfb4

Observation 765c349a-a682-4772-b025-c7490566712c · outbound

This paper cites The curse of depth in large language models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The curse of depth in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.187845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.187845Z digest=sha256:c595c6d5d545c6c3a9b135847b4cf658c67d466cadc81134789e1e424e50025d

Observation b71d0a8e-4860-43bd-b0ab-c7d4c60477ef · outbound

This paper cites Deepnet: Scaling transformers to 1,000 layers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Deepnet: Scaling transformers to 1,000 layers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.284486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:10.234776Z digest=sha256:8bdf0f1705c0c2edd0dabe0c3c9e959764e8bf74828d1562d8f80d88c35368a8

Observation 19ca331f-9c38-4f44-bd14-3e0f5396d1dc · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Cogview: Mastering text-to-image generation via transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.176152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:10.268765Z digest=sha256:858d8ca6f55b08b3e5f38e8f998508643abddc97f942c78ce266ec591be28952

Observation c04172f7-369d-4e02-a47b-9ef09d6682a4 · outbound

This paper cites Group normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Group normalization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:14.091570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:10.300584Z digest=sha256:ef279aead103bad034224b30e45bf500357cc73326dc489c6ae0b5c5e7a29846

Observation a22f81b1-924a-4493-a895-affa81442b78 · outbound

This paper cites Root mean square layer normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Root mean square layer normalization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.952220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:10.352396Z digest=sha256:48419d247164c30b2bb110e4301913007acd3f2ea5aa67f01a745cc8b3e8d3a0

Observation cce29769-20ca-4659-be46-5d740a038946 · outbound

This paper cites Understanding and Improving Layer Normalization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Understanding and Improving Layer Normalization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.438131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.438131Z digest=sha256:caa490d456417eb8620aaaf7335fbaa1632ea934c86a4eba123f682ac4c4f9fa

Observation 0b006660-0b36-4b59-949b-c47f4450d93a · outbound

This paper cites Transformers without normalization, 2025.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Transformers without normalization, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.515415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.515415Z digest=sha256:f66bab253135b844fa663e4213f7df47e181da3be6cf62bece3752e581876e21

Observation 214245db-f125-4b6d-b055-c67a9c129e3f · outbound

This paper cites On layer normalization in the transformer architecture.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling On layer normalization in the transformer architecture

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.865768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:10.573822Z digest=sha256:1eddf4f4f539bae110ebf4d8a375248c1a2c75f2c1bdb67d5221642c0b0cdf19

Observation accca324-9de0-47f8-b790-2b45b3eed285 · outbound

This paper cites B2t connection: Serving stability and performance in deep transformers.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling B2t connection: Serving stability and performance in deep transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.733130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:10.615392Z digest=sha256:18584c90fa6dbe2a5048c882d2b89f269b8160d0aaf0ea4eb9dbb9c0caf009aa

Observation 5ed40db9-0ac8-49d2-9536-daa3fb9bbafe · outbound

This paper cites Peri-LN: Revisiting Normalization Layer in the Transformer Architecture.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.661789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.661789Z digest=sha256:386ba0352d5fc388854c0519fd3d023230e70baca458949134c226c9d3520385

Observation 99d9f61f-7345-4669-8584-0e8a34e7dd7c · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.741333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.741333Z digest=sha256:9d1e75733429f710e95919515a1de925bd2c4da06c3a5aecdb780f246b880df8

Observation b3be30c0-e5f2-4ecc-b23e-b4811567bf73 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.769395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.769395Z digest=sha256:f1f96831a4a3ce02b0a346d81920a90320aeee23a6730e8bc15f2d691bea5a65

Observation 6d7f4837-08e2-4a40-b736-2a68eb5ec173 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.825974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.825974Z digest=sha256:fd5cfb717d731f16b5786f45e6581eb1ff0a7f04b7aa1c85d669cd298a73def9

Observation f6768497-b4eb-4e73-be14-ddcad43e1eb5 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.900239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.900239Z digest=sha256:dabc894684fdb9e2b0f587035f0c8b31b7a916a686f00da722ee126d8e45c0df

Observation 90c34a68-cf92-43ca-b45b-4a76e4068468 · outbound

This paper cites an unresolved cited work.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:10.954572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:10.954572Z digest=sha256:dcb5662c71b15915decca025dfb924df55f84e2aeabb2a2b417b43af4a037584

Observation ed815e55-47d4-48ac-8f61-9bb187ea90f2 · outbound

This paper cites Relora: High- rank training through low-rank updates.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Relora: High- rank training through low-rank updates

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.595864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.014246Z digest=sha256:ad3f2d7b18a0d6ea2c9aa06c131b2c75f816388f17111d23cf0cf4df2a387378

Observation 8a8621e3-d424-4be8-a038-6fa8409f93d5 · outbound

This paper cites Galore: Memory-efficient llm training by gradient low-rank projection.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Galore: Memory-efficient llm training by gradient low-rank projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.528482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.043414Z digest=sha256:685d52889766f44dd83306f9bb00e84c497e0cb70050eb8382398fa13688cb87

Observation 6472a69a-7795-4c60-86c3-14ae6c78ed73 · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Transformer-xl: Attentive language models beyond a fixed-length context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.357730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.118726Z digest=sha256:bf4122f47dd5c0f771181ea1157155bb67c804d57dafca23da716b3cdbb9b03d

Observation 740f9556-949b-4e22-a5cd-91292529a5ab · outbound

This paper cites Sandwich batch normal- ization: A drop-in replacement for feature distribution heterogeneity.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Sandwich batch normal- ization: A drop-in replacement for feature distribution heterogeneity

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.196929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.168935Z digest=sha256:2e8ec9785df38d6ef9b2a611f99fc9f2a88c5b37762b60130a812c33b3863fa2

Observation bc3ca830-ae9d-4668-b417-79cd9809d6bc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.222328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.222328Z digest=sha256:69e0aa806c859f62ee43faca1ae580a20ca4aca2326b22e042760bb8a12a38cb

Observation 5ac6186a-dee7-43ad-b92c-55d3030089bc · outbound

This paper cites GLU Variants Improve Transformer.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling GLU Variants Improve Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.265526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.265526Z digest=sha256:3a3880dc0db952c0bae4d967b5f26027a896e24114cf2298bd35f9b77af8ed29

Observation 992f5f56-b36b-417e-bdae-16dc89ca8156 · outbound

This paper cites Spike No More: Stabilizing the Pre-training of Large Language Models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Spike No More: Stabilizing the Pre-training of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.325852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.325852Z digest=sha256:5190549416ba4ebcd1aeb6eeda44c6f260d9f2b1dd693b3f3ee4d935efe463dc

Observation bfb1a337-b77a-4518-b7f5-97fef0e6e41c · outbound

This paper cites Adam: A method for stochastic optimization.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Adam: A method for stochastic optimization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:13.057360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.363870Z digest=sha256:b244d3d1c285402efea98afed7328714a7950fe35210406cd9b37b07864d036e

Observation 182ed576-34bf-4d7f-95f1-448bcf0a792b · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.394382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.394382Z digest=sha256:033091dce444c048f5a0b865d6dcc694eafc35eaa0f0c7b9491fc1eebd5a4d21

Observation 3d83aa1b-b6b4-4f13-9356-d47e3b1b86a0 · outbound

This paper cites Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Documenting large webtext corpora: A case study on the colossal clean crawled corpus, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:12.830440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.433172Z digest=sha256:acd1c4e63afd348d4762a15d06dfec83fc5cf5bd2f69dc61fd552b7b7b7fdf55

Observation 7be4cea1-d926-40e8-83b3-286c529a8def · outbound

This paper cites Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:12.697867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.518549Z digest=sha256:72f9cec79f9a1f0de64e36b9ec759ff2f198c08ff4c1ca2bdf9fe8954b010e37

Observation 8d26608b-4133-44f8-9cb5-2b2c2635f1f1 · outbound

This paper cites Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Lisa: Layerwise importance sampling for memory-efficient large language model fine-tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:21:12.512251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.584659Z digest=sha256:dd8eb49c291d19ec9d7ab993f7849e0ca0d94cb0b41a80e45495f4aee19c5713

Observation fb50ee85-7b04-441f-8740-cd7134cf8679 · outbound

This paper cites The language model evaluation harness, 07 2024.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling The language model evaluation harness, 07 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.650939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.650939Z digest=sha256:d26243a84b149f47cb1a0058dd00871ea911ce25a42585f2113f76a2814baba5

Observation f67d85ab-c70e-4247-817e-8c9c74823dcb · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models, 2023.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Llama 2: Open foundation and fine-tuned chat models, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.722895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:21:11.722895Z digest=sha256:324c4fa57bc60d56f7863b879bb61d084f0a4bef2109c0d155e46e0f5c1b4d9f

Observation bf37f784-a868-4bb5-9853-f70e99755612 · outbound

This paper cites an unresolved cited work.

GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation Scaling Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:21:12.211536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:21:11.749664Z digest=sha256:dbc81078973eee876e427e45d0e18068477ece300ff560aae199091213428e96

Pith citing papers

No inbound Pith citation observations are available.