Pith. sign in

Paper Citation Record · LEDGER

Reducing Activation Recomputation in Large Transformer Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2205.05198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.05198 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:56.938309Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

55
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23efadb5-39a3-4564-96fa-05b7a3431662 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Reducing Activation Recomputation in Large Transformer Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:19:46.710952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:005b79b43a4b68e6686155e9fb8826b1be1ede337f7cbcf379a9f5479e04eff9

Observation 5a993fea-6aa6-42ef-ab50-9fbbc7c9c94a · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Reducing Activation Recomputation in Large Transformer Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.812840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:d34fbfc9dacb9e9cad514b42ca6c75e6366db3bd5d77031909c031ee12e46111

Observation e3169a5d-bf93-40d0-b0f4-2af1cd38ed7b · inbound

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel cites this paper.

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel Reducing Activation Recomputation in Large Transformer Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:15:20.079382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:15:20.027659Z digest=sha256:c63c3dcf1a2f4a6763f4dcbf50ab5bf3b4d61bd5e7a14e5f43be30784b8df2f9

Observation cba71a44-9ef7-4c93-85f8-b271d07509da · inbound

Ring Attention with Blockwise Transformers for Near-Infinite Context cites this paper.

Ring Attention with Blockwise Transformers for Near-Infinite Context Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:28:28.258495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T19:28:28.201789Z digest=sha256:baf951ee3b2c33a6c2f09ba74fb51a3445fd8a5e779f3ef5d53f81033471909e

Observation 691d0b08-e397-4417-a090-0b2b67fe70d9 · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Reducing Activation Recomputation in Large Transformer Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:31:03.874186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:10ae87b4094a546114c121bcb2b0b5e635fee849fe7582f9d7cdf7e97a512963

Observation 6c13164e-123b-4bca-9aec-6c6c21527958 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Reducing Activation Recomputation in Large Transformer Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:07:14.441789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:ad3f240fedfce105960f86460649b0a05624366b378d1a253f651d25fd86b510

Observation c7c516d8-0213-4b7b-924e-62402ec17b46 · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Reducing Activation Recomputation in Large Transformer Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:15:09.565107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:f5849ed47861858f00b5410b16ac7e4ded76a7ea4412e3da615fd3c311aacd3b

Observation 6cf91668-8f56-47a2-9b89-8c4066a1a1d8 · inbound

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project cites this paper.

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project Reducing Activation Recomputation in Large Transformer Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:05:09.524204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T21:04:11.859563Z digest=sha256:06833a444ec16e0c89df3929106e1e0109bd5fbece96d9f3c4de0371125225da

Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.938309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.938309Z digest=sha256:949484ac7aca3a944516e6ad678dbf2a3f9b22fa33eb9b96e0306d0004c3eef3

Observation 56cb0289-22aa-40af-8685-0e92e1d25712 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Reducing Activation Recomputation in Large Transformer Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.846985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.846985Z digest=sha256:05dd9225024a003e545fad71f93f637d692284dbdbf46e101a55d5e5358c36dc

Observation ab079cb3-4ffd-4be9-96f4-13dc456ad6b2 · inbound

Photonic Fabric Platform for AI Accelerators cites this paper.

Photonic Fabric Platform for AI Accelerators Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:20:28.750375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:20:28.750375Z digest=sha256:6852906e85ef590f0c7fd62d84fedf1688aed4edf71ed9314b24eb8ca96a72ff

Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · inbound

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling cites this paper.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.758677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.758677Z digest=sha256:9d7883cf89acde22b9e574435286052abb4a0b9739d4297556ca04f8cfd51f10

Observation 2310f253-5f7a-4c4f-b35c-c2b3c75cacbb · inbound

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection cites this paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.661629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:723fc2c2fa11efbb87f9458e039bc793070bcc1e47477c4e6fb490361865ab4b

Observation 1599cf5d-4ed1-448c-9376-f64d6d84fb65 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:51:45.802088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:606a526b9579003c82e8ddba72c9980d6f69886e1dab3d510ebe042360f17d61

Observation e994574e-f6b7-406f-9d3a-6140912e951d · inbound

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training cites this paper.

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:06:27.277349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T14:04:31.017142Z digest=sha256:3f7672f2ccd160ce01cb44e41e051329f7844648ab0bf8f8d19788a657fa6cf3

Observation e0fd7fba-7f23-4a1d-9063-d8b7f7787471 · inbound

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs cites this paper.

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs Reducing Activation Recomputation in Large Transformer Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T22:28:46.339471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:28:46.339471Z digest=sha256:f30a675eca3212ea419c4c05c46fab7ef93a7d445b0b5e8adba17b910e39b54d

Observation 12782da5-748a-4f8c-b556-b6116eb53347 · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence Reducing Activation Recomputation in Large Transformer Models

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:40:42.615219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:fd3eb5e7ab7ed55f0c1cd47dbd7bb7004aeddc1438fa4e4f1f21a8f3e2540e7a

Observation 25597e34-7a25-4688-a37f-d0c9ac86d4fa · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.812534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.812534Z digest=sha256:9182167b1b40cb003e3952decc7d0e06ab21ec9e4bc029af231ec8eb6354a573

Observation 44a5279e-e3ff-4c8f-aa1e-c570a6561d63 · inbound

Efficient Scaling of LLM Training with Flexible Context Parallelism cites this paper.

Efficient Scaling of LLM Training with Flexible Context Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T20:59:24.724128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:59:24.724128Z digest=sha256:5dc1cd42e32dbd4d2e83b6225e038fc4e9f2e60fb3830bb9f15cc2b0601780f6

Observation b2543f4e-0d6d-47cd-9f31-2669ecf6a87b · inbound

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism cites this paper.

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:24:20.560853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:23:19.236803Z digest=sha256:47a0f2968f774779518971463b63f06423485d78f725ce4fe1e9fa7e2267db42

Observation 570aaedd-d074-469c-ab98-e26078be9db5 · inbound

Decoupled DiLoCo for Resilient Distributed Pre-training cites this paper.

Decoupled DiLoCo for Resilient Distributed Pre-training Reducing Activation Recomputation in Large Transformer Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:49:15.683024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T22:20:21.090246Z digest=sha256:38f0d9302401b3003bd2935b4423043dd9213f7e31ddb800cb9522207d73d81a

Observation 01b389ca-2465-4751-aeb1-731fb8ab9bfa · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe Reducing Activation Recomputation in Large Transformer Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.808957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:097925b6d84540ef178a0cf6cafdf0b1c9029abd05efd7e7337df57d106ebd0d

Observation 626393f0-cf06-4885-8f1f-e18a1ba5e337 · inbound

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism cites this paper.

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 40

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T17:51:07.974000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:04:02.418499Z digest=sha256:d66c54005f6b0bfa1418a257fdf3781d3aa9e219ef8174da93f214a563f4d481

Observation 8b587fa3-48b1-4433-8d58-456f31566783 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:06:15.053308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:348c36476de0763616641e35a142d147f9d4ea245444147aada44e69fa9b7cc8

Observation 99be7701-2fdb-4847-b917-fbca9bf0467d · inbound

Instant GPU Efficiency Visibility at Fleet Scale cites this paper.

Instant GPU Efficiency Visibility at Fleet Scale Reducing Activation Recomputation in Large Transformer Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:54.741443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T02:43:09.077121Z digest=sha256:af05f0970bb4cf3c70c33a1f0ca4ddc7164f63aea570d15af1a8a146687d938c

Observation f8d65840-193a-47fe-8a6c-495610f149ab · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.529353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T20:20:53.599064Z digest=sha256:0e942f9a6a6073275494f2316d8d0e7a648d18f29e08cdc55f47a0ec60992753

Observation 0c493078-284b-448d-af6a-03d51e9a3e57 · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T10:54:16.434702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:54:16.434702Z digest=sha256:d922d985d489a50b0c01f243b20547b731427b8a8af92bcdd3633f1ee9f000fd

Observation 8d91bf9b-23e5-4d5a-98c0-f507144fa1b9 · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T12:13:46.609563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:13:46.609563Z digest=sha256:38c08d6998bcd5b19be3db13d17104d8880c420d0eb92e524e6b2274c43555dc

Observation faa07337-c6ca-41dc-a16f-b1657d9e5bd8 · inbound

The Cost and Network Limits of Space-Based AI Compute cites this paper.

The Cost and Network Limits of Space-Based AI Compute Reducing Activation Recomputation in Large Transformer Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:46.028286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:46.028286Z digest=sha256:11e50325d4c4012f9b472dae1598e28a9f6806307cadad3f9f5e1aa7911043b0

Observation a31fe6a3-500d-467f-9f96-095c7d2e8fba · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.462772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.462772Z digest=sha256:4d83674c299ab38f619b9b13203c9f37e264c12874686e06676b99ee999d7349

Observation 58bacf8a-3c44-477c-ad9a-a7ebad505fcb · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Reducing Activation Recomputation in Large Transformer Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.263658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.263658Z digest=sha256:01ee7ae53289e8caa144a250339ddc313611b3bf6b1f8458123e231858fcc10d

Observation 2b777e7d-da11-4b27-8c90-c084795671c0 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reducing Activation Recomputation in Large Transformer Models

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:44.020572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:44.020572Z digest=sha256:4299c4ffbe7c76db1f7b477593bb130ebf6cd81f4d6fcf3352ed26b0230cba89