Pith. sign in

Paper Citation Record · LEDGER

Reducing Activation Recomputation in Large Transformer Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2205.05198.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.05198 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:30:56.938309Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

55
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23efadb5-39a3-4564-96fa-05b7a3431662 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Reducing Activation Recomputation in Large Transformer Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T23:19:46.710952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:b4d1e2f9d36cfa8285f039a6700fe007ed3e10d2f0943420ea07fbeea7482bcd

Observation 5a993fea-6aa6-42ef-ab50-9fbbc7c9c94a · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Reducing Activation Recomputation in Large Transformer Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.812840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:ebc414480444054da4895a90f0e1d355242c8e90aced5594f01d48d32762f703

Observation e3169a5d-bf93-40d0-b0f4-2af1cd38ed7b · inbound

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel cites this paper.

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel Reducing Activation Recomputation in Large Transformer Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:15:20.079382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:15:20.027659Z digest=sha256:4029851f1eb52459faca258ecb14532dc556b06423aa4096798de7982d132930

Observation cba71a44-9ef7-4c93-85f8-b271d07509da · inbound

Ring Attention with Blockwise Transformers for Near-Infinite Context cites this paper.

Ring Attention with Blockwise Transformers for Near-Infinite Context Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T19:28:28.258495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T19:28:28.201789Z digest=sha256:223763a2ce1213c27cec000016002f4f1c739c853486df64c0069846096b69f6

Observation 691d0b08-e397-4417-a090-0b2b67fe70d9 · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Reducing Activation Recomputation in Large Transformer Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:31:03.874186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:0fcef74e20759f9d7a7d949f89dd093e32d9ed11275230bbb6a17dedd3a216c2

Observation 6c13164e-123b-4bca-9aec-6c6c21527958 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Reducing Activation Recomputation in Large Transformer Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:07:14.441789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:7ea8adb7b490365f78e457b5e8e5a9afacec25731c0888f34aa0542c069063b6

Observation c7c516d8-0213-4b7b-924e-62402ec17b46 · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Reducing Activation Recomputation in Large Transformer Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T21:15:09.565107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:4888e51492ac4ea03ff465c9775d7dd56c33e206fbae8df77f838355135451c5

Observation 6cf91668-8f56-47a2-9b89-8c4066a1a1d8 · inbound

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project cites this paper.

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project Reducing Activation Recomputation in Large Transformer Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:05:09.524204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T21:04:11.859563Z digest=sha256:aedc7765ac059abe55e57a048c400565ac61cc1f1430a62f027597564bc5808f

Observation 05cf9cdb-0e18-47b3-8b67-145d5b4307f4 · inbound

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences cites this paper.

Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences Reducing Activation Recomputation in Large Transformer Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:56.938309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:56.938309Z digest=sha256:0d3fc58e6ad396b514af089148dba64a568aff23855945fdff38ee8441b3d5f8

Observation 56cb0289-22aa-40af-8685-0e92e1d25712 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Reducing Activation Recomputation in Large Transformer Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.846985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.846985Z digest=sha256:6938604c4b88658f0411c253abafa9ce42bcead6ce023495acd015aa819fe8db

Observation ab079cb3-4ffd-4be9-96f4-13dc456ad6b2 · inbound

Photonic Fabric Platform for AI Accelerators cites this paper.

Photonic Fabric Platform for AI Accelerators Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:20:28.750375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:20:28.750375Z digest=sha256:460c3baf182b2330eb81be6e3b9a8632c5c68ce5a1389cb1c8a3e2de38522d67

Observation 29feca7a-4ba7-4d1c-870a-f553e4fdb36f · inbound

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling cites this paper.

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling Reducing Activation Recomputation in Large Transformer Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.758677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.758677Z digest=sha256:654bd22e5bf2917cd6eac416829ffb382b50ee83353f0b70dfb43a3308b7aa04

Observation 2310f253-5f7a-4c4f-b35c-c2b3c75cacbb · inbound

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection cites this paper.

Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T20:41:50.661629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T20:39:32.123038Z digest=sha256:3d159ea4861851cdd30f7814e0e1a172cc8d09ae31f0d9d08afed015ebb9846e

Observation 1599cf5d-4ed1-448c-9376-f64d6d84fb65 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T18:51:45.802088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:a34c2314022685c9e608700da267b550230811f2fd66e9a0bc0742393bb2c677

Observation e994574e-f6b7-406f-9d3a-6140912e951d · inbound

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training cites this paper.

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training Reducing Activation Recomputation in Large Transformer Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:06:27.277349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T14:04:31.017142Z digest=sha256:61fa095df659239210396a5eed9f001638f8730bcf2de473a2bd42d0c5eb537d

Observation e0fd7fba-7f23-4a1d-9063-d8b7f7787471 · inbound

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs cites this paper.

Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs Reducing Activation Recomputation in Large Transformer Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T22:28:46.339471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:28:46.339471Z digest=sha256:217dee55b1bbc50014f6462b0a06b66aa162fe99482e7606ea7cc6254450e4b2

Observation 12782da5-748a-4f8c-b556-b6116eb53347 · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence Reducing Activation Recomputation in Large Transformer Models

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:40:42.615219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:def259e680fe2d7eb6e77f531cc732940fb4ee044e5619ae7bbf19e5111ef899

Observation 25597e34-7a25-4688-a37f-d0c9ac86d4fa · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking Reducing Activation Recomputation in Large Transformer Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:49.812534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:49.812534Z digest=sha256:1b14ba43fb2091817831298f722060b8a97cce6f894fa97d969f3f3615aaab09

Observation 44a5279e-e3ff-4c8f-aa1e-c570a6561d63 · inbound

Efficient Scaling of LLM Training with Flexible Context Parallelism cites this paper.

Efficient Scaling of LLM Training with Flexible Context Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T20:59:24.724128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:59:24.724128Z digest=sha256:583b09c8eaec1375022b742c7fd5d9dc9519683c9e5a564c3e6cf22f3689a8bf

Observation b2543f4e-0d6d-47cd-9f31-2669ecf6a87b · inbound

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism cites this paper.

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:24:20.560853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T10:23:19.236803Z digest=sha256:ab87dcc76d72c074264e6a14e08cd1ce595907cc6f70129936434aaabf4c5d31

Observation 570aaedd-d074-469c-ab98-e26078be9db5 · inbound

Decoupled DiLoCo for Resilient Distributed Pre-training cites this paper.

Decoupled DiLoCo for Resilient Distributed Pre-training Reducing Activation Recomputation in Large Transformer Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:49:15.683024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T22:20:21.090246Z digest=sha256:d1f5f8846010acfe82431ce7d437683d9b3862578d58b531cde19ae6bdfa1852

Observation 01b389ca-2465-4751-aeb1-731fb8ab9bfa · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe Reducing Activation Recomputation in Large Transformer Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.808957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:afba2583da1c8615504a2f6f65b9eaed932ee825eb351c36731bb7483b5aaba4

Observation 626393f0-cf06-4885-8f1f-e18a1ba5e337 · inbound

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism cites this paper.

Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism Reducing Activation Recomputation in Large Transformer Models

Reference 40

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T17:51:07.974000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:04:02.418499Z digest=sha256:fafb3597d832c9587f69b170231cd33e825c3687fd1a7a33391946a8da996914

Observation 8b587fa3-48b1-4433-8d58-456f31566783 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Reducing Activation Recomputation in Large Transformer Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:06:15.053308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:a3509f432e8d3ed0e9cf42b1154dc9a7182de5dcb8147f462bb9a5276e2534de

Observation 99be7701-2fdb-4847-b917-fbca9bf0467d · inbound

Instant GPU Efficiency Visibility at Fleet Scale cites this paper.

Instant GPU Efficiency Visibility at Fleet Scale Reducing Activation Recomputation in Large Transformer Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:54.741443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:43:09.077121Z digest=sha256:894532f806a1c96900762239fb29c536d4f371a0b35f7237d865c271d5722358

Observation f8d65840-193a-47fe-8a6c-495610f149ab · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.529353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T20:20:53.599064Z digest=sha256:83a6f6a18276779f5da053da2b6e930b30ddcb67d60be2d7cd6f14070c4ab892

Observation 0c493078-284b-448d-af6a-03d51e9a3e57 · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-15T10:54:16.434702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:54:16.434702Z digest=sha256:e8fe882c9ad404c22defcfa5121c951523592cec8f041febc4192467a29dcb11

Observation 8d91bf9b-23e5-4d5a-98c0-f507144fa1b9 · inbound

Explaining Data Mixing Scaling Laws cites this paper.

Explaining Data Mixing Scaling Laws Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T12:13:46.609563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:13:46.609563Z digest=sha256:b87887405160e29eba7c4383641403824e53571511c52dfff2854dcc65f8a945

Observation faa07337-c6ca-41dc-a16f-b1657d9e5bd8 · inbound

The Cost and Network Limits of Space-Based AI Compute cites this paper.

The Cost and Network Limits of Space-Based AI Compute Reducing Activation Recomputation in Large Transformer Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:46.028286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:46.028286Z digest=sha256:15b56326a63460d2062dfc310d561d3089bcb6ca5036474d5f25b92751242781

Observation a31fe6a3-500d-467f-9f96-095c7d2e8fba · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Reducing Activation Recomputation in Large Transformer Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.462772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.462772Z digest=sha256:ad8b961b6a5857baeeb3515f22c52f55024154763f1297bbfb83018fff98f5a9

Observation 58bacf8a-3c44-477c-ad9a-a7ebad505fcb · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix Reducing Activation Recomputation in Large Transformer Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.263658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.263658Z digest=sha256:f94a878ff140233cdead64b8cba4939f109a653f9ef58b40ac3fd9eeca4982cc

Observation 2b777e7d-da11-4b27-8c90-c084795671c0 · inbound

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization cites this paper.

QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization Reducing Activation Recomputation in Large Transformer Models

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-01T06:48:44.020572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T06:48:44.020572Z digest=sha256:0788216ef9d27c3e6a5860bbca00f0a65250647111d3e27612532b3204fff100