Pith. sign in

Paper Citation Record · LEDGER

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2509.03047.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03047 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:37:58.541062Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:19:41.302164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-09T06:35:39.100050Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5945f027-025d-45a8-a696-ee649e690a1b · outbound

This paper cites GPT-4 Technical Report.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.416767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.416767Z digest=sha256:e38a9fb0d01956445818aa79b63e9895b18de0e5aba6a6653c5621f6cb502208

Observation d416538b-b179-4ec1-9326-55b5a5d5d332 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.422548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.422548Z digest=sha256:b2919c98b402c6ad0a8c48e8cd7b44dc5805fa8850cb8323301fe3b68fdcdfed

Observation 23de1fb2-58d3-49ec-9c53-0be202d09fcd · outbound

This paper cites DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.427111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.427111Z digest=sha256:7439e004323a7bbc4b3ad1f118bad33e63be14b82f7d98396fdb3a7b60d10ed8

Observation 7394f22f-0879-4349-a63c-76fa634708cb · outbound

This paper cites Scaling Laws for Neural Language Models.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Scaling Laws for Neural Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.431587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.431587Z digest=sha256:a23cd4e6388c876e61a73da2a48df6a51aef9ecd8a7589459a0cfd248497132b

Observation 2a14e99e-cf18-4ad1-a406-8d315e75dd62 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Zero: Memory optimizations toward training trillion parameter models,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.436305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.436305Z digest=sha256:2e408e5de014a5d6ef9306a33ecdc608aafb0b3597e5c9f15da2abe456d39273

Observation cd7afc78-b4d5-40a2-b8c3-0aca8c42ff7b · outbound

This paper cites Colossal-ai: A unified deep learning system for large-scale parallel training,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Colossal-ai: A unified deep learning system for large-scale parallel training,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.440176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.440176Z digest=sha256:f6aefff2f76dc0469a16a0d2cea32bcf9720ff98bb0de8807f9d74937195ddc4

Observation 0bb3527d-0304-4d70-b0bc-1a87b6882214 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.444382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.444382Z digest=sha256:eef720106c1035627e95de4f418f0d0e0359144fd54519651d78798657e93b51

Observation 81eb126e-fcd0-4f55-99bb-2aee298e55fe · outbound

This paper cites Bloom: A 176b-parameter open-access multilingual language model,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Bloom: A 176b-parameter open-access multilingual language model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.930683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.449250Z digest=sha256:983e2a08806c2c9a8097514ca54fb903eb44c18777d9661294186c4378de7e80

Observation 697d7c74-9ef7-41db-baf2-9d65b3ef7695 · outbound

This paper cites Metaseq: Opt-175 logbook,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Metaseq: Opt-175 logbook,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.915447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.453379Z digest=sha256:f44c17589441c4a59a3bc55bd2c88e3d59f1081f034a7a17b100a87d0ebdfb8c

Observation 99ad5573-e904-4e12-a6dd-3d2cbd54a51b · outbound

This paper cites The Llama 3 Herd of Models.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.457341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.457341Z digest=sha256:f6a644c04000b942c3734ed0fe55d2ef397c31b5695fd0d07bacd3b226a257c7

Observation 4a8d17c5-fd21-4d1b-8e63-d7dd21deea3b · outbound

This paper cites Evaluating disaster recovery plans using the cloud,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Evaluating disaster recovery plans using the cloud,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.902569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.461286Z digest=sha256:0ddab439aefb32b4105255cce773efeb1f117814c40fe9cfdcd02e8080e2e3d5

Observation ffec4c98-d280-4001-ace9-a7874146a622 · outbound

This paper cites Just-in-time checkpointing: Low cost error recovery from deep learning training failures,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Just-in-time checkpointing: Low cost error recovery from deep learning training failures,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.889992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.465254Z digest=sha256:e651eca46018fe54d81b45d2f8f044a764f628d227ad5621aeecf8353b7ba777

Observation 365391e2-5239-4387-994b-9bdfdc4dc666 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.876966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.470233Z digest=sha256:9334373e098cd722e7cab91c5316e539a4f96bca44dd3e00bc2c91629980d87a

Observation 1f9c979b-47d3-45ee-bce1-34b477f45b6b · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Gpipe: Efficient training of giant neural networks using pipeline parallelism,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.863136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.474410Z digest=sha256:77d9269571f832a3875deba0901e58e20a0564d11199c79014b37b148e4b547a

Observation 753678ce-7c28-4060-a522-56537fec87eb · outbound

This paper cites Reducing activation recomputation in large transformer models,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Reducing activation recomputation in large transformer models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.478336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.478336Z digest=sha256:2a2d2825e81fe8c6b4eec1de3ae28f57b8e7cde4d0363d0a9eb8a7d15d139c9f

Observation a54368dd-c1f5-4f29-91a4-7f5e4f5482f9 · outbound

This paper cites TRANSOM: An Efficient Fault-Tolerant System for Training LLMs.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs TRANSOM: An Efficient Fault-Tolerant System for Training LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.482820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.482820Z digest=sha256:1f150ada33636b04030ad0823c343cbc79743cad1e91caa1890461685604d728

Observation f4186c31-4c0e-4fa4-9eae-8c48b4e53c22 · outbound

This paper cites Unicron: Economizing Self-Healing LLM Training at Scale.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Unicron: Economizing Self-Healing LLM Training at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.489021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.489021Z digest=sha256:e35886024cc0543f87acfd416d5f0b7cbc72e477a951b0b658f57803fd5388a1

Observation 9fd2580d-70aa-4df0-ba9f-ccb23c9aeb28 · outbound

This paper cites Megascale: Scaling large language model training to more than 10,000 gpus,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Megascale: Scaling large language model training to more than 10,000 gpus,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.842634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.494421Z digest=sha256:62804a43ad194edf34a9e61a299f63521822c10d93873dceacc83c17098d0ec6

Observation 870a91f6-2235-4c16-9ad9-63767c299acd · outbound

This paper cites MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-15T16:37:58.624673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.498533Z digest=sha256:e2c7f482753fe2f94bc63e5cd58ea1f7b6ed4891f5756f4305048ff753f45414

Observation 9d05668e-6d71-4f22-80a7-4fda4f38d769 · outbound

This paper cites Efficient Training of Large Language Models on Distributed Infrastructures: A Survey.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Efficient Training of Large Language Models on Distributed Infrastructures: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.502783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.502783Z digest=sha256:234acc5c01410e77bd4ec4f024107fa296ec6e034855ef7ded6b1191082a0aa6

Observation c6593d39-28fc-48a9-91ce-6c52a1f311d8 · outbound

This paper cites Datastates-llm: Lazy asynchronous checkpointing for large language models,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Datastates-llm: Lazy asynchronous checkpointing for large language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.828497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.506844Z digest=sha256:7aeffa72ce34f7d51eb818f630c9396151339d3182cae0b877457f72a9ab7fa9

Observation 3400d5c3-3a11-44cd-b6f3-bceb8f94e5a2 · outbound

This paper cites Checkfreq: Frequent, fine-grained dnn checkpointing,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Checkfreq: Frequent, fine-grained dnn checkpointing,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.815194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.511568Z digest=sha256:f7061834036c5c05a1fa549da1d0147e143316cbdf1dd249dc8364146ba63463

Observation 693b984d-08ca-408a-9685-53e4135faba7 · outbound

This paper cites A cost-efficient failure-tolerant scheme for distributed dnn training,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs A cost-efficient failure-tolerant scheme for distributed dnn training,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.796893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.515788Z digest=sha256:0acb5776c74a845e6a0357217319aad044b1166d66e02febf5a07b09e483fcae

Observation fcf4fbf8-d4a0-4694-9b70-a1390445ff05 · outbound

This paper cites Check-n-run: A checkpointing system for training deep learning recommendation models,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Check-n-run: A checkpointing system for training deep learning recommendation models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.780370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.519748Z digest=sha256:ad267548eaf44d9432c233ca2ec0aab3164a4e9e29375c720052064c9fc8f538

Observation 72e414b3-4563-4f4b-b95e-a6f5e742febb · outbound

This paper cites Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.765530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.523822Z digest=sha256:e13058f1118c6c502d28a16991db31dcd7c83b00a7de3206a2ee5b06c58b9cd3

Observation 1b9b73ab-1fcd-48c7-9497-36ee3aeaa9d4 · outbound

This paper cites Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.527618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.527618Z digest=sha256:379f1f3c84798d1ce53d9b60025d22c65e4ead6a454faf8ed5145dbffbbbbaa3

Observation 4f10ff25-b57c-45cf-a70a-35652420a2f8 · outbound

This paper cites ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:37:58.532754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:37:58.532754Z digest=sha256:f5996ecb98cd8c23290398880841e3ac8336d6883c1b3f4c5e19a90bfd28631e

Observation e09f494b-7704-4f35-a878-674f3cffa4ed · outbound

This paper cites Swift: Expedited failure recovery for large-scale dnn training,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Swift: Expedited failure recovery for large-scale dnn training,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.744392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.536768Z digest=sha256:c36877b4d5eb9ef205afb7c168c5934b3aeaeecf8911d914fa0d089324585762

Observation a7aeba1a-b462-4e97-9311-11c64d7ec529 · outbound

This paper cites Parcae: Proactive,liveput-optimized dnn training on preemptible instances,.

FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs Parcae: Proactive,liveput-optimized dnn training on preemptible instances,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:37:58.731224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T16:37:58.541062Z digest=sha256:dd1dd346ae1cb38188f6f8657e6e08767de11071e719667abb4bc38999c4c2b3

Pith citing papers

Observation 7ac5040a-21ae-43fc-a248-ea58525c51d2 · inbound

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning cites this paper.

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.101920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T18:19:41.302164Z digest=sha256:e9428361b9ef68bf06620976b780e950433e9a765af02398d7f5a099f5936e33