Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:46:02.443365Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2502.02581.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:46:02.443365Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T15:21:14.811340Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
47 of 47 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 6afbf8c7-ff13-46c3-8d7e-bcefef5cfd93 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fd32696-03fc-4b39-b23d-d10c3b821586 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Language Models are Few-Shot Learners
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5628e386-f491-4dce-a1eb-02f4c41bc096 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Synthesizing optimal collective algorithms
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57bd95d9-311d-4c25-ade8-559fbc328125 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Collective communication on architectures that support simultaneous communication over multiple links
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 977c2b10-19be-49a6-9600-a1cb15e195e1 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Nvidia collective communication library (nccl) documentation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85bca8e6-d6ea-4baa-8904-9a569818ea3f · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Nvidia nvlink
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1d0393a0-ab1e-4225-b2ca-1de84f175a7b · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Nvswitch: The world’s highest-bandwidth on-node switch
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e170693-cb70-4908-8c6b-e35806918ba8 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism nvbandwidth
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e629f57-3436-4116-a0a8-d787339c03b5 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism GC3: An Optimizing Compiler for GPU Collective Communication
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437bb618-352c-443c-887e-64bfc2c858b9 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cb877fab-4423-485c-990f-cf52b9e5368b · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bdb830ed-304a-4f78-ba3d-8d523f39e5f7 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d4d098-17aa-4400-86ae-951212e75ae6 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Gpt-3: Its nature, scope, limits, and consequences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0de21e-2c55-46f3-928f-df9205b76c01 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e91372-0037-4ca5-8ade-f6e98b1ee725 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Tictac: Accelerating distributed deep learning with communication scheduling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1f317f8c-19ed-4d31-a38c-bf082760d153 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09e6bfb6-5805-4f68-8dbc-7386387cbd19 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Rae, and Laurent Sifre
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58cd5668-5305-4f1a-941d-68138c166aa5 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Tutel: Adaptive Mixture-of-Experts at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8bdfe5-2b05-451c-ad6a-e5cc15fb3cba · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Adaptive mixtures of local experts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9c3419-385b-48e0-b497-5adb1eeb4692 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Scaling Laws for Neural Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5ba881-04a1-44ec-b48e-3440cc4ae403 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Tccl: Discovering better communication paths for pcie gpu clusters
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca30462b-e555-4888-808e-e855ee751eeb · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Adam: A Method for Stochastic Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da66578a-6978-463c-843a-ff2a382abbeb · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Breadth-first pipeline parallelism
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b2e9f84d-4126-408f-a717-4bdcbbd96b9d · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism A theoretical framework for back-propagation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 892ae4e6-1a52-447b-8e59-46a90d83b7c7 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15cb7034-8632-49a7-a5aa-702d4b59136f · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Accelerating Distributed MoE Training and Inference with Lina
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d50059-2a23-4d8f-befe-4779d71db8a5 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Pytorch distributed: experiences on accelerating data parallel training
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 601f99d2-cc61-46dc-b521-29152893b821 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Near-optimal sparse allreduce for distributed deep learning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 899e9e92-2a2a-4562-be46-d54d52cc1865 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Swin transformer: Hierarchical vision transformer using shifted windows
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5257b0-377a-4f9a-9114-83b429fb38ce · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Mixed Precision Training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d6eeb2b-336f-4b34-bd9b-d51f726bcd0d · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b4f2119-fa2c-436c-8165-478e6eb393b3 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism PyTorch: an imperative style, high-performance deep learning library
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ff58927-d358-4126-8217-cc52294f4b69 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Sparse gradient commu- nication with alltoall for accelerating distributed deep learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 576399ce-7b6c-4e33-aa5d-7bf577a6def2 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Zero: Memory optimizations toward training trillion parameter mod- els
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7a748b57-b433-419d-b6c3-9de9df80c68d · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Sparcml: High-performance sparse commu- nication for machine learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a7e94bb4-8652-484b-83c5-4683ec87ffb2 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Scaling vision with sparse mixture of experts
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 32ccdbec-524d-4193-9d47-20ceb15e5895 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , pages 593–612, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 60adf970-5084-420c-ba29-557373550f75 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8e3f55-e30d-48b1-9d4e-8c85a64ce0f9 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57ebab43-fcdc-4c37-b205-74058d648e6a · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b6788e-478f-45c4-acd5-b802f84f52b3 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Attention is all you need
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac189ba9-da3a-41c4-8173-30517377c4f2 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Blink: Fast and generic collec- tives for distributed ml
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d1f8c0d5-94af-4e65-9bf9-44645d9b6207 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61544aa1-fba5-48d2-8d4f-61d4b4df0d83 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Smartmoe: Efficiently training sparsely-activated models 13 through combining offline and online parallelization
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2285d163-d4da-4cb3-9249-8677ba2014b1 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Spardl: Distributed deep learning training with efficient sparse communication
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09d39c1d-e199-4f60-a730-25bb402af00d · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Pytorch fsdp: Experiences on scaling fully sharded data parallel
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3e444596-3129-40f2-a79a-8fdce25b2642 · outbound
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Unresolved cited work
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381dc057-b796-4990-8d4d-c637d87e2884 · inbound
Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU Platforms Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.