Pith. sign in

Paper Citation Record · LEDGER

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2502.02581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02581 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:46:02.443365Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T15:21:14.811340Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6afbf8c7-ff13-46c3-8d7e-bcefef5cfd93 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:01.990656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:01.990656Z digest=sha256:c9edaf5af24735581725f65dbf11d7eaa456f19f2923750bdc0ac6bd9fe184cc

Observation 9fd32696-03fc-4b39-b23d-d10c3b821586 · outbound

This paper cites Language Models are Few-Shot Learners.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Language Models are Few-Shot Learners

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:01.997188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:01.997188Z digest=sha256:412702d35cfc4aa08ff7e0cbd8932489df5a549317f32255457fccce5ac00c3e

Observation 5628e386-f491-4dce-a1eb-02f4c41bc096 · outbound

This paper cites Synthesizing optimal collective algorithms.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Synthesizing optimal collective algorithms

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.545534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.004525Z digest=sha256:0cbc9a81f7830cf5ae44f4d310eb2c2d36ddd0224a62b038ecb69c6975471675

Observation 57bd95d9-311d-4c25-ade8-559fbc328125 · outbound

This paper cites Collective communication on architectures that support simultaneous communication over multiple links.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Collective communication on architectures that support simultaneous communication over multiple links

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.529234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.009458Z digest=sha256:3d74f25780497df6294b710bfea33e1020f72a2895536225cd6a5e4097c57c21

Observation 977c2b10-19be-49a6-9600-a1cb15e195e1 · outbound

This paper cites Nvidia collective communication library (nccl) documentation.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Nvidia collective communication library (nccl) documentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.511967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.014155Z digest=sha256:5b5322792ad41295658f7626c8f21a47485154eee7b25f2908ebffead3d611fb

Observation 85bca8e6-d6ea-4baa-8904-9a569818ea3f · outbound

This paper cites Nvidia nvlink.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Nvidia nvlink

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.492736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.019109Z digest=sha256:5469ae173347498707208dd6348e2940d3b62b2b859e5060cc173d4c6b4b47f2

Observation 1d0393a0-ab1e-4225-b2ca-1de84f175a7b · outbound

This paper cites Nvswitch: The world’s highest-bandwidth on-node switch.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Nvswitch: The world’s highest-bandwidth on-node switch

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.475943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.025425Z digest=sha256:16ae1ebf13eba78ffc013fc408c6f9cb796fcaa3d8c5eeb870e301c2be71543b

Observation 4e170693-cb70-4908-8c6b-e35806918ba8 · outbound

This paper cites nvbandwidth.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism nvbandwidth

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.457963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.030330Z digest=sha256:937c88b37489a4a5705ebfbc1dc3af9200282bf4bd4f496ffbd58c154ca27a33

Observation 1e629f57-3436-4116-a0a8-d787339c03b5 · outbound

This paper cites GC3: An Optimizing Compiler for GPU Collective Communication.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism GC3: An Optimizing Compiler for GPU Collective Communication

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.035005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.035005Z digest=sha256:6f324779b4ac306599c44279907168f94cbd41081dcbbd0aa71aa32023420ed7

Observation 437bb618-352c-443c-887e-64bfc2c858b9 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.440544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.039997Z digest=sha256:ac0bb2556a8269a0340ec2ce7f72b47092c84eb13f46eb0ae1c29e52957cc57e

Observation cb877fab-4423-485c-990f-cf52b9e5368b · outbound

This paper cites an unresolved cited work.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:46:03.423636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.045180Z digest=sha256:78838e2963aa340f86de466c2ed29a7035ee8df8e1f57a8bd284fbce19cdbd22

Observation bdb830ed-304a-4f78-ba3d-8d523f39e5f7 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.050398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.050398Z digest=sha256:b40a2d0b1214f7333c220d8a3d9f45429a170220e585bc5ca897d7653ff2951a

Observation 64d4d098-17aa-4400-86ae-951212e75ae6 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Gpt-3: Its nature, scope, limits, and consequences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.055478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.055478Z digest=sha256:cd18a1913c67e269fce008c232a0d9b3318590a98cb7948598e0d9bc73c096b2

Observation 4e0de21e-2c55-46f3-928f-df9205b76c01 · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.060413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.060413Z digest=sha256:6d90cea1fa1fb957e3f3c9e8c01ae59879deeaea6cb11c8ee8934ce4a6a0de7f

Observation d7e91372-0037-4ca5-8ade-f6e98b1ee725 · outbound

This paper cites Tictac: Accelerating distributed deep learning with communication scheduling.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Tictac: Accelerating distributed deep learning with communication scheduling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.380538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.065369Z digest=sha256:7ddeadcf648423d84f04f43ef8e7a44f37997a9f7cc2d1e36390f902c151916a

Observation 1f317f8c-19ed-4d31-a38c-bf082760d153 · outbound

This paper cites Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.363283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.069913Z digest=sha256:90f54a9f9195cba97970d66eaf7511d429db656b45f4f52cd066682b86640ea0

Observation 09e6bfb6-5805-4f68-8dbc-7386387cbd19 · outbound

This paper cites Rae, and Laurent Sifre.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Rae, and Laurent Sifre

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.333659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.080238Z digest=sha256:507efbbe37b08890bd38300a9226d26422aebae06d3e1e943baf22d5e56cfd96

Observation 58cd5668-5305-4f1a-941d-68138c166aa5 · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Tutel: Adaptive Mixture-of-Experts at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.086074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.086074Z digest=sha256:52adec6a7d1789cae3077d4c5c98de4b13cc27353bceb67dddcd6b5e1d59c75d

Observation 2a8bdfe5-2b05-451c-ad6a-e5cc15fb3cba · outbound

This paper cites Adaptive mixtures of local experts.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Adaptive mixtures of local experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.091113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.091113Z digest=sha256:9153d61cfc5abb16c1b1ba1a5cd9aad59a915c468f3756d09da4670defd2cff6

Observation fc9c3419-385b-48e0-b497-5adb1eeb4692 · outbound

This paper cites Scaling Laws for Neural Language Models.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Scaling Laws for Neural Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.095641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.095641Z digest=sha256:17bb2a647172e7fe3412d00d9b2a82b74a1986c5ede4c51dfe2423182db86f38

Observation 8c5ba881-04a1-44ec-b48e-3440cc4ae403 · outbound

This paper cites Tccl: Discovering better communication paths for pcie gpu clusters.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Tccl: Discovering better communication paths for pcie gpu clusters

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.299475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.100541Z digest=sha256:4319149be97cec70d17bafab291a600ed3ce6f279984632e805c4c6a5672c4a9

Observation ca30462b-e555-4888-808e-e855ee751eeb · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Adam: A Method for Stochastic Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.105197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.105197Z digest=sha256:ac0f01d265d45425827562d26308de82b7eb2545d3bf5520e90be310346f17ab

Observation da66578a-6978-463c-843a-ff2a382abbeb · outbound

This paper cites Breadth-first pipeline parallelism.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Breadth-first pipeline parallelism

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.281816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.110528Z digest=sha256:f898f98915c599f80b29c0cef9c5b85f359a178241ec6a1411b18e4cac3d1190

Observation b2e9f84d-4126-408f-a717-4bdcbbd96b9d · outbound

This paper cites A theoretical framework for back-propagation.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism A theoretical framework for back-propagation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.263955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.115322Z digest=sha256:f9e05dfc21cc53884bccbd522d15ec30ac679949f2a52a4cf945f3e2ff87b1a2

Observation 892ae4e6-1a52-447b-8e59-46a90d83b7c7 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.120842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.120842Z digest=sha256:ed0112bd8b7c4589de97d575a3660d412ecc9ff9a7be72b1cc35e5644163409e

Observation 15cb7034-8632-49a7-a5aa-702d4b59136f · outbound

This paper cites Accelerating Distributed MoE Training and Inference with Lina.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Accelerating Distributed MoE Training and Inference with Lina

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.126145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.126145Z digest=sha256:420afcbfcfb2d0dd9b756795543fb7f05706cb0abf4876776446811ab6417ead

Observation 21d50059-2a23-4d8f-befe-4779d71db8a5 · outbound

This paper cites Pytorch distributed: experiences on accelerating data parallel training.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Pytorch distributed: experiences on accelerating data parallel training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.244639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.131223Z digest=sha256:ccd4281ebfd7859414cbd59b80348b2bf9d4b03c4457f14d9a93a7309ca76b42

Observation 601f99d2-cc61-46dc-b521-29152893b821 · outbound

This paper cites Near-optimal sparse allreduce for distributed deep learning.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Near-optimal sparse allreduce for distributed deep learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.226033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.136345Z digest=sha256:70def94e2624d12947acd787919853a099876e46ad49b0cd2205ab221fd0cee7

Observation 899e9e92-2a2a-4562-be46-d54d52cc1865 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Swin transformer: Hierarchical vision transformer using shifted windows

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.141382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.141382Z digest=sha256:e760d8e69f8cc7ed22ccdefca00e97ef95b6df66e1425e5f8ae2eac641a09a57

Observation 8a5257b0-377a-4f9a-9114-83b429fb38ce · outbound

This paper cites Mixed Precision Training.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Mixed Precision Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.146593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.146593Z digest=sha256:c81dd4424b999bcbd8f458751f74be13530909b2f5f4487c6b4f607e06f6e497

Observation 8d6eeb2b-336f-4b34-bd9b-d51f726bcd0d · outbound

This paper cites FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:46:02.756735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.152961Z digest=sha256:a9f7b5c66a2610e9c8c7f2a6cdfb54c000ac900de06c8c87b1dad614aeac9118

Observation 5b4f2119-fa2c-436c-8165-478e6eb393b3 · outbound

This paper cites PyTorch: an imperative style, high-performance deep learning library.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism PyTorch: an imperative style, high-performance deep learning library

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.195901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.158390Z digest=sha256:d114c1823f3d47a887018ca5a9e5cf65c81b3f3e120dee418751d46aa2b71a69

Observation 9ff58927-d358-4126-8217-cc52294f4b69 · outbound

This paper cites Sparse gradient commu- nication with alltoall for accelerating distributed deep learning.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Sparse gradient commu- nication with alltoall for accelerating distributed deep learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.178185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.163597Z digest=sha256:6b31cd78e57c45517c8e743fbe53722bda54e8c110a28e123b528efede93f455

Observation 576399ce-7b6c-4e33-aa5d-7bf577a6def2 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter mod- els.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Zero: Memory optimizations toward training trillion parameter mod- els

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.159640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.168760Z digest=sha256:1f1961f5e0cb4f2f62a8c7645d0c5fbb6f026aa742c5d568d780d4c380d9e585

Observation 7a748b57-b433-419d-b6c3-9de9df80c68d · outbound

This paper cites Sparcml: High-performance sparse commu- nication for machine learning.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Sparcml: High-performance sparse commu- nication for machine learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.140495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.173658Z digest=sha256:c95a3c6cb6d90acc018e727e0f8788baf8faca44cac0d4679b1426d80458c6d7

Observation a7e94bb4-8652-484b-83c5-4683ec87ffb2 · outbound

This paper cites Scaling vision with sparse mixture of experts.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Scaling vision with sparse mixture of experts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.121110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.178601Z digest=sha256:3d4fa987c99334ec1510b28685d128fa430e16b4b14ed53165aaceb40a163f76

Observation 32ccdbec-524d-4193-9d47-20ceb15e5895 · outbound

This paper cites In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , pages 593–612, 2023.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) , pages 593–612, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.102932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.183860Z digest=sha256:f9d7063ca89b5ff23ff6099531da78774d2e12725d2c0a791748e97ed9683c98

Observation 60adf970-5084-420c-ba29-557373550f75 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.200462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.200462Z digest=sha256:e9d6cb49e639784d528791cac21f434c2224432fafc526f7d2ab6defc53761c7

Observation bf8e3f55-e30d-48b1-9d4e-8c85a64ce0f9 · outbound

This paper cites A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.083615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.240697Z digest=sha256:8ec0c4591078ce17943d9b48f4619d262bc422fbda8c6d2d4c9dae23225b64e7

Observation 57ebab43-fcdc-4c37-b205-74058d648e6a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.319829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.319829Z digest=sha256:81a6c832d807e56dd6b6206a527a9ab42779dd211039dad64d734d914dd7f6c9

Observation 02b6788e-478f-45c4-acd5-b802f84f52b3 · outbound

This paper cites Attention is all you need.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Attention is all you need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.358203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.358203Z digest=sha256:bee4175f495f1fcb8a27bf51119f11942ae7f6efe6f497d6c6f78973ef20302e

Observation ac189ba9-da3a-41c4-8173-30517377c4f2 · outbound

This paper cites Blink: Fast and generic collec- tives for distributed ml.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Blink: Fast and generic collec- tives for distributed ml

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.041674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.389570Z digest=sha256:490c2e8d74424857b888122e6dc5265f86e7045d0894c9b8a333e3bf6d659768

Observation d1f8c0d5-94af-4e65-9bf9-44645d9b6207 · outbound

This paper cites Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.405892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.405892Z digest=sha256:69d2407919b87759dec0bfb96904165cb1b58384458601d0075af92161eb1d48

Observation 61544aa1-fba5-48d2-8d4f-61d4b4df0d83 · outbound

This paper cites Smartmoe: Efficiently training sparsely-activated models 13 through combining offline and online parallelization.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Smartmoe: Efficiently training sparsely-activated models 13 through combining offline and online parallelization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:03.017349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.433263Z digest=sha256:3277cf5a51c94e2e25f3ec28d24a15d596dd0ac99cc35850fb8a21897019cb36

Observation 2285d163-d4da-4cb3-9249-8677ba2014b1 · outbound

This paper cites Spardl: Distributed deep learning training with efficient sparse communication.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Spardl: Distributed deep learning training with efficient sparse communication

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:02.993685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.438636Z digest=sha256:df6eef82f88e9b473ae35539a42464ea96670474d4b906ebe146378af097d224

Observation 09d39c1d-e199-4f60-a730-25bb402af00d · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Pytorch fsdp: Experiences on scaling fully sharded data parallel

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:46:02.973573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:46:02.443365Z digest=sha256:b277d057f1980d01097fb5a1bbd40809298d804b8a135a9d2f0834b348fc548c

Observation 3e444596-3129-40f2-a79a-8fdce25b2642 · outbound

This paper cites an unresolved cited work.

Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T11:46:02.074884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:46:02.074884Z digest=sha256:9073e0c54112b3ef6962b0f94802271eeb4d467be0b78cace0ee13be6137e5b9

Pith citing papers

Observation 381dc057-b796-4990-8d4d-c637d87e2884 · inbound

Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU Platforms cites this paper.

Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU Platforms Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-11T15:28:05.291593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T15:21:14.811340Z digest=sha256:d83c0f8e7a53342a9031be914dd103efa12c254c5904ee89ac2f8e5cb9d9ea2b