Pith. sign in

Paper Citation Record · LEDGER

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2412.14335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14335 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:19.004428Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:27:43.845408Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:30:33.076250Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 890877d4-5f77-49ac-b521-e89d0ff47993 · outbound

This paper cites A bridging model for parallel computation,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines A bridging model for parallel computation,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.468135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.468135Z digest=sha256:1765085294030d3dc933de28ab9a1a4c5ed1f73f5f27ee7512d894488c408306

Observation d1a0b163-db15-42ec-8fb3-c4784c7ce237 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.500285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.500285Z digest=sha256:0f38d044d6db0aeed0e2c6ed506e269a9246e1c55a78a84149a511216e8e2dea

Observation 3c7b4e4f-8200-4234-ad70-3c5a017746e5 · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.596808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.596808Z digest=sha256:003fad1b73a098bb96bf80b4a8a76797043ed0806e20da3fe225116f69ecc6b8

Observation c32deaf5-bd94-4184-bdab-1ec72303a5b3 · outbound

This paper cites The Llama 3 Herd of Models.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.637910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.637910Z digest=sha256:3dbfb425ba9a9d7d1839873a7870c5dafcd847d6075c56d7517ed9fbcedbe2eb

Observation 1ae43603-382f-4fc8-968c-5d3e08f8b168 · outbound

This paper cites {ARK}:{GPU-driven} code execution for distributed deep learning,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines {ARK}:{GPU-driven} code execution for distributed deep learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.799105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.651544Z digest=sha256:9c1e3982c924339a19df76c76c0989f8d15714ff9a77964c704746a80ab9b77c

Observation e6ebe5e3-331a-4d90-98ad-863da059bdbe · outbound

This paper cites 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.783250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.658910Z digest=sha256:b3748476701c195b5b85d3628cfc32a7dcfda0170b94ae6af69572237f34274c

Observation 6a1f410e-90f4-4d51-b122-aa53bde69838 · outbound

This paper cites AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.772487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.670888Z digest=sha256:3302345f9816bb7cdf0e380305a32396ce9b425a8e51583ce16f24f1a645c4cc

Observation 1dd6be83-e635-4c87-9d31-354fad14baf0 · outbound

This paper cites The AMD CDNA™ 3 architecture,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The AMD CDNA™ 3 architecture,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.761505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.681382Z digest=sha256:b30f46444bb92027bc0057e5c40b61d16d931ecbce4d858e9fe39474939cb53a

Observation c75bc5a0-a566-4c90-846e-9ad51407dc64 · outbound

This paper cites HSA Runtime API and runtime for ROCm,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HSA Runtime API and runtime for ROCm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.750820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.690115Z digest=sha256:15d06368c982a4e03c31628a2a2cde25bffc3c0dd8e619ed3756c2f838cc0ff9

Observation ba5d4042-d53e-4154-9a86-8291b703c223 · outbound

This paper cites HIP: C++ Heterogeneous-Compute Interface for Portability,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HIP: C++ Heterogeneous-Compute Interface for Portability,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.739086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.698583Z digest=sha256:e91b6d9f6a3987118ef6fda4e2c274f50c3bf4fc3d888d212f27779fabcd65a7

Observation 243312cf-886f-423c-86bb-63307a79a35f · outbound

This paper cites Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.727367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.703200Z digest=sha256:01601161195604b4d62ace6528909388f3e941ce85d0d0f6ffca9fdea43cba34

Observation 35802812-fcae-49fc-96f9-f86b8d45709e · outbound

This paper cites AMD ROCm™ Software,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD ROCm™ Software,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.716251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.707350Z digest=sha256:b120584367710f5aa23ffa84085779dea18d1b16b3dac03908863be22652fd9e

Observation f858a095-6ea6-4c4f-8ac3-3a7610b2952d · outbound

This paper cites ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.703069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.711930Z digest=sha256:1c7b630910870ddeee508387e3ba65669cf2834d45e084f8b906e86f630f2d75

Observation 66cd9b53-7e39-4089-b23f-8874fe049d04 · outbound

This paper cites ROCm Communication Collectives Library (RCCL).

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm Communication Collectives Library (RCCL)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.690679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.716063Z digest=sha256:2f1880cc24f475164c6ca502d4476c59adfc40cfb1c598b1dad1ab921156eb51

Observation 6927e377-6619-403c-9cde-3c0acd2c1728 · outbound

This paper cites ROCm: HIPStream,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: HIPStream,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.679660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.720394Z digest=sha256:29a6c10a3d24d5151259c300d29a169ae1196e883fb07a5ac4f3b4992099db35

Observation 5cdb2bd9-e3db-4143-9029-4dceaab0961b · outbound

This paper cites rocprof — ROC Profiler Documentation,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines rocprof — ROC Profiler Documentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.668818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.724462Z digest=sha256:7886a52f9cb431bfe79a68a1e3a28919b194b527ea6dd63f8cccd1f2707a0e6e

Observation 2c930b0c-a81c-455d-aa56-f678e87621f8 · outbound

This paper cites AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.658083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.728339Z digest=sha256:904c75c3e6d0370ef97258a533c0b8f9a4cbbf2b48b238c536d7fc854fc27297

Observation a8f051b4-d58a-4f55-9839-7cdb3ccaf05a · outbound

This paper cites ROCm: ROCR-Runtime,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: ROCR-Runtime,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.647582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.732671Z digest=sha256:4a321c4cfefa1989fb31a9ab2e7008bd2faec03308efe0cf6da4f48112e625f0

Observation 19addbb6-178a-457f-b8f0-8dbceae946c7 · outbound

This paper cites T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.637138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.736987Z digest=sha256:ac942904cdc4532838e57e377068d79d43797b8ac10aa44bdcc7a7e30d343979

Observation f1ba0c67-19ca-41da-a15f-04e42cd467ac · outbound

This paper cites DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.741039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.741039Z digest=sha256:a6767c7a79e3f396b2920cdddd821bf4685364817e856052dd36d267573edb3a

Observation 918f3ca6-b710-4094-97a6-603550ca0fb0 · outbound

This paper cites TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.626380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.745563Z digest=sha256:5d8148fe94ca081f3ef57f51769b17ab1926e076ffbce5124be1538283f17baf

Observation 79aa09f4-c0dc-4557-bc57-6e0ea243859b · outbound

This paper cites MSC- CLang: Microsoft Collective Communication Language,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSC- CLang: Microsoft Collective Communication Language,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.615472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.749301Z digest=sha256:58b1586e276ab08a2c158c37afd5d6d9fb6a1d7f72acfc708554219d49a758aa

Observation b6541107-eac4-40d7-9ec4-7267ff685389 · outbound

This paper cites Available: https://github.com/NVIDIA/nccl.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Available: https://github.com/NVIDIA/nccl

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.604358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.752792Z digest=sha256:d6730206048d8ad1b6491be8d746532428148bb9cc605e3fe3d3fe6a82541fe0

Observation d7372df3-5cb9-4ae8-828a-430e2ab53fb6 · outbound

This paper cites TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.593400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.756341Z digest=sha256:4c1de2a00f8d88195353138a4d59ec7b146de111d19be00f9c207d14dda49495

Observation 5c965cdf-de6a-4e2e-a15c-201a20225f60 · outbound

This paper cites MSCCL++: A GPU-driven communication stack for scalable AI applications.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSCCL++: A GPU-driven communication stack for scalable AI applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.581937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.795376Z digest=sha256:2433312493b59b0e9eb6ce588a7f1eb64bb3d6c02954acc6e4e840100a02adbf

Observation 33a9f01b-36cd-481c-abf5-2b04b1c5de54 · outbound

This paper cites Deep Learning Recommendation Model for Personalization and Recommendation Systems.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Deep Learning Recommendation Model for Personalization and Recommendation Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.838572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.838572Z digest=sha256:1b36d18af73acca1370c4db7fffa26ec9b899bbe71260a30aa85e72dc72acc8e

Observation c3c9734d-8040-4f14-ac61-bcea09cbcb98 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.570814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.888958Z digest=sha256:c7fbc79afc9ce925c229a464a982a94ed67db2c50efdc8d5592061c749ffdf49

Observation 93de37ab-023d-4849-9713-109d1645eaee · outbound

This paper cites Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.917412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.917412Z digest=sha256:727017165cbc96d3b264f0bf61ff3b5629ef15006d0afa573d33e3ad90224a56

Observation febf812d-6ccd-40b8-b15f-3c2bfb32a186 · outbound

This paper cites The Case for GPGPU Spatial Multitasking,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Case for GPGPU Spatial Multitasking,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.559159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.936810Z digest=sha256:9aeb4108faa551e86d988acc0274b523fc10b95cd2634c13b2affa21e7337816

Observation f1ed9c9e-0e3f-42c4-9330-7a15f54fc25a · outbound

This paper cites Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.546408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.941351Z digest=sha256:83a16939246ea2997ee1b4b8cad2f33dfee22f67de9a08c251f3f33472e445a8

Observation 0de2b38d-dcb5-4768-81ea-ba7fe669f6cb · outbound

This paper cites Improving GPGPU Concurrency with Elastic Kernels,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU Concurrency with Elastic Kernels,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.533927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.945764Z digest=sha256:b55e1647c66c60e946451f13a5a0760fbc49dce98a9351b3d61ae49d2c4f7017

Observation 6ba035e1-6067-4d66-82c6-0c07a7857b84 · outbound

This paper cites Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.521377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.950054Z digest=sha256:8e0657a941e45e7c01db7428d2fe4e67c9b4af54c695308126fe6778eb07c4f1

Observation a3a604e8-475b-4bff-abd0-253409ab8303 · outbound

This paper cites Global Optimizations & Lightweight Dynamic Logic for Concurrency.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Global Optimizations & Lightweight Dynamic Logic for Concurrency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:23:19.250779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.954391Z digest=sha256:802663399f25b68525c3cfef05b59804d50136b90736fa7b5d44701de4ef3077

Observation fb53f232-3849-4ca5-ae25-0ac468d29a7e · outbound

This paper cites An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.508949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.966166Z digest=sha256:4b47f25efb09d9b0628e2c67bf36b1550b7c62251b229596bf764d709c89d185

Observation 1925a9da-25db-43c4-948d-e6befcaffe78 · outbound

This paper cites Introducing Async Tensor Parallelism in PyTorch,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Introducing Async Tensor Parallelism in PyTorch,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.496119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T12:23:18.970195Z digest=sha256:b49149131799993009fdeccff9f726a339f8a6ce065f7b4381afd1391dbbd714

Observation 0a16a08d-0815-4c5e-98c6-19f098eae38b · outbound

This paper cites Optimizing Distributed ML Communication with Fused Computation-Collective Operations.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Optimizing Distributed ML Communication with Fused Computation-Collective Operations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.976636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.976636Z digest=sha256:10783c604e4a575b9098dc5b2f4e05dabcab4e6a515f29c43ad1b27f511264fd

Observation 25d52610-b9b8-4db3-917c-45ac4402b6f3 · outbound

This paper cites Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.988507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.988507Z digest=sha256:22dfb5b3df80727f7b31081e35fdd48cfa27b363c76bdb3aa1a568b5b3dfaf92

Observation 54b67e22-4c11-4532-a4be-47ce55e0e45e · outbound

This paper cites Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:19.004428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:19.004428Z digest=sha256:c3542308fdbf2da7c9897234bad13a0a94c88c2acd0d8f70969b922e678a500c

Observation babe03a7-517b-4b39-9cf5-f1d3e8580eb1 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.559962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.559962Z digest=sha256:70a58f6c8410c52a2449a2c78b5df2b4dd19bb02ef1cab65b5c05ecc285aa1ab

Pith citing papers

Observation a6d06ad2-aa27-4edc-8a66-4da6323f8d25 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.078708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:2eab2da338e8471ec295eb56dd5ff6e9c6d87c7d5c97595658b49a3d6193bf46