Pith. sign in

Paper Citation Record · LEDGER

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2412.14335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14335 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:19.004428Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:27:43.845408Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:30:33.076250Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 890877d4-5f77-49ac-b521-e89d0ff47993 · outbound

This paper cites A bridging model for parallel computation,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines A bridging model for parallel computation,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.468135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.468135Z digest=sha256:280707a9b1d025bfe25d36eba21ccfd93527ec8b469ec5b517d8a0cc19d000cf

Observation d1a0b163-db15-42ec-8fb3-c4784c7ce237 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Pytorch fsdp: Experiences on scaling fully sharded data parallel,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.500285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.500285Z digest=sha256:ef90b5a33a372235b84ea19969f2cfe0cf9813ff3826496e490e89d39f7bf747

Observation 3c7b4e4f-8200-4234-ad70-3c5a017746e5 · outbound

This paper cites NanoFlow: Towards Optimal Large Language Model Serving Throughput.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines NanoFlow: Towards Optimal Large Language Model Serving Throughput

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.596808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.596808Z digest=sha256:e12d649632b039aaab1cbc1366782c0b08014af95c3f22ed1cc11ffc399f39bf

Observation c32deaf5-bd94-4184-bdab-1ec72303a5b3 · outbound

This paper cites The Llama 3 Herd of Models.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.637910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.637910Z digest=sha256:61811538a708aa9b6b6e96ac72ecb0055634806d67ce3467c742b809636f2c63

Observation 1ae43603-382f-4fc8-968c-5d3e08f8b168 · outbound

This paper cites {ARK}:{GPU-driven} code execution for distributed deep learning,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines {ARK}:{GPU-driven} code execution for distributed deep learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.799105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.651544Z digest=sha256:0c541153c43520dd303be3eb0db39d49128ed4f6248069141be684f37e06f588

Observation e6ebe5e3-331a-4d90-98ad-863da059bdbe · outbound

This paper cites 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.783250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.658910Z digest=sha256:a2234c12c0cb49e0701c2f8fbec0d47a426baa0b6bf5e1315898e6ca8beb5b04

Observation 6a1f410e-90f4-4d51-b122-aa53bde69838 · outbound

This paper cites AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.772487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.670888Z digest=sha256:fd4a61621effc326498b98d50da26e9a4358dbeca027e02935f0026ed011768c

Observation 1dd6be83-e635-4c87-9d31-354fad14baf0 · outbound

This paper cites The AMD CDNA™ 3 architecture,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The AMD CDNA™ 3 architecture,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.761505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.681382Z digest=sha256:8347fa82a3fef7c0f491108876a397bb98563e8fcaa235469892567e0cf413d9

Observation c75bc5a0-a566-4c90-846e-9ad51407dc64 · outbound

This paper cites HSA Runtime API and runtime for ROCm,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HSA Runtime API and runtime for ROCm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.750820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.690115Z digest=sha256:34d209714ae8254f257d7e27ce259a681cdadbd95dd7690f3bc6dbd977dc083e

Observation ba5d4042-d53e-4154-9a86-8291b703c223 · outbound

This paper cites HIP: C++ Heterogeneous-Compute Interface for Portability,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HIP: C++ Heterogeneous-Compute Interface for Portability,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.739086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.698583Z digest=sha256:bf794a25348ca6a76db7c6b385dc6e541938fbb8d1962d1aaf67e785eace8dc2

Observation 243312cf-886f-423c-86bb-63307a79a35f · outbound

This paper cites Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.727367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.703200Z digest=sha256:7731903baf819fae99c078924ee710dd31f39e35ffc0e86b6eebc331d827c2a5

Observation 35802812-fcae-49fc-96f9-f86b8d45709e · outbound

This paper cites AMD ROCm™ Software,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD ROCm™ Software,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.716251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.707350Z digest=sha256:3db4adbc0a6dd1794d9b49b283e1b5879e121a3839f15a44cadc7b6da1b6950a

Observation f858a095-6ea6-4c4f-8ac3-3a7610b2952d · outbound

This paper cites ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.703069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.711930Z digest=sha256:857d66f99cfce63a07266db417c0fc30b83be8179b97ffc598c4d916960b8658

Observation 66cd9b53-7e39-4089-b23f-8874fe049d04 · outbound

This paper cites ROCm Communication Collectives Library (RCCL).

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm Communication Collectives Library (RCCL)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.690679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.716063Z digest=sha256:a019ec4f476d9b51112c71826dbeed026a68ecd489979f05d6f0ce0b24c92853

Observation 6927e377-6619-403c-9cde-3c0acd2c1728 · outbound

This paper cites ROCm: HIPStream,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: HIPStream,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.679660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.720394Z digest=sha256:7b12f7b45eafe070e95e91ffe1a7f98a59b4ab49a616a8475020f8cae642c77c

Observation 5cdb2bd9-e3db-4143-9029-4dceaab0961b · outbound

This paper cites rocprof — ROC Profiler Documentation,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines rocprof — ROC Profiler Documentation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.668818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.724462Z digest=sha256:41e6764f4ffd1b62b2923e43e0bd80d96ef8ca97aa40333e8b1884969710967a

Observation 2c930b0c-a81c-455d-aa56-f678e87621f8 · outbound

This paper cites AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.658083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.728339Z digest=sha256:697b532e70845b252f3637e04e28967b84d95539f96c97c75848e9b0ad1577ed

Observation a8f051b4-d58a-4f55-9839-7cdb3ccaf05a · outbound

This paper cites ROCm: ROCR-Runtime,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: ROCR-Runtime,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.647582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.732671Z digest=sha256:d7fa40339218673cb0d001cfc7ac7933f4c15436ccb423449a016e0d810a4251

Observation 19addbb6-178a-457f-b8f0-8dbceae946c7 · outbound

This paper cites T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.637138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.736987Z digest=sha256:b91324736130a3379237d114b88b912c87d86aa3eafe58fe7abc03427beff9ac

Observation f1ba0c67-19ca-41da-a15f-04e42cd467ac · outbound

This paper cites DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.741039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.741039Z digest=sha256:b32ae0965bcdbe29b19bc710e8d9e0546e9945b28d1f10fb6788ecd09624cf0a

Observation 918f3ca6-b710-4094-97a6-603550ca0fb0 · outbound

This paper cites TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.626380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.745563Z digest=sha256:f3804d48cc5d8e521496be59ed851ecc4cec34beff6348b8c0d5386ade29305c

Observation 79aa09f4-c0dc-4557-bc57-6e0ea243859b · outbound

This paper cites MSC- CLang: Microsoft Collective Communication Language,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSC- CLang: Microsoft Collective Communication Language,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.615472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.749301Z digest=sha256:e280b4b51c249e0afaa22d9ba9fa43ed79c9bb7173f8c0e8c552852c4b34bd71

Observation b6541107-eac4-40d7-9ec4-7267ff685389 · outbound

This paper cites Available: https://github.com/NVIDIA/nccl.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Available: https://github.com/NVIDIA/nccl

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.604358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.752792Z digest=sha256:c4d49ecf4bd13fb1a8633a26faa408a8d971a6918ade9ffb3fb07099b13ff215

Observation d7372df3-5cb9-4ae8-828a-430e2ab53fb6 · outbound

This paper cites TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.593400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.756341Z digest=sha256:510c7f400d3e94cd88cd88c3008a3dd5c0dd50a1396f0defe7e009ce3012d847

Observation 5c965cdf-de6a-4e2e-a15c-201a20225f60 · outbound

This paper cites MSCCL++: A GPU-driven communication stack for scalable AI applications.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSCCL++: A GPU-driven communication stack for scalable AI applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.581937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.795376Z digest=sha256:6a8ebeb0726309c2b48528e350147f5d42c97d31ec5fa5193f35bad7ef3e7916

Observation 33a9f01b-36cd-481c-abf5-2b04b1c5de54 · outbound

This paper cites Deep Learning Recommendation Model for Personalization and Recommendation Systems.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Deep Learning Recommendation Model for Personalization and Recommendation Systems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.838572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.838572Z digest=sha256:92ac2fface6b455bd244be839ea61e0a9b3e3c93c940a1a6d425ff74a7020222

Observation c3c9734d-8040-4f14-ac61-bcea09cbcb98 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.570814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.888958Z digest=sha256:62c3a1618e4f84ce5ce5e21fe764a17dd27eb5d5d4b1893435037931fd8f0427

Observation 93de37ab-023d-4849-9713-109d1645eaee · outbound

This paper cites Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.917412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.917412Z digest=sha256:a20323c961881e5cfc6a1b602dc2dcbbe17878fc6cd96fa2be005b9aeb036e1e

Observation febf812d-6ccd-40b8-b15f-3c2bfb32a186 · outbound

This paper cites The Case for GPGPU Spatial Multitasking,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Case for GPGPU Spatial Multitasking,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.559159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.936810Z digest=sha256:de9c8be2dcf51229fc2dfb516b93a9f318a15b791b89e47289ecabccd1a49b71

Observation f1ed9c9e-0e3f-42c4-9330-7a15f54fc25a · outbound

This paper cites Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.546408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.941351Z digest=sha256:977a2d6c386948a88f6709a9276c2508ae5aea1610d1a5cded5a088901e7c58c

Observation 0de2b38d-dcb5-4768-81ea-ba7fe669f6cb · outbound

This paper cites Improving GPGPU Concurrency with Elastic Kernels,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU Concurrency with Elastic Kernels,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.533927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.945764Z digest=sha256:2c9715155e2def11fb2527ac1a78d380ce4835309a461d1b96f0e9c9f3bc361b

Observation 6ba035e1-6067-4d66-82c6-0c07a7857b84 · outbound

This paper cites Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.521377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.950054Z digest=sha256:1f4a26e1fce116c8c873ea946cf063dcd5fc248b62877bc87813b1b126c5d947

Observation a3a604e8-475b-4bff-abd0-253409ab8303 · outbound

This paper cites Global Optimizations & Lightweight Dynamic Logic for Concurrency.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Global Optimizations & Lightweight Dynamic Logic for Concurrency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:23:19.250779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.954391Z digest=sha256:f2b48bdc998806678e178e25ce4ea6fe733ce57772798436e1b5db95ba981d4b

Observation fb53f232-3849-4ca5-ae25-0ac468d29a7e · outbound

This paper cites An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.508949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.966166Z digest=sha256:78974dfea01ad5e9b98aa3469902a7f09ed1e2dbdf0e6ff49a4c0b6da73ac22b

Observation 1925a9da-25db-43c4-948d-e6befcaffe78 · outbound

This paper cites Introducing Async Tensor Parallelism in PyTorch,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Introducing Async Tensor Parallelism in PyTorch,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:19.496119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T12:23:18.970195Z digest=sha256:3eaa303b6c021ac8b6151db062aee9394561cbe9a1ae5c88e64456b7ef6bdc19

Observation 0a16a08d-0815-4c5e-98c6-19f098eae38b · outbound

This paper cites Optimizing Distributed ML Communication with Fused Computation-Collective Operations.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Optimizing Distributed ML Communication with Fused Computation-Collective Operations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.976636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.976636Z digest=sha256:30eb0228e5a1c65f6ac5f59c4f8d984924017d3535dc4cc707b9d79262c4fd89

Observation 25d52610-b9b8-4db3-917c-45ac4402b6f3 · outbound

This paper cites Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.988507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.988507Z digest=sha256:a8e45f339963a1ea9155e233d580e07afd74c197332708733db11ae7f89793bd

Observation 54b67e22-4c11-4532-a4be-47ce55e0e45e · outbound

This paper cites Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:19.004428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:19.004428Z digest=sha256:cdf4bf1b0631aaeb7dd2b0c96b4b976fdb9381005d52fcead17889aa23942b61

Observation babe03a7-517b-4b39-9cf5-f1d3e8580eb1 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Optimizing ML Concurrent Computation and Communication with GPU DMA Engines PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:18.559962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:18.559962Z digest=sha256:7cb32fd696ef8271337faa7505e0cc18cbef05a12cc6a3087a5aa6fcd48eb695

Pith citing papers

Observation a6d06ad2-aa27-4edc-8a66-4da6323f8d25 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Optimizing ML Concurrent Computation and Communication with GPU DMA Engines

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.078708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:5f2430880ae7db6af6e53094ceb90b2430d7e28af1479c0c897c81209b9d53f6