Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:19.004428Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2412.14335.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:19.004428Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-18T00:27:43.845408Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T00:30:33.076250Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 890877d4-5f77-49ac-b521-e89d0ff47993 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines A bridging model for parallel computation,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1a0b163-db15-42ec-8fb3-c4784c7ce237 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Pytorch fsdp: Experiences on scaling fully sharded data parallel,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c7b4e4f-8200-4234-ad70-3c5a017746e5 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines NanoFlow: Towards Optimal Large Language Model Serving Throughput
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c32deaf5-bd94-4184-bdab-1ec72303a5b3 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae43603-382f-4fc8-968c-5d3e08f8b168 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines {ARK}:{GPU-driven} code execution for distributed deep learning,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6ebe5e3-331a-4d90-98ad-863da059bdbe · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines 11.1 amd instincttm mi300 series modular chiplet package – hpc and ai accelerator for exa-class systems,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6a1f410e-90f4-4d51-b122-aa53bde69838 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1dd6be83-e635-4c87-9d31-354fad14baf0 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The AMD CDNA™ 3 architecture,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c75bc5a0-a566-4c90-846e-9ad51407dc64 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HSA Runtime API and runtime for ROCm,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba5d4042-d53e-4154-9a86-8291b703c223 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines HIP: C++ Heterogeneous-Compute Interface for Portability,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 243312cf-886f-423c-86bb-63307a79a35f · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Tale of two cs: Computation vs. communication scaling for future transformers on future hardware,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35802812-fcae-49fc-96f9-f86b8d45709e · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD ROCm™ Software,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f858a095-6ea6-4c4f-8ac3-3a7610b2952d · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm/rocBLAS: Next generation BLAS implementation for ROCm platform,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 66cd9b53-7e39-4089-b23f-8874fe049d04 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm Communication Collectives Library (RCCL)
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6927e377-6619-403c-9cde-3c0acd2c1728 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: HIPStream,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5cdb2bd9-e3db-4143-9029-4dceaab0961b · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines rocprof — ROC Profiler Documentation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2c930b0c-a81c-455d-aa56-f678e87621f8 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines AMD Instinct™ MI300X Accelerator Perfor- mance Validation Guide,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a8f051b4-d58a-4f55-9839-7cdb3ccaf05a · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines ROCm: ROCR-Runtime,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 19addbb6-178a-457f-b8f0-8dbceae946c7 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines T3: Transparent tracking & triggering for fine-grained overlap of compute & collectives,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f1ba0c67-19ca-41da-a15f-04e42cd467ac · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 918f3ca6-b710-4094-97a6-603550ca0fb0 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TimeGraph: GPU Scheduling for Real-Time Multi-Tasking Environments,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 79aa09f4-c0dc-4557-bc57-6e0ea243859b · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSC- CLang: Microsoft Collective Communication Language,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b6541107-eac4-40d7-9ec4-7267ff685389 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Available: https://github.com/NVIDIA/nccl
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d7372df3-5cb9-4ae8-828a-430e2ab53fb6 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5c965cdf-de6a-4e2e-a15c-201a20225f60 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines MSCCL++: A GPU-driven communication stack for scalable AI applications
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 33a9f01b-36cd-481c-abf5-2b04b1c5de54 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Deep Learning Recommendation Model for Personalization and Recommendation Systems
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c9734d-8040-4f14-ac61-bcea09cbcb98 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Megatron-LM: Training Multi-Billion Parameter Language Mod- els Using Model Parallelism,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 93de37ab-023d-4849-9713-109d1645eaee · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation febf812d-6ccd-40b8-b15f-3c2bfb32a186 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines The Case for GPGPU Spatial Multitasking,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f1ed9c9e-0e3f-42c4-9330-7a15f54fc25a · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU energy-efficiency through concurrent kernel execution and DVFS,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0de2b38d-dcb5-4768-81ea-ba7fe669f6cb · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Improving GPGPU Concurrency with Elastic Kernels,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6ba035e1-6067-4d66-82c6-0c07a7857b84 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3a604e8-475b-4bff-abd0-253409ab8303 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Global Optimizations & Lightweight Dynamic Logic for Concurrency
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fb53f232-3849-4ca5-ae25-0ac468d29a7e · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines An In-Network Architecture for Accelerating Shared-Memory Multiprocessor Collec- tives,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1925a9da-25db-43c4-948d-e6befcaffe78 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Introducing Async Tensor Parallelism in PyTorch,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0a16a08d-0815-4c5e-98c6-19f098eae38b · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Optimizing Distributed ML Communication with Fused Computation-Collective Operations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d52610-b9b8-4db3-917c-45ac4402b6f3 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Breaking the Computation and Communication Abstraction Barrier in Distributed Machine Learning Workloads,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b67e22-4c11-4532-a4be-47ce55e0e45e · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation babe03a7-517b-4b39-9cf5-f1d3e8580eb1 · outbound
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d06ad2-aa27-4edc-8a66-4da6323f8d25 · inbound
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.