Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:56:49.328497Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 19 inbound Pith citation observations for arXiv:2504.19442.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:56:49.328497Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:59:12.728206Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T07:37:44.599691Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 280aa325-bb9e-4cdc-8e10-5c6138b0dd2f · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Rocm communication collectives library.https://github.com/ROCm/rccl, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 127fb14b-5856-473d-a6ef-61f67030d5bc · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Triton-linalg for mlu, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 61a959a8-99c4-4bde-8b1a-b9843fb1a340 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4545f50-1458-4902-8758-003982b1e6a0 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95176c07-1386-489d-93ec-982034aa0fa5 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4de3f634-be53-4d82-9227-e8b2c4548141 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Flash-decoding for long-context inference, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2d4d291c-61bd-41f8-ac47-f62e212393da · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09037f36-59a6-4499-8417-c1878dc4679c · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Amarasinghe
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570e5e45-e4c6-41e8-8d3d-43d673f4a14a · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Tensorir: An abstraction for automatic tensorized program optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b02ffc-8353-4305-9502-653f13eeacbf · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Pallas, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88e805f9-b77e-44cb-aeeb-83a1e4283deb · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Goumas, Aristidis Sotiropoulos, and Nectarios Koziris
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bb7b60b1-5bb2-40ad-9fcd-38e227c12600 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Le, Yonghui Wu, and Zhifeng Chen
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b461d3f4-7c39-42d1-b64b-e3f2b9865071 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18ef61c-c336-4f2c-99ce-e633e503922e · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Amarasinghe
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a358689-3327-4dd7-8b63-385cb452b9f7 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler MLIR: A Compiler Infrastructure for the End of Moore's Law
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b92840d-7cbf-4158-81b3-18f21b4f151c · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler MPI+ULT: overlapping communication and computation with user-level threads
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d13acd64-9b70-4604-9dc8-fa42aa0ff9f6 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Overlapping communication and computation by using a hybrid mpi/smpss approach
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 423d82e4-f95a-4c2f-8597-e112f664db17 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Efficient large-scale language model training on GPU clusters using megatron-lm
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdd5752-b79d-41cf-9dcb-e4dc1c62d1c1 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler cuBLAS, 2022
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b86c41-65fa-427d-a5d6-5c2c5566f1f7 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Cutlass, 2022
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c738b87b-77d6-4d51-8734-bd21167d064b · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Nvidia collective communications library.https://developer.nvidia.com/nccl, 2024
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6db03254-fdbb-453e-a9da-7a1376eadf86 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Chatgpt, 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3e90ca48-5469-4994-b8dc-1bf55851efab · outbound
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fef45f2d-9a39-4fba-88dc-3cc1f36b3489 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Addendum to gpt-4o system card: Native image generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1838869a-a4fb-45c0-a81a-402a506adb29 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Optimizing Distributed ML Communication with Fused Computation-Collective Operations
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435cc7a7-7fc7-4a06-8113-044342fd4f49 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Ama- rasinghe
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e84b639-e9c2-43b7-adb5-abd50fd34541 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c041d9b8-aa40-48b7-b7e7-259ad28ad882 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Doubao-1.5-pro, 2025
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0a6d506d-385a-43bd-9477-485012cf5744 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32de1d5-71f4-4db8-94ca-7e2ea2122b0d · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Qwen2.5 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f62dcb12-73db-41cb-ac32-831a4ae792cc · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Tilelang, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bfbca8d8-e425-4c3f-88f5-f2f03994a0b5 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c116ba-7342-4b6d-a9ce-5d028f1d1ccb · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler DISTAL: the distributed tensor algebra compiler
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de2bafe-49c9-4bdc-8068-1705086537ab · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 119a3061-deba-4248-ba42-f6a0de4bcd83 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4624ed39-1996-4001-919e-8174a131875d · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Deepep: an efficient expert-parallel communication library.https://github.com/deepseek-ai/ DeepEP, 2025
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ac5655-3770-4591-983e-72b32ce89ddc · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Gonzalez, and Ion Stoica
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 13a322e2-d6fe-4bef-8bfd-1498b8130e82 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4802ef6-fa99-4088-a61a-48a073fe92bd · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01651ee0-471d-4be5-92ba-a611d337b7fb · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec862e63-71b6-490e-8a17-737d3373e240 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler URLhttps://doi.org/10.1109/IPDPS.2001.924976
Reference 2001
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9278b5ba-0d9a-4f40-8094-f814fbddf8d8 · outbound
Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler URLhttps://doi.org/10.1145/3458817.3476209
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b82ecf43-30b1-4223-95f0-bf96799fa64f · inbound
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 22004933-2a11-4577-8921-b843a28ddd72 · inbound
Robix: A Unified Model for Robot Interaction, Reasoning and Planning Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6022ff37-35f4-49d2-8ed7-d9d7a4d24545 · inbound
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5f4c9bf8-c59f-4d2c-819f-3ee88f462063 · inbound
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aad259e2-f4d6-4b9d-a755-929a04214a70 · inbound
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12a7ce50-5cc3-4baa-9d83-7e4967d02403 · inbound
Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb172f6c-e390-4cab-ad04-4f2e76452a46 · inbound
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4df9561-beba-4fdf-938e-bef5ded9aed7 · inbound
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3ad8e02c-cd63-41e6-a347-8849af033675 · inbound
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 504f7409-3dba-4dc3-89b8-4e17ce7042d4 · inbound
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2014761c-7057-48e6-950b-90133d74b0ac · inbound
Eliminating Hidden Serialization in Multi-Node Megakernel Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3cf7948d-804f-48db-afe3-8846d520fcb2 · inbound
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d631532b-9e71-4ecb-b590-2f9b335e9624 · inbound
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f7e1ebf8-4eba-42e4-8548-ac37829603ee · inbound
HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9b15cc1c-a649-46f9-a1de-332a5b2edf7b · inbound
HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6fe09191-b2fe-407e-b7ca-9fd37eccf6a8 · inbound
HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f7a96141-20b8-4940-b17d-16c900ec5a1e · inbound
Coalesced Matrix-Free Finite Elements in Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3b2296a-e9f8-4a83-89c6-0fc88688d802 · inbound
Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 448dfd36-bda8-4f71-b366-2421e9937c5a · inbound
SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.