Pith. sign in

Paper Citation Record · LEDGER

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

As of 16 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 19 inbound Pith citation observations for arXiv:2504.19442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19442 v3

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:56:49.328497Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:59:12.728206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T07:37:44.599691Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 280aa325-bb9e-4cdc-8e10-5c6138b0dd2f · outbound

This paper cites Rocm communication collectives library.https://github.com/ROCm/rccl, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Rocm communication collectives library.https://github.com/ROCm/rccl, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.918412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.102645Z digest=sha256:ad66bed0a0ca361b0716b647a3d225a08ade22ea0f5f9a958f43420bfd36ca80

Observation 127fb14b-5856-473d-a6ef-61f67030d5bc · outbound

This paper cites Triton-linalg for mlu, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Triton-linalg for mlu, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.902492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.109503Z digest=sha256:22b30a52a44607a59ba3a321e5e7ee66612f5fcc2969e1dd6660e01be29abdc7

Observation 61a959a8-99c4-4bde-8b1a-b9843fb1a340 · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.115053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.115053Z digest=sha256:96b8e141984cb330e1ec26d6bc1e8e1f1e92e399420938f0ed9072a8aaf77980

Observation a4545f50-1458-4902-8758-003982b1e6a0 · outbound

This paper cites Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.122499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.122499Z digest=sha256:cc13ac56714461586aa4a2c4ec1e0d51e44ef9ab868053b2caf0524d872b1d08

Observation 95176c07-1386-489d-93ec-982034aa0fa5 · outbound

This paper cites Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.128875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.128875Z digest=sha256:94c13a183a0b5831102f5f0c33e02cb48d54d8b33c73dcd8d94e4bba2226f179

Observation 4de3f634-be53-4d82-9227-e8b2c4548141 · outbound

This paper cites Flash-decoding for long-context inference, 2023.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Flash-decoding for long-context inference, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.871794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.135780Z digest=sha256:b8e8698ae120eb400ef708635956d4354df3ce4f5b12ab9ed149a23d6ae7a145

Observation 2d4d291c-61bd-41f8-ac47-f62e212393da · outbound

This paper cites DeepSeek-V3 Technical Report.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.142017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.142017Z digest=sha256:1944dd40da3e3e1bf130894717074c8f5ddde231dc4e6ec906b90ff08ef89942

Observation 09037f36-59a6-4499-8417-c1878dc4679c · outbound

This paper cites Amarasinghe.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Amarasinghe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.148268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.148268Z digest=sha256:fd20d0739bbe6dc893d67d5626308b25e586e10e5be2c1fb1b3a374c5e086413

Observation 570e5e45-e4c6-41e8-8d3d-43d673f4a14a · outbound

This paper cites Tensorir: An abstraction for automatic tensorized program optimization.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Tensorir: An abstraction for automatic tensorized program optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.153156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.153156Z digest=sha256:77ea62df5f9d1c18f559f77e06d648d1e25f9353aca2ffa52ed427a05a3cb174

Observation 75b02ffc-8353-4305-9502-653f13eeacbf · outbound

This paper cites Pallas, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Pallas, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.158699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.158699Z digest=sha256:b61bb7129dd467f0bab243cfeb9294322941c26186dfa426348591307c5356c7

Observation 88e805f9-b77e-44cb-aeeb-83a1e4283deb · outbound

This paper cites Goumas, Aristidis Sotiropoulos, and Nectarios Koziris.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Goumas, Aristidis Sotiropoulos, and Nectarios Koziris

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.843104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.166183Z digest=sha256:dfdce7abe1b903283f85894b2fd18d3263975ec338acf02d9ac4a707369d2e20

Observation bb7b60b1-5bb2-40ad-9fcd-38e227c12600 · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Le, Yonghui Wu, and Zhifeng Chen

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.826886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.179643Z digest=sha256:7f7c1b059feed1a8197ed907111db645cd3f174b5005750805de72155568a24a

Observation b461d3f4-7c39-42d1-b64b-e3f2b9865071 · outbound

This paper cites Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.184567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.184567Z digest=sha256:c02096742f1a8922d7708f6b47810205ccae3db56d03f0df28fb0ee48958d287

Observation a18ef61c-c336-4f2c-99ce-e633e503922e · outbound

This paper cites Amarasinghe.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Amarasinghe

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.189431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.189431Z digest=sha256:33a50a11d06da9d022d37b0a95c6c929d38127531db0c716c702a6b71475d664

Observation 5a358689-3327-4dd7-8b63-385cb452b9f7 · outbound

This paper cites MLIR: A Compiler Infrastructure for the End of Moore's Law.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler MLIR: A Compiler Infrastructure for the End of Moore's Law

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.195787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.195787Z digest=sha256:c60d2c3ed7491d0066617858ef0da5898f6a26f2c1021204c3d42b3e23a9390b

Observation 2b92840d-7cbf-4158-81b3-18f21b4f151c · outbound

This paper cites MPI+ULT: overlapping communication and computation with user-level threads.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler MPI+ULT: overlapping communication and computation with user-level threads

Reference 16

Resolution
verified exact
doi, observed 2026-08-16T05:56:49.539932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.201061Z digest=sha256:572515bb4659cfc8e316eda2ae42e8d4261382120fda95337904cf5970120ade

Observation d13acd64-9b70-4604-9dc8-fa42aa0ff9f6 · outbound

This paper cites Overlapping communication and computation by using a hybrid mpi/smpss approach.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Overlapping communication and computation by using a hybrid mpi/smpss approach

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.206394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.206394Z digest=sha256:db2f009d3d1fca83f66a2341df318d38c76ff09b63772747892824a1eecb3c60

Observation 423d82e4-f95a-4c2f-8597-e112f664db17 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using megatron-lm.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Efficient large-scale language model training on GPU clusters using megatron-lm

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.211788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.211788Z digest=sha256:a95ccf96397b11257d36bf51b906d7e64674ea9d172c6352d1d565cf87ef8bff

Observation 1fdd5752-b79d-41cf-9dcb-e4dc1c62d1c1 · outbound

This paper cites cuBLAS, 2022.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler cuBLAS, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.223756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.223756Z digest=sha256:78337c2004fbd3f92b8ae51d0a0918e52a8fdbea12e2bc9fced5e1f39885342b

Observation 11b86c41-65fa-427d-a5d6-5c2c5566f1f7 · outbound

This paper cites Cutlass, 2022.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Cutlass, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.228550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.228550Z digest=sha256:26c328205351b8fe9e68200d20d7e278ea877c0e6a1268ffb2ec3fd55fec880e

Observation c738b87b-77d6-4d51-8734-bd21167d064b · outbound

This paper cites Nvidia collective communications library.https://developer.nvidia.com/nccl, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Nvidia collective communications library.https://developer.nvidia.com/nccl, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.233580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.233580Z digest=sha256:6925f5d98e4edfdb2dca4e3e3d542309b4115a637600be00f1466981fae962ee

Observation 6db03254-fdbb-453e-a9da-7a1376eadf86 · outbound

This paper cites Chatgpt, 2022.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Chatgpt, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.764451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.238383Z digest=sha256:f7b98fd4c90ca76f5d4ed5a4ccae943b033c487ec468ed9f577a03d1ea5d3725

Observation 3e90ca48-5469-4994-b8dc-1bf55851efab · outbound

This paper cites Sora, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Sora, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.741562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.242972Z digest=sha256:abd243c0826f456dcbc24d589f072cef4f5a1412c2600f46c3d759bb13f87830

Observation fef45f2d-9a39-4fba-88dc-3cc1f36b3489 · outbound

This paper cites Addendum to gpt-4o system card: Native image generation.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Addendum to gpt-4o system card: Native image generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.723878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.247596Z digest=sha256:254362a045e0f36091cb4b8519ca7093b55c31c68ae920fbf491ded91afe8991

Observation 1838869a-a4fb-45c0-a81a-402a506adb29 · outbound

This paper cites Optimizing Distributed ML Communication with Fused Computation-Collective Operations.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Optimizing Distributed ML Communication with Fused Computation-Collective Operations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.252665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.252665Z digest=sha256:e579a203722d34aecb931170541461de42796f68b164bb13bc54768ec0dacc07

Observation 435cc7a7-7fc7-4a06-8113-044342fd4f49 · outbound

This paper cites Ama- rasinghe.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Ama- rasinghe

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.259665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.259665Z digest=sha256:00148a4d972b5c6d65f2b78bf844110a62fe61b501fa01dc80924f439a36b067

Observation 0e84b639-e9c2-43b7-adb5-abd50fd34541 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.264286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.264286Z digest=sha256:6ecce221c50d5ac0db86a367c373c3adcdf8f35715da7ee8a897c6fac657de3f

Observation c041d9b8-aa40-48b7-b7e7-259ad28ad882 · outbound

This paper cites Doubao-1.5-pro, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Doubao-1.5-pro, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.707581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.269604Z digest=sha256:4e56ffe034cbfce997d7b18e738dd1a4f2d54c6aa9faf90ea1806ea17ab7b1b8

Observation 0a6d506d-385a-43bd-9477-485012cf5744 · outbound

This paper cites an unresolved cited work.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.274264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.274264Z digest=sha256:73b4f9fa767c94e3588c3a3c4c692d36f696cfd54f5a4b519841696cd7e86659

Observation b32de1d5-71f4-4db8-94ca-7e2ea2122b0d · outbound

This paper cites Qwen2.5 Technical Report.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Qwen2.5 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.279097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.279097Z digest=sha256:84f68fd77ffa3a9b06ccdbd0ecffefb65e5f9c95be0e270bdf2357575f6e3cb4

Observation f62dcb12-73db-41cb-ac32-831a4ae792cc · outbound

This paper cites Tilelang, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Tilelang, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.679881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.284292Z digest=sha256:309f2abb7e2c68341470d33d7340f19072ceff7bc96f8a0fd49f2b13808d18c0

Observation bfbca8d8-e425-4c3f-88f5-f2f03994a0b5 · outbound

This paper cites an unresolved cited work.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.289496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.289496Z digest=sha256:764556d83d959d23955fad52bc51538c963512654d92c653ca2ab5a4b8df4135

Observation 57c116ba-7342-4b6d-a9ce-5d028f1d1ccb · outbound

This paper cites DISTAL: the distributed tensor algebra compiler.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler DISTAL: the distributed tensor algebra compiler

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.293976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.293976Z digest=sha256:87d8c17f9282095641600a622f16bb6a6dee16424762bb620777b8518e2407d7

Observation 9de2bafe-49c9-4bdc-8068-1705086537ab · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.299130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.299130Z digest=sha256:3eab0ec84885b1a94179781d78cf889229650b39822d02dd0e72079f97a7026e

Observation 119a3061-deba-4248-ba42-f6a0de4bcd83 · outbound

This paper cites Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.304091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.304091Z digest=sha256:9971e083907e919bc1a71043d95325973dccd54941dd0078e9f2b876d70d9112

Observation 4624ed39-1996-4001-919e-8174a131875d · outbound

This paper cites Deepep: an efficient expert-parallel communication library.https://github.com/deepseek-ai/ DeepEP, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Deepep: an efficient expert-parallel communication library.https://github.com/deepseek-ai/ DeepEP, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.309237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.309237Z digest=sha256:0695c96bc32d4677871704d881f9bc1e78806d9b9654f3fff5ba30b514683340

Observation 80ac5655-3770-4591-983e-72b32ce89ddc · outbound

This paper cites Gonzalez, and Ion Stoica.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Gonzalez, and Ion Stoica

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.650966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.313906Z digest=sha256:83f345954a4c0c3c1a64fb8159e54cacae72a4a7f0232f7c3b2a95ac96ec157a

Observation 13a322e2-d6fe-4bef-8bfd-1498b8130e82 · outbound

This paper cites Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.318601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.318601Z digest=sha256:72837e39163e044f5d678c009b72fc12d933f5f1d5d54659db221362befa450e

Observation e4802ef6-fa99-4088-a61a-48a073fe92bd · outbound

This paper cites AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.324143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.324143Z digest=sha256:83c4aaaa58d3e20e6337274946c240f14d27c3a0cff2fe7f0c8b8c28bbbeccb4

Observation 01651ee0-471d-4be5-92ba-a611d337b7fb · outbound

This paper cites TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.328497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.328497Z digest=sha256:0c301a5233adc18a064e2457b5bf55fcd46c761f4fff987acf42946c5165e95a

Observation ec862e63-71b6-490e-8a17-737d3373e240 · outbound

This paper cites URLhttps://doi.org/10.1109/IPDPS.2001.924976.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler URLhttps://doi.org/10.1109/IPDPS.2001.924976

Reference 2001

Resolution
metadata mismatch
raw_fallback, observed 2026-08-16T05:56:50.441937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.174198Z digest=sha256:0e1ad99fe621e5bb76e2ce478ee50f39abc33977e130406aa76f142f9c7f855a

Observation 9278b5ba-0d9a-4f40-8094-f814fbddf8d8 · outbound

This paper cites URLhttps://doi.org/10.1145/3458817.3476209.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler URLhttps://doi.org/10.1145/3458817.3476209

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.218663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.218663Z digest=sha256:bfeda047bfa253d4fc805b24a00a706c3d0257184213a2b07ff6929b30d03567

Pith citing papers

Observation b82ecf43-30b1-4223-95f0-bf96799fa64f · inbound

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference cites this paper.

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:26:40.426398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T14:25:14.150581Z digest=sha256:2743f12be090d8dcf5e5a195fbbed82225abdbacf348cdbb95b69a827af4533c

Observation 22004933-2a11-4577-8921-b843a28ddd72 · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:12.728206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:12.728206Z digest=sha256:358f438493dabe1fb86a60c0f845292cb4e781cf2b7b9ee8856d3a012baa7f06

Observation 6022ff37-35f4-49d2-8ed7-d9d7a4d24545 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.083551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:ffab57411e84dd92aacf3a749ce3897b37e114f2027dca96d48c46f9d7ed07ad

Observation 5f4c9bf8-c59f-4d2c-819f-3ee88f462063 · inbound

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap cites this paper.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:35.012041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:35.012041Z digest=sha256:d2202044451acee39c4e54cd194f9b83c6a7f77d181e71d32ef40de5a6c8d3cb

Observation aad259e2-f4d6-4b9d-a755-929a04214a70 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:07:43.209953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T10:06:41.150046Z digest=sha256:f89c85b19b00a0f32693f9a1e6037594e71767832768d27f5c30904a6ccaecba

Observation 12a7ce50-5cc3-4baa-9d83-7e4967d02403 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T07:21:23.181366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:21:23.181366Z digest=sha256:1283f6b5b8c25e5d809cdc6f04fa5514d90fbe109ef8ff3f1423344ba4405a66

Observation bb172f6c-e390-4cab-ad04-4f2e76452a46 · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:51.981196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:ea043879d42a190add2c77cd7f480cf71ed1f3f6e6556dddd855825ba45ccb3c

Observation f4df9561-beba-4fdf-938e-bef5ded9aed7 · inbound

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel cites this paper.

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:36:02.757348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:45:53.346024Z digest=sha256:224be2de62378592c46b3d7eed08c63f42d9ee14e010fda8dc797ddaa925718e

Observation 3ad8e02c-cd63-41e6-a347-8849af033675 · inbound

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training cites this paper.

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.137902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T02:20:00.625923Z digest=sha256:871a4bc1fc0d6216a2beead6a5229f1cbe34732b545dfd76fba50668f9da12a1

Observation 504f7409-3dba-4dc3-89b8-4e17ce7042d4 · inbound

FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training cites this paper.

FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.339124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:14:18.241481Z digest=sha256:683fb81dcb09c16dfce5a277e1aac1c5022d26a41fde1739ee57f79288e11a49

Observation 2014761c-7057-48e6-950b-90133d74b0ac · inbound

Eliminating Hidden Serialization in Multi-Node Megakernel Communication cites this paper.

Eliminating Hidden Serialization in Multi-Node Megakernel Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:11:09.153963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T18:35:12.493334Z digest=sha256:3353a754c3ba4c044b4779dc538d93255d4fe1a67b0276387e65d30110d3c281

Observation 3cf7948d-804f-48db-afe3-8846d520fcb2 · inbound

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding cites this paper.

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:53:54.557117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T01:51:38.671365Z digest=sha256:3b35a15b4bcaf8c5b5e60752c54092c11fa2734922951badf8fb29c573e79166

Observation d631532b-9e71-4ecb-b590-2f9b335e9624 · inbound

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling cites this paper.

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:36:17.082595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T08:34:47.923752Z digest=sha256:d53acc57e365bd91dd65ae0615425aed594f5e4ca4e3b288956b576d7ef12ee5

Observation f7e1ebf8-4eba-42e4-8548-ac37829603ee · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.750739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T02:49:09.109990Z digest=sha256:64701504b104ce23a67781ba319d0750b415e850d473c93af716093a06386199

Observation 9b15cc1c-a649-46f9-a1de-332a5b2edf7b · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:14:46.917726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T15:10:25.071087Z digest=sha256:79325e2141a5099c3776c8c413c4e062d189d4539e6e229b90f8e0b52bdad3cb

Observation 6fe09191-b2fe-407e-b7ca-9fd37eccf6a8 · inbound

HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters cites this paper.

HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.227181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T20:16:40.966177Z digest=sha256:072dd1114b766cf5239a6de2d16c7d6946a3e48f33bd24dfd812e44c09dc5b78

Observation f7a96141-20b8-4940-b17d-16c900ec5a1e · inbound

Coalesced Matrix-Free Finite Elements in Cell-Wise Storage cites this paper.

Coalesced Matrix-Free Finite Elements in Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:37:44.601363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-03T07:37:40.688370Z digest=sha256:33e59f24f2a45bdc78e696fb9042402d3a0edcde44a73111a6f66c74339ad011

Observation a3b2296a-e9f8-4a83-89c6-0fc88688d802 · inbound

Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage cites this paper.

Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-12T02:40:50.976808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:40:50.976808Z digest=sha256:741ad5dc41a114d2a90054dca4fff1c89291e77743528c3795b110ba958ab45e

Observation 448dfd36-bda8-4f71-b366-2421e9937c5a · inbound

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference cites this paper.

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T01:18:48.061500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:18:48.061500Z digest=sha256:42d0488a85ff8e660e83cb4ca4374834dd8a4b4a636ec99f74f5b0a7e302e913