Pith. sign in

Paper Citation Record · LEDGER

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

As of 17 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 19 inbound Pith citation observations for arXiv:2504.19442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19442 v3

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:56:49.328497Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:59:12.728206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T07:37:44.599691Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 280aa325-bb9e-4cdc-8e10-5c6138b0dd2f · outbound

This paper cites Rocm communication collectives library.https://github.com/ROCm/rccl, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Rocm communication collectives library.https://github.com/ROCm/rccl, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.918412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.102645Z digest=sha256:2e50781c72a71b408ee5b6148857de915407875c17d6ffb02b6b49898929f1c2

Observation 127fb14b-5856-473d-a6ef-61f67030d5bc · outbound

This paper cites Triton-linalg for mlu, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Triton-linalg for mlu, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.902492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.109503Z digest=sha256:6d29d316ac649c6af2f8e6bdf7b17be0c196abe767bbf806ae29e78ffe6f91f9

Observation 61a959a8-99c4-4bde-8b1a-b9843fb1a340 · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.115053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.115053Z digest=sha256:4393a8a699e888e8945bd031b8e8e8a07ede8e2c4d382dd4604c197a45f50a8f

Observation a4545f50-1458-4902-8758-003982b1e6a0 · outbound

This paper cites Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.122499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.122499Z digest=sha256:ffe8382fd3d0a23298529edabc372abb37f5935e287eef871821849358cf2835

Observation 95176c07-1386-489d-93ec-982034aa0fa5 · outbound

This paper cites Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.128875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.128875Z digest=sha256:e296c8253cbd913beef51e4a0a4bbc1fadfa9d667d564b379bdec1c761811c2a

Observation 4de3f634-be53-4d82-9227-e8b2c4548141 · outbound

This paper cites Flash-decoding for long-context inference, 2023.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Flash-decoding for long-context inference, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.871794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.135780Z digest=sha256:da437aae99c33fbc32126c35f4509406314e179a6f510d3d87223d25022560a3

Observation 2d4d291c-61bd-41f8-ac47-f62e212393da · outbound

This paper cites DeepSeek-V3 Technical Report.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.142017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.142017Z digest=sha256:9b9f62a3e11d5ee8dde1a87cd2b09f73fa6ef8e85979237edaca2c256ec14ef5

Observation 09037f36-59a6-4499-8417-c1878dc4679c · outbound

This paper cites Amarasinghe.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Amarasinghe

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.148268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.148268Z digest=sha256:be5f246453aaf5ce37762d9ccb851d20cc6bc78c8fe47f126b6b6ef83ffa5242

Observation 570e5e45-e4c6-41e8-8d3d-43d673f4a14a · outbound

This paper cites Tensorir: An abstraction for automatic tensorized program optimization.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Tensorir: An abstraction for automatic tensorized program optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.153156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.153156Z digest=sha256:c9b49adcdb5b70933f193d77ebe39183048d07c8ef0c69572d12da18ad2ea105

Observation 75b02ffc-8353-4305-9502-653f13eeacbf · outbound

This paper cites Pallas, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Pallas, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.158699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.158699Z digest=sha256:c943ef78d473fd1fbbfdb067833edab4cb69963b73bd9900cceb52dd6835317e

Observation 88e805f9-b77e-44cb-aeeb-83a1e4283deb · outbound

This paper cites Goumas, Aristidis Sotiropoulos, and Nectarios Koziris.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Goumas, Aristidis Sotiropoulos, and Nectarios Koziris

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.843104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.166183Z digest=sha256:646714da8e12ab84647da4818f28008d10ed8bd631f40206ab260b47316d7dfc

Observation bb7b60b1-5bb2-40ad-9fcd-38e227c12600 · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Le, Yonghui Wu, and Zhifeng Chen

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.826886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.179643Z digest=sha256:27d2e7183ec891ca10c84cc8abb53f79ed5e44a0fd43a061a666e7f31c85ae0f

Observation b461d3f4-7c39-42d1-b64b-e3f2b9865071 · outbound

This paper cites Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.184567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.184567Z digest=sha256:66910e1c356f7dc3423bbe5bc803158eb0df2a73ea312c3882c7fd3a57d80d9c

Observation a18ef61c-c336-4f2c-99ce-e633e503922e · outbound

This paper cites Amarasinghe.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Amarasinghe

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.189431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.189431Z digest=sha256:ccc26490bda0d5d46a492955c97dbcbc7ef329dab4b355baddd2e5956c1854d8

Observation 5a358689-3327-4dd7-8b63-385cb452b9f7 · outbound

This paper cites MLIR: A Compiler Infrastructure for the End of Moore's Law.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler MLIR: A Compiler Infrastructure for the End of Moore's Law

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.195787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.195787Z digest=sha256:c8557d49b04653db8972f20da71148715cb8a89cce49c66362b6ef38d504b573

Observation 2b92840d-7cbf-4158-81b3-18f21b4f151c · outbound

This paper cites MPI+ULT: overlapping communication and computation with user-level threads.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler MPI+ULT: overlapping communication and computation with user-level threads

Reference 16

Resolution
verified exact
doi, observed 2026-08-16T05:56:49.539932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.201061Z digest=sha256:fb97720644e2a483d3de2e0655e3642a4da584229bc37a86139df8b3e6de42ca

Observation d13acd64-9b70-4604-9dc8-fa42aa0ff9f6 · outbound

This paper cites Overlapping communication and computation by using a hybrid mpi/smpss approach.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Overlapping communication and computation by using a hybrid mpi/smpss approach

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.206394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.206394Z digest=sha256:4cd4eee62b8e6716a713f2ed86bf5ae247cd4420a7ebfee335b47c78075b791a

Observation 423d82e4-f95a-4c2f-8597-e112f664db17 · outbound

This paper cites Efficient large-scale language model training on GPU clusters using megatron-lm.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Efficient large-scale language model training on GPU clusters using megatron-lm

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.211788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.211788Z digest=sha256:962b6e2c28b267b17c570607e16e1d6c356625f8086deb6d8b6395c1223fac25

Observation 1fdd5752-b79d-41cf-9dcb-e4dc1c62d1c1 · outbound

This paper cites cuBLAS, 2022.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler cuBLAS, 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.223756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.223756Z digest=sha256:57d0b5e7333d3f854fc74cf3c81c4fe77b155ab58b9f6034c9fb5749ec19ad20

Observation 11b86c41-65fa-427d-a5d6-5c2c5566f1f7 · outbound

This paper cites Cutlass, 2022.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Cutlass, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.228550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.228550Z digest=sha256:dd4b81bfcc9b0055d18b54b1ee8f6832d6e072142a6270378bf8f743f3860153

Observation c738b87b-77d6-4d51-8734-bd21167d064b · outbound

This paper cites Nvidia collective communications library.https://developer.nvidia.com/nccl, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Nvidia collective communications library.https://developer.nvidia.com/nccl, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.233580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.233580Z digest=sha256:bd58c10df8a80b9bbe2b41446f4f9645390df341c2aa2d755eaae035af8589e2

Observation 6db03254-fdbb-453e-a9da-7a1376eadf86 · outbound

This paper cites Chatgpt, 2022.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Chatgpt, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.764451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.238383Z digest=sha256:a9cde44a67c711acf3b8f40b636512466229ea08d6ce406e28c139285ef8c706

Observation 3e90ca48-5469-4994-b8dc-1bf55851efab · outbound

This paper cites Sora, 2024.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Sora, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.741562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.242972Z digest=sha256:ae19e517ca11abf87ef74a0ac1125e7258d7076345fdb84807f4261cedf06542

Observation fef45f2d-9a39-4fba-88dc-3cc1f36b3489 · outbound

This paper cites Addendum to gpt-4o system card: Native image generation.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Addendum to gpt-4o system card: Native image generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.723878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.247596Z digest=sha256:79009d8f0e49b3cbe663fe9dffe0d91fbe874446cbcf69b850764c0b91981e61

Observation 1838869a-a4fb-45c0-a81a-402a506adb29 · outbound

This paper cites Optimizing Distributed ML Communication with Fused Computation-Collective Operations.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Optimizing Distributed ML Communication with Fused Computation-Collective Operations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.252665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.252665Z digest=sha256:24c982f43589dda9e41f28f4c9b2b5c5e90a9a1e459b17accca4fe7fd479e999

Observation 435cc7a7-7fc7-4a06-8113-044342fd4f49 · outbound

This paper cites Ama- rasinghe.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Ama- rasinghe

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.259665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.259665Z digest=sha256:1599deec67657f17e8a86e9a72d1da13d6cc24a3ccb36a81d246d1a46fa514a9

Observation 0e84b639-e9c2-43b7-adb5-abd50fd34541 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.264286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.264286Z digest=sha256:a8235f5784332dfead8246b2612a8b7d194f54d56d94ded7bd178a936c016b24

Observation c041d9b8-aa40-48b7-b7e7-259ad28ad882 · outbound

This paper cites Doubao-1.5-pro, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Doubao-1.5-pro, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.707581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.269604Z digest=sha256:8c73729891bb82eb56b3d443629e8fa8c831df9624cdeb21e9bd1994d8cdbd64

Observation 0a6d506d-385a-43bd-9477-485012cf5744 · outbound

This paper cites an unresolved cited work.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.274264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.274264Z digest=sha256:30f44ad67ada8a949ea250e0f230bc0eef84e4c322db3714123066169d786bd5

Observation b32de1d5-71f4-4db8-94ca-7e2ea2122b0d · outbound

This paper cites Qwen2.5 Technical Report.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Qwen2.5 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.279097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.279097Z digest=sha256:d316b9eb7c38af7883c5f795559556c0f0a83c77110961fd68235693e4438d07

Observation f62dcb12-73db-41cb-ac32-831a4ae792cc · outbound

This paper cites Tilelang, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Tilelang, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.679881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.284292Z digest=sha256:3f7011488b644587c9c48f7282c53bb9893fd6c64b04f14425929b0eccfc5813

Observation bfbca8d8-e425-4c3f-88f5-f2f03994a0b5 · outbound

This paper cites an unresolved cited work.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.289496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.289496Z digest=sha256:b43b77305ee13d1268dfc0569884ad8da7ab0b5728427d847f9f2b2cc1f7fb6e

Observation 57c116ba-7342-4b6d-a9ce-5d028f1d1ccb · outbound

This paper cites DISTAL: the distributed tensor algebra compiler.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler DISTAL: the distributed tensor algebra compiler

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.293976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.293976Z digest=sha256:8927d06c74ed6d300169f41df69bdea48a6f810d9d609cfc49ae4ae762940db2

Observation 9de2bafe-49c9-4bdc-8068-1705086537ab · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.299130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.299130Z digest=sha256:ee8322f03cedcedb284470ac1c7ecd67fb4c5dd8461738ab1f2420ce7159bd7d

Observation 119a3061-deba-4248-ba42-f6a0de4bcd83 · outbound

This paper cites Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.304091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.304091Z digest=sha256:187b7c04eaf0526a3f6120c3ab5d34a5bc24f4ebe837014543e2b22939c3547f

Observation 4624ed39-1996-4001-919e-8174a131875d · outbound

This paper cites Deepep: an efficient expert-parallel communication library.https://github.com/deepseek-ai/ DeepEP, 2025.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Deepep: an efficient expert-parallel communication library.https://github.com/deepseek-ai/ DeepEP, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.309237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.309237Z digest=sha256:5b487cf6090a4d85cb77027747725762d00c7ba66b6db6a4357bfd9316c8b0c9

Observation 80ac5655-3770-4591-983e-72b32ce89ddc · outbound

This paper cites Gonzalez, and Ion Stoica.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Gonzalez, and Ion Stoica

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:56:50.650966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.313906Z digest=sha256:2ffa3c7498b3de073360061a719063b7ae3eae5f2e45db7ac860925b1dd56af5

Observation 13a322e2-d6fe-4bef-8bfd-1498b8130e82 · outbound

This paper cites Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler Flextensor: An automatic schedule exploration and optimization framework for tensor computation on heterogeneous system

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.318601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.318601Z digest=sha256:5355b5961bf201d0347dc3c5a8b958b08558d622dc2d5676acca1a77428e4a8f

Observation e4802ef6-fa99-4088-a61a-48a073fe92bd · outbound

This paper cites AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstraction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.324143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.324143Z digest=sha256:8813574c12d622beb685aca1da0b4862ca800a7ce0f75839045e6a40abeab6db

Observation 01651ee0-471d-4be5-92ba-a611d337b7fb · outbound

This paper cites TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.328497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.328497Z digest=sha256:3814ebcc39772b475ea236f2eac358aa3d6fdad632bf24490b700e810acb36aa

Observation ec862e63-71b6-490e-8a17-737d3373e240 · outbound

This paper cites URLhttps://doi.org/10.1109/IPDPS.2001.924976.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler URLhttps://doi.org/10.1109/IPDPS.2001.924976

Reference 2001

Resolution
metadata mismatch
raw_fallback, observed 2026-08-16T05:56:50.441937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T05:56:49.174198Z digest=sha256:0ac881cfeda52e08bf2e7f60667b1448fdf34c7c13cee36780b310fac967357d

Observation 9278b5ba-0d9a-4f40-8094-f814fbddf8d8 · outbound

This paper cites URLhttps://doi.org/10.1145/3458817.3476209.

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler URLhttps://doi.org/10.1145/3458817.3476209

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T05:56:49.218663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:56:49.218663Z digest=sha256:a2b75084805d446fd54f024f755b81019755b52b098605391673ff8a3595ae55

Pith citing papers

Observation b82ecf43-30b1-4223-95f0-bf96799fa64f · inbound

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference cites this paper.

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:26:40.426398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T14:25:14.150581Z digest=sha256:0a79cfc2afd1058b7b02b2cf7296f9b678e105b73e16195e0840b7e0105f048f

Observation 22004933-2a11-4577-8921-b843a28ddd72 · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:12.728206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:12.728206Z digest=sha256:381e998a5d7e6babb23dd324045433be117bb5783fd1007477d7891569074ecf

Observation 6022ff37-35f4-49d2-8ed7-d9d7a4d24545 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.083551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:5c4426b47c8352f991a07450f148e8caa2f59bee58f7dea84314f6335c5662fb

Observation 5f4c9bf8-c59f-4d2c-819f-3ee88f462063 · inbound

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap cites this paper.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:35.012041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:35.012041Z digest=sha256:eb152404fbe35bfa4ac35b3ec7b660f68a0d1131f5636c6f90f9e3be4578aff9

Observation aad259e2-f4d6-4b9d-a755-929a04214a70 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:07:43.209953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T10:06:41.150046Z digest=sha256:13189a554c666ed3c216445eb971bb140942ec0867572a7ceb97265af94e563d

Observation 12a7ce50-5cc3-4baa-9d83-7e4967d02403 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T07:21:23.181366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:21:23.181366Z digest=sha256:d276ab6845217b94d7b98504c7cf48ad83b4b08739a5db185090808285488326

Observation bb172f6c-e390-4cab-ad04-4f2e76452a46 · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:51.981196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:e2e5536677167e73c7fe0264647746c5ce50a47f1b026ec8814db4a8f07ec2e2

Observation f4df9561-beba-4fdf-938e-bef5ded9aed7 · inbound

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel cites this paper.

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:36:02.757348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:45:53.346024Z digest=sha256:8fa2f318b3f2cead160715efb4b61ec35d4339f22ed602b87fe7f550a5e26085

Observation 3ad8e02c-cd63-41e6-a347-8849af033675 · inbound

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training cites this paper.

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.137902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T02:20:00.625923Z digest=sha256:c8a4879a76ac178ffb38dfa4faddaf871a3f00067a9dab97703988b2abd07ac6

Observation 504f7409-3dba-4dc3-89b8-4e17ce7042d4 · inbound

FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training cites this paper.

FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.339124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T01:14:18.241481Z digest=sha256:725cbc4aa5c9984986f11847569aa669c9d0ee3bd745238b3e07cb00efd66736

Observation 2014761c-7057-48e6-950b-90133d74b0ac · inbound

Eliminating Hidden Serialization in Multi-Node Megakernel Communication cites this paper.

Eliminating Hidden Serialization in Multi-Node Megakernel Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:11:09.153963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T18:35:12.493334Z digest=sha256:f2d03cf8fe73c715a3446bc441f24b7b8f13a37c4a3896a88588dad8a1897c43

Observation 3cf7948d-804f-48db-afe3-8846d520fcb2 · inbound

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding cites this paper.

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:53:54.557117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T01:51:38.671365Z digest=sha256:bc9b6b2f4e90e6e0aeca0cc7736fd7b333cc8a9a8c2057e13e95c5f4bce8fdbd

Observation d631532b-9e71-4ecb-b590-2f9b335e9624 · inbound

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling cites this paper.

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:36:17.082595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T08:34:47.923752Z digest=sha256:0954d6dc8e9f82a30455f7d67eb9ceec0117849b3738cfec4681cd24eae57aab

Observation f7e1ebf8-4eba-42e4-8548-ac37829603ee · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.750739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T02:49:09.109990Z digest=sha256:8c2d5f145d6ec86765b02a4c586538f64f8980fcd6e1bbdb5916c8ffcac5eb7d

Observation 9b15cc1c-a649-46f9-a1de-332a5b2edf7b · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:14:46.917726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T15:10:25.071087Z digest=sha256:64f22df6ded12888e259dfedef9e740b769fbe2a2729ae9e83d6905e763a9989

Observation 6fe09191-b2fe-407e-b7ca-9fd37eccf6a8 · inbound

HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters cites this paper.

HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.227181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T20:16:40.966177Z digest=sha256:7c0d4ba2e5a03bc62b28e34d63398c30b5570dcd1ab596bd455bbdde058ba966

Observation f7a96141-20b8-4940-b17d-16c900ec5a1e · inbound

Coalesced Matrix-Free Finite Elements in Cell-Wise Storage cites this paper.

Coalesced Matrix-Free Finite Elements in Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:37:44.601363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-03T07:37:40.688370Z digest=sha256:76704d34ec6b82491ba92a419ae2eedf1becf461690f1ead014fb254540725dd

Observation a3b2296a-e9f8-4a83-89c6-0fc88688d802 · inbound

Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage cites this paper.

Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-12T02:40:50.976808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:40:50.976808Z digest=sha256:2cbb564daa715825a96c8079115d1e4cddc506ddb9fd864804097ee99e3f75bf

Observation 448dfd36-bda8-4f71-b366-2421e9937c5a · inbound

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference cites this paper.

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T01:18:48.061500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:18:48.061500Z digest=sha256:ad44f675e6c6f08d3759f309981f902b7c2695255d0d2e09cae5acd4d28a02a1