Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:14:18.740009Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.03537.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:14:18.740009Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de83941b-3105-4d33-8c4f-3e8f4b789b38 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures TensorFlow: A system for Large-Scale machine learning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f8a1a920-6bfd-4beb-acb7-5ba9610ede9a · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Learning to op- timize halide with tree search and random programs,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dbbd22f3-545c-4526-8c50-63f458e096f7 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a731703-094b-4bfc-b007-dbb1f94e2176 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures TVM: An automated end-to-end optimizing compiler for deep learning,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58231797-7d19-4de7-a11b-dcebafd6cc81 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Learning to optimize tensor programs,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d1a62d32-53c3-4fb2-b398-29dad4f513d3 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Evt: Accelerating deep learning training with epilogue visitor tree,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 95ca7f43-86fd-4cc4-ac48-9c0896e07727 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures cuDNN: Efficient Primitives for Deep Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51f1f166-1332-4459-969e-2676d49c715e · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Flashattention-2: Faster attention with better parallelism and work partitioning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6af0f6c4-efec-4282-818d-dde8d198105e · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Analyzing the impact of kernel fusion on gpu tensor operation performance: A systematic performance study,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1ebde18d-88a7-4062-ad49-e34abe702834 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6078042a-77b1-4e35-841c-37f57c13b633 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mixed-input matrix multiplication performance optimiza- tions,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7fa5ce43-b93d-4440-af7d-cc99e82c7e42 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Fireiron: A data-movement-aware scheduling language for gpus,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb018d17-0471-4be0-9f06-6950717cd1fc · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Making deep learning go brrrr from first principles,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d5a46cd-8399-4fc1-8dd7-111352935560 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Data movement is all you need: A case study on optimizing transformers,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6697361a-1a6d-40cb-9f81-84b8fbb683c8 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures In-datacenter performance analysis of a tensor processing unit,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a01e35-2706-42a8-b9ab-8ee24c562ece · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures onednn graph compiler: A hybrid approach for high-performance deep learning compilation,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b93095-fb03-43d5-aac3-1699d63f0319 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Deep Learning Recommendation Model for Personalization and Recommendation Systems
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e3e51c-3f46-4c85-86e8-8cbfa62bab2d · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1901e4c-4a20-4c97-b9b9-409eb397507a · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Cutlass: Fast linear algebra in cuda c++,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca35fd91-cbdb-48fb-a424-ec2e0d94829d · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures NVIDIA TensorRT: An sdk for high-performance deep learn- ing inference,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 89097831-03eb-41e5-b98f-24a82aea9918 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures cuBLAS: The nvidia cuda basic linear algebra subroutines library,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a209196-fa8c-452d-b4d0-d75d7bf85ba5 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Cuda c++ programming guide,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 686478c4-0146-4fbc-b377-e36874ff1810 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Triton: An open-source programming language for writing highly efficient gpu code,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a9ee1b1-8a15-42b4-8ec1-b9bbe5dc9b10 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Automatic kernel fusion for image processing dsls,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 312e487d-0668-4644-8a4f-af5fe7ce1592 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Tensor program optimization with probabilistic programs,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ffe88e2c-271f-4114-aea2-50edce9eb9ac · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Astra: Exploiting predictability to optimize deep learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 863ded5a-d359-412e-9574-d3d5d3b30703 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures XLA: Optimizing compiler for machine learning,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4fa07624-d39a-45ce-9125-a17a8d8aeea1 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a844f1b8-d2d4-49c5-95a5-de50376987a9 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Attention is all you need,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f83c22-51b8-4517-8890-71eb6f4cf4a1 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdffdb4f-b987-41bd-ac8e-7b6ce2e49f4d · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mirage: A{Multi-Level}superoptimizer for tensor programs,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 36b85fd2-7fc5-45ab-b66d-1afe3a4597ef · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures {PluS}: Highly efficient and expandable{ML}compiler with pluggable graph schedules,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 928d8a25-e863-44a8-90f3-64c05f261e48 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f0a2b04-2e16-4526-9ce1-76b6c58dd2a8 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Demystifying tensor cores to optimize half-precision matrix multiply,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2129f888-f11e-474b-b1e0-4ac00c7a7479 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eac4fa19-6532-4ecc-9fc5-12b560288995 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mcfuser: High- performance and rapid fusion of memory-bound compute-intensive operators,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7286c7c0-384d-4940-984f-f803229c4991 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Apollo: Automatic partition-based operator fusion through layer by layer optimization,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8988c86d-05de-410c-92a2-c50d5273d328 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Operator fusion scheduling optimization for tvm deep learning compilers,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3151d97-b989-4776-adf9-6845a8625ce6 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Ansor: Generating{High-Performance}tensor programs for deep learning,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df1ebaf6-9996-4733-8e75-88d7bbf85b24 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Chimera: An analytical optimizing framework for effective 12 compute-intensive operators fusion,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1fd8d8ba-539e-47e2-8964-c73a52e0752b · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Astitch: enabling a new multi-dimensional optimization space for memory-intensive ml training and inference on modern simt architectures,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6c2a8ee3-907b-4f38-bae2-b48a41f44e22 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64845e81-c353-4684-8de4-7f7da5983a75 · outbound
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Deep interest network for click-through rate prediction,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.