Pith. sign in

Paper Citation Record · LEDGER

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.03537.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03537 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:14:18.740009Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de83941b-3105-4d33-8c4f-3e8f4b789b38 · outbound

This paper cites TensorFlow: A system for Large-Scale machine learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures TensorFlow: A system for Large-Scale machine learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.269781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.150954Z digest=sha256:3f62bfca21525c166714a3f208c54146d19d0f20c53e3706d54ed21177189e60

Observation f8a1a920-6bfd-4beb-acb7-5ba9610ede9a · outbound

This paper cites Learning to op- timize halide with tree search and random programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Learning to op- timize halide with tree search and random programs,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.260061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.176970Z digest=sha256:26b238b15b51d448fef28f95ac49d2cf0bfcaa92eb5500aa7f6cc85b112a60dd

Observation dbbd22f3-545c-4526-8c50-63f458e096f7 · outbound

This paper cites Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:15.249607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:15.249607Z digest=sha256:c0c04e838d7fe89babb17d5077e4e86a0cb7e3ad6583ac37b044826c4e63735a

Observation 8a731703-094b-4bfc-b007-dbb1f94e2176 · outbound

This paper cites TVM: An automated end-to-end optimizing compiler for deep learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures TVM: An automated end-to-end optimizing compiler for deep learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.245075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.316409Z digest=sha256:5f822e66b753bda7b463d3e78a124335e5de9a89b1cd050d185204ff52b5381a

Observation 58231797-7d19-4de7-a11b-dcebafd6cc81 · outbound

This paper cites Learning to optimize tensor programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Learning to optimize tensor programs,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.235551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.395958Z digest=sha256:ebc5dd858cb999934b2f72486528d96447e0385074afa81b181b5f86b4fc0546

Observation d1a62d32-53c3-4fb2-b398-29dad4f513d3 · outbound

This paper cites Evt: Accelerating deep learning training with epilogue visitor tree,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Evt: Accelerating deep learning training with epilogue visitor tree,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.226744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.470126Z digest=sha256:b4b9dd9a803687fc63aafe7fb15e6613d4d0fe74388d48d13b1f5cb42841067b

Observation 95ca7f43-86fd-4cc4-ac48-9c0896e07727 · outbound

This paper cites cuDNN: Efficient Primitives for Deep Learning.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures cuDNN: Efficient Primitives for Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:15.535544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:15.535544Z digest=sha256:6a757d2ba0c9925c2dfd1b94d03830f27439b714355e48b7f9cb89f8b6ed33bd

Observation 51f1f166-1332-4459-969e-2676d49c715e · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Flashattention-2: Faster attention with better parallelism and work partitioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.217549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.581087Z digest=sha256:0f730eb6b31f045c68645f6faba7978841d8cbf2ef7f089e369cb80e20f54515

Observation 6af0f6c4-efec-4282-818d-dde8d198105e · outbound

This paper cites Analyzing the impact of kernel fusion on gpu tensor operation performance: A systematic performance study,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Analyzing the impact of kernel fusion on gpu tensor operation performance: A systematic performance study,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.207763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.658022Z digest=sha256:ac55b966ff5f3ad16ebd8c885e97a15e4031b57ee7a5f92909e66246185bb28a

Observation 1ebde18d-88a7-4062-ad49-e34abe702834 · outbound

This paper cites CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:15.729983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:15.729983Z digest=sha256:a131d7962d0510b41a1540fb66268805a90f6fdce9084bbabae82edeb8c59da0

Observation 6078042a-77b1-4e35-841c-37f57c13b633 · outbound

This paper cites Mixed-input matrix multiplication performance optimiza- tions,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mixed-input matrix multiplication performance optimiza- tions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.197930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.790178Z digest=sha256:b23a1112d8a3851728e60f8bf0b7c362f39e0048d62b118752f5a2b0285c9a16

Observation 7fa5ce43-b93d-4440-af7d-cc99e82c7e42 · outbound

This paper cites Fireiron: A data-movement-aware scheduling language for gpus,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Fireiron: A data-movement-aware scheduling language for gpus,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.188148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.843159Z digest=sha256:f7a12cdbf59077adbd844372ac7d48932940ad520a4ef02d5288595a566c1ad2

Observation fb018d17-0471-4be0-9f06-6950717cd1fc · outbound

This paper cites Making deep learning go brrrr from first principles,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Making deep learning go brrrr from first principles,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.178532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.881797Z digest=sha256:527dbf3bdba846d61f4fc502834321a3b8fc1bbc9074f0a60d4c3e74f14ca367

Observation 9d5a46cd-8399-4fc1-8dd7-111352935560 · outbound

This paper cites Data movement is all you need: A case study on optimizing transformers,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Data movement is all you need: A case study on optimizing transformers,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.169517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:15.944166Z digest=sha256:60a1dcb22d227c03196b43d98382614236d4d0eea514c8a635b6d733c66d2dce

Observation 6697361a-1a6d-40cb-9f81-84b8fbb683c8 · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures In-datacenter performance analysis of a tensor processing unit,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:16.004122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:16.004122Z digest=sha256:1400197df6e37085d446baa3f11b546b88bd6ed8fd0348e3eef63df3f52d257e

Observation 05a01e35-2706-42a8-b9ab-8ee24c562ece · outbound

This paper cites onednn graph compiler: A hybrid approach for high-performance deep learning compilation,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures onednn graph compiler: A hybrid approach for high-performance deep learning compilation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:16.083133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:16.083133Z digest=sha256:a235f492db6fd332eb1dec22964b3665fb5c8bc73c2d42525fb83dd4380e22e2

Observation 23b93095-fb03-43d5-aac3-1699d63f0319 · outbound

This paper cites Deep Learning Recommendation Model for Personalization and Recommendation Systems.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Deep Learning Recommendation Model for Personalization and Recommendation Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:16.231025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:16.231025Z digest=sha256:6617a700970e71dd77115117e433e7cbab98f0c6245ebcb289b531d5fa80afc8

Observation f1e3e51c-3f46-4c85-86e8-8cbfa62bab2d · outbound

This paper cites Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Dnnfusion: accelerating deep neural networks execution with advanced operator fusion,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.149447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.305681Z digest=sha256:baa59c94f3d054199298733afb8fc532bd58bce7fccc171720eeb3aa905f45a0

Observation f1901e4c-4a20-4c97-b9b9-409eb397507a · outbound

This paper cites Cutlass: Fast linear algebra in cuda c++,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Cutlass: Fast linear algebra in cuda c++,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.139488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.385907Z digest=sha256:5646e0b230e3b46af7f5bfa0f706e0c99b101eed323709bc9d5b9144c266de4b

Observation ca35fd91-cbdb-48fb-a424-ec2e0d94829d · outbound

This paper cites NVIDIA TensorRT: An sdk for high-performance deep learn- ing inference,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures NVIDIA TensorRT: An sdk for high-performance deep learn- ing inference,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.129913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.440457Z digest=sha256:215e44c13009fe01fdac5e1029be5b715cb0662b364766893320e06cf64c5e55

Observation 89097831-03eb-41e5-b98f-24a82aea9918 · outbound

This paper cites cuBLAS: The nvidia cuda basic linear algebra subroutines library,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures cuBLAS: The nvidia cuda basic linear algebra subroutines library,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.120105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.547526Z digest=sha256:b6149eeb0b487605c0adc920fa7182c1f8a576274692835a6eb0997313569ac0

Observation 5a209196-fa8c-452d-b4d0-d75d7bf85ba5 · outbound

This paper cites Cuda c++ programming guide,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Cuda c++ programming guide,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.109012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.641554Z digest=sha256:af9a945d491babf900c9ade4240749ab6803315940effbbb5f1672843babfab5

Observation 686478c4-0146-4fbc-b377-e36874ff1810 · outbound

This paper cites Triton: An open-source programming language for writing highly efficient gpu code,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Triton: An open-source programming language for writing highly efficient gpu code,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.098694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.692826Z digest=sha256:4c8e40798ab425c380639fbdf1e4ae3eaa086af6823a2fa46df1c13447746bcd

Observation 0a9ee1b1-8a15-42b4-8ec1-b9bbe5dc9b10 · outbound

This paper cites Automatic kernel fusion for image processing dsls,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Automatic kernel fusion for image processing dsls,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.089869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.771487Z digest=sha256:dfcf0eda10a392b4a41350ca1d07fd364042dde37552c9dc5be62145dc7470ac

Observation 312e487d-0668-4644-8a4f-af5fe7ce1592 · outbound

This paper cites Tensor program optimization with probabilistic programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Tensor program optimization with probabilistic programs,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.079505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.872430Z digest=sha256:48298da5946c46835432046477cc48672472ef1c24eb26969d4ba6923c951db8

Observation ffe88e2c-271f-4114-aea2-50edce9eb9ac · outbound

This paper cites Astra: Exploiting predictability to optimize deep learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Astra: Exploiting predictability to optimize deep learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.068888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:16.938136Z digest=sha256:dd44dd70fefa3037ef9bf86b6b8f9392e28be2566912238a8c2a688b71d0fdeb

Observation 863ded5a-d359-412e-9574-d3d5d3b30703 · outbound

This paper cites XLA: Optimizing compiler for machine learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures XLA: Optimizing compiler for machine learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.059093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.003096Z digest=sha256:09d554234a23e64d8ac44658a16f528bdbe916cc662d5d683b4f7c1e1a73dee8

Observation 4fa07624-d39a-45ce-9125-a17a8d8aeea1 · outbound

This paper cites Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:17.006441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:17.006441Z digest=sha256:f2eb84e08d37b0bac478b4aed1addecca45fefdb9d149282e4e694fa06fa86ac

Observation a844f1b8-d2d4-49c5-95a5-de50376987a9 · outbound

This paper cites Attention is all you need,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Attention is all you need,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:17.033024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:17.033024Z digest=sha256:8d3e3564fb3061f40f718a9d5a3afa347bdbf67686c2eb7bbab2abad27bb1ad0

Observation f6f83c22-51b8-4517-8890-71eb6f4cf4a1 · outbound

This paper cites Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Memory Is All You Need: An Overview of Compute-in-Memory Architectures for Accelerating Large Language Model Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:14:17.119250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:14:17.119250Z digest=sha256:77dcf6a53f1992cf456b8984c86d695f8374d1e377a6fee3f834e388ece646aa

Observation cdffdb4f-b987-41bd-ac8e-7b6ce2e49f4d · outbound

This paper cites Mirage: A{Multi-Level}superoptimizer for tensor programs,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mirage: A{Multi-Level}superoptimizer for tensor programs,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.042563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.204354Z digest=sha256:630d98f5587a82df51314319af8250edb29d7f8c5ee2fcf21315f3ad70a746c9

Observation 36b85fd2-7fc5-45ab-b66d-1afe3a4597ef · outbound

This paper cites {PluS}: Highly efficient and expandable{ML}compiler with pluggable graph schedules,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures {PluS}: Highly efficient and expandable{ML}compiler with pluggable graph schedules,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.033790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.291030Z digest=sha256:aaf16f48190eb9c5da01148367b00026231cf92c3df1f3d68857638359d24b2e

Observation 928d8a25-e863-44a8-90f3-64c05f261e48 · outbound

This paper cites Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Bolt: Bridg- ing the gap between auto-tuners and hardware-native performance,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.024271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.409790Z digest=sha256:39a45718ed214604cc345aaaae8463dbf6cdc061c9341716d799bfdbb6f82776

Observation 5f0a2b04-2e16-4526-9ce1-76b6c58dd2a8 · outbound

This paper cites Demystifying tensor cores to optimize half-precision matrix multiply,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Demystifying tensor cores to optimize half-precision matrix multiply,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.015551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.590768Z digest=sha256:ea1ea88293fcc23c9f3d961b9caff132479733e22061825168c72f57b3f031c3

Observation 2129f888-f11e-474b-b1e0-4ac00c7a7479 · outbound

This paper cites Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:14:19.222523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.683649Z digest=sha256:fa1dae8a4dce634b47583c27ff3b7c53d82818beae98c43a0e38c3b1029153ee

Observation eac4fa19-6532-4ecc-9fc5-12b560288995 · outbound

This paper cites Mcfuser: High- performance and rapid fusion of memory-bound compute-intensive operators,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Mcfuser: High- performance and rapid fusion of memory-bound compute-intensive operators,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:20.006601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.813122Z digest=sha256:6f7284ec1d91591eca6458cc6e71be0f7f7cc6ea96fb8812c5c1b4f901f4280b

Observation 7286c7c0-384d-4940-984f-f803229c4991 · outbound

This paper cites Apollo: Automatic partition-based operator fusion through layer by layer optimization,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Apollo: Automatic partition-based operator fusion through layer by layer optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.997973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:17.906380Z digest=sha256:9f27dc3a157ba2f85b59f81dcdc2fd706e9e6513889eab49ae704c9f7cd632f8

Observation 8988c86d-05de-410c-92a2-c50d5273d328 · outbound

This paper cites Operator fusion scheduling optimization for tvm deep learning compilers,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Operator fusion scheduling optimization for tvm deep learning compilers,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.988750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:18.030327Z digest=sha256:cb04e8df91f115a52357fdc02c0a5e4c55b54a1529b61414ad135a0acd9e35ff

Observation c3151d97-b989-4776-adf9-6845a8625ce6 · outbound

This paper cites Ansor: Generating{High-Performance}tensor programs for deep learning,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Ansor: Generating{High-Performance}tensor programs for deep learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.979920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:18.157773Z digest=sha256:28067214179f329c562ec43e8cbd19c833ec9184347d6a3a173ea1bd8ff71f8a

Observation df1ebaf6-9996-4733-8e75-88d7bbf85b24 · outbound

This paper cites Chimera: An analytical optimizing framework for effective 12 compute-intensive operators fusion,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Chimera: An analytical optimizing framework for effective 12 compute-intensive operators fusion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.970526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:18.299365Z digest=sha256:2408bb5bcc277af8f45a1b7c34eaf6af6dbe78b1613ded90cd68b4a04d5cab67

Observation 1fd8d8ba-539e-47e2-8964-c73a52e0752b · outbound

This paper cites Astitch: enabling a new multi-dimensional optimization space for memory-intensive ml training and inference on modern simt architectures,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Astitch: enabling a new multi-dimensional optimization space for memory-intensive ml training and inference on modern simt architectures,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.768992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:18.390911Z digest=sha256:dc5a18a589e47853d213959d4311b8aef15f84c8b1968f9965f4c0d21347cfce

Observation 6c2a8ee3-907b-4f38-bae2-b48a41f44e22 · outbound

This paper cites FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-05T17:14:18.939842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:18.523453Z digest=sha256:3b30ce68644816775ccc1faf51e1a82ed9bf447157487827d3317ef1c352d0e4

Observation 64845e81-c353-4684-8de4-7f7da5983a75 · outbound

This paper cites Deep interest network for click-through rate prediction,.

ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures Deep interest network for click-through rate prediction,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:14:19.549972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T17:14:18.740009Z digest=sha256:10623723fde03b0727d89cab504399a21bd401873c02d096319214fc4b515bef

Pith citing papers

No inbound Pith citation observations are available.