Pith. sign in

Paper Citation Record · LEDGER

FluidML: Fast and Memory Efficient Inference Optimization

As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2411.09242.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09242 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:56:15.620948Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c46a14b-1056-44cc-81f3-b66999537102 · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.352187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.352187Z digest=sha256:71ac26c747dcee6c941ef84964eb9bcc4f850533354fa0c6fffd194e94942b12

Observation 67ddb99e-1cbf-4df8-9ae1-2784b73a2ecf · outbound

This paper cites GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch , 9 2023.

FluidML: Fast and Memory Efficient Inference Optimization GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch , 9 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.878701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.358037Z digest=sha256:6a5091ff48e540e3717132b16dfaf6539de2e759c108bb2354941737d0af24a5

Observation 4a371ee6-a733-4e5e-b3d7-2a65a34d81a2 · outbound

This paper cites High Performance Code Generation in MLIR: An Early Case Study with GEMM.

FluidML: Fast and Memory Efficient Inference Optimization High Performance Code Generation in MLIR: An Early Case Study with GEMM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.363489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.363489Z digest=sha256:af5448c91a0fd3be1b6cd2dcd3a5e21151a7519ca934581133301031fbdcb2de

Observation f0d28679-13a3-4408-b167-bf29d096d115 · outbound

This paper cites The slab allocator: an object-caching kernel memory allocator.

FluidML: Fast and Memory Efficient Inference Optimization The slab allocator: an object-caching kernel memory allocator

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.857233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.370423Z digest=sha256:27514c66c02fc61c3c21199ad645f127deb1a99b61d420990d9f8276f2fa0110

Observation ba640ada-42f8-4be8-bdcd-296794f8006e · outbound

This paper cites J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q.

FluidML: Fast and Memory Efficient Inference Optimization J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., Vander P las, J., Wanderman- M ilne, S., and Zhang, Q

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.375817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.375817Z digest=sha256:1002e88c6323bf546f0c2dfe0f978586bc3b4d89bbc39dabfae840b9fb873ec4

Observation 35733c97-50cf-4b60-b5e6-f7310f69d9d3 · outbound

This paper cites TVM: An Automated End-to-End Optimizing Compiler for Deep Learning.

FluidML: Fast and Memory Efficient Inference Optimization TVM: An Automated End-to-End Optimizing Compiler for Deep Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.381198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.381198Z digest=sha256:6ef85938db43f575bcf33cd915e91817ec675a0a05ddf8a5bcf3c1a3274d1bc8

Observation e2cac0c7-83e4-447c-9e2a-aa4d963e85af · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.387520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.387520Z digest=sha256:653d7d9e85ad2c32e8de7caf60df44506a2dff1f3dce0d7a6f89a7b64094029c

Observation 60b706b7-3ec3-4e60-80e3-3799153b1ffa · outbound

This paper cites Intel(r) math kernel library for deep neural networks (intel(r) mkl-dnn).

FluidML: Fast and Memory Efficient Inference Optimization Intel(r) math kernel library for deep neural networks (intel(r) mkl-dnn)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.820009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.392551Z digest=sha256:edf6ee092c20fdc674db53b930fb85b7ff380d98511b5ba9ac5f0a71dfa359d4

Observation 8ff634b7-3931-496e-b56f-49dffeac24c3 · outbound

This paper cites TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems.

FluidML: Fast and Memory Efficient Inference Optimization TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.397594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.397594Z digest=sha256:6e43e2f7fea1a77c8baf09543a9d9d9364f1413624b2bb34cc8fdfa467282e85

Observation 9e0dda72-379b-451d-9e03-21bb2eb33b8b · outbound

This paper cites Tensorflow lite micro: Embedded machine learning for tinyml systems.

FluidML: Fast and Memory Efficient Inference Optimization Tensorflow lite micro: Embedded machine learning for tinyml systems

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.802748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.402724Z digest=sha256:42e2bd16694ae070ec5271c402aa5ab59eba4e32603d01196c1192a3a7b050a6

Observation 6d102845-b61d-4a27-a197-b8df7e7fb21f · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:56:16.782712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.408043Z digest=sha256:d500c98524111a5660abc1d8dfd1093b44aa5b11e28d287734923d65436a3b1d

Observation 7bb98dc0-72c8-4adf-a419-8b276598b219 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

FluidML: Fast and Memory Efficient Inference Optimization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.412985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.412985Z digest=sha256:b89dce0ebfa94b10a64d0399fab58bdf4cbb10413bd144dd911617b066f015f6

Observation a7921ea0-f8ae-4312-a12f-90cf00bbc586 · outbound

This paper cites Algorithms for compile-time memory optimization.

FluidML: Fast and Memory Efficient Inference Optimization Algorithms for compile-time memory optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.763065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.418619Z digest=sha256:520d68fa3807265bed302367d22cb8030ea61530eb1aa5c2bbb1828a890a9e6c

Observation 8a9c8a04-cd6f-4927-a43d-cb41d951dbcb · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:56:16.743662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.423361Z digest=sha256:7475d3a1918252eb014aca0a95897a87f6f6474f34dc20ac690156fb45bd371f

Observation ab60ea0e-a3ba-4910-a915-cdc994f06cc6 · outbound

This paper cites and Geijn, R.

FluidML: Fast and Memory Efficient Inference Optimization and Geijn, R

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.428402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.428402Z digest=sha256:bc41a10aea900bad0bd3008a35be411ba0d04906e2a4bd824bf4f597fa434498

Observation c6efccda-0494-45be-95ab-ca700345e725 · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-12T20:56:15.676860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.433277Z digest=sha256:be49ffd5923b7c72c61e0d4e96033609d8ffad76716ba46b8a86119b09131d58

Observation e013fa76-f195-47ba-8175-c2b45a6a39ca · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

FluidML: Fast and Memory Efficient Inference Optimization Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.438510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.438510Z digest=sha256:0a530aefcca03d8a88241f054e9072c011eb9e988a01c3400b96572dc48fd74c

Observation 330c517a-6c1b-4ce8-8a1b-4e1dfa0b060e · outbound

This paper cites Distilling the Knowledge in a Neural Network.

FluidML: Fast and Memory Efficient Inference Optimization Distilling the Knowledge in a Neural Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.443882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.443882Z digest=sha256:ca6744b815d486610de224367d6cdf76dcbebe89755453a6381da6da1ac138b2

Observation ec94146e-6de9-4959-97f2-88dbe04eb24c · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:56:16.725128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.450501Z digest=sha256:66ac24ee4651b1202f73ea82c29b99d7001ed69c2c9b2ae24e4c1ac1f0379a0c

Observation 47a0e427-e4dc-4539-9b3a-4b2a3a4a5a33 · outbound

This paper cites Openvino.

FluidML: Fast and Memory Efficient Inference Optimization Openvino

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.708624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.455286Z digest=sha256:1af24c792a5fc97760ec4d3afd56b772640459c4e031917f04292a9fb6930893

Observation 54873ebe-6de1-4abc-afe2-11d86425316a · outbound

This paper cites ConvBERT: Improving BERT with Span-based Dynamic Convolution.

FluidML: Fast and Memory Efficient Inference Optimization ConvBERT: Improving BERT with Span-based Dynamic Convolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.460527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.460527Z digest=sha256:c06f70e4c03c1ad418a09c05ad7c5228f36598fe2c12c680bde244aaea70bf4e

Observation d3266c7c-d358-4813-83ac-b72e9fd23524 · outbound

This paper cites FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference.

FluidML: Fast and Memory Efficient Inference Optimization FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.466664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.466664Z digest=sha256:7238d500c45e97897061c1a02483e33f7e207acd1e663f96bde5d45f0a5fe0cc

Observation b0647653-e7a9-4bb6-aef9-ce4697c75ab2 · outbound

This paper cites W., and Keutzer, K.

FluidML: Fast and Memory Efficient Inference Optimization W., and Keutzer, K

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.692107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.471845Z digest=sha256:61b8b57d7e5cb83e434f2857a66360533025c8f64ce108141ce7eb1c39cdc84e

Observation 0204436d-8b31-4502-83f9-59772d9e8992 · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:56:16.674612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.476719Z digest=sha256:d8a6fb09d613390b21603d1d80d9ffe7b01b5a7d3b0c948aa567218590714fbc

Observation 29196095-6593-4a14-8ddd-c66b900bb85b · outbound

This paper cites CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs.

FluidML: Fast and Memory Efficient Inference Optimization CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.481470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.481470Z digest=sha256:77d6406ed1d8d7c0af83b3a79cc1676585a8f025cfb1d268767b990fb49f4ea3

Observation 75fe1ef6-a324-49c9-8f69-17b1186b8030 · outbound

This paper cites MLIR: A Compiler Infrastructure for the End of Moore's Law.

FluidML: Fast and Memory Efficient Inference Optimization MLIR: A Compiler Infrastructure for the End of Moore's Law

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.486087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.486087Z digest=sha256:4d0d06b96175170f72d64da308f4310679b78a96263a0fbf3c0292b1234e0f90

Observation 8e905d47-f3a4-4035-837d-b946e51f15c1 · outbound

This paper cites MLIR : Scaling compiler infrastructure for domain specific computation.

FluidML: Fast and Memory Efficient Inference Optimization MLIR : Scaling compiler infrastructure for domain specific computation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.491457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.491457Z digest=sha256:68cbf8c8cbb4d85f2783c5dd209a0ab7492ae944b81db59767cdcb3eea7c0e97

Observation 867c6d68-25de-4b0d-9dcd-11f0f7f0c093 · outbound

This paper cites Compiling ONNX Neural Network Models Using MLIR.

FluidML: Fast and Memory Efficient Inference Optimization Compiling ONNX Neural Network Models Using MLIR

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.495894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.495894Z digest=sha256:7ca2231a93bd2297a06b2ecf5edc9b436e62d954213bf4a8836513b438a7700e

Observation 2a0f5617-e628-427f-8c59-a7f86c1579a8 · outbound

This paper cites On-Device Neural Net Inference with Mobile GPUs.

FluidML: Fast and Memory Efficient Inference Optimization On-Device Neural Net Inference with Mobile GPUs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.500865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.500865Z digest=sha256:478a83ac190773807bd1dd7e252701efcec1d5cb18d0a3d20859841f80fe5604

Observation fd8231fb-ed7e-4a60-917d-9c73b56192eb · outbound

This paper cites MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning.

FluidML: Fast and Memory Efficient Inference Optimization MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.506081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.506081Z digest=sha256:ca2fc03d280fe61b8cda25e9ca570821a8ff1e0612d2cb2aedb21e974e48dd6b

Observation dc4f2966-590c-4e7e-aaf0-b1f1d064f4ad · outbound

This paper cites and Deng, W.

FluidML: Fast and Memory Efficient Inference Optimization and Deng, W

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.511170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.511170Z digest=sha256:3a15badaec7028974071b6a0935a2a8d2592f0a0420519d723868daf749613ca

Observation 9f86e7b1-706e-4333-960f-05d662e9c11b · outbound

This paper cites Optimizing CNN model inference on CPUs.

FluidML: Fast and Memory Efficient Inference Optimization Optimizing CNN model inference on CPUs

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.657661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.516050Z digest=sha256:f908cd5369bd9bc638b7795ce09ff01d846f1f6e4c3b5f2c5688c8c961faac6f

Observation 4b5caed0-8a29-4675-bb2f-cfc91d4af710 · outbound

This paper cites Rammer: Enabling holistic deep learning compiler optimizations with rTasks.

FluidML: Fast and Memory Efficient Inference Optimization Rammer: Enabling holistic deep learning compiler optimizations with rTasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.520957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.520957Z digest=sha256:fd009ec60c2ad2ccc00f326d5590f6c1032b504e11f17be634d9f640ddff5b0a

Observation fa0c3e2f-d18e-493c-ba61-766a99ef7dae · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:56:16.628490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.525990Z digest=sha256:6473cee185cb8c7d610f6382a942a81e988c9b2bf67f85837a0107f6e72bcb92

Observation 11a49a10-119a-4226-bb22-260ccd0aad17 · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 35

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T20:56:16.063895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.531651Z digest=sha256:cbcee31d6a5222ecacfaee211495af9f8e346cad5728fe19d60d46aa5ff1f1ee

Observation a3969b7f-7fca-457e-95a5-da03fa5cc8e7 · outbound

This paper cites Tensorrt.

FluidML: Fast and Memory Efficient Inference Optimization Tensorrt

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.611754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.536508Z digest=sha256:22559af31f958a6a50c586219c401502baa9778058bffa209d74959f1e6e320f

Observation 3a37fea0-d617-4a59-9ac2-5226e0d6a1d4 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

FluidML: Fast and Memory Efficient Inference Optimization PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.541507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.541507Z digest=sha256:7dc388129917b2c0150124b9bc6cbddcea40d5a70ab2d658dbc2ae10b1c868a0

Observation f52dbcd3-e8e0-4f12-a81e-859755eea8d2 · outbound

This paper cites Efficient Memory Management for Deep Neural Net Inference.

FluidML: Fast and Memory Efficient Inference Optimization Efficient Memory Management for Deep Neural Net Inference

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-12T20:56:15.973118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.546977Z digest=sha256:abef31f010d615ad54191e540f116fe98c73c2089534c1777963efb4ec7a4d58

Observation b8ccd284-3882-4be1-899b-067239ed600e · outbound

This paper cites Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines.

FluidML: Fast and Memory Efficient Inference Optimization Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.552456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.552456Z digest=sha256:200b9206edb9a2811b38b236a63b69e1fe77b89c023ff31b54da21f4d5e99f72

Observation 61996e79-8239-4cac-94e3-19e8a18c9c0d · outbound

This paper cites Glow: Graph Lowering Compiler Techniques for Neural Networks.

FluidML: Fast and Memory Efficient Inference Optimization Glow: Graph Lowering Compiler Techniques for Neural Networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.558547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.558547Z digest=sha256:375a42cdac0fa82d81d28e0764c117fa75c2305e67669d40747875c54b899623

Observation af7e3904-88dc-4e36-9328-81398812238f · outbound

This paper cites Xla : Compiling machine learning for peak performance, 2020.

FluidML: Fast and Memory Efficient Inference Optimization Xla : Compiling machine learning for peak performance, 2020

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.594812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.563876Z digest=sha256:f7064f149af3c89abc2c64e42aebf0ecc1689765f711c8be32d410d7e3c9e15f

Observation c954bb67-6291-4847-b3c3-0945b98bf92e · outbound

This paper cites Efficient Transformers: A Survey.

FluidML: Fast and Memory Efficient Inference Optimization Efficient Transformers: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.569068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.569068Z digest=sha256:80aa2fae29f7f2b79083b8297ff03ca74716650756cd73e8c6d72c6204a0e79d

Observation a777bded-2662-4b06-8289-855dd326dab3 · outbound

This paper cites IREE , September 2019.

FluidML: Fast and Memory Efficient Inference Optimization IREE , September 2019

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.579026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.574702Z digest=sha256:acc5d990c0cd7a3152133820632029b67b976d46bf4c07610fb91f85cab3b5fa

Observation 363f2bf0-1ae2-457f-82bf-79e47cab334d · outbound

This paper cites Attention Is All You Need.

FluidML: Fast and Memory Efficient Inference Optimization Attention Is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.579575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.579575Z digest=sha256:4d0267ffcd3d1ed1639d6e94dc8d43a92b0c194f12a9ceef6ab8632fc9de5e17

Observation c7855c6d-940a-4e13-a323-193a0f8ade81 · outbound

This paper cites Augem: Automatically generate high performance dense linear algebra kernels on x86 cpus.

FluidML: Fast and Memory Efficient Inference Optimization Augem: Automatically generate high performance dense linear algebra kernels on x86 cpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.584899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.584899Z digest=sha256:c1595218033130441b6ead1da2a39952188d69f3534b7b90f4aa0326e26d8842

Observation 51ffc36d-2cdd-4f0f-ac66-1aa7e8a9dca7 · outbound

This paper cites an unresolved cited work.

FluidML: Fast and Memory Efficient Inference Optimization Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T20:56:16.563719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.589890Z digest=sha256:67f8a1c6c69d93332bf1f8b38c4d3c06ba560e2ae690c1f43852289b76f232b9

Observation 67f7592a-28c4-406a-909f-8afcd4bb9f33 · outbound

This paper cites Model-driven level 3 blas performance optimization on loongson 3a processor.

FluidML: Fast and Memory Efficient Inference Optimization Model-driven level 3 blas performance optimization on loongson 3a processor

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.595781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.595781Z digest=sha256:2b8d6180c28daf7c1017baa92fe84f18b9659385e851726342e717152aeb7abb

Observation 9a66df72-f32a-41b1-ba77-4ab3357bfea5 · outbound

This paper cites DeepCPU : Serving RNN-based deep learning models 10x faster.

FluidML: Fast and Memory Efficient Inference Optimization DeepCPU : Serving RNN-based deep learning models 10x faster

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.547933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.601544Z digest=sha256:4b56405b3b7a93cf4c102a1f9b989e2629df5243883a5edf4b3af81eaa72bfff

Observation 9656f4ed-f265-46c7-9bdf-6434388e5888 · outbound

This paper cites H., Haj-Ali, A., Wang, Y., Yang, J., Zhuo, D., Sen, K., Gonzalez, J.

FluidML: Fast and Memory Efficient Inference Optimization H., Haj-Ali, A., Wang, Y., Yang, J., Zhuo, D., Sen, K., Gonzalez, J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T20:56:16.531220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.605965Z digest=sha256:afd857e95d3f668d10c42605d818fac6adcaa999ddbdfa6d708c40e47baa2b5f

Observation 0f583c32-1c31-45a0-8511-a2fd3c574681 · outbound

This paper cites vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs.

FluidML: Fast and Memory Efficient Inference Optimization vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-12T20:56:15.730113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-12T20:56:15.610777Z digest=sha256:fad2495e76cf5631463ef4f3a64929a19e062db3cbc084f37d370d68bcd8228e

Observation 1ee2c716-5ce3-4f87-8c01-a730e8241a36 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

FluidML: Fast and Memory Efficient Inference Optimization Neural Architecture Search with Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.615628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.615628Z digest=sha256:c2e93e0e566253cc0d5092c4397641af86a8480b48bcdd8f83601b808666cb74

Observation e4ac0616-10e2-4c9c-9291-206640762832 · outbound

This paper cites write newline.

FluidML: Fast and Memory Efficient Inference Optimization write newline

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T20:56:15.620948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:56:15.620948Z digest=sha256:d1b7b4d44a9d688d27e80f719213e804a74c334949f855b5a710fc69db0508d2

Pith citing papers

No inbound Pith citation observations are available.