Pith. sign in

Paper Citation Record · LEDGER

LLM Inference Unveiled: Survey and Roofline Model Insights

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2402.16363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.16363 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:22:24.427154Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.001945Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 90806b5a-d669-4a41-ae24-75184098e22d · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:49:33.794422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:e171c8078bf03f50a2ceecf3541f05bf9ab31e75d168366c2ef5de8e58be0a8d

Observation 456ea09f-255f-4c50-a111-d7d38c867900 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.159051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:8a3977c1ae46c9b7b669deadc2ce00bedd9fa7734c726c7df516a74ae1b45d98

Observation a7a73652-6c22-41ac-92fe-6bfd2dedddae · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.959988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:a1c6411366c64ca7a5003858b7321963542ca6ef9504ddae5c4e860080014318

Observation 16c57e80-98ba-423f-8a67-c216f57374ef · inbound

Genetic AI: Evolutionary Games for ab initio dynamic Multi-Objective Optimization cites this paper.

Genetic AI: Evolutionary Games for ab initio dynamic Multi-Objective Optimization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T21:22:24.427154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:22:24.427154Z digest=sha256:6e47c821b7a37942b5357804b4882c69e9cb2837fe04dc285bbfce27b44e47b1

Observation a73d2dac-4e2a-42e9-b7c5-149bf4ab8b80 · inbound

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization cites this paper.

MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T19:12:56.309924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:12:56.309924Z digest=sha256:c146f6cb1b697adb6d2392dacf4e4b07aff000dc9dfcf1fec29359f0a99654f4

Observation ecfd4f90-ce9d-497b-a883-24e8b28f6620 · inbound

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation cites this paper.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.815570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.815570Z digest=sha256:e9112e38ef8aa9a926d12e4193b2123e9f5935f18f6dc29cfd12fb93394735bd

Observation 2164ac5a-ad2a-4c46-afae-d0fb53ae3d2d · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.852245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.852245Z digest=sha256:dba394fb3a5cdd0369f818e88b7c5b2bd961aac6eb76d57be05713ff03f5ace9

Observation 6ed6c039-3b37-495c-8cbe-b021ae7d85a0 · inbound

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees cites this paper.

Memory Offloading for Large Language Model Inference with Latency SLO Guarantees LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T10:10:23.113427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:10:23.113427Z digest=sha256:0c506c7656d81d9c35116259a1de3327409e678f7e19bb66e6ff6d81edee8c4f

Observation 70f92a0a-1dbc-4fd9-840c-8d0f1db5a590 · inbound

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation cites this paper.

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:12:11.621335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:07:37.906916Z digest=sha256:28be0d246a55cebc7fb6639638a72d135dc42aec1ea1d2bd1ed8ed8d3794eaf0

Observation ecca2898-dbd4-46c9-b25a-38c76e1393d0 · inbound

Scaling Law for Quantization-Aware Training cites this paper.

Scaling Law for Quantization-Aware Training LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:09.214394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:09.214394Z digest=sha256:cd7f196cb592707fdf54924b5004828da5fbaf11e522852853703b4d3290d690

Observation 97d55eaf-e1ea-4233-aa9f-aea58d0b400a · inbound

How to keep pushing ML accelerator performance? Know your rooflines! cites this paper.

How to keep pushing ML accelerator performance? Know your rooflines! LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:29.707975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:29.707975Z digest=sha256:14849f24e4658f4aaef5fd5aa6d244b5f3d085c7200af07639298dcac99451a3

Observation b93a57aa-0a94-4728-8bc8-8f5214d49e12 · inbound

LatentLLM: Attention-Aware Joint Tensor Compression cites this paper.

LatentLLM: Attention-Aware Joint Tensor Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:36:28.556122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:36:28.556122Z digest=sha256:554267b9508ae8ae8e4312077fbd8d773e8b9f513f977c1b393fe830ae6be869

Observation 50a59906-288a-4c38-990d-e47abb93f1d2 · inbound

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts cites this paper.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.984961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.984961Z digest=sha256:c6c0f320605a41087d789fc5365b764dab7817e8ce5c94cbfe18d06d55256cd0

Observation 2da04287-b6d6-4094-9ce7-7ba8329aa926 · inbound

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators cites this paper.

FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:26.256211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:26.256211Z digest=sha256:89dadade441c160282b60c60bad5e8392183f3576429f59c078805f696231fba

Observation bb364619-0240-43b7-917f-9b5b2c3920ad · inbound

DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model cites this paper.

DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline Model LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:58.446351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:58.446351Z digest=sha256:695817fb503c697b250cd8a29f945378aad3409d01433792a8c4bc497a69f012

Observation dcfdf893-6348-4cbd-95bc-e83b1ef50f6a · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.430496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.430496Z digest=sha256:c804916ab5ea075410a778c8497ae3a6d1e5be7b6ffe89e2eef36b7a7d85b495

Observation 1d8ad172-a5ce-46a3-ad55-399e14c4ea5e · inbound

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques cites this paper.

Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:56.770100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:56:56.770100Z digest=sha256:8a5592e4636882000f354c7683283488755a5ff9b6c896d6dae70f4f1f4ad524

Observation d9042ddc-7027-424a-9e3a-ca490e83a41c · inbound

RAILS: Retrieval-Augmented Intelligence for Learning Software Development cites this paper.

RAILS: Retrieval-Augmented Intelligence for Learning Software Development LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:04:05.650998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:04:05.650998Z digest=sha256:ed938bb584c34f995cb36585b5f7462d409c252b48e74f1c28598f41a3fbe47c

Observation ca666bfe-4770-47a5-a596-1b322e529a10 · inbound

Teaching LLMs to Speak Spectroscopy cites this paper.

Teaching LLMs to Speak Spectroscopy LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T20:46:50.618756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:46:50.618756Z digest=sha256:173d359608cfcf569b2a174fe069f6357590b1bb989dc4d26d46dd40dfac2285

Observation 84d11a4a-1084-4bf8-812f-086dfa615dee · inbound

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization cites this paper.

XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:37:15.475893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:37:15.475893Z digest=sha256:309ad9ea81ceec4df99aca552cfa6256d1e2de293ac3a4d6fbdcfe19049de93a

Observation 50dbda44-73ad-4c8c-b464-a047058de65d · inbound

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration cites this paper.

APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:03:35.382312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:03:35.382312Z digest=sha256:dc556a236f8c0ababf22751f8bcc4135af7986c88d05149634586d34a7008f06

Observation 56688505-a8b2-4bcc-9b9b-e778f36b9c4e · inbound

HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference cites this paper.

HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T16:35:33.896184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:35:33.896184Z digest=sha256:d0b49017d887324360c41e2b65876139de4d85f6a8413b81bb932a428d78cce3

Observation 159f7077-0ace-467e-af00-344d9027a091 · inbound

Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads cites this paper.

Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:57:43.030400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:53:13.057587Z digest=sha256:aa437dd3dc70cb9710ce0cb5dff169966b55b989d145492e323b48a9515feb49

Observation 0d18adca-0792-44d3-890d-59c722e37061 · inbound

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference cites this paper.

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:29.606715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:05:56.380048Z digest=sha256:45a2fdaa9446cd64bd4eb808ca3a44110ded4fa3fc1ad21981937613a73c2c95

Observation 147cae87-5bfc-4a14-beb9-9ecd4747d509 · inbound

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs cites this paper.

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:58:03.122102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:56:38.684104Z digest=sha256:5a53874893b5048573970f5f21b4ba615b08fa2c8dc51efbddba714781dedad6

Observation f5ff9e9a-1a6d-4c80-ac42-0fe94c2f005e · inbound

Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling cites this paper.

Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:38:02.718186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:36:06.003661Z digest=sha256:558cdf5771ae7e18e77046f7634260418d4601359387fdf01007c165ccb3168c

Observation a65b5d76-97c6-4b4b-a227-02f450f67153 · inbound

Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics cites this paper.

Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:19.616843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:30:40.021376Z digest=sha256:d5a3267febc42a19b4c5e963dae0a26f5228a08e50e679c8be59a463680a0630

Observation 045be78f-94a9-4672-9169-9c9e6102a303 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.463875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:abf222a783316214cfd1334b3cf38b746211e54a69f1fae2a20d030b6d472470

Observation 056bba68-557b-4834-9c14-b507ffce08df · inbound

A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks cites this paper.

A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:12.821198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:11:26.862363Z digest=sha256:8ca97ab33dedc6998d907a027c576df39eb56a448edd43365dff6c7a48e01052

Observation 8ef061b4-f73b-4e58-80d3-c1a271d5a89c · inbound

Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding cites this paper.

Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:05.348415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T19:53:40.791974Z digest=sha256:b5e5a8c3f417193e3d1518b29af98e2ad582331b008d461ef707096b7ed4691b

Observation 06e377ad-dcda-403a-bccc-88435008d2aa · inbound

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis cites this paper.

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:26.967126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T18:43:06.803528Z digest=sha256:bf91285a26a46ea76abdab58c82be66b26c0d39e5be8f51c2049389b7ed71818

Observation e530f9f2-06aa-4692-912e-9a96b177b4c0 · inbound

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis cites this paper.

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T15:09:44.098075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:09:44.098075Z digest=sha256:20e6ab2ca98a7e09e64ab46515ca05d5683c3eed6d0115ebf17a181f23dffd59

Observation e9c5676f-0207-4e02-bdfb-cf1ed7537661 · inbound

Gated Subspace Inference for Transformer Acceleration cites this paper.

Gated Subspace Inference for Transformer Acceleration LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:41:31.916138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:26:08.510348Z digest=sha256:35967a8897f9f36035de2b2aafa4a01a95e96ead96ce5346da55d4b9cabfd2fc

Observation 42e9d4d3-7ed4-41d9-a55b-429b60591205 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:30:44.191274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T18:23:14.935801Z digest=sha256:83dad4cdc87a66d3ca8a04c128253a043535f530ee787fa52843868e904e39f3

Observation b24b8616-dbc5-43d5-a0c1-d1b4ed63a515 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.177694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:59:00.997742Z digest=sha256:ebde4021a34089672b94870c5fec8582ef23478d20beb028d52f4a8a4b759d9f

Observation 8246754b-f481-416b-9f79-16f07c5cff9c · inbound

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference cites this paper.

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.959019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:13:46.098095Z digest=sha256:4cac89f22c85502f9e7df5b0faeec5a14f606fcd63bdca011f4d528fbe08148c

Observation 1bd1c73a-6b8a-45a0-a48f-b7e1d9ba5b0c · inbound

On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection cites this paper.

On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.308427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T07:51:13.367200Z digest=sha256:a42a360e0d5442c13c7dc0a4929c1c104bf8831086a159b2b511650d3388af28

Observation 275b708e-8721-476b-a29a-1fd21f16047b · inbound

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture cites this paper.

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:09.263898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:12:13.114405Z digest=sha256:c4cb6fb4426c030383da0cbb9a7d1a47caff8f6ca2e07f4ee37fc61c8097b93f

Observation 5bb01164-4430-424f-8618-806e1500196c · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.711487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:8ed0cd6a908fc361a996612580e4a808f701fa578527af328d8a5d1f8d83af73

Observation edaa94ed-94a3-4f1c-85da-db9cf5bf72a7 · inbound

Operator Fusion for LLM Inference on the Tensix Architecture cites this paper.

Operator Fusion for LLM Inference on the Tensix Architecture LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.969993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:01:15.392433Z digest=sha256:98fd515db0dc418426d1bd3d6bedf4c81f80d3ecbff28f1605a3260e971f8067

Observation d58a23c3-f7e6-4a53-997d-a141d21b67c9 · inbound

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts cites this paper.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.177818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:2c82b15a088253cabe5eb3c9c7fe427895b44261e0a00b979c840df5b10f3563

Observation 0c1ecd48-3c85-496f-86e8-8da5a2c8fcfb · inbound

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding cites this paper.

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.985869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:11:58.636939Z digest=sha256:1dab22df1c47bdd1c3f3d5591e4ab8e41bf03ee93f7520f9f04ede58124b3bef

Observation 71735d02-751c-4462-b4a0-537cf837c9c6 · inbound

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression cites this paper.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:22.036003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T01:23:46.175022Z digest=sha256:c641fb2077be6f708c469807d3640346d917e8f91277567b64a71f6030890d36

Observation 335dbaa4-29f2-468a-a406-8b2e36a447af · inbound

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression cites this paper.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:96f546b5de85e5a91b50221be8d527682282137b820a1dfdef4a7664b6b6d9a1

Observation e4ae8fa4-9ab3-4f41-9429-8aca7ece18b7 · inbound

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems cites this paper.

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 136

Resolution
unresolved
no resolver link, observed 2026-07-12T11:05:56.233115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T11:05:56.233115Z digest=sha256:149a4b9dfa7918df0df9f7f7538cfe6098a00f851a963373360c2213de97e21e

Observation 3df85082-7be3-42d5-bf95-d3b761ddf1b2 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.136167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:69a4cbc4836a83470794f338ac6e79dbac10bb2bb6bebc01b858a9f0f8e09200

Observation a0405de9-45c6-4168-a446-b473c3262421 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.036518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:4bfb754ca30dc7a7defa8671d9eab515f6dfd091e71abea2abe14eb68f19da46

Observation 8ecd1f7a-97e0-4398-8016-4c8f0275c8f9 · inbound

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing cites this paper.

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:12.343498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:12.343498Z digest=sha256:1113d735e315177bc9cf04be6412764f0a17f19c9f850917818736847ff5b3d9

Observation 9168e695-d063-4ad4-9634-4fa85f8a9393 · inbound

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding cites this paper.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.181845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.181845Z digest=sha256:9e98dfe006272f54865c773964a8d06a4ebb3db50503f3bae2340e4023357a52

Observation 7af833aa-6c1b-444c-9024-198ac481d754 · inbound

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters cites this paper.

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T04:49:48.043261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:49:48.043261Z digest=sha256:d0d886caebd926fea8528b4e9fdfa7d96a64be2837b5d352cda6406fca3f1f60

Observation 9d2d8d7c-48a5-4710-a506-4a869d174421 · inbound

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement cites this paper.

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:00:24.412238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:00:24.412238Z digest=sha256:20abd13946b9a4db350bed17720fbf84edde2947a00ec5a330290e84f0a5308c

Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · inbound

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving cites this paper.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.202838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.202838Z digest=sha256:b443929bdc9b94d973d2db95ac1e8dbf4f9c12b5512aa62ce28b235650fe844b