Pith. sign in

Paper Citation Record · LEDGER

Scaling On-Device GPU Inference for Large Generative Models

As of 19 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2505.00232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00232 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:51:31.541760Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:57:30.372775Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:57:30.733875Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8a6ea8cf-7f83-451e-b584-6b4e92d8fb08 · outbound

This paper cites The Khronos Group Inc., 2019.

Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.226757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.326199Z digest=sha256:c7f6e9dbf4dfcefa24de9dff40a88130e2b1924c0911c488a1b2c159f1d28a90

Observation 84050aef-ddba-42d5-b04b-ecb0c3593d90 · outbound

This paper cites The Khronos Group Inc., 2025.

Scaling On-Device GPU Inference for Large Generative Models The Khronos Group Inc., 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.216060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.330309Z digest=sha256:b4b4ae1a11768f380c64f9a1e5107fe448b2b80762fb2e284c7825f4f329ac72

Observation 64e918cc-7901-4b36-ac95-2d2349573ec9 · outbound

This paper cites AMD ROCm Soft- ware.

Scaling On-Device GPU Inference for Large Generative Models AMD ROCm Soft- ware

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.205450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.333808Z digest=sha256:b1e5a1990e25648494cf7776f2a107a53cb9cacba36dbbc03ba2939ab5d0b4f5

Observation 35619c9e-5e6b-4bf3-80c4-cf0862e6fb69 · outbound

This paper cites LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024.

Scaling On-Device GPU Inference for Large Generative Models LLM in a flash: Effi- cient Large Language Model Inference with Limited Mem- ory, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.195166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.337188Z digest=sha256:657231b64636c2bcba1a5f80e19cd2e62bc6b16d4cd1d01ee82c87a5e543b8e9

Observation 50f83bf4-3208-4146-9e8f-0d8a1022d446 · outbound

This paper cites Core ML Stable Diffusion.

Scaling On-Device GPU Inference for Large Generative Models Core ML Stable Diffusion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.184617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.340853Z digest=sha256:173740a625b3a1ba863e9ecee3a461f9886bac560002770e0aedfe8be1a241f2

Observation 1600da46-e464-4788-8bf6-58925191c0e6 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.174494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.343987Z digest=sha256:17feb474375ad79dd3f449699cc916329720e4e2dd9ea154f700ed9f9986eff0

Observation 03b334a5-6335-4ae8-a223-c2118fe97de6 · outbound

This paper cites Compute Library.

Scaling On-Device GPU Inference for Large Generative Models Compute Library

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.163667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.347636Z digest=sha256:52ea428bbadd83b33503689e50c3624b469e6dd8449ab4ab161938e08ed0b657

Observation 3a2aac5d-634a-4b7f-9068-6a346dbb38e6 · outbound

This paper cites TVM: An automated End-to-End optimizing com- piler for deep learning.

Scaling On-Device GPU Inference for Large Generative Models TVM: An automated End-to-End optimizing com- piler for deep learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.153440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.350756Z digest=sha256:438aff2e68a55b7ed79c890697c74392f7540863135676f83520b55b9b1a8b67

Observation ea88d8f3-31d9-4065-b0a5-4af7c03d8324 · outbound

This paper cites Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023.

Scaling On-Device GPU Inference for Large Generative Models Speed Is All You Need: On-Device Acceleration of Large Diffusion Models via GPU-Aware Optimizations, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.143000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.353895Z digest=sha256:c2ec0611ad40bd7d09e346ec8e32091bb18e7dbccaa7684333395c4d04063fd6

Observation 1bb45ed1-17a1-458f-8cd6-cb1221f0cc8f · outbound

This paper cites LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM.

Scaling On-Device GPU Inference for Large Generative Models LMDeploy: A Toolkit for Com- pressing, Deploying, and Serving LLM

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.132514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.356758Z digest=sha256:688574c32ea74696cdd65828f50b3830d8790bdd940eab4a7993f17651bb394c

Observation a02154a5-1f7c-46cc-85ef-d54b8eef9de5 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Scaling On-Device GPU Inference for Large Generative Models Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.122766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.359871Z digest=sha256:e192e62862d5d7854c44a3f0a2d91e9e5c42579bd09aaa6bd356ba11b08d68f8

Observation d1287604-d0ae-4cfc-aa76-73f52a5097b6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Scaling On-Device GPU Inference for Large Generative Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.363141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.363141Z digest=sha256:1c9687a7fde35da599c9f59410a56174a8ed796748f4cc636ca10b7c4e51454c

Observation 677ff3df-9209-451b-a302-346cbdb2a438 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023.

Scaling On-Device GPU Inference for Large Generative Models GPTQ: Accurate Post-Training Quantization for Gen- erative Pre-trained Transformers, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.111893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.367086Z digest=sha256:7c7229a8b570acabe376cd094427c3969784683e195e68e4a91a790a8d8bfdfe

Observation 02d036d3-af8f-4eee-b43d-c27a107cd2d7 · outbound

This paper cites Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024.

Scaling On-Device GPU Inference for Large Generative Models Gemma 2: Improving Open Language Mod- els at a Practical Size, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.099149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.370650Z digest=sha256:46adb7f41230b4cf3e586fd022faa48dfba561400dfff122b9677daa78c6d00f

Observation c7015efd-80bd-4314-9f32-01401939d2b7 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology, 2024.

Scaling On-Device GPU Inference for Large Generative Models Gemma: Open Models Based on Gemini Research and Technology, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.085958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.374881Z digest=sha256:9c63cc2943aafce4905538ea9731a3b3ea00a6d40e4e0d728bd480058bd99fef

Observation 0aea38be-4609-4085-89ad-90ae866ed3ee · outbound

This paper cites llama.cpp.

Scaling On-Device GPU Inference for Large Generative Models llama.cpp

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.075012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.379167Z digest=sha256:9ad91159cb529015a664054d57053fb8df747f47a3a1bc4287d0a0641fb4db70

Observation 6e8b1186-d337-49c2-8e95-ddfdba6a5788 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.064032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.383046Z digest=sha256:bf0474f66b4f65a5588d3cdbcb538a50295bfe76f655018dac1472b9634c2f7b

Observation be793fe5-21cb-4c21-858d-104d7958d003 · outbound

This paper cites LiteRT Overview.

Scaling On-Device GPU Inference for Large Generative Models LiteRT Overview

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.052536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.387319Z digest=sha256:5f4521f1481fba4870dc192aba1d7eddd13de71d22300307b628305c5c271c54

Observation 4f96f6ab-f0fe-43a4-a47c-bdc7b9ba331f · outbound

This paper cites Huawei HiAI.

Scaling On-Device GPU Inference for Large Generative Models Huawei HiAI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.038995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.392090Z digest=sha256:1388dd011fc9990d185afe9d372e680927ff7c18dd90718f21faf170c0ccce7d

Observation 4c372e43-57e6-4f94-a62a-e31759ed6572 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:32.029100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.396993Z digest=sha256:d11e84d8c0f5aea67405dce705d68ff828a602ad012d4a68f4e7813f3200db07

Observation a88a2194-250a-488e-9509-80693c5284fe · outbound

This paper cites OpenVINO.

Scaling On-Device GPU Inference for Large Generative Models OpenVINO

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.018800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.401664Z digest=sha256:b91b802995692fdf79ad1cee0285a27905b433bc1519ba6acc39a842e49dc70d

Observation 9ac07930-d5b7-4318-8149-b9713f0d4c83 · outbound

This paper cites Intel Core Ultra Series 2 Media Deck.

Scaling On-Device GPU Inference for Large Generative Models Intel Core Ultra Series 2 Media Deck

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:32.007819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.405781Z digest=sha256:2dddc0ee147842c1be2cb6adf648b992014672061bd2ffb16e513ad982fcc890

Observation 64a14253-3f9a-4fb7-ba60-b5f0216bb40d · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.996993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.411970Z digest=sha256:7cb726fc48981bc520460682525bfede45b92ef1f858b91fe3087ac820c561e0

Observation 5171003c-2e67-4fdb-8567-d34c505fe892 · outbound

This paper cites MNN: A Universal and Efficient Inference Engine.

Scaling On-Device GPU Inference for Large Generative Models MNN: A Universal and Efficient Inference Engine

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.415946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.415946Z digest=sha256:e5aaebf974d7fd9a767de6bedfd70208b914e77af4d9994aff4f854c5a199c4f

Observation f28ca99f-b57a-461b-8dde-39ac0ab86c1f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.985623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.420135Z digest=sha256:7d13ee0ce4e9304154a4f5b59a97520c9c0dd75ff73179f5ee496a191107ed06

Observation 01410587-29d5-48ee-a2fe-8319d24df45f · outbound

This paper cites On-Device Neu- ral Net Inference with Mobile GPUs.

Scaling On-Device GPU Inference for Large Generative Models On-Device Neu- ral Net Inference with Mobile GPUs

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.974737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.424276Z digest=sha256:21e6241abf1250ec7decdb0fac5da5084bcc915212f244544e4cfdc81a08ff1e

Observation 6ac4ea48-9622-4f7b-92d0-760a9b383d46 · outbound

This paper cites OpenGL ES Version 3.1.

Scaling On-Device GPU Inference for Large Generative Models OpenGL ES Version 3.1

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.963044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.427827Z digest=sha256:e66e8659327f1cfb4f4e25353aa49b869d3c9c54624b5af93f3e4166d80da7bd

Observation b7b87274-5c91-466f-82f3-535ef56f05ce · outbound

This paper cites Fast In- ference from Transformers via Speculative Decoding, 2023.

Scaling On-Device GPU Inference for Large Generative Models Fast In- ference from Transformers via Speculative Decoding, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.951919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.431461Z digest=sha256:b6c88759d7ccc2894aba2748cb0538b00d5191ac128493d3a787f32b98c22602

Observation 48103cb2-cd68-40b7-8954-9ca3ef371738 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion.

Scaling On-Device GPU Inference for Large Generative Models AWQ: Activation-aware Weight Quantization for LLM Compression and Accelera- tion

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.942289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.435413Z digest=sha256:cc22d59a739bdf8ea9e01bd8b7f4c4b0fafbc89077ea787f677f357103674860

Observation d5a778d6-56d8-4ef9-88ce-5d00b8bba3a9 · outbound

This paper cites The Llama 3 Herd of Models, 2024.

Scaling On-Device GPU Inference for Large Generative Models The Llama 3 Herd of Models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.932892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.438744Z digest=sha256:306560813b1dd66c2c2575a5bb3fd0eaa2348d88102bb2587e892013e06f12f3

Observation 39867c08-cad7-418b-bd15-3f6fba21ecca · outbound

This paper cites DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications.

Scaling On-Device GPU Inference for Large Generative Models DeepMon: Mobile GPU-based Deep Learning Framework for Continuous Vision Applications

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.921260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.442284Z digest=sha256:769050ae4618bf9e3f2ac38525d9683f6e9306dd6a26a90d67971818fd78678a

Observation bee3c6e9-0c36-4cf2-8a12-31c50c1489ac · outbound

This paper cites NeuroPilot.

Scaling On-Device GPU Inference for Large Generative Models NeuroPilot

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.909923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.445790Z digest=sha256:8b2901dc8656cbe54a08127fca3eddf1935e48894f7cbb8f6c528f3dc20e59d4

Observation c3ef5adf-bf8c-4145-8251-dab241f281a6 · outbound

This paper cites ExecuTorch.

Scaling On-Device GPU Inference for Large Generative Models ExecuTorch

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.898276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.449046Z digest=sha256:b8a691861e004d5eacc588f6dea59a3e7dcae1522c65ecc11eab61fdd46f39b4

Observation ac7b1861-7567-46ab-9388-0652468d864f · outbound

This paper cites Llama 3.2: Revolutionizing edge AI and vision with open, customizable models.

Scaling On-Device GPU Inference for Large Generative Models Llama 3.2: Revolutionizing edge AI and vision with open, customizable models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.887304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.452475Z digest=sha256:ea3b0ca1df1536bbfdb8a66bc07a761ebe6ae8c13975f15e667a6e26c1dd671c

Observation c0838d42-45f4-4b9e-a3a1-48effb96e14d · outbound

This paper cites DirectML Overview.

Scaling On-Device GPU Inference for Large Generative Models DirectML Overview

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.876047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.456091Z digest=sha256:55598ad090d1e57c6dcedbddbd91bc625f97901343e64884e3945c8b06dc9080

Observation af0ad225-015b-4b94-b423-7c10975a72c6 · outbound

This paper cites Get started with ONNX Run- time Mobile.

Scaling On-Device GPU Inference for Large Generative Models Get started with ONNX Run- time Mobile

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.864756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.459461Z digest=sha256:51d6b4d95d7da9eeda83fda19216bb3e323c87c78347b69822952ffb49d6be35

Observation d0268f61-6ea7-4c56-9c13-91e0b67c767b · outbound

This paper cites Stable Diffusion Op- timization with DirectML.

Scaling On-Device GPU Inference for Large Generative Models Stable Diffusion Op- timization with DirectML

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.854183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.463215Z digest=sha256:650a328583a3e97740fbe143789e1a9f007c4a65ab9fa45857d2acc5b90ee4c5

Observation bfa471c6-016f-43ce-96da-0bb3acfc7576 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.842022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.466325Z digest=sha256:e3d10d93d660d88ec1677978003fd1aed3d417e87f09c7a80c6a0e63ba480300

Observation 67262328-6b4d-4379-a936-70c921ad655a · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.828786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.469475Z digest=sha256:23b42821117b4d232d4cea2853a861d38897446d4d0d922ab6c461d6032f397e

Observation 96c5db1d-d03b-4196-b9ad-5167ee423445 · outbound

This paper cites NVIDIA Tensor Cores.

Scaling On-Device GPU Inference for Large Generative Models NVIDIA Tensor Cores

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.816678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.473074Z digest=sha256:09e78107b0e45571240d9ef9b2804b0ba699d31a3a34f0fe4313c34c80ea4cb2

Observation f4b38987-3aeb-4e27-a84d-b1d2b3938805 · outbound

This paper cites TensorRT.

Scaling On-Device GPU Inference for Large Generative Models TensorRT

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.803472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.475989Z digest=sha256:db471fbcee1732aa444a58053915bd10d27508a4303ce331abef7d904ccf7b27

Observation 17bc6ceb-1dde-4cc4-9ee9-81596a10cf25 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.792711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.479341Z digest=sha256:cd534a6df79f86f9eac228aad2fad76266eec09673b2a3ab98b4c2d8469d1ff9

Observation 2572a02f-e45d-4571-9175-ec554570753e · outbound

This paper cites Efficient Memory Manage- ment for Deep Neural Net Inference.

Scaling On-Device GPU Inference for Large Generative Models Efficient Memory Manage- ment for Deep Neural Net Inference

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.782707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.482345Z digest=sha256:240d6175f5f03609a47b936a4a96c412bea6f0cc3c789e9a9a79ad4157195eda

Observation c1f17b16-a16d-4582-99ac-55d254b36740 · outbound

This paper cites Snapdragon Neural Processing Engine SDK.

Scaling On-Device GPU Inference for Large Generative Models Snapdragon Neural Processing Engine SDK

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.771094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.485334Z digest=sha256:2f49206b08cf794bf1901803049a3fbaefebaf9fb6beecdd017007a95cd059e6

Observation cfda3ff7-68cc-4da5-914f-db7745bd842e · outbound

This paper cites QualComm AI Hub Llama-v3.2-3B-Chat.

Scaling On-Device GPU Inference for Large Generative Models QualComm AI Hub Llama-v3.2-3B-Chat

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.760222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.488945Z digest=sha256:d814b33756a995b0836dcf5215555216ce82a35c02d22952a955fef88f880022

Observation 5e70183f-a43b-4b95-81a6-02a97d5caadb · outbound

This paper cites World’s first on-device demonstration of Stable Diffusion on an Android phone.

Scaling On-Device GPU Inference for Large Generative Models World’s first on-device demonstration of Stable Diffusion on an Android phone

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.749703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.492697Z digest=sha256:4ae130b6d02dc4d5e9b3cfd31e55a0f297ae68766143bfd4630ae015c61360d7

Observation 9ed90683-f398-487f-a7d3-b88063eda807 · outbound

This paper cites Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen.

Scaling On-Device GPU Inference for Large Generative Models Ruan, Yucheng Qin, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, and Tianqi Chen

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.737091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.496210Z digest=sha256:f793844f3ea01f2dfdd3fc455f92e215a9302122f88050094c21f3f18985fad4

Observation f547c7bc-5890-481e-b1d7-8b8c3decf77e · outbound

This paper cites XLA: Compiling Machine Learning for Peak Performance, 2020.

Scaling On-Device GPU Inference for Large Generative Models XLA: Compiling Machine Learning for Peak Performance, 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.725717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.499934Z digest=sha256:bdf33d354e49a483fd1c78ada3489efad2be781ac8f5037707c407bf802abed3

Observation 10aef569-a863-4b62-a5e4-f42a03f003c5 · outbound

This paper cites Introducing Stable Diffusion 3.5.

Scaling On-Device GPU Inference for Large Generative Models Introducing Stable Diffusion 3.5

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.713687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.503769Z digest=sha256:49a78ff01f684f63212969828a6d79b0e29c2bfc0794b7dd00920ac0a30d8cbe

Observation 7fa999f2-5162-44ad-a18e-57d302ee4093 · outbound

This paper cites Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile.

Scaling On-Device GPU Inference for Large Generative Models Introducing torchchat: Accelerating Local LLM Inference on Laptop, Desktop and Mobile

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.703132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.507604Z digest=sha256:9952c1a42c967a1a943b0700d5e1af5ab6e91cb4953737e05b5180c948f774da

Observation 8524c447-e6d9-4edc-9d07-816c3c206f37 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.691885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.511990Z digest=sha256:2a4f684b0b1d0829440c47c66e3f7f6a0ff7e3144278c4c9191e2249906c38ed

Observation 6adc8945-0d67-42f6-8bbd-27b2fb14f743 · outbound

This paper cites Dawn, a WebGPU implemen- tation.

Scaling On-Device GPU Inference for Large Generative Models Dawn, a WebGPU implemen- tation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.680108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.515581Z digest=sha256:21198dddb48791094ebd3fe77730d2b361580f84c892a0e9c7f93829578204ff

Observation 050b6cd3-07d8-4802-aa26-7d3b9057637f · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.668575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.519818Z digest=sha256:c27a6ed5670bb6f30ba07ef9b57a4404010c28912f13af83977807712cdeefc8

Observation 3b55a910-0d4c-4677-8cf2-78534dd35c7e · outbound

This paper cites SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els.

Scaling On-Device GPU Inference for Large Generative Models SmoothQuant: Accurate and Effi- cient Post-Training Quantization for Large Language Mod- els

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.655869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.523550Z digest=sha256:6aac88040e51ecc6ffc7df4ec60b001ae51d970362673f3bf7b9ab7518db7de1

Observation 4ebfc85a-c516-44d0-8f83-2390397b2c97 · outbound

This paper cites an unresolved cited work.

Scaling On-Device GPU Inference for Large Generative Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:51:31.643405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.527151Z digest=sha256:5b658598b976945d7d1805c063474c0cc6b1523ac4ea73701b819b5cf56ac0a8

Observation 4c0350f0-bd82-43d1-a93e-54649577265e · outbound

This paper cites LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023.

Scaling On-Device GPU Inference for Large Generative Models LLMCad: Fast and Scalable On-device Large Language Model Inference, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.632711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.530744Z digest=sha256:db2263177cc86483fa2da0295da6d558daf2d7053380fe727976acc244820c7d

Observation d73f38fa-cffa-4fa2-b165-fbec16e9cc78 · outbound

This paper cites Fast On-device LLM Inference with NPUs.

Scaling On-Device GPU Inference for Large Generative Models Fast On-device LLM Inference with NPUs

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.621573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.534059Z digest=sha256:d5750d0d50ad64d4de1632c3b0b393e6687be0a33077989f949dae034b2ab6f9

Observation 339db099-dd19-46ef-be92-964ead0a53a3 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

Scaling On-Device GPU Inference for Large Generative Models PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:51:31.537714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:51:31.537714Z digest=sha256:95b1cfa2810d11b3307136897722dbc3415b82b5c72d75577826f57fd2b34d3c

Observation 984377af-19f7-4a48-9783-9de59c1de3fc · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

Scaling On-Device GPU Inference for Large Generative Models Gonzalez, Clark Barrett, and Ying Sheng

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:51:31.609190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T04:51:31.541760Z digest=sha256:74c5eace02687299f42fb6bb19d3be1a2e312c1367f0132f5b4d517865b7b15b

Pith citing papers

Observation 035f8d65-6761-4107-896c-35502c5af71f · inbound

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions cites this paper.

EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions Scaling On-Device GPU Inference for Large Generative Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:57:30.740928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:57:30.372775Z digest=sha256:357dde76fb6300f02a331c95c238a93fe16ae5d5b1214ccb6f1084135c752ac3