Pith. sign in

Paper Citation Record · LEDGER

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference

As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2605.27435.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27435 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T15:03:31.289211Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact7
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72c2fc1c-2789-47f1-a691-d88637bc627a · outbound

This paper cites A Survey of LLM Inference Systems.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference A Survey of LLM Inference Systems

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.172024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:2a503c6eb33f7063787963ef9fb4f2f127fe5c50be902dae5858fdc3c01ad2bd

Observation 4cd305fd-a733-450c-b296-270d1c1580d9 · outbound

This paper cites Empowering edge intelligence: A comprehensive survey on on-device AI models,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Empowering edge intelligence: A comprehensive survey on on-device AI models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.630511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:187c2301c1f16ed3879880f2452a461c02e0ce2939336f59c5a7b9e788ef22f0

Observation 5474a78b-3d68-46c3-81d5-f5794b57b03b · outbound

This paper cites LLaMA 3.2: Vision and edge-optimized models for multi- modal and mobile AI,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference LLaMA 3.2: Vision and edge-optimized models for multi- modal and mobile AI,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.628414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:93a4ad060e55b7f764ac38508c372c88778a440558b32e4b7b26ca06047dbe64

Observation 3a8cdc4c-220d-4b2d-95c9-f8ebc3c82caf · outbound

This paper cites Qwen3 Technical Report.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Qwen3 Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:04:46.185045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:1037f30d6595f8c84bf2d23790221568763cb03d0cd541f2e77ce75c509e26ff

Observation 01f58251-7609-4ccc-bd02-7a78e73df3e5 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:04:46.174674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:460c428ea3ee3235d69c23dcaabd6445112add5ccb430cf12384f5b416855a95

Observation 729cc128-fa7a-4ba3-b2e9-575a32ef2bff · outbound

This paper cites AWQ: Activation-aware weight quantiza- tion for on-device LLM compression and acceleration,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference AWQ: Activation-aware weight quantiza- tion for on-device LLM compression and acceleration,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.633922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:3af82d35b2f995d3f6aa65c8abf727377182248f97e36f4dc2d4dd9129f9a72a

Observation 889cde39-ecb6-414b-b0ff-bda5ba117faa · outbound

This paper cites SmoothQuant: Accurate and efficient post-training quantization for large language models,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.625882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:b0256f380cf38cf0f204919d6cbf177ba7f2c95868f9c59693f3ff35d47a09e1

Observation d56d4330-afff-4cf2-b55b-6864e4d2789b · outbound

This paper cites Scaling Laws for Precision.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Scaling Laws for Precision

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T15:04:46.177292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:5d1cc56728a2a9de298c05d89d56bc0596c2a6167bb69c32c0f8d143e10cb820

Observation 21935cb9-172e-44a0-b139-cc05bb947124 · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Efficient memory management for large language model serving with PagedAttention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.622158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:043237b2571d6eccadd1d34a60413f7b4973c11e335d80f47b1237c75c23dcdd

Observation 16c7ed9f-6b6e-4285-a022-2b2a5177799d · outbound

This paper cites MLC-LLM: Universal llm deployment engine with ml com- pilation,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference MLC-LLM: Universal llm deployment engine with ml com- pilation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.639236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:bf0dded7cb887b0669640e56d4c1b59399fe0250a485f0b56449731f7493212f

Observation 193de604-d52b-485a-9bf1-930be28d0665 · outbound

This paper cites Fast on-device LLM inference with NPUs,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Fast on-device LLM inference with NPUs,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.620264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:f6354809ee642ba6ea5c696ebd242d8bc654e2502f5d5258b0cc4f36fb9a4071

Observation e6e60fec-694c-40db-a75a-e7b687591a5c · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.182427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:2842c03b74367aab8b361c9073cfbf6738c80ef48a8410ababc40a5bf0bc7ebb

Observation f1bbf492-22d1-4b26-89be-f67b6f343da2 · outbound

This paper cites Flightllm: Efficient large language model inference with a complete mapping flow on FPGAs,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Flightllm: Efficient large language model inference with a complete mapping flow on FPGAs,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.624019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:1f53e6933194b5e9b485dc33459359986d8b03ee03ca037e2ffbd4f8d1905465

Observation f0ba2ee1-1d48-4343-abd1-db478141dbaf · outbound

This paper cites MLPerf Mobile Inference Benchmark.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference MLPerf Mobile Inference Benchmark

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.179771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:c10c9ad050310c85226ef6f1b98e31494c1e9623e26dc7d28cc2e9cdadb1a264

Observation 5d1dd023-4a81-4d4f-8038-6656acde83c4 · outbound

This paper cites Characterizing mobile SoC for accelerating heterogeneous LLM inference,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Characterizing mobile SoC for accelerating heterogeneous LLM inference,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.642901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:a37061e69927589f684f4d3b73ea76da66bd2424dd48269c507a575080bbce95

Observation 579fe8a0-dfef-46b8-ac71-1e6ab1fd8449 · outbound

This paper cites Scaling llm test-time compute with mobile npu on smart- phones.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Scaling llm test-time compute with mobile npu on smart- phones

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.187806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:4ce37feadaa775932e62d0fcc68c07c03adcd18adac7bb656f9a0e00f0aac9be

Observation 6613b2e4-7add-4248-9952-8a083720d293 · outbound

This paper cites Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.190798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:6663a2d4c10a4db7f9596275e5e7d7c4afdfd8e1eab187013ae4d87dec2d37ff

Observation 6df9f1df-40f1-4179-8587-ec2d85c833c9 · outbound

This paper cites llama.cpp: Llm inference in c/c++,.

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference llama.cpp: Llm inference in c/c++,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T18:45:27.641047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T15:03:31.289211Z digest=sha256:4ae4b0473567286a72327435b6bff874d83928013bf70c2033c3ed516f962824

Pith citing papers

No inbound Pith citation observations are available.