Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T15:03:31.289211Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2605.27435.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T15:03:31.289211Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 72c2fc1c-2789-47f1-a691-d88637bc627a · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference A Survey of LLM Inference Systems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4cd305fd-a733-450c-b296-270d1c1580d9 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Empowering edge intelligence: A comprehensive survey on on-device AI models,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5474a78b-3d68-46c3-81d5-f5794b57b03b · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference LLaMA 3.2: Vision and edge-optimized models for multi- modal and mobile AI,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a8cdc4c-220d-4b2d-95c9-f8ebc3c82caf · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Qwen3 Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 01f58251-7609-4ccc-bd02-7a78e73df3e5 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 729cc128-fa7a-4ba3-b2e9-575a32ef2bff · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference AWQ: Activation-aware weight quantiza- tion for on-device LLM compression and acceleration,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 889cde39-ecb6-414b-b0ff-bda5ba117faa · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference SmoothQuant: Accurate and efficient post-training quantization for large language models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d56d4330-afff-4cf2-b55b-6864e4d2789b · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Scaling Laws for Precision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 21935cb9-172e-44a0-b139-cc05bb947124 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Efficient memory management for large language model serving with PagedAttention,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 16c7ed9f-6b6e-4285-a022-2b2a5177799d · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference MLC-LLM: Universal llm deployment engine with ml com- pilation,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 193de604-d52b-485a-9bf1-930be28d0665 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Fast on-device LLM inference with NPUs,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6e60fec-694c-40db-a75a-e7b687591a5c · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1bbf492-22d1-4b26-89be-f67b6f343da2 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Flightllm: Efficient large language model inference with a complete mapping flow on FPGAs,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f0ba2ee1-1d48-4343-abd1-db478141dbaf · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference MLPerf Mobile Inference Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5d1dd023-4a81-4d4f-8038-6656acde83c4 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Characterizing mobile SoC for accelerating heterogeneous LLM inference,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 579fe8a0-dfef-46b8-ac71-1e6ab1fd8449 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Scaling llm test-time compute with mobile npu on smart- phones
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6613b2e4-7add-4248-9952-8a083720d293 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6df9f1df-40f1-4179-8587-ec2d85c833c9 · outbound
When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference llama.cpp: Llm inference in c/c++,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
No inbound Pith citation observations are available.