Pith. sign in

Paper Citation Record · LEDGER

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding

As of 22 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2508.12590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12590 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:31:39.052107Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13ffb03c-64ba-4c7f-8177-5843e13104ef · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.461160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.002774Z digest=sha256:6ae65e905830ac2de4529a75a46c5d033409650993bb8eec01e91dc3ee3ef07a

Observation e68eeda1-f9d8-4850-a7e5-c2f12fb57fdd · outbound

This paper cites Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.005967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.005967Z digest=sha256:55e729ced294e3865381b9ee6c33c386c7fc5426f6a6446f163212a8d8eb9f8f

Observation 90689d9d-85e8-4f72-83f0-02512697cc30 · outbound

This paper cites Hybrid slm and llm for edge-cloud collaborative inference,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Hybrid slm and llm for edge-cloud collaborative inference,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.451366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.009391Z digest=sha256:6a734d3814f77f65e044acd52b730e4901dae03de3ba95bdd4333ea64ffe98ca

Observation fbdc6fd9-d6e9-4907-ab36-08da220991de · outbound

This paper cites Fast inference from trans- formers via speculative decoding,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Fast inference from trans- formers via speculative decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.441705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.012982Z digest=sha256:c5b534e2bf91f902762f5dba37634e9a3af7d0f89e0fd26ae5fae9c10d28e607

Observation 81bbfd8e-ab94-4f7e-8c4a-830767f993fe · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.016413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.016413Z digest=sha256:7014201fd32031e410f1c75013e65aa7b5981c6a53efd9967ac3e0a05bb0b2d1

Observation 4d0d7295-cda8-4254-a264-dfb25f134a61 · outbound

This paper cites Uncertainty-aware hybrid inference with on-device small and remote large language models,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Uncertainty-aware hybrid inference with on-device small and remote large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.431631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.019737Z digest=sha256:bb67ed6e3512bb35f715d48860175e7980c5d33c7c4a27db2164a06e2d328f1a

Observation dbfd4561-579e-4f76-b464-79dd46aef187 · outbound

This paper cites Uncertainty-aware opportunistic hybrid language model in wireless robotic systems,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Uncertainty-aware opportunistic hybrid language model in wireless robotic systems,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.411228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.022856Z digest=sha256:b1d0a5c9d0fb72d37db85b82954f75a91fb64b018c0db4a00d2da8123bd4ecf1

Observation f0b85060-eb4e-4bc8-9d85-93284ee55516 · outbound

This paper cites Seeing far and clearly: Mitigating halluci- nations in mllms with attention causal decoding,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Seeing far and clearly: Mitigating halluci- nations in mllms with attention causal decoding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.397924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.025859Z digest=sha256:107ee580a1a396b55b059797ec1628cfa385a7496577da47761c1a7f76d77b16

Observation 3c9b5960-3ac4-42fd-9706-84f200110ca4 · outbound

This paper cites Understanding the metropolis-hastings algorithm,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Understanding the metropolis-hastings algorithm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.378874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.028602Z digest=sha256:f4c5c8e8e7ae6f6901b6c31f105cf02eb91e87883bfa7fdd5bec9606635014d2

Observation 130c4ab6-96f5-4c1c-a11f-e393be730b73 · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.031313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.031313Z digest=sha256:a87029b078b41f5014a0705f051ce7051481106639f64c3c77d2199e50603540

Observation 9b3dc1b0-98b5-458e-a22b-c7ebd21d54ec · outbound

This paper cites Latte: Low-precision approximate attention with head-wise trainable threshold for efficient transformer,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Latte: Low-precision approximate attention with head-wise trainable threshold for efficient transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.341197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.034625Z digest=sha256:d987c8d11a5f3f54c6ada417149e58d1dce5c784843628ec4f53651d47b44833

Observation afd689fe-6809-4571-98fd-d42752e95ac7 · outbound

This paper cites Energon: Toward efficient acceleration of transformers using dynamic sparse attention,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Energon: Toward efficient acceleration of transformers using dynamic sparse attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.259478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.037367Z digest=sha256:a04374bd163fc7c9849763e53aa30d82ffddf4f862c687c31dd5005b76e3589f

Observation fae26586-119d-4841-ba5f-6bd02c280899 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding TinyLlama: An Open-Source Small Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.039834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.039834Z digest=sha256:d3441c349120f6333ac74ce783a22f8c5a9a735af945eb854196df49b094c9f3

Observation 079bf107-d6d2-426b-9f3b-5d8852fbc3a6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.043735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.043735Z digest=sha256:8d629a748f2e8117740014862d78fad081f975e2ed2891fdbaeda77d13e23ea5

Observation 966da510-63e7-4b40-a262-12ba3c6448d5 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Stanford alpaca: An instruction-following llama model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.046483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.046483Z digest=sha256:74d508e732c12e2c703f5dff5123a699f1394f98637cacd6ec13285135dc1fee

Observation ff3167e8-2ec7-4f99-bb73-feafd452aaf7 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding BERTScore: Evaluating Text Generation with BERT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.049209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.049209Z digest=sha256:868d9aa76f14a7252e8e54035853f81ca7558142240417e1659d253b87b55c77

Observation 36761e8f-6eb2-4193-8387-555625992248 · outbound

This paper cites Joint computation and communication cooperation for energy-efficient mobile edge comput- ing,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Joint computation and communication cooperation for energy-efficient mobile edge comput- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.162658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T19:31:39.052107Z digest=sha256:ebf0e323a8b477e74aa39fdc3a86b56921b0407929f0c044be0dcfa162896b27

Pith citing papers

No inbound Pith citation observations are available.