Pith. sign in

Paper Citation Record · LEDGER

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2508.12590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12590 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:31:39.052107Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13ffb03c-64ba-4c7f-8177-5843e13104ef · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.461160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.002774Z digest=sha256:63518af72f774547cf24c7607a592dc8a8a03682ed4767add42f263154db701b

Observation e68eeda1-f9d8-4850-a7e5-c2f12fb57fdd · outbound

This paper cites Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.005967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.005967Z digest=sha256:17ae04e1aa7ef4795e76efc975550774c32e3f52256cd94f4011df1988ca6a37

Observation 90689d9d-85e8-4f72-83f0-02512697cc30 · outbound

This paper cites Hybrid slm and llm for edge-cloud collaborative inference,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Hybrid slm and llm for edge-cloud collaborative inference,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.451366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.009391Z digest=sha256:2c36176db8ce3cc599a017f24f379ceede55872889863840b0f9922bff33d7b0

Observation fbdc6fd9-d6e9-4907-ab36-08da220991de · outbound

This paper cites Fast inference from trans- formers via speculative decoding,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Fast inference from trans- formers via speculative decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.441705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.012982Z digest=sha256:dbbb07b1a79961fa13ffe2c4f3d775d8999326c731a490c3a4fec2586f8a659a

Observation 81bbfd8e-ab94-4f7e-8c4a-830767f993fe · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.016413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.016413Z digest=sha256:4c766006792a2faabb97bf20a7e468d5aad9b207d1081c6ad4c5b5d04be23c85

Observation 4d0d7295-cda8-4254-a264-dfb25f134a61 · outbound

This paper cites Uncertainty-aware hybrid inference with on-device small and remote large language models,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Uncertainty-aware hybrid inference with on-device small and remote large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.431631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.019737Z digest=sha256:a1ab35038ec3be60ede2ce6a73699492d34d87d8bebffdd595d085d848df5203

Observation dbfd4561-579e-4f76-b464-79dd46aef187 · outbound

This paper cites Uncertainty-aware opportunistic hybrid language model in wireless robotic systems,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Uncertainty-aware opportunistic hybrid language model in wireless robotic systems,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.411228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.022856Z digest=sha256:a1eeccf03bd4d4ade064603c5484b7f2252dd59e5ec18b9d9fcbbd4ea9257464

Observation f0b85060-eb4e-4bc8-9d85-93284ee55516 · outbound

This paper cites Seeing far and clearly: Mitigating halluci- nations in mllms with attention causal decoding,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Seeing far and clearly: Mitigating halluci- nations in mllms with attention causal decoding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.397924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.025859Z digest=sha256:9ae5ca370ec76614646a5e981061aa38b4987fa0aa60b81eea95a69f8de86d15

Observation 3c9b5960-3ac4-42fd-9706-84f200110ca4 · outbound

This paper cites Understanding the metropolis-hastings algorithm,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Understanding the metropolis-hastings algorithm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.378874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.028602Z digest=sha256:d0cf14e1356ef1fbc3bba148b31f558468866acabbab06c98a6f85435ca98a05

Observation 130c4ab6-96f5-4c1c-a11f-e393be730b73 · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.031313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.031313Z digest=sha256:21389752a9d95e27ee89a6a36972d3fb56f8c4fe4ad063daabf37640e357ac3d

Observation 9b3dc1b0-98b5-458e-a22b-c7ebd21d54ec · outbound

This paper cites Latte: Low-precision approximate attention with head-wise trainable threshold for efficient transformer,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Latte: Low-precision approximate attention with head-wise trainable threshold for efficient transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.341197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.034625Z digest=sha256:24e7118a33ec88751b0525fa10e95f0b3c63f4a41de41751647832da52545972

Observation afd689fe-6809-4571-98fd-d42752e95ac7 · outbound

This paper cites Energon: Toward efficient acceleration of transformers using dynamic sparse attention,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Energon: Toward efficient acceleration of transformers using dynamic sparse attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.259478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.037367Z digest=sha256:987a3d6f35f28a67d938eed28d01da4611b45ce98b99b76e02f612a78a3d018d

Observation fae26586-119d-4841-ba5f-6bd02c280899 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding TinyLlama: An Open-Source Small Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.039834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.039834Z digest=sha256:7e1eaa99b4f584bbad55973672d9ab6eb354024ea9debb221712505a27d9182f

Observation 079bf107-d6d2-426b-9f3b-5d8852fbc3a6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.043735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.043735Z digest=sha256:470697d7c91e8cecc3c0d543da6730cfcb2966357ca4276f06319f4751f95200

Observation 966da510-63e7-4b40-a262-12ba3c6448d5 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Stanford alpaca: An instruction-following llama model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.046483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.046483Z digest=sha256:94d972fba60e924fb7a2ba0d60e1a41167109aa9ca92cfdcee7d6e797d9bc757

Observation ff3167e8-2ec7-4f99-bb73-feafd452aaf7 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding BERTScore: Evaluating Text Generation with BERT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.049209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.049209Z digest=sha256:63110d8f31827aaacdde9dfaa89536d3b513a9e5a917d8ab8f0937648b31ac4a

Observation 36761e8f-6eb2-4193-8387-555625992248 · outbound

This paper cites Joint computation and communication cooperation for energy-efficient mobile edge comput- ing,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Joint computation and communication cooperation for energy-efficient mobile edge comput- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.162658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:31:39.052107Z digest=sha256:69fd533f7058ebd647b4ecede09688b2364e67c1b3010fdc0443a586d2d83f60

Pith citing papers

No inbound Pith citation observations are available.