Pith. sign in

Paper Citation Record · LEDGER

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2312.12456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12456 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:26:17.238480Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.594987Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0f5232b-ef16-4177-bc9e-752bc21b198a · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 274

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.808663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:850f9bd505d25d9e78e6ee487816cfb784769f8acce81c3db91d5c0d8ab594cf

Observation 9bfc00a5-a9a7-4688-9c6e-dacdc19a8d58 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.393963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:d724e2f46fe502c0223b57f89ea69d2a327f4a53fde63ea593207cee610d4dde

Observation 025dc3cd-c52b-4e23-91cd-2d17303e6d9e · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.891930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:35ace58faf238726ba7f170f2ae0d4c7b81c3944ec9ef318559545d23c736a01

Observation 69b87f83-9793-4573-9514-e99b11547a90 · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.238480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.238480Z digest=sha256:5572e1143ac86c12c6b841cdc1ff38bd93a3a767a9500d2c2726f1b6c0dc60cd

Observation 8865263a-8f4e-44a8-b067-b61b5be100ea · inbound

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts cites this paper.

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:17.159468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:17.159468Z digest=sha256:2af7374c31924b18c4a2e20705b3d5fc79defb4b946a94ad95a59d923f33c2a7

Observation 5568e98f-9628-4a2d-b939-7a04aeaab1fb · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.173785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.173785Z digest=sha256:53e4edb56a7307266d475d47fbf714a000dd4fe0dc2bae0dcae540799a734c33

Observation 0aeb7da2-92c1-4501-bf32-4b3b6ad58c68 · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.384466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.384466Z digest=sha256:499b2e0baa77541f58d6b4a12103102812288989be1d3a34062cb7ecbbafdc10

Observation 5e860b4a-708f-43e3-aad1-8ed942bdf3a0 · inbound

Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage cites this paper.

Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:57:31.381924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:57:31.381924Z digest=sha256:63285b51ea290c845288952a3cf28f458b975d94e6e2892c14ddaa58614fde6e

Observation 2950b28d-ac6b-4d06-9bbf-fcdf52566cca · inbound

A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering cites this paper.

A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:17:30.724777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:17:30.724777Z digest=sha256:b08797f1b29786a3567093d0d23cfe2b4e1fa7332c3028781c0137422b5625e0

Observation 4c185bc3-56b6-4eaa-a43f-8583afd93821 · inbound

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy cites this paper.

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.434711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:46:52.811371Z digest=sha256:865237b4d87e45565319749fcb6decf963240ae66fa0d5c0e298330693994846

Observation f469afaf-17f4-40c8-abdf-4732a5164f8e · inbound

Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models cites this paper.

Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:27.716071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T10:05:31.211127Z digest=sha256:26f429ea59a19528300222feae4713c6ca11004434588915a93ccd4070f04756

Observation b3e29c5b-c69b-4ebc-8251-d4c5128ee048 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.689914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T08:30:35.264804Z digest=sha256:8f32edb499d6e7deb54ba37b378c57417362e58a9507ae3168ba6d0dc6e0b6e8

Observation 0edf9054-3d03-4b0f-b0e9-96ef98c2bd5f · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:34.733096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T08:54:41.845820Z digest=sha256:c25164086880b455dd001bb22083338be18cb60ae044369faee759eab44fa205

Observation 257da112-d13a-4153-bd07-b2d47146e05d · inbound

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories cites this paper.

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:13.435370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:11:57.929452Z digest=sha256:30683097bc412a12462111eed2c71fa5d50def7cb6519e5d4f02c90fad7a0920

Observation 0f38a8d6-31a3-4311-a2a7-6456ae94e70b · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.596383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:baad633df4b23be3a259f32f6537d347ea8f417a32788007c8c936c99a9aba28

Observation ef95162c-6e8d-4289-8c31-7a1c19ea98ea · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:f22be60966e96994c5f01ea2cda70001003940cb065f44a586173af1ad5a39ed

Observation 6698b789-2425-45fc-bdd7-130a55fb1c32 · inbound

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference cites this paper.

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T03:14:02.219678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:14:02.219678Z digest=sha256:4f6da0a6644573fa8a13cd51b66374c156143814e56ce5e75288f506bc0e0cd7

Observation e6d75a5c-45f6-461b-82ac-cd4a75d5742f · inbound

Transition-Aware Backend Dispatch for Edge LLM Inference cites this paper.

Transition-Aware Backend Dispatch for Edge LLM Inference PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:57.394979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:57.394979Z digest=sha256:a9d4be9141bfc09936b9985042eb3086b33d959c885c2f3bf242959d1f035221

Observation dd1d5176-b30e-48b7-8aa1-d750a2da225a · inbound

A CXL Memory Rack for Multi-Turn LLM Serving cites this paper.

A CXL Memory Rack for Multi-Turn LLM Serving PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T15:58:29.909466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:58:29.909466Z digest=sha256:0e55f93d1e44a993a7eb1cd118f9704c664f63c795391f033337b46be00cd5cc