Pith. sign in

Paper Citation Record · LEDGER

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.12456.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.12456 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:31:35.806995Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.594987Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0f5232b-ef16-4177-bc9e-752bc21b198a · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 274

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.808663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:f30d3d7d73a2c64098b32ab687f596dcf36de6f386d4f5fd1e1f777b867746e4

Observation 9bfc00a5-a9a7-4688-9c6e-dacdc19a8d58 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.393963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:2422403d4bedcd3f83fde19404e19b9dbd92c9c127b04a41c5ccb29eab8bd7d7

Observation 025dc3cd-c52b-4e23-91cd-2d17303e6d9e · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.891930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:d5358abaa1c6264a2db03291c4d33f5dc7a8d53853be104e2e67270434374e9a

Observation c8978d1f-3d8f-4ff8-b6ed-06c6ef73bddf · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.806995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.806995Z digest=sha256:5ffc0dfcbec6f58a67e5a304119948daabc5fc5abd4a0e3737f9d1ecc47b629f

Observation 69b87f83-9793-4573-9514-e99b11547a90 · inbound

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models cites this paper.

CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for Comprehensive Acceleration of Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:17.238480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:17.238480Z digest=sha256:5572e1143ac86c12c6b841cdc1ff38bd93a3a767a9500d2c2726f1b6c0dc60cd

Observation 8865263a-8f4e-44a8-b067-b61b5be100ea · inbound

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts cites this paper.

DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:17.159468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:17.159468Z digest=sha256:bbbdd65f2aaf15af6672d00fe422d8f3dd920ec0d05bc5ba6f7fa31223c80a5f

Observation 5568e98f-9628-4a2d-b939-7a04aeaab1fb · inbound

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU cites this paper.

Breaking the Boundaries of Long-Context LLM Inference: Adaptive KV Management on a Single Commodity GPU PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:11.173785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:11.173785Z digest=sha256:53e4edb56a7307266d475d47fbf714a000dd4fe0dc2bae0dcae540799a734c33

Observation 0aeb7da2-92c1-4501-bf32-4b3b6ad58c68 · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.384466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.384466Z digest=sha256:499b2e0baa77541f58d6b4a12103102812288989be1d3a34062cb7ecbbafdc10

Observation 5e860b4a-708f-43e3-aad1-8ed942bdf3a0 · inbound

Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage cites this paper.

Toward Efficient SpMV in Sparse LLMs via Block Extraction and Compressed Storage PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:57:31.381924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:57:31.381924Z digest=sha256:63285b51ea290c845288952a3cf28f458b975d94e6e2892c14ddaa58614fde6e

Observation 2950b28d-ac6b-4d06-9bbf-fcdf52566cca · inbound

A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering cites this paper.

A Sparsity Predicting Approach for Large Language Models via Activation Pattern Clustering PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:17:30.724777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:17:30.724777Z digest=sha256:dd39fdb169482c09ff6b7f7c8834d7a555c0de1c5e42f0ef77bb2e0561218d9d

Observation 4c185bc3-56b6-4eaa-a43f-8583afd93821 · inbound

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy cites this paper.

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.434711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T06:46:52.811371Z digest=sha256:e0d077a2e7647a92f861a008cf39408e69c347e25dca84d94dab2fd2fc3638f0

Observation f469afaf-17f4-40c8-abdf-4732a5164f8e · inbound

Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models cites this paper.

Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:27.716071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T10:05:31.211127Z digest=sha256:d09c53775f7c1bedba4a7c0fc2499e6b87081539a22c307bac5fb2d04bdd467b

Observation b3e29c5b-c69b-4ebc-8251-d4c5128ee048 · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.689914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T08:30:35.264804Z digest=sha256:421e98ce2d8ef9a0e1d93f5243d28833cf8f03cfe9f2e444c47aacc08a37fba4

Observation 0edf9054-3d03-4b0f-b0e9-96ef98c2bd5f · inbound

EdgeFM: Efficient Edge Inference for Vision-Language Models cites this paper.

EdgeFM: Efficient Edge Inference for Vision-Language Models PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:34.733096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:54:41.845820Z digest=sha256:e9cc0e878b92ebe71de7720ccbf384c3f53994f2f054bd95e15168837d81476e

Observation 257da112-d13a-4153-bd07-b2d47146e05d · inbound

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories cites this paper.

ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:13.435370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:11:57.929452Z digest=sha256:ede5ddc28b9dbdd4a44b52014a32c7cfca0f0480dcca7ff5394023fa73c74ce5

Observation 0f38a8d6-31a3-4311-a2a7-6456ae94e70b · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.596383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:bd0e1ed017b08a60601a700cb4b78489dbc7a4595f4e522a0284e04ed9ffdc1f

Observation ef95162c-6e8d-4289-8c31-7a1c19ea98ea · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:f22be60966e96994c5f01ea2cda70001003940cb065f44a586173af1ad5a39ed

Observation 6698b789-2425-45fc-bdd7-130a55fb1c32 · inbound

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference cites this paper.

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T03:14:02.219678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:14:02.219678Z digest=sha256:4f6da0a6644573fa8a13cd51b66374c156143814e56ce5e75288f506bc0e0cd7

Observation e6d75a5c-45f6-461b-82ac-cd4a75d5742f · inbound

Transition-Aware Backend Dispatch for Edge LLM Inference cites this paper.

Transition-Aware Backend Dispatch for Edge LLM Inference PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:57.394979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:57.394979Z digest=sha256:315b517982e86e9f0a732dbc285f13285750fc9083ce7ca808328b7d8241ca9e

Observation dd1d5176-b30e-48b7-8aa1-d750a2da225a · inbound

A CXL Memory Rack for Multi-Turn LLM Serving cites this paper.

A CXL Memory Rack for Multi-Turn LLM Serving PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T15:58:29.909466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:58:29.909466Z digest=sha256:5da3f2dc72d40ecfc85a05f878fe9e39795c619ddd37d0417c6c3830412558d2