Pith. sign in

Paper Citation Record · LEDGER

FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2403.11421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.11421 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:37.472671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:30:17.948710Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1aabc17c-9418-4901-9a3f-caef676b27bc · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 252

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:39:33.397582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:9a8e5971aa9929677d1ccc6ffa98519389579d094facf352bd566e2dcc1a107d

Observation 25305e44-e40b-428b-8584-0e4fcba07bfe · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.224262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.224262Z digest=sha256:966973dce92a2695c3bcd11d178dfc1128b7e24524926f93f7a85dd23403cbc5

Observation 1db91454-2a3d-4e31-908e-ff723db9d091 · inbound

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation cites this paper.

Beyond the Buzz: A Pragmatic Take on Inference Disaggregation FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:32.741822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:32.741822Z digest=sha256:c6ecf152981d8d39b5934223852b92e05042813034d9a8ac3a2a568c4537ce89

Observation 1b096e79-904e-4e7a-9c35-dab5e3ecdf9f · inbound

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts cites this paper.

MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:52.754758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:37:52.754758Z digest=sha256:2c63bb9f80975ba9970a149804154012caf781a8d5af993fb58d9bb0b11729c5

Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · inbound

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding cites this paper.

Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:38:37.472671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:38:37.472671Z digest=sha256:18f3fb5d854eef32c1c99f941da4cadbded076f1dac614258fcc8e241b9fd4f4

Observation 5737006c-aef0-4827-b38a-d548350d40ee · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:32.334492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:32.334492Z digest=sha256:c2def34ae9d29a8f642d61ca9ba8db5d39a5a43fc5116b1e70135c7daa7dc57b

Observation 3104cf80-a1c0-4618-84aa-fcf5c0feb196 · inbound

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding cites this paper.

KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:20:26.333389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:20:26.333389Z digest=sha256:2a75362c612d720bbb88be118cfe2fe509a4bf9e1301e68c0af67fb9e11e40f2

Observation 950360d6-18a2-4927-9922-c48f3179ece1 · inbound

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs cites this paper.

HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:06:37.907873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:06:37.907873Z digest=sha256:e0e10dcc9452fda08e201ca94c4f42d52dd97b21917201a38fd8c8bacec6d94a

Observation 3e8225a2-deb0-4c9f-b07a-c421677e2b82 · inbound

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism cites this paper.

Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T20:53:31.847951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:53:31.847951Z digest=sha256:cde397d52a9e1bbc676f88eed4c66750967f91a92a8667f3e6f23aadd4c5ee82

Observation 2f6b9d54-a29c-4d28-973a-c09b00bd5bb0 · inbound

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips cites this paper.

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:30:17.950957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T15:26:01.283448Z digest=sha256:00ad5750f7efca030c2894ca19eef45e47c08419a77e5145ae9b9dfd64cc7881

Observation 90749329-89f8-4a3d-ba2b-73467a013d77 · inbound

Understanding Rate-Distortion Performance in Distributed Transformer Inference cites this paper.

Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:02:42.597700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:01:53.390831Z digest=sha256:bb6b167f5e2be94cee00b6bb122263db91b042d21c42179a08c0c3c690c22a28

Observation 8a7c8d74-18bd-448f-848f-31ac9a830ebc · inbound

Understanding Rate-Distortion Performance in Distributed Transformer Inference cites this paper.

Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T06:51:13.983136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:51:13.983136Z digest=sha256:10ba1b6bff72223cad0f7143121806b19ae1835155d1c6d099ae27e0ff409bba

Observation 0a05612f-965f-40c6-acb2-8d82e7d567af · inbound

Understanding Rate-Distortion Performance in Distributed Transformer Inference cites this paper.

Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T06:23:13.506645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:23:13.506645Z digest=sha256:b3eddf668faaf2c3a00ac50d373766920c99e59d322b11f20c6f69c383d76287

Observation 63dae347-e822-4f8a-80ad-d723fa99cf1c · inbound

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch cites this paper.

DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:06:54.210093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:06:54.210093Z digest=sha256:69a861ee2ff489a4b9cab55178235ed88352a5940756b35d658f8c192b379ba2

Observation e3cec5c8-1849-4e2e-89b7-01434f9dcc5b · inbound

Training-Free Hashing-Based Attention via Binary Principal Components cites this paper.

Training-Free Hashing-Based Attention via Binary Principal Components FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:48.467163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:45:48.467163Z digest=sha256:cd3a38d807f0cc4cf5c273de78a28c16ad1e5fee3d0b4d4e603d3162a0874cb7