Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2403.11421.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:38:37.472671Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T15:30:17.948710Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 1aabc17c-9418-4901-9a3f-caef676b27bc · inbound
A Survey on Efficient Inference for Large Language Models FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 252
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25305e44-e40b-428b-8584-0e4fcba07bfe · inbound
Kinetics: Rethinking Test-Time Scaling Laws FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db91454-2a3d-4e31-908e-ff723db9d091 · inbound
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b096e79-904e-4e7a-9c35-dab5e3ecdf9f · inbound
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9c97e9d-4fe5-4fba-ba4d-0a63006dc951 · inbound
Learn from the Past: Fast Sparse Indexing for Large Language Model Decoding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5737006c-aef0-4827-b38a-d548350d40ee · inbound
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3104cf80-a1c0-4618-84aa-fcf5c0feb196 · inbound
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 950360d6-18a2-4927-9922-c48f3179ece1 · inbound
HCAttention: Extreme KV Cache Compression via Heterogeneous Attention Computing for LLMs FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8225a2-deb0-4c9f-b07a-c421677e2b82 · inbound
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f6b9d54-a29c-4d28-973a-c09b00bd5bb0 · inbound
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 90749329-89f8-4a3d-ba2b-73467a013d77 · inbound
Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a7c8d74-18bd-448f-848f-31ac9a830ebc · inbound
Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a05612f-965f-40c6-acb2-8d82e7d567af · inbound
Understanding Rate-Distortion Performance in Distributed Transformer Inference FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63dae347-e822-4f8a-80ad-d723fa99cf1c · inbound
DualDecoder: Accelerate Long Context LLM Inference by Predictive Prefetch FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3cec5c8-1849-4e2e-89b7-01434f9dcc5b · inbound
Training-Free Hashing-Based Attention via Binary Principal Components FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.