Pith. sign in

Paper Citation Record · LEDGER

FlashDecoding++: Faster Large Language Model Inference on GPUs

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2311.01282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.01282 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:21:54.206928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.798667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ef22293-e0aa-4262-89e0-326b6b8ee658 · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.799232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:9cd5c49f76787b35a2ef2887092ab79eb3d43642092146cc2e94176bdd655312

Observation ddec9caf-42b6-4a52-9843-c3892b099d91 · inbound

WaferLLM: Large Language Model Inference at Wafer Scale cites this paper.

WaferLLM: Large Language Model Inference at Wafer Scale FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.206928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.206928Z digest=sha256:9dc4cda45235cb2513232fdc2cb5efc71de04b401e6f358a95a6ca0c2d048ec7

Observation 751b4ea8-832f-4923-9adf-d3417765cea8 · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.679828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.679828Z digest=sha256:9d3c02be9173303ada1e4dc293236657902b412f7115b6bbfe162b8c10d38b7f

Observation ac17d14e-2ccb-4c22-adc6-cccc746d29ca · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.857737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.857737Z digest=sha256:90d1f7ee2b0196dbc588970ebb57f62eb0008202ff0be07e4c9ffa70b960f45c

Observation 67557a58-4e61-4f7d-9e52-63de85969a1b · inbound

QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm cites this paper.

QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:50.245539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:55:50.245539Z digest=sha256:ac0e5bbe906163292308dc5acd44fa94e5fa87d4937c6e734c44a108bb0ae32d

Observation a4678885-969e-4a5f-86b7-82f6600a50de · inbound

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations cites this paper.

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:28.045319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:28.045319Z digest=sha256:7a96a1b4384677bd0b5a9836b9c7ca4f4bfbef8bef73342f999b5cb253bf9393

Observation 08228ed5-dff6-4ef2-b535-85988aa99550 · inbound

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure cites this paper.

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:52.895202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:52.895202Z digest=sha256:d0b674e2498545870441394754e9e2442ecf86f6ef2db9c1e01afbe25fa7ffd4

Observation 5b123bcb-f408-452d-add1-7d409451bda2 · inbound

Past-Future Scheduler for LLM Serving under SLA Guarantees cites this paper.

Past-Future Scheduler for LLM Serving under SLA Guarantees FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.797553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.797553Z digest=sha256:e7cbdda4405feae46df759db094d0fb91cef11b31ba00b65aa6eafaad66850d0

Observation cffde461-3275-4a7e-b093-b1caaca2907a · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.142783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:f575dc1bd6e45010535eb10db3bc211cc323ea4a994598d1dfe44083008f6ccd

Observation 18b815b4-1dd2-4022-a073-4f091de5632a · inbound

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models cites this paper.

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:02.030920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:02.030920Z digest=sha256:5b6fe2b1f63aad5e6b85296e4908765c4e96f7a24d0d6fd3d0f76ca08442184b

Observation f319ba07-7273-4021-8f80-619847606956 · inbound

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation cites this paper.

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:00.032774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:10:40.858525Z digest=sha256:50ad4002805063170a2d04d9e810550cf16eaa939850956486872c597890c4b0

Observation 20f99678-05a1-4b93-a6aa-fdece8d46f93 · inbound

Prism: Symbolic Superoptimization of Tensor Programs cites this paper.

Prism: Symbolic Superoptimization of Tensor Programs FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:24.949228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:01:38.783602Z digest=sha256:76c1ad6923b863730636d4ca7c9a74a961994f94ebd34b613961d658df28d9f7

Observation ffda947e-5625-4717-b167-f7e1f580643f · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.233012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:21d10b2229fd19f30fb3ea8c28b4cefc41d91a97a6f5a111f3ca112547140479

Observation 3dccb04e-b102-4f83-801d-e8eac0eec47f · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.405629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:5ccf70d0cc8601560a12682977f97d0633d3aef48891e4a43d86db588c46518a

Observation 50cd2131-44bf-4af8-9404-d4bb2279ba9f · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.800083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:1f5118632a201c5eb9145e51644e4e35318140a4590c24ecddf2c1f8a12f8a88

Observation f9d0b40c-a746-425d-a2eb-54a225b0ff58 · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.543190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.543190Z digest=sha256:5b664184a630552b53578354745d959579ff60146623bf39de02f2f4dd6bfbbf