Pith. sign in

Paper Citation Record · LEDGER

FlashDecoding++: Faster Large Language Model Inference on GPUs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2311.01282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.01282 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:21:54.206928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.798667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ef22293-e0aa-4262-89e0-326b6b8ee658 · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.799232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:0450c494f4e78ebc16e9c41dbbf31ec29003d26aa49448e798dda14c5e283389

Observation ddec9caf-42b6-4a52-9843-c3892b099d91 · inbound

WaferLLM: Large Language Model Inference at Wafer Scale cites this paper.

WaferLLM: Large Language Model Inference at Wafer Scale FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:21:54.206928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:21:54.206928Z digest=sha256:28efaf1ae26578aa9d0cc93386a4216d902b4e9494bf133b0a26aff53e48dd3a

Observation 751b4ea8-832f-4923-9adf-d3417765cea8 · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.679828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.679828Z digest=sha256:2230f22acf76b9e27d2c372c23895ee09bd3f5ee6423635508b542dc23863332

Observation ac17d14e-2ccb-4c22-adc6-cccc746d29ca · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.857737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.857737Z digest=sha256:f7fcb0e29c97b6f69874a8ad523c813439714ddc06efc374a7762089e1379820

Observation 67557a58-4e61-4f7d-9e52-63de85969a1b · inbound

QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm cites this paper.

QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:50.245539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:55:50.245539Z digest=sha256:a36f9eba63cbea710d7e3b67c3c91e2ae48bf3b69e31cea3b19402ac200b7707

Observation a4678885-969e-4a5f-86b7-82f6600a50de · inbound

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations cites this paper.

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:28.045319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:28.045319Z digest=sha256:f5624f6b8984fd5e6ee2b0375416b06297349e858daa552f5fbc88012dcddaa2

Observation 08228ed5-dff6-4ef2-b535-85988aa99550 · inbound

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure cites this paper.

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:52.895202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:52.895202Z digest=sha256:c65bfef50fc35de15c9f197704baeb2f198cdbbf33a9224124b17ca61ca35cc4

Observation 5b123bcb-f408-452d-add1-7d409451bda2 · inbound

Past-Future Scheduler for LLM Serving under SLA Guarantees cites this paper.

Past-Future Scheduler for LLM Serving under SLA Guarantees FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.797553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.797553Z digest=sha256:cf0e56ded0c067d8546e0fae89e982299fe8da2f209cd2f70bd04bc3fe495816

Observation cffde461-3275-4a7e-b093-b1caaca2907a · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.142783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:c62c1654f12f36727f8dc0ef640c7e65c716c8b3c30d689b377cecd14966c79f

Observation 18b815b4-1dd2-4022-a073-4f091de5632a · inbound

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models cites this paper.

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:02.030920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:02.030920Z digest=sha256:3d4c8b421a8da78f39136d3aeb94911d0b80613c9896ad6676bd38f70f6e5fed

Observation f319ba07-7273-4021-8f80-619847606956 · inbound

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation cites this paper.

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:00.032774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:10:40.858525Z digest=sha256:9a9409ff38c0d09d2f15211f4eec70e6611ceaac26991bf5197b2fffaa58797b

Observation 20f99678-05a1-4b93-a6aa-fdece8d46f93 · inbound

Prism: Symbolic Superoptimization of Tensor Programs cites this paper.

Prism: Symbolic Superoptimization of Tensor Programs FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:24.949228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:01:38.783602Z digest=sha256:e59b9a59c72ea76bd6b1be99b526eb0aa739a2f62cd08fdd27a6d41318c9620b

Observation ffda947e-5625-4717-b167-f7e1f580643f · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.233012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:ae59294664a26b41d4cd463a58d5e0375426eb0e858ae1437c632470e8d57205

Observation 3dccb04e-b102-4f83-801d-e8eac0eec47f · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.405629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:24cc5d52560cc1d82538d92bd40381e7f779c7d304a2ca9743ada42fa67032cf

Observation 50cd2131-44bf-4af8-9404-d4bb2279ba9f · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.800083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:31ee4442237acebae0e521c6c25793a8b8ebed74e3a5087ade9fd2e400e42e18

Observation f9d0b40c-a746-425d-a2eb-54a225b0ff58 · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.543190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.543190Z digest=sha256:e23e74ce333e18507c3abe569e12397f47d4c95591d2c7def10311677d746d8d