Pith. sign in

Paper Citation Record · LEDGER

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels

As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2412.18106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18106 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:05:22.596292Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:37:20.774251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:36:15.531550Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47cd4147-f3cc-47dd-ac60-4559b03ac1b3 · outbound

This paper cites https://pytorch.org/blog/acceleratin g-llama3/?hss_channel=lcp-78618366/.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://pytorch.org/blog/acceleratin g-llama3/?hss_channel=lcp-78618366/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.331066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.392066Z digest=sha256:c2d6523ef0bbbe2469e00327be3dc628703dcd9dcb446d23266f2275001129ba

Observation 83a42cb8-8a3e-474b-a87f-8c1bf358989e · outbound

This paper cites https://flashinfer.ai/ 2024/02/02/introduce-flashinfer.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://flashinfer.ai/ 2024/02/02/introduce-flashinfer.html

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.319515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.396611Z digest=sha256:1c2c848375b05a584c626f20a952490d3ccd10ce71b1fc1ef6c456197e374185

Observation 22947872-6e5a-40b3-a0f1-7ad9f52a1758 · outbound

This paper cites https://pytorch.org/blog/cutlass-p ing-pong-gemm-kernel/.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://pytorch.org/blog/cutlass-p ing-pong-gemm-kernel/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.305150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.401104Z digest=sha256:d653e3585ed9351bd47f347d40278c9e24a24885872e60a86b48111ef2ca78c5

Observation a18affcc-3f87-4cba-b693-1485f0841d30 · outbound

This paper cites https://huggingface.co/blog/layerskip.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://huggingface.co/blog/layerskip

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.291470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.405940Z digest=sha256:89b8a89471b7ca4344f29331f3d587d31ee1e81e97c36517668f526bf082f423

Observation 452c3868-d977-4ffc-a866-f90d86387830 · outbound

This paper cites https: //pytorch.org/blog/flash-decoding/.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https: //pytorch.org/blog/flash-decoding/

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.277652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.410870Z digest=sha256:cd03b18579ab3e12e545c8330b5874619fa82e152f8d80b7858c5d52c12193a7

Observation 64bc80c1-1cfc-46b9-aa7a-daf962464996 · outbound

This paper cites Mirror of https://gitee.com/ascend/pytorch.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mirror of https://gitee.com/ascend/pytorch

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.263812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.415024Z digest=sha256:99019ad4adff3d0daf417951b15ef80d2541ccecd19ee6083848c5755c2cfd00

Observation dd28661e-e65c-4227-9cc1-8c37acf99bae · outbound

This paper cites https://github.com/p ybind/pybind11.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com/p ybind/pybind11

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.250528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.419398Z digest=sha256:2a6c9287df10e26716b9f9cb9860decf09bc6f68409177909bccc6643b28a38c

Observation 5edfda8b-cdd9-4bac-a27d-9644b9eb8474 · outbound

This paper cites https://docs.vllm.ai/e n/latest/automatic_prefix_caching/apc.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://docs.vllm.ai/e n/latest/automatic_prefix_caching/apc.html

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.232533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.423136Z digest=sha256:70c007f6d2c7275b1faeb00ba3f1adf31c36842878b28ae87f8b2f20267c4ec9

Observation 83cefb15-b204-4231-81d8-6abdf8fcb6ae · outbound

This paper cites https://www.hiascend.com/docum ent/detail/en/canncommercial/700/modeldevpt /ptmigr/ptaoplist_000006.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/docum ent/detail/en/canncommercial/700/modeldevpt /ptmigr/ptaoplist_000006.html

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.219157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.426276Z digest=sha256:c3303f7965a50b54e86605d36a7b03ad446a381d702e72935fbcb0cfed49e753

Observation 111339f5-494f-42d3-99f4-e97c7f379713 · outbound

This paper cites https://www.hiascend.com/doc_center/source /zh/Pytorch/60RC2/apiref/apilist/ptaoplist _000787.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/doc_center/source /zh/Pytorch/60RC2/apiref/apilist/ptaoplist _000787.html

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.206340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.429593Z digest=sha256:17069339d8c4c7f26fcf7c118829c36c54cca6bee1dbef8ab358a3d8e8720d1d

Observation 9e869776-9d22-46e9-b9b5-733fc3a691ac · outbound

This paper cites https://www.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.194461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.434016Z digest=sha256:18cc467a4c1c4e1e7e654adce62b5b49e19247e0c24fa5466f9ce7487bfc3c35

Observation f37ccd31-1e6d-4ad5-a445-3b98d387190f · outbound

This paper cites https://www.hiascend.com/doc_center/sour ce/zh/CANNCommunityEdition/80RC1alpha001/ap iref/fmkadptapi/ptaoplist_000142.html.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://www.hiascend.com/doc_center/sour ce/zh/CANNCommunityEdition/80RC1alpha001/ap iref/fmkadptapi/ptaoplist_000142.html

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.183067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.437889Z digest=sha256:67b98a874b88846ef2398dac70f34a1e7f043c212490c8616d58feab0ef454c2

Observation 0bbb9e4b-c871-4b17-9262-7d6c6e8aa49e · outbound

This paper cites https://github.com/v llm-project/vllm/tree/main/.buildkite/nig htly-benchmarks.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com/v llm-project/vllm/tree/main/.buildkite/nig htly-benchmarks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.169997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.441379Z digest=sha256:1c5a937ec16fd9313d6ddffa7570da9f3036e9747c5ed0fb06c33ccd3562eeaf

Observation f099ea11-6cec-4a5b-a15d-cbcf46a99a58 · outbound

This paper cites https://github.com /vllm-project/vllm/pull/8054.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://github.com /vllm-project/vllm/pull/8054

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.155090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.445080Z digest=sha256:f22f913e9eade99d624a795bc04ce91b61f0ea69352e4b8d88e021de94a26847

Observation 6e76a983-1087-48a1-bdbd-d0c58b75f3c2 · outbound

This paper cites https://developer.nvidia.com/blog/optimi zing-compute-shaders-for-l2-locality-using -thread-group-id-swizzling/ , July 2020.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels https://developer.nvidia.com/blog/optimi zing-compute-shaders-for-l2-locality-using -thread-group-id-swizzling/ , July 2020

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.139298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.448986Z digest=sha256:9792a0e332299310ff61f2be38a89153f0c362e697cc0f167082a1ec53dfb6ef

Observation 6b08a2df-99af-4b74-a82c-04bb52b56086 · outbound

This paper cites Mnemosyne: Parallelization strategies for efficiently serving multi-million context length llm in- ference requests without approximations.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mnemosyne: Parallelization strategies for efficiently serving multi-million context length llm in- ference requests without approximations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.453803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.453803Z digest=sha256:fe1801827e44316a5bea76094e91ab7039430852b3812ae5934af9fc737bfb97

Observation 2d69371e-fbdb-400e-bda8-e7a15a12b8f2 · outbound

This paper cites Taming throughput- latency tradeoff in llm inference with sarathi-serve.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Taming throughput- latency tradeoff in llm inference with sarathi-serve

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.125471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.458077Z digest=sha256:f94449ff387c8aea6aecb7d974d17b8b281c09d0c1138fef04a401bdc1d6aa3e

Observation 36bab499-f5d5-491e-8201-f8db57f5b7b7 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.462957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.462957Z digest=sha256:40421f3192ceb0b5a5203d5f6aa95452cfa773a421795e1b906836a84210f2eb

Observation 0d67471d-a83d-4795-9981-14a0fe5af7a9 · outbound

This paper cites an unresolved cited work.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:05:23.112597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.468538Z digest=sha256:6aecd88fd4abe6fe5f56ae3b979b5d03e7e3b4fa895d5b5bbc60383f41fbcfdb

Observation 75bfec98-92e3-4eb8-8d29-2d43e65f6b38 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.473831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.473831Z digest=sha256:4ce5fa37279d019f548e49db1d96fb4e22d3b22a20594f7220a8c3ec3820c7e7

Observation 5e8a7e0a-e400-4e19-943b-a7ad5eea69ee · outbound

This paper cites End-to-end object detection with transform- ers.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels End-to-end object detection with transform- ers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.478313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.478313Z digest=sha256:61ceaf77f23bdccf2db05d4a40aa1520c00f523da1fbe787c2c770b78dec4a42

Observation 733cd9e4-6420-4366-b0d2-338805b786a5 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Accelerating Large Language Model Decoding with Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.482679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.482679Z digest=sha256:a74a4739073a7c82584441277c1855657118e7a0c78dd35d4fa208bc6cce5718

Observation d0b6a347-0352-4d1d-8e03-58225f182926 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.486482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.486482Z digest=sha256:f618cda0c522434cbfc822e9346304ad0aa3c826e7c60aa11bc0c7e896888b29

Observation be8e9371-9f9c-4480-b7d6-75c73f083ede · outbound

This paper cites Flashattention: Fast and memory- efficient exact attention with io-awareness.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Flashattention: Fast and memory- efficient exact attention with io-awareness

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.091824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.490605Z digest=sha256:aaf332739c2fd3a0b26d190891e2be0167a0a3c0263d739cefdc0dbcc196b108

Observation a0bb6f57-3fb0-440d-a1ef-f3fa95263ba5 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels An image is worth 16x16 words: Transformers for image recognition at scale

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.078223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.494447Z digest=sha256:9fa971f2c9176beb25195adb11fb9dd4d7ae87f1059322abea760fc8748e8868

Observation 5c526a61-35fc-4073-9207-c18d2d746ee0 · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.498348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.498348Z digest=sha256:f8399eea34a9ab40049b31a2db9120bc3357f52c1d5eb3c5764b68b507de5659

Observation ac97fea9-8f50-4cd6-a9b7-c55ae0c2bb51 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.503397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.503397Z digest=sha256:9115d62fa724fdcdafffb62096ebf320f54d7b804c935b9eb4a7c43e2fd0ea1c

Observation 9857202e-31fb-4f66-baee-4a8f4072702d · outbound

This paper cites P/D-Serve: Serving Disaggregated Large Language Model at Scale.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels P/D-Serve: Serving Disaggregated Large Language Model at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.508469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.508469Z digest=sha256:0c4965e01a903b4f3f52510659eabcf55b6d6f7e34177644bb7e510273d057e8

Observation b7800d6c-7386-4453-a1c9-cd79dd7fe34b · outbound

This paper cites POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:05:22.775442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.514760Z digest=sha256:369bd78e473b6f4325afe13abbe0c245c0203e16f2cb0ecaa6e98b2f9c0d8661

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.519070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.519070Z digest=sha256:095ef8579ec7660c0b6f78093b831e4d1fca0f2982fffb50291206b512a9cfdb

Observation bb016a02-9b7e-42f5-84f8-f03d8a495ba2 · outbound

This paper cites Efficient memory man- agement for large language model serving with page- dattention.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Efficient memory man- agement for large language model serving with page- dattention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.523327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.523327Z digest=sha256:180dca8ba2fbc05bd493d30c0d5a7a0d1d53468352c5b3f32bef047f164c4688

Observation 81fc6725-0bf6-436b-b36f-2d510aa3f4b4 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Fast inference from transformers via speculative decoding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.527297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.527297Z digest=sha256:089460733907d21d224907597e0c6ab6236fcb6d956cfcc43eda5a62a2cd884b

Observation b1aec2a8-ca3b-4f6d-9fb9-70510bc93d1e · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.531342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.531342Z digest=sha256:0dceffb6ed425ed438168bd252a5b765d3b67d35917c80772e590cf31f6d5d17

Observation 106ad724-c428-45ed-a10e-76a854dcaa36 · outbound

This paper cites Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.047071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.535557Z digest=sha256:727a6317b3e815cd758aacf413c4103c14610da5a87956bc499348d73bf7aafc

Observation 07c0b696-aace-4bae-87a8-e6a60384a8ef · outbound

This paper cites Davinci: A scalable architecture for neural network com- puting.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Davinci: A scalable architecture for neural network com- puting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.033583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.539706Z digest=sha256:460b3773e9f627a9f5a5637f2220bd2730d0c1e3758478a67e7606c8afc59b6a

Observation bfec7400-f702-40ed-9ca1-0d7ab176d340 · outbound

This paper cites FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.543764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.543764Z digest=sha256:4b715ca4f70c209ada2c01c3a644e32ad586224ee298299af4a169dbe24c71ae

Observation bd49f4a0-a2f4-4f8f-9570-d8fc9bfaabf2 · outbound

This paper cites TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.548112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.548112Z digest=sha256:d2fa9886b52034eceea9438fe370c196f68969b704ec91ec18e0f3108d270844

Observation e8651dbe-77c5-4aa3-9ffd-a011c769b9e4 · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Specinfer: Accelerating large language model serving with tree-based speculative inference and verification

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.018073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.553178Z digest=sha256:f081d7ef62234ccc08d870fe7be193133a11b0a22f7a1bc787c12ccd64dec15d

Observation 3fb187e3-ee08-4c7d-8e13-2b351b54ce6c · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.557261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.557261Z digest=sha256:5fc2ceb9c12e3f59e1eee707e242a9c42a59f0117df3ea994502946f5c2c3bfe

Observation b8c36028-98ec-4624-a5cc-0d7b426b321d · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.561973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.561973Z digest=sha256:5073925251cc408f4303a1370c7b5d9d7ea7d4427034a772e0f464cd76a3c892

Observation 1f0dc711-fbd7-432d-be49-72d79524c17a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.567647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.567647Z digest=sha256:86b5f857ec9f1f6194c09d649247ec76171b4f9f8e6dee0d819b05f3df60fa6c

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.572111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.572111Z digest=sha256:252744195222a18a44c651b8bfc5ff1f34193b1dd6d899c37b0d8bfee95f0386

Observation 76612cb5-28c6-48c6-951e-65d09d21bbce · outbound

This paper cites ChunkAt- tention: Efficient self-attention with prefix-aware KV cache and two-phase partition.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels ChunkAt- tention: Efficient self-attention with prefix-aware KV cache and two-phase partition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:23.002281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.576332Z digest=sha256:9b8b3028f3a99546de47f4f371f80db4cf6cb50ae0d57c5ee38eb353e0a6659a

Observation d24a8bee-8606-495d-9c90-1a43c4577a8b · outbound

This paper cites Orca: A distributed serving system for transformer-based generative mod- els.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Orca: A distributed serving system for transformer-based generative mod- els

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:22.987665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.580400Z digest=sha256:e839519ebee277bba7f9100c602cec2c941c642204ebe7a855a637eeafc16887

Observation 8058edb3-d53b-41ee-a168-34bf101eebb0 · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.584734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.584734Z digest=sha256:16f89bcaaacb8fdf7e1f07139492c0b7f510e9d03dde70b4e904494fd6b5b5d1

Observation 7878538c-d35e-4ba6-93b8-b14f9a978001 · outbound

This paper cites Lookahead: An inference acceleration frame- work for large language model with lossless generation accuracy.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Lookahead: An inference acceleration frame- work for large language model with lossless generation accuracy

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:22.973937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.589029Z digest=sha256:30a8bb720ed780fc5456997100f7f3394521b97a3306f01d76670341c85b5236

Observation 794db7ca-1f52-48d4-af1f-04c218142f6b · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels SGLang: Efficient Execution of Structured Language Model Programs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T05:05:22.592491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:05:22.592491Z digest=sha256:38943979f22b4cae38d3c4546cc65ff18d5b3c9ae8ba5c0200192c7791414cbf

Observation c41de5c9-3c40-480d-8740-319d8f0be7ee · outbound

This paper cites Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving.

Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels Dist- serve: Disaggregating prefill and decoding for goodput- optimized large language model serving

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:05:22.960244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T05:05:22.596292Z digest=sha256:beb2fcef4ee340c16f25d15987a4226afb7b30a9001f4e6925dd3f417db045e8

Pith citing papers

Observation 948b4573-1367-45f1-adf1-ed2a3aaa1f87 · inbound

AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments cites this paper.

AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments Tackling the Dynamicity in a Production LLM Serving System with SOTA Optimizations via Hybrid Prefill/Decode/Verify Scheduling on Efficient Meta-kernels

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:15.533193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T16:37:20.774251Z digest=sha256:bd88770077ca39689de207146756e126fb9d7159fecdbb50e8c39c9cc7c09be3