Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:56.846403Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.01651.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:56.846403Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6329b19f-2260-42fd-b9bf-2ab9f8c1b84a · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c3e815c-d5ac-46f1-91fd-671471c8f9f9 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b358d3a9-81ac-414e-9fe2-f908aaa5f5a0 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d1821c-8c83-48db-91bd-4d22083b50dd · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Medusa: Simple LLM inference acceleration framework with multiple decoding heads,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d56c3e1-8fa1-4da9-a50a-4006741a1276 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating large language model decoding with speculative sampling,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 41c3e34a-2038-4378-b496-0fa3be60e862 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a81927f-d5c7-4993-a494-1ce773e71ff2 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 296a7cc9-540f-4ea6-8f4a-de259765a09a · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Better & faster large language models via multi-token prediction,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f03a369e-2f3c-4956-b816-dc933ee702c4 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94009cd4-5c80-4777-aab6-25484e9187ca · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d65ce328-f1ed-48ce-9452-d1ba3ef61784 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Bridging draft policy misalignment: Group tree optimization for speculative decoding,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fa33ab29-b983-483f-aaff-9bd1176265f0 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 819522c8-b345-4da7-b4b4-19d448e8f0f9 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pimba: A processing-in-memory acceleration for post-transformer large language model serving,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b98fde6-42c6-4d26-b3d5-2e0de4616326 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient memory management for large language model serving with pagedattention,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a41fce-23c2-4e55-a210-5bfde570bdf1 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Fast inference from transformers via speculative decoding,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 510c84a3-2a10-4a5c-8067-64be9b2152fd · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8490ba93-2d01-485a-a550-cf3627768e51 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models EAGLE-2: Faster inference of language models with dynamic draft trees,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 343c301c-3e3d-45d1-80b2-3e1d6974f9be · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Eagle-3: Scaling up inference acceleration of large language models via training-time test,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b0a1622-72f0-44b1-8409-f9b8015de48a · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 541cab22-df40-40a4-9be6-1db16051545a · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Speculative decoding: Performance or illusion?
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef9c02c0-d2bf-4737-bb93-dbf8600041bc · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3460b5a7-58fa-4aa3-ad95-8654d9c6bff3 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Agentix: An efficient serving engine for LLM agents as general programs,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7650ec64-9a17-41cc-887e-007fdee0c498 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c00efa0f-bb4c-44b4-a9ea-f5bb1b31479f · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3102521f-27a9-443d-aa13-b1c30c48ac49 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient large-scale language model training on gpu clusters using megatron-lm,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cbac254-604b-47ce-bfb8-e22dd2af5948 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CUDA Programming Guide: CUDA Graphs,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4c92c93e-63c6-4bb4-8c3f-29821dd13de7 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Compute Documentation,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ad36ac27-bea9-41e2-a4f1-d6857fd8fe50 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Systems User Guide,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 74e51260-d55d-4a1d-95e1-14039000aba4 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Marconi: Prefix caching for the era of hybrid llms,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 52486329-8c91-4d4b-84bd-4b94310c1a63 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 145fe7ed-7026-4bc8-a77f-92bca7747bd8 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Qwen3.5: Towards native multimodal agents,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ed74e44a-f1bc-4d3d-a066-6b49d9fe8740 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Get to the point: Summarization with pointer-generator networks,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation adb0e956-5cf9-4cc3-b540-22e1d91385f2 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models GitHub - sgl-project/sglang at release/v0.5.12 — github.com,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9eca6b96-d966-434e-b9cc-5227e861a25c · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27902834-7cb9-4369-9541-b26c00a13516 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llumnix: Dynamic scheduling for large language model serving,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a5a3837-1995-4e42-b918-6191772cb2a2 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Kimi Linear: An Expressive, Efficient Attention Architecture
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74088bc6-074f-4dba-a5bf-651ef49861b7 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Triton: an intermediate language and compiler for tiled neural network computations,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f23af17-10ee-4ed3-ae8d-8f8dffa7159b · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9f2b2780-522c-4994-a2a1-e88fda4ec9ff · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Openhands: An open platform for ai software developers as generalist agents,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 34f30b30-2b7c-4c4f-97c4-f776e3459680 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Roofline: an insightful visual performance model for multicore architectures,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22984f73-11fc-4377-bf95-5170c7303300 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Stree: Speculative tree decoding for hybrid state space models,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation defa3908-f307-47a4-8665-d722fd90f93b · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Strata: Hierarchical context caching for long context language model serving,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation da079084-1e09-4c0a-b498-92afe391c455 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Gated delta networks: Improving mamba2 with delta rule,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9febaefa-f8f9-4399-89a0-a73dc2045386 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 37f2872d-6750-46d3-9319-69a66e55f7c5 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Flashinfer: Efficient and customizable attention engine for llm inference serving,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 880a4d15-aa6e-4a15-9f48-a3ee80593255 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llmcompass: Enabling efficient hardware design for large language model inference,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dd1f0d1-d306-4d0a-a222-e9585658d688 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fa610865-6a1f-44f2-800d-443f12480b41 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Atom: Low-bit quantization for efficient and accurate llm serving,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4851dad1-8f2d-455d-be6b-1ab59ce296f0 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sglang: Efficient execution of structured language model programs,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b65ea5f1-7791-4bf5-bc26-0800be44c666 · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NanoFlow: Towards optimal large language model serving throughput,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 74e19205-6bed-40de-8329-6fe6d609e5ff · outbound
Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating Large Language Model Decoding with Speculative Sampling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.