Pith. sign in

Paper Citation Record · LEDGER

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

As of 11 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.01651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01651 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:56.846403Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6329b19f-2260-42fd-b9bf-2ab9f8c1b84a · outbound

This paper cites Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.124106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.695540Z digest=sha256:06bc4890f666cdf33ac9ae1bee4d1041108e68ade5bf535ddce53a06129bf736

Observation 9c3e815c-d5ac-46f1-91fd-671471c8f9f9 · outbound

This paper cites Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-04T23:34:57.854572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.698868Z digest=sha256:fb2ba96407b09ef4279fa39d2b63cc68095e1855f596cd1b4d4c7d8ac40e88a3

Observation b358d3a9-81ac-414e-9fe2-f908aaa5f5a0 · outbound

This paper cites Program Synthesis with Large Language Models.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.701873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.701873Z digest=sha256:a793bc91ae00187c86f93bd1137a96fd1cddfafb25e863008b866596f334cd2e

Observation a2d1821c-8c83-48db-91bd-4d22083b50dd · outbound

This paper cites Medusa: Simple LLM inference acceleration framework with multiple decoding heads,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Medusa: Simple LLM inference acceleration framework with multiple decoding heads,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.116183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.704970Z digest=sha256:9e327c22e19634ec9f1aeb4de7cdcb313f53c3fc7f1060ce7134f2024d919ae4

Observation 5d56c3e1-8fa1-4da9-a50a-4006741a1276 · outbound

This paper cites Accelerating large language model decoding with speculative sampling,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating large language model decoding with speculative sampling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.108000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.708145Z digest=sha256:6cd09f969695a4c6ba5892c74a643dfcd5c5abff436976e75678bd6b7c671eb5

Observation 41c3e34a-2038-4378-b496-0fa3be60e862 · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.714038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.714038Z digest=sha256:0d7ea2c5d972beb6232425baeb03730a5e863da5e3d506b793f813d5927a76a5

Observation 3a81927f-d5c7-4993-a494-1ce773e71ff2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.717585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.717585Z digest=sha256:98fe5e361fdfdf3dd970f4cf0fe2a8451c30d4e1e743528a0474186c44d8c39f

Observation 296a7cc9-540f-4ea6-8f4a-de259765a09a · outbound

This paper cites Better & faster large language models via multi-token prediction,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Better & faster large language models via multi-token prediction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.099691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.720709Z digest=sha256:3c023e87a459a85514cd24f039760deb3173b8a3ca8617cea0f581a115a2b581

Observation f03a369e-2f3c-4956-b816-dc933ee702c4 · outbound

This paper cites Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.092451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.723608Z digest=sha256:89452cdc77f211a0fdbae357a561bbe6491250465debd4c6ed5798ae940c1e80

Observation 94009cd4-5c80-4777-aab6-25484e9187ca · outbound

This paper cites Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.726380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.726380Z digest=sha256:1aa9eb06a26a853ef8decae4a71019c52957412e5f00cfc355c59a22a7309822

Observation d65ce328-f1ed-48ce-9452-d1ba3ef61784 · outbound

This paper cites Bridging draft policy misalignment: Group tree optimization for speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Bridging draft policy misalignment: Group tree optimization for speculative decoding,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.084724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.729247Z digest=sha256:3d89a1bd0ca3ec0f64cc1f1fb246ed14bc17ddb25d949b41fd5c7284c085a19b

Observation fa33ab29-b983-483f-aaff-9bd1176265f0 · outbound

This paper cites Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.734525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.734525Z digest=sha256:ed9c1f74df4d321d6912e1f57bfa3a73bea95a9e17eea321819aea54d8b94424

Observation 819522c8-b345-4da7-b4b4-19d448e8f0f9 · outbound

This paper cites Pimba: A processing-in-memory acceleration for post-transformer large language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pimba: A processing-in-memory acceleration for post-transformer large language model serving,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.737160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.737160Z digest=sha256:0ac2d820311808d60c10c965d2f5a89073209164d1c30c5fa042760a93dc180d

Observation 4b98fde6-42c6-4d26-b3d5-2e0de4616326 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient memory management for large language model serving with pagedattention,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.739878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.739878Z digest=sha256:38bdf544f2240b680d87e26f0cf039d0c8f72b3e7cfcce00aa211ac955eb9fd2

Observation a3a41fce-23c2-4e55-a210-5bfde570bdf1 · outbound

This paper cites Fast inference from transformers via speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Fast inference from transformers via speculative decoding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.076505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.742647Z digest=sha256:847ce16eedf923b1096bd4fe161c1e0aabec0ee7dc30fe117065f7b8dd6ac390

Observation 510c84a3-2a10-4a5c-8067-64be9b2152fd · outbound

This paper cites Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.748666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.748666Z digest=sha256:7db9949fbb09a95906dfc5c23ac418bcbd453d772bdcbe208d58d783b7f464f9

Observation 8490ba93-2d01-485a-a550-cf3627768e51 · outbound

This paper cites EAGLE-2: Faster inference of language models with dynamic draft trees,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models EAGLE-2: Faster inference of language models with dynamic draft trees,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.068928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.751295Z digest=sha256:e9dcdedf83d7fc56ca5057aaa74c3f4a11da51192dc27e31eb542da60aa876df

Observation 343c301c-3e3d-45d1-80b2-3e1d6974f9be · outbound

This paper cites Eagle-3: Scaling up inference acceleration of large language models via training-time test,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Eagle-3: Scaling up inference acceleration of large language models via training-time test,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.061487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.753852Z digest=sha256:4149da6cb749f6c047cd5de5d6450ec59cc8f266b55e82d7a9ce14d1da08af98

Observation 7b0a1622-72f0-44b1-8409-f9b8015de48a · outbound

This paper cites Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.756424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.756424Z digest=sha256:d22a7b7bcb1865d6e07e61942fe6ae1993ec85ff3090189dd65c393175ca6be4

Observation 541cab22-df40-40a4-9be6-1db16051545a · outbound

This paper cites Speculative decoding: Performance or illusion?.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Speculative decoding: Performance or illusion?

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.052560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.759847Z digest=sha256:c930729a281681fdb13faa3c87715152c8c6a5f0b0fc19f3d2d52b89450d8074

Observation ef9c02c0-d2bf-4737-bb93-dbf8600041bc · outbound

This paper cites CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.045015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.762499Z digest=sha256:c1c34d557d53df7f58068eb0ab04f41ef755de9a913c1179bfcbb72151e03927

Observation 3460b5a7-58fa-4aa3-ad95-8654d9c6bff3 · outbound

This paper cites Agentix: An efficient serving engine for LLM agents as general programs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Agentix: An efficient serving engine for LLM agents as general programs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.036706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.765160Z digest=sha256:5824bf00ec7d09f712c49a77f39c26b0afa808377be864e65f29b829238d53a1

Observation 7650ec64-9a17-41cc-887e-007fdee0c498 · outbound

This paper cites No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.029402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.767778Z digest=sha256:48c8bc5c6fcb064431782ec531d0a89aad1a5fd89e37877175e22562288647f0

Observation c00efa0f-bb4c-44b4-a9ea-f5bb1b31479f · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.770410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.770410Z digest=sha256:3a804a23ccfd99a92551728b528fe367ce5a5222d0db93a69ae4aa541fcc0db6

Observation 3102521f-27a9-443d-aa13-b1c30c48ac49 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.773049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.773049Z digest=sha256:6c4da1c7441fd4c347956f0797e146c6cf80bd2f78969610c1fc4e6d763a1267

Observation 1cbac254-604b-47ce-bfb8-e22dd2af5948 · outbound

This paper cites CUDA Programming Guide: CUDA Graphs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CUDA Programming Guide: CUDA Graphs,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.020532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.775582Z digest=sha256:0e48f5b082d9c1cc29e71a3afaf1c1d7390b63765330181e0fa10649a654686f

Observation 4c92c93e-63c6-4bb4-8c3f-29821dd13de7 · outbound

This paper cites NVIDIA Nsight Compute Documentation,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Compute Documentation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.012710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.778082Z digest=sha256:d4ae291bfc3ee53ba33fa784cb3f55018f2ce131d98d246a1733f174525037ed

Observation ad36ac27-bea9-41e2-a4f1-d6857fd8fe50 · outbound

This paper cites NVIDIA Nsight Systems User Guide,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Systems User Guide,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.005213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.780615Z digest=sha256:d48da869110d41c2ce81a9c121d77ab2e0df591f54487541cc3ffeec1035c604

Observation 74e51260-d55d-4a1d-95e1-14039000aba4 · outbound

This paper cites Marconi: Prefix caching for the era of hybrid llms,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Marconi: Prefix caching for the era of hybrid llms,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.998021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.783043Z digest=sha256:6fe086de970e7b592937a668db88b5063485582b2205e5d4e17a8befecbce04d

Observation 52486329-8c91-4d4b-84bd-4b94310c1a63 · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.990529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.788834Z digest=sha256:93ab6b5dc0c34f691843ab22254b01a968ba76371d0febacc7f7234581670c5d

Observation 145fe7ed-7026-4bc8-a77f-92bca7747bd8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Qwen3.5: Towards native multimodal agents,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.982799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.791474Z digest=sha256:2297e82651901fbf880850535f39951b5ce75bfb8427ed965e055d6df2426c12

Observation ed74e44a-f1bc-4d3d-a066-6b49d9fe8740 · outbound

This paper cites Get to the point: Summarization with pointer-generator networks,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Get to the point: Summarization with pointer-generator networks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.974060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.793960Z digest=sha256:295901d5ba316258af8c3983c91cb8e9312665f8718bb23125feb9e1246acd05

Observation adb0e956-5cf9-4cc3-b540-22e1d91385f2 · outbound

This paper cites GitHub - sgl-project/sglang at release/v0.5.12 — github.com,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models GitHub - sgl-project/sglang at release/v0.5.12 — github.com,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.965305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.796907Z digest=sha256:77e3b20562b70968a2871a15a759ef62c30759c8b6d0d5a23407fc7f5d66ece4

Observation 9eca6b96-d966-434e-b9cc-5227e861a25c · outbound

This paper cites anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.956695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.799558Z digest=sha256:c68f0ae75afd03b6440b01288fe2ae719b736cc8b1e17be0b9f01f310ceb3271

Observation 27902834-7cb9-4369-9541-b26c00a13516 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llumnix: Dynamic scheduling for large language model serving,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.948954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.801941Z digest=sha256:25f3784a95afbe4c299ae57a7e5201a1bd2373128f565ff49369d558ffd7f9fa

Observation 2a5a3837-1995-4e42-b918-6191772cb2a2 · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.804544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.804544Z digest=sha256:3979633e70a24ce263ff5957fe604cfd60bde6ac48fa63785f633b78419c03db

Observation 74088bc6-074f-4dba-a5bf-651ef49861b7 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Triton: an intermediate language and compiler for tiled neural network computations,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.807486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.807486Z digest=sha256:ce7a1b1c5edd87c53ceb3812958c51005bcc3ee81d4110c3e2733460ba829d01

Observation 9f23af17-10ee-4ed3-ae8d-8f8dffa7159b · outbound

This paper cites Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.941600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.810674Z digest=sha256:1d1ca781242e7daa20a1cf167eb7ac1e70d604dd253fea4d6545f5b7bc7b74a1

Observation 9f2b2780-522c-4994-a2a1-e88fda4ec9ff · outbound

This paper cites Openhands: An open platform for ai software developers as generalist agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Openhands: An open platform for ai software developers as generalist agents,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.933584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.813367Z digest=sha256:aa859320731d853e4e387e0dfb553ac87423283182412143e98d9d4f15ce551c

Observation 34f30b30-2b7c-4c4f-97c4-f776e3459680 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Roofline: an insightful visual performance model for multicore architectures,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.817103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.817103Z digest=sha256:79b2957a3c2ea60565cd92c0598bbf67b2d1d17268978338a457398bfcaabe80

Observation 22984f73-11fc-4377-bf95-5170c7303300 · outbound

This paper cites Stree: Speculative tree decoding for hybrid state space models,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Stree: Speculative tree decoding for hybrid state space models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.924591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.819686Z digest=sha256:78dd0b9d069142fb8407f71d9532f76d6c8aabffca9b189966c20a6deccdb475

Observation defa3908-f307-47a4-8665-d722fd90f93b · outbound

This paper cites Strata: Hierarchical context caching for long context language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Strata: Hierarchical context caching for long context language model serving,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.915181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.822338Z digest=sha256:138dbb70ce2c8440f462c9bf2ab1948fbe3058e997a151b20c8d49d3c91d1de3

Observation da079084-1e09-4c0a-b498-92afe391c455 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Gated delta networks: Improving mamba2 with delta rule,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.906315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.825091Z digest=sha256:e70af7610661e486f57999a3a77b2d5c4e296053d830b776f5e8c351fbe81dd4

Observation 9febaefa-f8f9-4399-89a0-a73dc2045386 · outbound

This paper cites Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.896471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.827799Z digest=sha256:c04d74b4f2821bfbbfc63d4f01caf7a19558aee8337e06ef1c9f3d7d32caab9a

Observation 37f2872d-6750-46d3-9319-69a66e55f7c5 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Flashinfer: Efficient and customizable attention engine for llm inference serving,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.887251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.830769Z digest=sha256:74282d5182371cc1c4af44ef4ec7de413721ba1c96e28278c851fc580b68a995

Observation 880a4d15-aa6e-4a15-9f48-a3ee80593255 · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.836278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.836278Z digest=sha256:322d84453af28041882d779ce15142df3ef497dc7009772dc4cad48adadf0934

Observation 2dd1f0d1-d306-4d0a-a222-e9585658d688 · outbound

This paper cites Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,

Reference 51

Resolution
metadata mismatch
raw_fallback, observed 2026-08-04T23:34:56.952275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.838809Z digest=sha256:fb321afa9f526fa95997821a7210721a3800f605dedfab7812e8dd5cdeee7d6a

Observation fa610865-6a1f-44f2-800d-443f12480b41 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.841217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.841217Z digest=sha256:958bb7515344d5a2ff9196f682b8d0c365c3a555bbc25f67365b84c48837edd2

Observation 4851dad1-8f2d-455d-be6b-1ab59ce296f0 · outbound

This paper cites Sglang: Efficient execution of structured language model programs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sglang: Efficient execution of structured language model programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.873171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.843854Z digest=sha256:bdef2901ef39aed827d2dbff9ce7f0ddcf86b98941288f2912befff200b2b594

Observation b65ea5f1-7791-4bf5-bc26-0800be44c666 · outbound

This paper cites NanoFlow: Towards optimal large language model serving throughput,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NanoFlow: Towards optimal large language model serving throughput,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.862730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-04T23:34:56.846403Z digest=sha256:5b5568c193c5f959abbfeda71dfed8225bb3c4b05b12a21f5df39cd44b33edfb

Observation 74e19205-6bed-40de-8329-6fe6d609e5ff · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.711032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.711032Z digest=sha256:62426f88dd22fb00992e3ac079df1139f02d6a23def57ca9e4c8b27614e581d9

Pith citing papers

No inbound Pith citation observations are available.