Pith. sign in

Paper Citation Record · LEDGER

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2608.01651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01651 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:34:56.846403Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6329b19f-2260-42fd-b9bf-2ab9f8c1b84a · outbound

This paper cites Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.124106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.695540Z digest=sha256:333ef4c6c8114b7ffc0d400d19d9e812d1dd4cad01036d0952f2fedd6c9c89e1

Observation 9c3e815c-d5ac-46f1-91fd-671471c8f9f9 · outbound

This paper cites Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Open-swe-traces: Advancing dual-mode multilingual distillation for software engineering agents,

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-04T23:34:57.854572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.698868Z digest=sha256:571dbb32ec937ab436d1cc10d6f33e93b34115edc0667e182a3a39ae1473fd99

Observation b358d3a9-81ac-414e-9fe2-f908aaa5f5a0 · outbound

This paper cites Program Synthesis with Large Language Models.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.701873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.701873Z digest=sha256:a793bc91ae00187c86f93bd1137a96fd1cddfafb25e863008b866596f334cd2e

Observation a2d1821c-8c83-48db-91bd-4d22083b50dd · outbound

This paper cites Medusa: Simple LLM inference acceleration framework with multiple decoding heads,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Medusa: Simple LLM inference acceleration framework with multiple decoding heads,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.116183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.704970Z digest=sha256:12b35eaf061b94ff9fec8a7bebadeee19d11647d9c95436ce13ad778ae5f1bd3

Observation 5d56c3e1-8fa1-4da9-a50a-4006741a1276 · outbound

This paper cites Accelerating large language model decoding with speculative sampling,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating large language model decoding with speculative sampling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.108000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.708145Z digest=sha256:68a1e919de6b587a17388e6d5bdea3b9d75ce3e798eefbdbd318dcf27b9cffb5

Observation 41c3e34a-2038-4378-b496-0fa3be60e862 · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.714038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.714038Z digest=sha256:0d7ea2c5d972beb6232425baeb03730a5e863da5e3d506b793f813d5927a76a5

Observation 3a81927f-d5c7-4993-a494-1ce773e71ff2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.717585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.717585Z digest=sha256:98fe5e361fdfdf3dd970f4cf0fe2a8451c30d4e1e743528a0474186c44d8c39f

Observation 296a7cc9-540f-4ea6-8f4a-de259765a09a · outbound

This paper cites Better & faster large language models via multi-token prediction,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Better & faster large language models via multi-token prediction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.099691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.720709Z digest=sha256:03012f3de8b8a8aa74fecaf5f4693ef7f564b2b2c1d00b9072aa77cdbd6910df

Observation f03a369e-2f3c-4956-b816-dc933ee702c4 · outbound

This paper cites Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Yggdrasil: Bridging dynamic speculation and static runtime for latency-optimal tree- based LLM decoding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.092451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.723608Z digest=sha256:18b3bf86f9e96ef57325a4c1dd7c87de437dc7c6c5e526774f8b0b99bc7494ec

Observation 94009cd4-5c80-4777-aab6-25484e9187ca · outbound

This paper cites Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Papi: Exploiting dynamic parallelism in large language model decoding with a processing-in-memory-enabled computing system,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.726380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.726380Z digest=sha256:1aa9eb06a26a853ef8decae4a71019c52957412e5f00cfc355c59a22a7309822

Observation d65ce328-f1ed-48ce-9452-d1ba3ef61784 · outbound

This paper cites Bridging draft policy misalignment: Group tree optimization for speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Bridging draft policy misalignment: Group tree optimization for speculative decoding,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.084724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.729247Z digest=sha256:47f0e2ad8f452748ecd17dddce441514e7ed7ccaa590637a7e7b0f865d1f423d

Observation fa33ab29-b983-483f-aaff-9bd1176265f0 · outbound

This paper cites Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pod-attention: Unlocking full prefill-decode overlap for faster llm inference,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.734525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.734525Z digest=sha256:ed9c1f74df4d321d6912e1f57bfa3a73bea95a9e17eea321819aea54d8b94424

Observation 819522c8-b345-4da7-b4b4-19d448e8f0f9 · outbound

This paper cites Pimba: A processing-in-memory acceleration for post-transformer large language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Pimba: A processing-in-memory acceleration for post-transformer large language model serving,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.737160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.737160Z digest=sha256:0ac2d820311808d60c10c965d2f5a89073209164d1c30c5fa042760a93dc180d

Observation 4b98fde6-42c6-4d26-b3d5-2e0de4616326 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient memory management for large language model serving with pagedattention,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.739878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.739878Z digest=sha256:38bdf544f2240b680d87e26f0cf039d0c8f72b3e7cfcce00aa211ac955eb9fd2

Observation a3a41fce-23c2-4e55-a210-5bfde570bdf1 · outbound

This paper cites Fast inference from transformers via speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Fast inference from transformers via speculative decoding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.076505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.742647Z digest=sha256:3276b4a9a4c53429e0c1a7155262ce6959540aeaf0e5874f386dd4c923e1af0e

Observation 510c84a3-2a10-4a5c-8067-64be9b2152fd · outbound

This paper cites Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Orches: Orchestrated test-time-compute-based llm reasoning on collaborative gpu-pim heterogeneous system,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.748666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.748666Z digest=sha256:7db9949fbb09a95906dfc5c23ac418bcbd453d772bdcbe208d58d783b7f464f9

Observation 8490ba93-2d01-485a-a550-cf3627768e51 · outbound

This paper cites EAGLE-2: Faster inference of language models with dynamic draft trees,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models EAGLE-2: Faster inference of language models with dynamic draft trees,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.068928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.751295Z digest=sha256:953602bc8003f5b16f1b091c35ac482f71f0ebf032437876e5183749c6ac5be8

Observation 343c301c-3e3d-45d1-80b2-3e1d6974f9be · outbound

This paper cites Eagle-3: Scaling up inference acceleration of large language models via training-time test,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Eagle-3: Scaling up inference acceleration of large language models via training-time test,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.061487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.753852Z digest=sha256:f4e86482d1116c2e5bb5a12ceadb5053ac7b000c17436e6f2335e10c28004d0b

Observation 7b0a1622-72f0-44b1-8409-f9b8015de48a · outbound

This paper cites Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adaserve: Accelerating multi-slo llm serving with slo-customized speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.756424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.756424Z digest=sha256:d22a7b7bcb1865d6e07e61942fe6ae1993ec85ff3090189dd65c393175ca6be4

Observation 541cab22-df40-40a4-9be6-1db16051545a · outbound

This paper cites Speculative decoding: Performance or illusion?.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Speculative decoding: Performance or illusion?

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.052560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.759847Z digest=sha256:d74015949e3a7e32cbccbea9f8b87f54332905c863a1bdb4a8b31205f689f5d1

Observation ef9c02c0-d2bf-4737-bb93-dbf8600041bc · outbound

This paper cites CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CacheSlide: Unlocking cross Position-Aware KV cache reuse for accelerating LLM serving,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.045015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.762499Z digest=sha256:fba35634e1e56292bd69b8cf7362db5906c1cbadeb4d755061581b90c9bb433c

Observation 3460b5a7-58fa-4aa3-ad95-8654d9c6bff3 · outbound

This paper cites Agentix: An efficient serving engine for LLM agents as general programs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Agentix: An efficient serving engine for LLM agents as general programs,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.036706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.765160Z digest=sha256:63e0b16f64ec83bea0f5bff3d2efb493bf924070aed045c20978ff5528fce60d

Observation 7650ec64-9a17-41cc-887e-007fdee0c498 · outbound

This paper cites No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models No buffer, no bottleneck: Efficient Zero-Copy KV cache offloading for Long-Context LLMs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.029402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.767778Z digest=sha256:f1e91b90195cc8d85d042d21ea364cff2266e2a922224a611754180e6a79cd64

Observation c00efa0f-bb4c-44b4-a9ea-f5bb1b31479f · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Specinfer: Accelerating large language model serving with tree-based speculative inference and verification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.770410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.770410Z digest=sha256:3a804a23ccfd99a92551728b528fe367ce5a5222d0db93a69ae4aa541fcc0db6

Observation 3102521f-27a9-443d-aa13-b1c30c48ac49 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Efficient large-scale language model training on gpu clusters using megatron-lm,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.773049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.773049Z digest=sha256:6c4da1c7441fd4c347956f0797e146c6cf80bd2f78969610c1fc4e6d763a1267

Observation 1cbac254-604b-47ce-bfb8-e22dd2af5948 · outbound

This paper cites CUDA Programming Guide: CUDA Graphs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models CUDA Programming Guide: CUDA Graphs,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.020532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.775582Z digest=sha256:f44735bc0279891ee74fa297a4d1528f2b3652e19a632a3ac3aad3621eab9c35

Observation 4c92c93e-63c6-4bb4-8c3f-29821dd13de7 · outbound

This paper cites NVIDIA Nsight Compute Documentation,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Compute Documentation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.012710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.778082Z digest=sha256:f4e918f9d27e372083e66db787de85e7b5684fc47a3918dbbfbf72bcf31a56ca

Observation ad36ac27-bea9-41e2-a4f1-d6857fd8fe50 · outbound

This paper cites NVIDIA Nsight Systems User Guide,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NVIDIA Nsight Systems User Guide,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:58.005213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.780615Z digest=sha256:27ef37b1eef5bff3fe7c19642e1b471d684187d6b964eac11914b92ef3758a7e

Observation 74e51260-d55d-4a1d-95e1-14039000aba4 · outbound

This paper cites Marconi: Prefix caching for the era of hybrid llms,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Marconi: Prefix caching for the era of hybrid llms,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.998021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.783043Z digest=sha256:27877d1d933485209802e5222514f57e2fee5b250d9cefc42007b3659eb2c8bf

Observation 52486329-8c91-4d4b-84bd-4b94310c1a63 · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.990529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.788834Z digest=sha256:1cc6e9ecfd3519841bfb6c5b5d13096bd73828027a2bde445b874e8986f4a02e

Observation 145fe7ed-7026-4bc8-a77f-92bca7747bd8 · outbound

This paper cites Qwen3.5: Towards native multimodal agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Qwen3.5: Towards native multimodal agents,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.982799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.791474Z digest=sha256:1b2381a266b9ae069bcab809b49517ca19dd072db43f18cb5cbfa3fbd29a4c00

Observation ed74e44a-f1bc-4d3d-a066-6b49d9fe8740 · outbound

This paper cites Get to the point: Summarization with pointer-generator networks,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Get to the point: Summarization with pointer-generator networks,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.974060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.793960Z digest=sha256:f75d5cd28f6d8284b19121b36422349e5791a860122ce0a763858b13df9f794a

Observation adb0e956-5cf9-4cc3-b540-22e1d91385f2 · outbound

This paper cites GitHub - sgl-project/sglang at release/v0.5.12 — github.com,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models GitHub - sgl-project/sglang at release/v0.5.12 — github.com,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.965305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.796907Z digest=sha256:0b61dbac9ef42969425caad93e2bc39df2173c18c8f3e7ae1ca0033bb086b3d7

Observation 9eca6b96-d966-434e-b9cc-5227e861a25c · outbound

This paper cites anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models anon8231489123/ShareGPT Vicuna unfiltered · Datasets at Hugging Face — huggingface.co,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.956695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.799558Z digest=sha256:976bee1217858630c70e734f7e5a965f23f09c61c5fa6c592c2cb5eeb89b28e4

Observation 27902834-7cb9-4369-9541-b26c00a13516 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llumnix: Dynamic scheduling for large language model serving,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.948954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.801941Z digest=sha256:050365820deacef805293b883ddd3e2bd42ad05840ee2623ee83754fb2e628ab

Observation 2a5a3837-1995-4e42-b918-6191772cb2a2 · outbound

This paper cites Kimi Linear: An Expressive, Efficient Attention Architecture.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Kimi Linear: An Expressive, Efficient Attention Architecture

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.804544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.804544Z digest=sha256:3979633e70a24ce263ff5957fe604cfd60bde6ac48fa63785f633b78419c03db

Observation 74088bc6-074f-4dba-a5bf-651ef49861b7 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Triton: an intermediate language and compiler for tiled neural network computations,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.807486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.807486Z digest=sha256:ce7a1b1c5edd87c53ceb3812958c51005bcc3ee81d4110c3e2733460ba829d01

Observation 9f23af17-10ee-4ed3-ae8d-8f8dffa7159b · outbound

This paper cites Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Adap- tive draft sequence length: Enhancing speculative decoding throughput on pim-enabled systems,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.941600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.810674Z digest=sha256:36a45b027f72e75e55bca6b85ee34dc07e3d67451af0401f4edb9078d1c43311

Observation 9f2b2780-522c-4994-a2a1-e88fda4ec9ff · outbound

This paper cites Openhands: An open platform for ai software developers as generalist agents,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Openhands: An open platform for ai software developers as generalist agents,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.933584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.813367Z digest=sha256:20e5002f06452c060963aa2fd987e9931ab4d2059584ce7361c1d596e6c68b7d

Observation 34f30b30-2b7c-4c4f-97c4-f776e3459680 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Roofline: an insightful visual performance model for multicore architectures,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.817103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.817103Z digest=sha256:79b2957a3c2ea60565cd92c0598bbf67b2d1d17268978338a457398bfcaabe80

Observation 22984f73-11fc-4377-bf95-5170c7303300 · outbound

This paper cites Stree: Speculative tree decoding for hybrid state space models,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Stree: Speculative tree decoding for hybrid state space models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.924591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.819686Z digest=sha256:365bffdcf023b17c04962020fd3b4e4ec2e9965c08b7f9ca00723eb801e9528e

Observation defa3908-f307-47a4-8665-d722fd90f93b · outbound

This paper cites Strata: Hierarchical context caching for long context language model serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Strata: Hierarchical context caching for long context language model serving,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.915181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.822338Z digest=sha256:a56a4ab244772cd772c2c1f4e991f4161c47a6156f8e12b55d0b8d4bb8eb90bd

Observation da079084-1e09-4c0a-b498-92afe391c455 · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Gated delta networks: Improving mamba2 with delta rule,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.906315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.825091Z digest=sha256:d58fc61c9376067ddeffccce527a3498cfc30bcc2760e935a79bf64de0738850

Observation 9febaefa-f8f9-4399-89a0-a73dc2045386 · outbound

This paper cites Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Deft: Decoding with flash tree-attention for efficient tree-structured llm inference,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.896471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.827799Z digest=sha256:f5c431a08dddfbac27c02646e35c948e0db6c7a603628f752a9af3bafe5b06c5

Observation 37f2872d-6750-46d3-9319-69a66e55f7c5 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Flashinfer: Efficient and customizable attention engine for llm inference serving,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.887251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.830769Z digest=sha256:414a7746d21909b47017280b6aa1ee21100654abd16ff0538b5602be5586ce4e

Observation 880a4d15-aa6e-4a15-9f48-a3ee80593255 · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.836278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.836278Z digest=sha256:322d84453af28041882d779ce15142df3ef497dc7009772dc4cad48adadf0934

Observation 2dd1f0d1-d306-4d0a-a222-e9585658d688 · outbound

This paper cites Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Swiftspec: Disaggregated speculative decoding and fused kernels for low-latency llm inference,

Reference 51

Resolution
metadata mismatch
raw_fallback, observed 2026-08-04T23:34:56.952275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.838809Z digest=sha256:0b71fa07dbebb90d54b4ccc609a16058e46400f97658e7cf987cb4ec371db966

Observation fa610865-6a1f-44f2-800d-443f12480b41 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Atom: Low-bit quantization for efficient and accurate llm serving,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.841217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.841217Z digest=sha256:958bb7515344d5a2ff9196f682b8d0c365c3a555bbc25f67365b84c48837edd2

Observation 4851dad1-8f2d-455d-be6b-1ab59ce296f0 · outbound

This paper cites Sglang: Efficient execution of structured language model programs,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Sglang: Efficient execution of structured language model programs,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.873171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.843854Z digest=sha256:410260bb1e2e06892a010fde242109317d3f5eb0bd883010e2d1e9a8cbdcd376

Observation b65ea5f1-7791-4bf5-bc26-0800be44c666 · outbound

This paper cites NanoFlow: Towards optimal large language model serving throughput,.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models NanoFlow: Towards optimal large language model serving throughput,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:34:57.862730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-04T23:34:56.846403Z digest=sha256:13d368eb41b14da2a89efc7a72f34b91f43f6000cd11d03242a9e7252e2c1bb1

Observation 74e19205-6bed-40de-8329-6fe6d609e5ff · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T23:34:56.711032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:34:56.711032Z digest=sha256:aa8baebed697ba2b64c95345c09925a430520f5f8a3ca7e94467d8a533b288ec

Pith citing papers

No inbound Pith citation observations are available.