Pith. sign in

Paper Citation Record · LEDGER

Speculative Streaming: Fast LLM Inference without Auxiliary Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2402.11131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.11131 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T20:43:27.532056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T05:52:37.526395Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a02f28c2-4398-40ae-b504-4adb62a11546 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.531501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:e614ca819047569e1a9d9c5d70605aa37af8140fcb7c51ff9fe7b049f57cacc8

Observation 55c28658-f78e-4556-8a76-5f94d086f314 · inbound

Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment cites this paper.

Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T20:43:27.532056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:43:27.532056Z digest=sha256:c414e1599cbac23f74dd7772b8c454468df89939ccdff4a312f8e21c11c02218

Observation 2abcd8ac-4e7f-44c2-a606-6e08b53cde8c · inbound

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing cites this paper.

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T13:58:44.876800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:58:44.876800Z digest=sha256:b809efe91157fc590e097b001932a33c1e6608846aff5d8bd989d227a0ca067b

Observation 0d7e73a3-b19b-4020-8b21-4c906ee9a86a · inbound

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference cites this paper.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.409002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.409002Z digest=sha256:19fd203d088b8738829dc6b3f9a3b88039d1eb4c9174cd5651011ca23cdf24c8

Observation 4c04e46d-7b85-4a4a-89a6-1e6ba3cd289f · inbound

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache cites this paper.

QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T04:28:03.898221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:28:03.898221Z digest=sha256:9f573ede0f82881207d58b43a32fce8dec3b3bd08b466da85fd992b1784567e8

Observation 31054bfa-44b0-4869-8a8e-e466a8545f10 · inbound

Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential cites this paper.

Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T17:06:07.458918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:06:07.458918Z digest=sha256:3eae8b9876ef8df4b9ab6d441e4153a23f1c02e33fc3e30fe5c6bff74cd59fdc

Observation 4a3dc938-8149-45cf-8772-de4aea22ac62 · inbound

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations cites this paper.

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:55.566889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:55:51.126210Z digest=sha256:9eaf5d47cd85738ec6d7f4ce42dcf66ea5a213af5f86bda170c12c351b8b675e

Observation 529d8384-7001-4f9c-b503-3d4eb6bbac29 · inbound

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing cites this paper.

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T09:35:52.548132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:35:52.548132Z digest=sha256:11b100426c3bbbd3659552cab480617c3c99083e0078f2e08b36904ff6d08e89