Pith. sign in

Paper Citation Record · LEDGER

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation

As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2502.02789.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02789 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:11:17.839813Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:05.263720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:33.364894Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 825cc507-4d35-4288-a66c-0e4333a50bbd · outbound

This paper cites write newline.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.602177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.602177Z digest=sha256:addbf27bd3398d225c3ed301670f8ec20de5304b14682a3b01c628e99fa93398

Observation c35c8163-c947-43dc-ac37-4657bc64e693 · outbound

This paper cites Program Synthesis with Large Language Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.607441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.607441Z digest=sha256:9dbbe21e8186b769126e636edb7b30634f5df226442fb5325f48dc4c1b45a012

Observation f0bb1da9-e101-4c37-a710-7353aed9f5c4 · outbound

This paper cites L ong B ench: A bilingual, multitask benchmark for long context understanding.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation L ong B ench: A bilingual, multitask benchmark for long context understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.611991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.611991Z digest=sha256:7cd63349f772fb521050151ac0b22a6268696481f7ae60025947a1420043eca7

Observation 8ba76ba8-24d4-46dd-9062-7b6f7dbcd9c7 · outbound

This paper cites Longformer: The Long-Document Transformer.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.615898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.615898Z digest=sha256:d2acabf3a3c18b053f73527cb0382349fbd86c63e2f8097c352ec22c34b766f6

Observation 40efdeaf-b4f3-4037-a16a-495ea9432db2 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.620625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.620625Z digest=sha256:3224b5509bd174c1b12d61f677b4d4400bce1bf4663bf4e0a5b971489c7b137d

Observation 233f7931-9bf8-4029-9d07-a205c366d6e2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.624690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.624690Z digest=sha256:af25f9876059383a0bdb9cbeacf653d8e04e8e7f2740f04484d5729011f4f700

Observation 67435126-6772-4b69-b780-3fc8020cd76b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.628946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.628946Z digest=sha256:8d952d18074345d6e56340e3bbdc9e87af897e3e6beee9fe7acfe50d60f5ff82

Observation 25149e5d-9d65-43a2-9266-6357f6084a6d · outbound

This paper cites Training verifiers to solve math word problems, 2021.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Training verifiers to solve math word problems, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:11:31.010258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:11:17.633140Z digest=sha256:85cc2640d1d4fe44e777665742339eec7cddb72acc1aa4d07408d5b779c6cf68

Observation db173266-a6eb-45d5-9fe3-470492b91798 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.636804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.636804Z digest=sha256:92ac58cdff852ea422e481403ffddf8b7c39892fe48279042791722d5013af5d

Observation 82b64e1e-fb99-41bc-9b80-58758ec30ea8 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.640537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.640537Z digest=sha256:10e97beb0cf266dcc7046c5c8f72e23069de98aa641d4bc4b3d4e08fa1bf8197

Observation da49d82b-1f3f-4462-a0a1-d40ea8354f05 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.644345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.644345Z digest=sha256:5279d21d2db1f6335dada33ba2e236c413df2d553f75c4cdd7f1ef0a71f6cdd6

Observation c793fdf5-84a3-4a65-8b54-3bb3a0111926 · outbound

This paper cites Layerskip: Enabling early exit inference and self-speculative decoding.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Layerskip: Enabling early exit inference and self-speculative decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.648359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.648359Z digest=sha256:34f9ec8c82a22bbfc8c85ca82fa0640ca5799e48280b8242d4278d2ba81db65e

Observation 1afb1e5c-3e4f-4a51-9223-9c80ed6df62e · outbound

This paper cites How Far Are We From AGI: Are LLMs All We Need?.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation How Far Are We From AGI: Are LLMs All We Need?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.652509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.652509Z digest=sha256:9d00abcb3c2b13f8a37cb7ef04557f81c47d04e9ff6ae52830e4649a3921e87f

Observation 0a4600c3-a957-4af5-8d74-fe913fb18020 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024 a.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation A framework for few-shot language model evaluation, 07 2024 a

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.656292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.656292Z digest=sha256:2769e1ec1a91b30ccfa6717641e112d27ac36e4aacf2c6fb6c50f00194ba669a

Observation 8cc5a2cf-5255-4c68-91c9-bd138ecd0c98 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.659971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.659971Z digest=sha256:902ce90300f37b34fb6c99d65f194f3c97d0e1b230480e3a14200191553868c3

Observation db3c62e1-e760-4347-9604-a1e22c7a2580 · outbound

This paper cites The Llama 3 Herd of Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.664015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.664015Z digest=sha256:9eef8e79ab8b40d3bc6563e9edaa318af18def96e18435646cb369ca8394d34b

Observation f41cc89b-f9d8-47f3-8744-968c435497e9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.669337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.669337Z digest=sha256:98faaf517dd8638806ed9e76943f9d4fc291a5b4a0e4d10c06af9ee6eb01defb

Observation b1d47dfe-223c-4c48-a082-8e450037c5a3 · outbound

This paper cites KV Prediction for Improved Time to First Token.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation KV Prediction for Improved Time to First Token

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.673384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.673384Z digest=sha256:09aaaa3a4bd9bd8d6fd696b4dff47e739f1c8c6ae637377aa73421dd0228e653

Observation 37a9dd95-f2aa-43f0-a172-06288745b923 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.677430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.677430Z digest=sha256:9788d66bd19c5c0bbf8907319eb4a2a2574369723d7a7b12dacfc326c4547343

Observation 6ee53d13-3743-46d1-9a53-d46014661c0b · outbound

This paper cites Mistral 7B.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.681510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.681510Z digest=sha256:231a42eadf95ae1d01a240b5501737030e88816c8578bbab5d0b5d0d441f7257

Observation c0805d18-4818-4e42-8e52-cf6f52f50a36 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.685933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.685933Z digest=sha256:14fc0546281f5e80fda74b95af52014827c4c3c29c61e935606d530a4a418982

Observation 81687733-b76f-4519-b52d-389e12613747 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.689930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.689930Z digest=sha256:3e07c9312f8dc67519b5b08170c6e3118fd5cad850fc9c08691f0c5fa27c688e

Observation e98294fb-c605-48be-9e27-8efe5c153e0c · outbound

This paper cites LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.694065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.694065Z digest=sha256:220af2d1be8ae4c0582d593e74719dace07684ea8d046646a0b95fdee726211e

Observation 40f0d174-2e08-400d-aa92-67080801e89e · outbound

This paper cites H., Gonzalez, J.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation H., Gonzalez, J

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.698253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.698253Z digest=sha256:1a2c86e6de1d33188e860e7a4dc3a6a18da154340d4813bc2c7d9bf0da4e56d8

Observation 730aa723-d591-417c-832f-5400fda1cef9 · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.702371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.702371Z digest=sha256:af4a88c82e2eaaeb73b75969374e62831c3e8969bff676cf30edb023fef54c38

Observation e97223e2-fc57-465d-8165-a03c84f75dbd · outbound

This paper cites InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.706860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.706860Z digest=sha256:b40620c51b2de87a8ba80d35b8f137131694a09851ce06097f2cc9833f0de0d8

Observation 67b9bca9-d6bf-4eb9-834a-6212459ce350 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Fast Inference from Transformers via Speculative Decoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.711098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.711098Z digest=sha256:e0dfedeca0dde042d1d4de15905d8c35093e7d5681a10b7ffcf71ed4dc7d00a4

Observation 02d31b63-ea18-4cd1-9788-c6119ad433e5 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.714944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.714944Z digest=sha256:8f741fc11e0d9af9a5b42a26dcacec51535944669bb9931314fa2c44293651f8

Observation c1aa0ade-cbc0-4d9f-8a4d-5969cbca6a80 · outbound

This paper cites Compressing Context to Enhance Inference Efficiency of Large Language Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Compressing Context to Enhance Inference Efficiency of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.718608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.718608Z digest=sha256:2d347fd9719e0d143238142dbead1691278d2cb99dc99f4f5eeebbd32f981d02

Observation 73f6391c-4166-407d-973e-5850a203fe1f · outbound

This paper cites SCBench: A KV Cache-Centric Analysis of Long-Context Methods.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.722512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.722512Z digest=sha256:3266a3df813d670d84d86359d127ebfb06e75518d738f17ff09ec5e6cf6829bb

Observation 5662039a-e072-439e-866d-c89b879a0ff1 · outbound

This paper cites Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.726660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.726660Z digest=sha256:8dd973e03c1c425ebf88d8478d9f34f109b112a0c7c376887c1847e8f1d0660b

Observation 2df015bd-3b85-4f0d-bc87-472d78c5de49 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.730999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.730999Z digest=sha256:7feafcc21114adb0027e1482e63ef0f0a9056553694178d61c69698a0c91a5fb

Observation 45bce6ca-a1cb-4663-b39f-64869f22f964 · outbound

This paper cites S., Wang, Y., and Zhang, L.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation S., Wang, Y., and Zhang, L

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.735194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.735194Z digest=sha256:bdce4fd9d55a48b9958e73db0be351a2e2eb4f6d50dc4c430f21cce85e2d4ce2

Observation 12d1fef0-7650-4007-95a7-1d84b3e70cb5 · outbound

This paper cites CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.739768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.739768Z digest=sha256:d26c2a971eaa705adcdc27b0ee3b11fe3bd163c488c8306722b97503af6c3cc2

Observation 4476ab2b-8dbc-4bc6-bb4e-ff108b27ed32 · outbound

This paper cites NLTK: The Natural Language Toolkit.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation NLTK: The Natural Language Toolkit

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.743641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.743641Z digest=sha256:fb5c783baf7321aebc2fa83c207aebd0dd3dcf3dd5ef1a7717aa2571629458ac

Observation be5abe06-8e50-4a74-b148-10464d34f148 · outbound

This paper cites CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.748351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.748351Z digest=sha256:d822277ed1abaa228f30ca75621257109e7b4180325586e85cbacb6b05dd67dd

Observation 56fc0fe0-bdde-48b2-82e9-b1931cd5c1a5 · outbound

This paper cites Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.752885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.752885Z digest=sha256:3b63522df7aac397e3f1b9539210add120a7be4fbead6115d6f6791743ce61eb

Observation 229521fa-32a6-4c05-a7a2-1ccd4e4db97d · outbound

This paper cites GPT-4 Technical Report.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation GPT-4 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.756918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.756918Z digest=sha256:f92a17b1aa6c992073a8bf68dec2014d91c5f2a89efca02e3074043f61aa4dcb

Observation 42e8049c-d875-46f0-8f10-659c77c0504c · outbound

This paper cites SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.760837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.760837Z digest=sha256:abad9814c43d9d9178de3536055c61d145687c8a13a94c0cbaacfb657a955cf2

Observation 02c128c2-b44d-4c5a-a9b5-205f2c17f685 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.764820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.764820Z digest=sha256:31366a0a972d1416b890485b9afc2201edc9fb8a18081a17a4ecd30c373e7ed7

Observation e2d7906d-5500-4d0d-aca2-5cfcc5490e8e · outbound

This paper cites L., Stickland, A.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation L., Stickland, A

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.768911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.768911Z digest=sha256:611510cfbdee1a4a38e252565b67296c503bf15af996ad1263af289d5f5a4e9e

Observation 0e6dc874-1a7f-4625-9aea-9f54564b78d3 · outbound

This paper cites Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.772628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.772628Z digest=sha256:16c9aabd63a4d9867f8d7942b8a2c72297035382853ee75d2d4a44b716e9818e

Observation 1bc901b6-2acf-4280-a233-9b90ced92faf · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.776294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.776294Z digest=sha256:0928e212b46d08c7072188ca0f3b013181706b337d3d81d81b4a42f51f8ea4f6

Observation 858ec0b8-1b25-42de-a93e-597ba1aa812a · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.779908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.779908Z digest=sha256:253c58402af3ba97de7af93f27d4388e0a157c6e4d8e64c4bfb87c7e4a8857f9

Observation fbe6a1a6-a70f-4a57-83b8-747c165b2d18 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Gemini: A Family of Highly Capable Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.783386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.783386Z digest=sha256:2eebdb13732227bc701c4393d6620de570462d7f53d8faad8b2e95343d42cb38

Observation 73bf3f2f-7191-4dfb-b041-4b4823b62010 · outbound

This paper cites LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.788093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.788093Z digest=sha256:ddf1bab067da5a75008b368ce93b75b66597d70bbfc9a31d4e9a9301bb0b8a47

Observation e4b23d10-0562-4206-9311-5db5d06a788b · outbound

This paper cites Emergent Abilities of Large Language Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Emergent Abilities of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.791959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.791959Z digest=sha256:a1d23bcb783af63256db1b5792b3ba9b171af116f94fa53ee726e89b338a3ee7

Observation 771287dc-540e-4263-a3ca-95401869bafc · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.795746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.795746Z digest=sha256:4118fa91b63ff93468249f6f8fe76196686082c9eea4474373793baec18003f9

Observation 3fbfa1cb-6a3e-4d12-81c3-d6c13e579e7d · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.799746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.799746Z digest=sha256:f604d6459628f1cdcfa81400089d567bf3a9304ba1b0111b0ff97b51b2d0826a

Observation ab3ff739-7f95-4d28-933e-196791966058 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Efficient Streaming Language Models with Attention Sinks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.803412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.803412Z digest=sha256:8f5c8ffb2f7cde4b197c2f292e5e367eeb45f9cbf042449d85a060c971f670cc

Observation f80f3215-5206-4997-ac88-d9ea0cd44b21 · outbound

This paper cites Effective Long-Context Scaling of Foundation Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Effective Long-Context Scaling of Foundation Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.807536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.807536Z digest=sha256:55e7d7dcac969dd7878ab7a403f730ed1957d2350b95ae127417cd5af40941f7

Observation 98b2bab2-103c-456d-90cf-932a6058636a · outbound

This paper cites Qwen2 Technical Report.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.811606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.811606Z digest=sha256:7365c0a1e0124b77bcc20c2fbf40c8bb8cb2e55b603fee6328bcfef781111d9d

Observation ecfd4f90-ce9d-497b-a883-24e8b28f6620 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.815570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.815570Z digest=sha256:e9112e38ef8aa9a926d12e4193b2123e9f5935f18f6dc29cfd12fb93394735bd

Observation fec84020-9065-4149-b785-b80c0267ae27 · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.819360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.819360Z digest=sha256:be95faefaa1858ddbcf9ba639e156966613f264fbf14fa22fab82a0edcea4dbb

Observation e183dc25-5ec7-4f8a-b359-10b893767903 · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.823352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.823352Z digest=sha256:304b9f06778d6fd5d76f2bdecbb092f41f5dd55f2d78469ec3d74b43fa7393f8

Observation 4df35ab7-1b69-492f-9ad8-f440963f86a7 · outbound

This paper cites Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:11:30.978098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T11:11:17.827379Z digest=sha256:ad26131eaef02e1d433d6b764510fa462d0c4e71d4bf76ab379d3f258e9361a1

Observation 3cc57133-f199-4d1a-ab98-0d5a9cc4a8f4 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SGLang: Efficient Execution of Structured Language Model Programs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.832003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.832003Z digest=sha256:40558599159ee55ba574d9b613ef7492d7bee695bf6fed41f1855e477b3ff602

Observation 9f5f087b-b25b-45c8-9ae1-ec36d0fe213d · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Instruction-Following Evaluation for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.835920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.835920Z digest=sha256:a3d1dd2b5a17ef47fc3e1fb4ecb79c96d22dcf0f05318596fef3020ad71e93b9

Observation e00000cd-e7e5-4938-b067-d560714c60f6 · outbound

This paper cites SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention.

Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T11:11:17.839813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:11:17.839813Z digest=sha256:09e51400e86a041ccc7af9ee1207932d587c857b05f6df5de187f1bd134d2267

Pith citing papers

Observation 68f5277a-7d78-4afb-813e-57c80128c8d0 · inbound

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling cites this paper.

SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:05.263720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:05.263720Z digest=sha256:d328187d7b236133c132ab3ef57636fa67adde0aaf721e5fa4662dc54d923287

Observation 559406c1-ef58-46ce-b620-3a64915e8109 · inbound

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models cites this paper.

Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:33.366365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:53:51.279880Z digest=sha256:41b5afee8856e51f39e43ae3804b06f9ec97eab37043530da1927379476a933e

Observation bdf29ee7-e41d-4710-93d0-0ab622e72a68 · inbound

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards cites this paper.

Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T02:01:48.669846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:01:48.669846Z digest=sha256:667f37b5f98c6db9e521584b8cf568c03112922d851cc4114b12bb8f82cfb618