Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:11:17.839813Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2502.02789.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T11:11:17.839813Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:39:05.263720Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:48:33.364894Z
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 825cc507-4d35-4288-a66c-0e4333a50bbd · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35c8163-c947-43dc-ac37-4657bc64e693 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bb1da9-e101-4c37-a710-7353aed9f5c4 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation L ong B ench: A bilingual, multitask benchmark for long context understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba76ba8-24d4-46dd-9062-7b6f7dbcd9c7 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Longformer: The Long-Document Transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40efdeaf-b4f3-4037-a16a-495ea9432db2 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233f7931-9bf8-4029-9d07-a205c366d6e2 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67435126-6772-4b69-b780-3fc8020cd76b · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25149e5d-9d65-43a2-9266-6357f6084a6d · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Training verifiers to solve math word problems, 2021
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db173266-a6eb-45d5-9fe3-470492b91798 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b64e1e-fb99-41bc-9b80-58758ec30ea8 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da49d82b-1f3f-4462-a0a1-d40ea8354f05 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c793fdf5-84a3-4a65-8b54-3bb3a0111926 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Layerskip: Enabling early exit inference and self-speculative decoding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afb1e5c-3e4f-4a51-9223-9c80ed6df62e · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation How Far Are We From AGI: Are LLMs All We Need?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a4600c3-a957-4af5-8d74-fe913fb18020 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation A framework for few-shot language model evaluation, 07 2024 a
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc5a2cf-5255-4c68-91c9-bd138ecd0c98 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db3c62e1-e760-4347-9604-a1e22c7a2580 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41cc89b-f9d8-47f3-8744-968c435497e9 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Measuring Massive Multitask Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d47dfe-223c-4c48-a082-8e450037c5a3 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation KV Prediction for Improved Time to First Token
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a9dd95-f2aa-43f0-a172-06288745b923 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee53d13-3743-46d1-9a53-d46014661c0b · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Mistral 7B
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0805d18-4818-4e42-8e52-cf6f52f50a36 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81687733-b76f-4519-b52d-389e12613747 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e98294fb-c605-48be-9e27-8efe5c153e0c · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f0d174-2e08-400d-aa92-67080801e89e · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation H., Gonzalez, J
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730aa723-d591-417c-832f-5400fda1cef9 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e97223e2-fc57-465d-8165-a03c84f75dbd · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b9bca9-d6bf-4eb9-834a-6212459ce350 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Fast Inference from Transformers via Speculative Decoding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d31b63-ea18-4cd1-9788-c6119ad433e5 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1aa0ade-cbc0-4d9f-8a4d-5969cbca6a80 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Compressing Context to Enhance Inference Efficiency of Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f6391c-4166-407d-973e-5850a203fe1f · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SCBench: A KV Cache-Centric Analysis of Long-Context Methods
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5662039a-e072-439e-866d-c89b879a0ff1 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2df015bd-3b85-4f0d-bc87-472d78c5de49 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bce6ca-a1cb-4663-b39f-64869f22f964 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation S., Wang, Y., and Zhang, L
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d1fef0-7650-4007-95a7-1d84b3e70cb5 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4476ab2b-8dbc-4bc6-bb4e-ff108b27ed32 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation NLTK: The Natural Language Toolkit
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5abe06-8e50-4a74-b148-10464d34f148 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56fc0fe0-bdde-48b2-82e9-b1931cd5c1a5 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229521fa-32a6-4c05-a7a2-1ccd4e4db97d · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation GPT-4 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e8049c-d875-46f0-8f10-659c77c0504c · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c128c2-b44d-4c5a-a9b5-205f2c17f685 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d7906d-5500-4d0d-aca2-5cfcc5490e8e · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation L., Stickland, A
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6dc874-1a7f-4625-9aea-9f54564b78d3 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc901b6-2acf-4280-a233-9b90ced92faf · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858ec0b8-1b25-42de-a93e-597ba1aa812a · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbe6a1a6-a70f-4a57-83b8-747c165b2d18 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Gemini: A Family of Highly Capable Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bf3f2f-7191-4dfb-b041-4b4823b62010 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b23d10-0562-4206-9311-5db5d06a788b · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Emergent Abilities of Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771287dc-540e-4263-a3ca-95401869bafc · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbfa1cb-6a3e-4d12-81c3-d6c13e579e7d · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab3ff739-7f95-4d28-933e-196791966058 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Efficient Streaming Language Models with Attention Sinks
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f80f3215-5206-4997-ac88-d9ea0cd44b21 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Effective Long-Context Scaling of Foundation Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98b2bab2-103c-456d-90cf-932a6058636a · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Qwen2 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecfd4f90-ce9d-497b-a883-24e8b28f6620 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation LLM Inference Unveiled: Survey and Roofline Model Insights
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec84020-9065-4149-b785-b80c0267ae27 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e183dc25-5ec7-4f8a-b359-10b893767903 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df35ab7-1b69-492f-9ad8-f440963f86a7 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Q-hitter: A better token oracle for efficient llm inference via sparse-quantized kv cache
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3cc57133-f199-4d1a-ab98-0d5a9cc4a8f4 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SGLang: Efficient Execution of Structured Language Model Programs
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5f087b-b25b-45c8-9ae1-ec36d0fe213d · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation Instruction-Following Evaluation for Large Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00000cd-e7e5-4938-b067-d560714c60f6 · outbound
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68f5277a-7d78-4afb-813e-57c80128c8d0 · inbound
SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559406c1-ef58-46ce-b620-3a64915e8109 · inbound
Multi-Modal Agents for Power Distribution Defect Detection: An Evaluation of Foundation Models Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bdf29ee7-e41d-4710-93d0-0ab622e72a68 · inbound
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.