Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:19:37.955444Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2502.08910.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:19:37.955444Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:00:43.304573Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T05:53:04.753054Z
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5da89157-b546-42a5-9486-966ef03a9ef9 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c36de9-0f72-4e09-95b3-1422237f243f · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1376c85a-8125-4214-8b04-6808982abe62 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU NTK - Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., June 2023
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1bb8716c-3377-4043-beac-34b16c36cf78 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a7e9fd-9fcf-4035-8590-4d2117e44b35 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Flash-decoding for long-context inference, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3545b50a-5b7b-4782-8228-be7689c494c1 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f4ec7d-d063-45b8-b78d-a5cdf8372526 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e4bc0d-81d0-4354-bc7d-ea55a0ebc2c7 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf36afdf-078a-4509-a810-d3dd43910de1 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Gemma 2: Improving Open Language Models at a Practical Size
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42f8bdbd-f073-407f-b83e-93d7f6cc8a34 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 238d4a53-ec19-4a44-95f3-f91b2843baa0 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c5451d0-0d44-40d1-9102-6d299b70deb3 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Mistral 7B
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ed5391c-3d37-4a5c-be11-bce0d9996892 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e5821af-de8f-4a27-ba7b-f7971d4cedff · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d3eec75-e510-4ba7-8f31-ff15cca3b029 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d086886-eac4-406b-a35b-8fc72910edbb · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 383c3a34-8b05-4958-ac3b-0fea11b3874e · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1abf58e2-42c7-4a8f-906b-62e7cc0517d9 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 407523f4-dad4-40ea-be25-5fa859bf4d05 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU EXAONE 3.0 7
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57decbb3-1e36-47ba-9f02-3d67d5c653fa · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU EXAONE 3.5: Series of Large Language Models for Real -world Use Cases , December 2024 b
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b7e9af-e392-4be5-b046-466c5b80d87f · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU SnapKV: LLM Knows What You are Looking for Before Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c57521d-a509-44c7-ab93-8bc7d2dc9658 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU The Llama 3 Herd of Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e6b6c2-8183-4301-8eab-e3804da67195 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Transformers are Multi-State RNNs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b388a196-ed47-4fe7-9652-fd5cc6970cb2 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Code Llama: Open Foundation Models for Code
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08b7145-6752-4f1f-8046-99752f72a97f · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU RoFormer: Enhanced Transformer with Rotary Position Embedding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c1a1c4-8617-42aa-bfba-9be1fd2a9751 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63b3f5e-75b9-4dbc-9aeb-73b44ba9cc58 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87ec3dcb-ead0-453b-8e13-e762b07b64cf · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Attention Is All You Need
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48941070-f497-4cd2-986e-3dffe4170ca6 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Training-Free Exponential Context Extension via Cascading KV Cache
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32ef12c0-7a4f-4fc5-b005-9a28ffecec08 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81550cf-d9b9-4065-82bd-9852e424cfb3 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Efficient Streaming Language Models with Attention Sinks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df593e2-e375-404e-be22-b30d877a80c0 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d1d90c1-3251-4737-b299-aaed6e22a2c2 · outbound
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce79bcc8-386d-4afd-af53-3b4a5f0ec7d0 · inbound
FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be78e69b-d60a-44a1-b951-354d1b35ac41 · inbound
Controllably Efficient Language Models InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3836c4a1-58cd-44df-a149-fa7a038473eb · inbound
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.