Pith. sign in

Paper Citation Record · LEDGER

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2502.08910.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08910 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:19:37.955444Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:00:43.304573Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T05:53:04.753054Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5da89157-b546-42a5-9486-966ef03a9ef9 · outbound

This paper cites write newline.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.781866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.781866Z digest=sha256:5778d967298ee5bb0b2287fd9cbf2657a3a9506067583229dbbe0574e121efc6

Observation 66c36de9-0f72-4e09-95b3-1422237f243f · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.788191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.788191Z digest=sha256:f5ae1c0e4c9cf9d639b9302b4a40ef214ce03ec1cae621ca570cbf2c516f3edc

Observation 1376c85a-8125-4214-8b04-6808982abe62 · outbound

This paper cites NTK - Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., June 2023.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU NTK - Aware Scaled RoPE allows LLaMA models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation., June 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:19:39.112315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.794159Z digest=sha256:1659710ec69e735d888cbc7cee44f9b1d3aa80f81d58181d103663d5ed07cd53

Observation 1bb8716c-3377-4043-beac-34b16c36cf78 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.800257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.800257Z digest=sha256:77d5ebb62372b7537b73494823c6fe3854e850398d5f7a703c953468c6a055d9

Observation 26a7e9fd-9fcf-4035-8590-4d2117e44b35 · outbound

This paper cites Flash-decoding for long-context inference, 2023.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Flash-decoding for long-context inference, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:19:39.063296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.805815Z digest=sha256:5fabf608778742d0b004488cb3bcabc98372dfd79650360096b5bdb82322581a

Observation 3545b50a-5b7b-4782-8228-be7689c494c1 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.811005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.811005Z digest=sha256:d3ed5e1395e95119314a90b95aa6844f4b02a10badcb61001b1ee50bd3685830

Observation 56f4ec7d-d063-45b8-b78d-a5cdf8372526 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.816306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.816306Z digest=sha256:d5e6438f7c501421c3a3e27b6851a6679e70728d47c4aacd9452cf67f6652d40

Observation f1e4bc0d-81d0-4354-bc7d-ea55a0ebc2c7 · outbound

This paper cites LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.822448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.822448Z digest=sha256:07f48553438f95a72cada45c74e767212a96ecd5c1a8cc5e808dfc2a3e7e61b6

Observation cf36afdf-078a-4509-a810-d3dd43910de1 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Gemma 2: Improving Open Language Models at a Practical Size

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.827717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.827717Z digest=sha256:fd612264f216f0c0533c821836696358556dbb8f29b1fc9dd8077bf1bbaf9da4

Observation 42f8bdbd-f073-407f-b83e-93d7f6cc8a34 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.832723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.832723Z digest=sha256:a266473ae832d4729e81e7df746b0745c34dac3a625c8eb0b5014e982344bb2c

Observation 238d4a53-ec19-4a44-95f3-f91b2843baa0 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.838100Z digest=sha256:5386f596bb7c44714009f1b980d91721f41d998d53f597b34ed4522b8dd6dc08

Observation 1c5451d0-0d44-40d1-9102-6d299b70deb3 · outbound

This paper cites Mistral 7B.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Mistral 7B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.843510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.843510Z digest=sha256:e10e9ffa136416e7fed9a0f2621cad4630465bcb21be7f01244dfc9eff4462bb

Observation 6ed5391c-3d37-4a5c-be11-bce0d9996892 · outbound

This paper cites MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.848294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.848294Z digest=sha256:c2cbddf547eb239de6ee5e0c7c0654b6ca719c1e4f3f398ee7a7d68678395cd5

Observation 5e5821af-de8f-4a27-ba7b-f7971d4cedff · outbound

This paper cites LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.854054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.854054Z digest=sha256:85f0abef958c9888b11f22b3ee5284265919b79885371d7280b955c9afe3c325

Observation 9d3eec75-e510-4ba7-8f31-ff15cca3b029 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.859198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.859198Z digest=sha256:4064592372722582ae331f245027e60f295f4608532a2e9926f661737b10656c

Observation 6d086886-eac4-406b-a35b-8fc72910edbb · outbound

This paper cites an unresolved cited work.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:19:38.997502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.864013Z digest=sha256:056f04c20e0a943c441c5f8b2a7f3cd14e3433794c44ec4c9844b119c13f714d

Observation 383c3a34-8b05-4958-ac3b-0fea11b3874e · outbound

This paper cites an unresolved cited work.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:19:38.894823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.872254Z digest=sha256:55be9758eb81305ad30b448e1b764529dee0800fd3c5a626e07a24ad9a7990de

Observation 1abf58e2-42c7-4a8f-906b-62e7cc0517d9 · outbound

This paper cites A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:19:38.625454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.877479Z digest=sha256:da0a134d08aa8572d396b31b7d6e394461ac0c9c3fab8462243da066a6c8a665

Observation 407523f4-dad4-40ea-be25-5fa859bf4d05 · outbound

This paper cites EXAONE 3.0 7.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU EXAONE 3.0 7

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.883272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.883272Z digest=sha256:de29c40dc9aa15e17dc89b5d4c5a5290dd36f31d81f5e74d7e988b56302cba94

Observation 57decbb3-1e36-47ba-9f02-3d67d5c653fa · outbound

This paper cites EXAONE 3.5: Series of Large Language Models for Real -world Use Cases , December 2024 b.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU EXAONE 3.5: Series of Large Language Models for Real -world Use Cases , December 2024 b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.888459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.888459Z digest=sha256:4d39b6d7a5771b4dfe91dbd662260b078266f84dcfa27032f302bcfba5a7eca8

Observation c6b7e9af-e392-4be5-b046-466c5b80d87f · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU SnapKV: LLM Knows What You are Looking for Before Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.892901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.892901Z digest=sha256:55625d4c14d97a4e7324c92da4eadeb01313bfcb16386cad7eedf4853dc4c3b2

Observation 6c57521d-a509-44c7-ab93-8bc7d2dc9658 · outbound

This paper cites The Llama 3 Herd of Models.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.897728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.897728Z digest=sha256:1e1c8e7e4b559394e9b90af838596f0927eab38aede73540f9ff7be6b74473ce

Observation 92e6b6c2-8183-4301-8eab-e3804da67195 · outbound

This paper cites Transformers are Multi-State RNNs.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Transformers are Multi-State RNNs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.902581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.902581Z digest=sha256:6b9c429ffd6cadca801ddc7e7e166969bddd21b8ba491365be8a1c8b8832e19a

Observation b388a196-ed47-4fe7-9652-fd5cc6970cb2 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Code Llama: Open Foundation Models for Code

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.908106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.908106Z digest=sha256:30172004cea21e2c2a8054663a3edc77dce72328fedced31692168cba3eec1cc

Observation c08b7145-6752-4f1f-8046-99752f72a97f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.913396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.913396Z digest=sha256:c62945201e25de90815b6f03e0eaeaf81299f76bf4a26759c7129cc80168bbbe

Observation 22c1a1c4-8617-42aa-bfba-9be1fd2a9751 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.918446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.918446Z digest=sha256:6e3d85f7c44e1d2a53e4f9173e77a560d1a89042847f9999c806fd6ea28c716b

Observation d63b3f5e-75b9-4dbc-9aeb-73b44ba9cc58 · outbound

This paper cites an unresolved cited work.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:19:38.834399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.924374Z digest=sha256:9bd95ac8bc6c725df4dcc936de48b1b8952f77481c0189789b1bd8153b2a1f3a

Observation 87ec3dcb-ead0-453b-8e13-e762b07b64cf · outbound

This paper cites Attention Is All You Need.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Attention Is All You Need

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.928806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.928806Z digest=sha256:658dd7e3e6db8c202397d3f5fd84e1cb1fa2c38debc2d97800f8eaddc801bef3

Observation 48941070-f497-4cd2-986e-3dffe4170ca6 · outbound

This paper cites Training-Free Exponential Context Extension via Cascading KV Cache.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Training-Free Exponential Context Extension via Cascading KV Cache

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:19:38.069832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T23:19:37.933737Z digest=sha256:b114975c56774e7cd5a7099b46fbb7320a5d6a459aa764f7bfc1c6c6e5435411

Observation 32ef12c0-7a4f-4fc5-b005-9a28ffecec08 · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.938758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.938758Z digest=sha256:e816b57a5634f4c27a0177fc8f58442ce5b0b55f55e7f6ffd3e57c1ffee4ef2d

Observation d81550cf-d9b9-4065-82bd-9852e424cfb3 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU Efficient Streaming Language Models with Attention Sinks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.944290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.944290Z digest=sha256:2054b666c761e2aa8fca40029b9cb054d3900c86ee8c3195a2deb33d0f09cbc0

Observation 7df593e2-e375-404e-be22-b30d877a80c0 · outbound

This paper cites $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU $\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.949529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.949529Z digest=sha256:d93b0810eec020a2da237f8901403d00b5052ed9e408702c3ef55ec687c0d0a4

Observation 7d1d90c1-3251-4737-b299-aaed6e22a2c2 · outbound

This paper cites H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:19:37.955444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:19:37.955444Z digest=sha256:cd51e2a10ded05449ef779d1e62a00eed15f58fce5a32f34e9227dfceff49ab5

Pith citing papers

Observation ce79bcc8-386d-4afd-af53-3b4a5f0ec7d0 · inbound

FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression cites this paper.

FAEDKV: Infinite-Window Fourier Transform for Unbiased KV Cache Compression InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:00:43.304573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:00:43.304573Z digest=sha256:4b2100a5fdee1e3f73bfed07dc1ec117c395e8cf01de6988f3c97efc6e81e911

Observation be78e69b-d60a-44a1-b951-354d1b35ac41 · inbound

Controllably Efficient Language Models cites this paper.

Controllably Efficient Language Models InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:57.667935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:33:57.667935Z digest=sha256:ffe1eb970380112be98c3545e0fef803946fb424ec44a6cdda17f443051f9a78

Observation 3836c4a1-58cd-44df-a149-fa7a038473eb · inbound

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents cites this paper.

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.754570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:50:14.266278Z digest=sha256:04045c5a8da0a68d8b0ca1f7cfc9023f27d2da34e2af7e7b854c6f574b004cc9