Pith. sign in

Paper Citation Record · LEDGER

Star Attention: Efficient LLM Inference over Long Sequences

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 4 inbound Pith citation observations for arXiv:2411.17116.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17116 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:36:41.350934Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:14:05.493385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:58:45.186838Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 301f6a37-141f-4bc5-b4f6-c97a85f30ef6 · outbound

This paper cites Longformer: The Long-Document Transformer.

Star Attention: Efficient LLM Inference over Long Sequences Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.267658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.267658Z digest=sha256:4e800c7bfca5bd78ba8d1f55e71f8ca8ed9c33a61c520ddf49317259fb7c49d8

Observation cd71722d-799d-43bd-b7e4-decb6ed1cdf0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Star Attention: Efficient LLM Inference over Long Sequences Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.277456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.277456Z digest=sha256:5ad50bb39a273fa9d42c02e27d25884e84085987a5d2161195a472cd109b65df

Observation bf781c6d-c133-4af4-9fec-7deb74e88ab0 · outbound

This paper cites Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y ., Ji, H., and Wang, S.

Star Attention: Efficient LLM Inference over Long Sequences Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y ., Ji, H., and Wang, S

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:36:41.701500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:36:41.282198Z digest=sha256:0f29784405f047511569c1c9b55776a251d28cc9d54cc859cb0bbb833ae95353

Observation 4dcd3aeb-e5c3-4068-a020-9f3e56f0dae4 · outbound

This paper cites BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack.

Star Attention: Efficient LLM Inference over Long Sequences BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.286575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.286575Z digest=sha256:f05b0f6e4ea394dbc2570d1c50dabfbc01a27e4212fec9a840b80c1c6bc6fd46

Observation 46bd833c-382d-4288-b642-f030dae59106 · outbound

This paper cites Online normalizer calculation for softmax.

Star Attention: Efficient LLM Inference over Long Sequences Online normalizer calculation for softmax

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.302149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.302149Z digest=sha256:c09d856d32ea3a7d345ed161ef6aa319acc575777020a26f3909bced9c1eee41

Observation 69d28db5-5bf0-469c-843c-c6b71a31694f · outbound

This paper cites Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models.

Star Attention: Efficient LLM Inference over Long Sequences Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.312148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.312148Z digest=sha256:48cd3974c8dafd5e95576a46a185c2579e76ad199b96a712cebd66a56ba7bfbe

Observation 0a9b9bb7-b475-411a-b7e7-ff870ece27f8 · outbound

This paper cites Qwen2.5 Technical Report.

Star Attention: Efficient LLM Inference over Long Sequences Qwen2.5 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.317759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.317759Z digest=sha256:78a39be62199822ce6803d9c7c86800faa203f67675a88b42008bb4c2bca22f6

Observation eb18a738-74ae-43ce-b22f-a3927428df7a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Star Attention: Efficient LLM Inference over Long Sequences Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.325530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.325530Z digest=sha256:8e610b71f63a878d5f58dc633808b93af1b02dfa2b75870431026c3832e017d3

Observation 5b7f8d72-611b-43a4-9635-6f890c34e254 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

Star Attention: Efficient LLM Inference over Long Sequences You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.334337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.334337Z digest=sha256:cd4f381b3ebc7847ab44aced88f61bad0a21ad0a9f0f5699d8b745cdc0b4350c

Observation 903c8dad-2c6c-4a2d-a563-e28baf4019d2 · outbound

This paper cites Transformers: State-of- the-art natural language processing.

Star Attention: Efficient LLM Inference over Long Sequences Transformers: State-of- the-art natural language processing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:36:41.669401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:36:41.338270Z digest=sha256:8e6e9b2d50e4e9ce9cc01a7a606a0f2a0e3886dc60cb3f2f612681c1f0d37986

Observation c94eed87-a0f5-4e72-a50c-2b594d8acd4a · outbound

This paper cites SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation.

Star Attention: Efficient LLM Inference over Long Sequences SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.342234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.342234Z digest=sha256:f59d4d438e1c72bc6e206d814311bf3c396b1076f5af75a243acadfe9c8cfb24

Observation 4b6c21d5-17ad-4e62-9f47-74640ea7368b · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Star Attention: Efficient LLM Inference over Long Sequences InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.346014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.346014Z digest=sha256:eddf1a800c997ef086db542b26a06cf17e5975ba673ab071b759375b453ae821

Observation 648563e6-e92e-4f43-afc1-c39e26a21738 · outbound

This paper cites Vanilla autoregressive generation encounters out-of-memory (OOM) at 128K sequence length.

Star Attention: Efficient LLM Inference over Long Sequences Vanilla autoregressive generation encounters out-of-memory (OOM) at 128K sequence length

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T12:36:41.652437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:36:41.350934Z digest=sha256:04af166a7312397d118d394c38d68790e1d1aa33dcb2af74fd55fe97e4c04403

Observation 11b0b415-67f1-482f-9a5d-6836d72b1c69 · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

Star Attention: Efficient LLM Inference over Long Sequences Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.306215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.306215Z digest=sha256:2141317d0bca7360342bc89aa8854e20ff7039611602e1af503dcf01191549c9

Observation 9aba7698-778a-497d-99a9-9a11d8d01e13 · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

Star Attention: Efficient LLM Inference over Long Sequences Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.329878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.329878Z digest=sha256:f116a15f7a15e22a8e7d94d5b0f3db74fd58f9f234ddba2a6ac43e5db822ee87

Observation 8de72b6c-7c7e-4006-a997-4f562f9c9dbb · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Star Attention: Efficient LLM Inference over Long Sequences Generating Long Sequences with Sparse Transformers

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.272002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.272002Z digest=sha256:1199170ce419a0de0240eb24a72ed03c7fece296bd9fbaa8239360c665399515

Observation 32ccf8b2-e46e-4f02-bd3c-6bd7de36715c · outbound

This paper cites fb.com/2021/07/15/open-source/fsdp/.

Star Attention: Efficient LLM Inference over Long Sequences fb.com/2021/07/15/open-source/fsdp/

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:36:41.685551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T12:36:41.298298Z digest=sha256:0a1ac6822b4613b48ff1ccb1f03e5d9d1f7ff1b53d72a70d96d980a3fd9e7dbd

Observation b5052458-e772-4053-be1d-3c580fef9266 · outbound

This paper cites E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning.

Star Attention: Efficient LLM Inference over Long Sequences E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.294074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.294074Z digest=sha256:d0fc578928706fdb33754edf668aa0361f4358aeb9d32cad4d7dd02b8e421a98

Observation c3b5315a-6d1b-4f98-a07e-80b483d7d2c8 · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Star Attention: Efficient LLM Inference over Long Sequences Titans: Learning to Memorize at Test Time

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.261501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.261501Z digest=sha256:c235f51d22857f8df91a207f28242233a0e1a80cd3b84bc4c4130682ee8f1ac2

Observation bad9a822-c269-4bf7-9ff2-2381dd7717ec · outbound

This paper cites Writing in the Margins: Better Inference Pattern for Long Context Retrieval.

Star Attention: Efficient LLM Inference over Long Sequences Writing in the Margins: Better Inference Pattern for Long Context Retrieval

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.321647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.321647Z digest=sha256:380c551ef7c7ef7db4bbbb2a4e65149a0e2c72a670403c0d6159901b8723242f

Pith citing papers

Observation 18bfbac8-1a26-4962-a053-34f6e6c8b3d8 · inbound

SCBench: A KV Cache-Centric Analysis of Long-Context Methods cites this paper.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Star Attention: Efficient LLM Inference over Long Sequences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.493385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.493385Z digest=sha256:20ab21d51077e6c6fd8984cfe4a5a81c1c70193c4c981715974d577e7408bfda

Observation 6a8cd906-a9f6-4fec-8372-60b7704116aa · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Star Attention: Efficient LLM Inference over Long Sequences

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.189007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:07a389c68ed5a6531fa46cdc4fee20e6d71d108c887ff16a1c044998654fe727

Observation 1cfa2c0e-7705-4dd1-9ec3-16a67d9e115a · inbound

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models cites this paper.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Star Attention: Efficient LLM Inference over Long Sequences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.963422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.963422Z digest=sha256:d01df11d5638ddabc1ed91fec4717dc344cb7ad9cd011d6df1164fc030f93561

Observation 3843c861-26e9-4101-8fae-6ba88205790d · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Star Attention: Efficient LLM Inference over Long Sequences

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.867587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:b3a276e1ca540d353286b8e56fc48792b5dd9ed80274cb6eaeebc493add3d3b2