Pith. sign in

Paper Citation Record · LEDGER

Star Attention: Efficient LLM Inference over Long Sequences

As of 21 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 7 inbound Pith citation observations for arXiv:2411.17116.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17116 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:36:41.350934Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:15:26.587682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T20:58:45.186838Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 301f6a37-141f-4bc5-b4f6-c97a85f30ef6 · outbound

This paper cites Longformer: The Long-Document Transformer.

Star Attention: Efficient LLM Inference over Long Sequences Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.267658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.267658Z digest=sha256:ad6a6ceca1791390a92b05c6cebd0ce3f17c15f0b3774fe0785f779e97f165f2

Observation cd71722d-799d-43bd-b7e4-decb6ed1cdf0 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Star Attention: Efficient LLM Inference over Long Sequences Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.277456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.277456Z digest=sha256:e7d19c7b7c7e1ab11738bc87753607fe75680c57668c60dc2ea0ceb4af46c9f8

Observation bf781c6d-c133-4af4-9fec-7deb74e88ab0 · outbound

This paper cites Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y ., Ji, H., and Wang, S.

Star Attention: Efficient LLM Inference over Long Sequences Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y ., Ji, H., and Wang, S

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:36:41.701500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:36:41.282198Z digest=sha256:7c1f731ee6d5a0c56879915034fc34b4b917b2bd8f8f3fe8869e9934a19c17cc

Observation 4dcd3aeb-e5c3-4068-a020-9f3e56f0dae4 · outbound

This paper cites BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack.

Star Attention: Efficient LLM Inference over Long Sequences BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.286575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.286575Z digest=sha256:3afbc3ddf06e2de8fdd6e57062450647c7078241ebc7ac20522bc9fce4a536b5

Observation 46bd833c-382d-4288-b642-f030dae59106 · outbound

This paper cites Online normalizer calculation for softmax.

Star Attention: Efficient LLM Inference over Long Sequences Online normalizer calculation for softmax

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.302149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.302149Z digest=sha256:f57a06c19fb009efd43e9283ca68b4f5b1bd74807e2606bb206555a6dd495cc5

Observation 69d28db5-5bf0-469c-843c-c6b71a31694f · outbound

This paper cites Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models.

Star Attention: Efficient LLM Inference over Long Sequences Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.312148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.312148Z digest=sha256:c7904b5485c0bb65601220c9406cb1f3f4fc1741ec74bb1f061d92211fbe8b58

Observation 0a9b9bb7-b475-411a-b7e7-ff870ece27f8 · outbound

This paper cites Qwen2.5 Technical Report.

Star Attention: Efficient LLM Inference over Long Sequences Qwen2.5 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.317759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.317759Z digest=sha256:d3561f95fd365dc7ae01951e045e7d196e6b67570b7c4386c4c88bb863871764

Observation eb18a738-74ae-43ce-b22f-a3927428df7a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Star Attention: Efficient LLM Inference over Long Sequences Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.325530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.325530Z digest=sha256:23db5df4539361383c0173afc78b5d67c26c54d2cdcc5fe284dd267e91741d38

Observation 5b7f8d72-611b-43a4-9635-6f890c34e254 · outbound

This paper cites You Only Cache Once: Decoder-Decoder Architectures for Language Models.

Star Attention: Efficient LLM Inference over Long Sequences You Only Cache Once: Decoder-Decoder Architectures for Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.334337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.334337Z digest=sha256:d8c043e1f7c7c41a45df5ca45e2afacd20f8dc73220e52520cd026e8632772a0

Observation 903c8dad-2c6c-4a2d-a563-e28baf4019d2 · outbound

This paper cites Transformers: State-of- the-art natural language processing.

Star Attention: Efficient LLM Inference over Long Sequences Transformers: State-of- the-art natural language processing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:36:41.669401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:36:41.338270Z digest=sha256:df1d707f45c6b3d62cd088b9a0f9d6223d0fe6e3090b30dcddf45ed703d9fa3e

Observation c94eed87-a0f5-4e72-a50c-2b594d8acd4a · outbound

This paper cites SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation.

Star Attention: Efficient LLM Inference over Long Sequences SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.342234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.342234Z digest=sha256:97764ca0726823783712ffd10fce791b21092b31b5d76087a8308a8a7804a851

Observation 4b6c21d5-17ad-4e62-9f47-74640ea7368b · outbound

This paper cites InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory.

Star Attention: Efficient LLM Inference over Long Sequences InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.346014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.346014Z digest=sha256:e321df13225952a0c212d0141a4cfc9f68bad900b100d31e6847571a2445b53b

Observation 648563e6-e92e-4f43-afc1-c39e26a21738 · outbound

This paper cites Vanilla autoregressive generation encounters out-of-memory (OOM) at 128K sequence length.

Star Attention: Efficient LLM Inference over Long Sequences Vanilla autoregressive generation encounters out-of-memory (OOM) at 128K sequence length

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T12:36:41.652437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:36:41.350934Z digest=sha256:d6b02c3a6cdc4bb5e59ab3687b9ee3322fb1c22d69ad2ccb3448c60e89881318

Observation 11b0b415-67f1-482f-9a5d-6836d72b1c69 · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

Star Attention: Efficient LLM Inference over Long Sequences Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.306215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.306215Z digest=sha256:20881e5993d6d071ff35cef7be1a230aa16c3e1a4d7351e97eb1a0cae7e778fc

Observation 9aba7698-778a-497d-99a9-9a11d8d01e13 · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

Star Attention: Efficient LLM Inference over Long Sequences Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.329878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.329878Z digest=sha256:14ea74bae612c4cdc60259e255f252f7cdbcbd3da80719c6f1773af03e902512

Observation 8de72b6c-7c7e-4006-a997-4f562f9c9dbb · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Star Attention: Efficient LLM Inference over Long Sequences Generating Long Sequences with Sparse Transformers

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.272002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.272002Z digest=sha256:5b9c37a46b488476bd56681d4ea1a22eddf87ecce12329522a1d7d7b0f8ea961

Observation 32ccf8b2-e46e-4f02-bd3c-6bd7de36715c · outbound

This paper cites fb.com/2021/07/15/open-source/fsdp/.

Star Attention: Efficient LLM Inference over Long Sequences fb.com/2021/07/15/open-source/fsdp/

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:36:41.685551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T12:36:41.298298Z digest=sha256:dd3eb8a5e9f58e5364e63e017512a35f710eafd4e0c8ef210c59d007323f1c06

Observation b5052458-e772-4053-be1d-3c580fef9266 · outbound

This paper cites E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning.

Star Attention: Efficient LLM Inference over Long Sequences E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.294074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.294074Z digest=sha256:4a96fb944b68f46cac4f63f0901097e080b3ee132ac514732e317b17b5f9e250

Observation c3b5315a-6d1b-4f98-a07e-80b483d7d2c8 · outbound

This paper cites Titans: Learning to Memorize at Test Time.

Star Attention: Efficient LLM Inference over Long Sequences Titans: Learning to Memorize at Test Time

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.261501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.261501Z digest=sha256:49f01c817526488fb6479eb41829b505cbd63cd9171210a71d0ee0331dbe9d01

Observation bad9a822-c269-4bf7-9ff2-2381dd7717ec · outbound

This paper cites Writing in the Margins: Better Inference Pattern for Long Context Retrieval.

Star Attention: Efficient LLM Inference over Long Sequences Writing in the Margins: Better Inference Pattern for Long Context Retrieval

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.321647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.321647Z digest=sha256:e7cf28a67cc5d90611e279e16ba7352673173b01fc4ceaf2001fa300a9fdcd5f

Pith citing papers

Observation 18bfbac8-1a26-4962-a053-34f6e6c8b3d8 · inbound

SCBench: A KV Cache-Centric Analysis of Long-Context Methods cites this paper.

SCBench: A KV Cache-Centric Analysis of Long-Context Methods Star Attention: Efficient LLM Inference over Long Sequences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:14:05.493385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:14:05.493385Z digest=sha256:3e30bc4c69418d24e19612dca8736115b33db1f38aa79dd24b855543e4d6999d

Observation f9b35de8-bd10-4615-9827-3803d1ede451 · inbound

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention cites this paper.

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention Star Attention: Efficient LLM Inference over Long Sequences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:15:26.587682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:15:26.587682Z digest=sha256:0fea961aaa3fab13805c19124c18a18d6ccc3665e7cafdd46123565af19ea1e5

Observation 38c6e429-e0e5-4138-951d-fbd762d539dd · inbound

Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data cites this paper.

Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data Star Attention: Efficient LLM Inference over Long Sequences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T01:07:10.282629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T01:07:10.282629Z digest=sha256:3b449c29f2efa9d780f82f9266deb15841e489e2a8392697d6f75fd5f9b9f3ca

Observation 6a8cd906-a9f6-4fec-8372-60b7704116aa · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Star Attention: Efficient LLM Inference over Long Sequences

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.189007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:f169c6ad516a5116f53a241c89738a32c34462c10155734723dd3ea13c74c141

Observation 1cfa2c0e-7705-4dd1-9ec3-16a67d9e115a · inbound

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models cites this paper.

FlashSVD: Memory-Efficient Inference with Streaming for Low-Rank Models Star Attention: Efficient LLM Inference over Long Sequences

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:47.963422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:40:47.963422Z digest=sha256:c56d2fda89b40cd9a362b6eba05590164f72cd91fd9d9a17d493b38465935ea2

Observation d2a5b67c-64f8-4def-9291-00163c139acd · inbound

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning cites this paper.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Star Attention: Efficient LLM Inference over Long Sequences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.546665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.546665Z digest=sha256:940d7fb45c1d7bc02f2d82df4aa5228e20ca85d902adbf8ea31a3a502d41376b

Observation 3843c861-26e9-4101-8fae-6ba88205790d · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Star Attention: Efficient LLM Inference over Long Sequences

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.867587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:270de0e3405871da93d70c57797be4c888dedca4673058b2f27488dc9e504d62