Pith. sign in

Paper Citation Record · LEDGER

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.01659.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01659 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:58:14.675856Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:30:14.955831Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:30:15.054198Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ad0472da-339c-4524-80d8-8a306225c2c6 · outbound

This paper cites A Survey of Large Language Models.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.552267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.552267Z digest=sha256:b6021fa052e688c818b8ef548304cb53cd346e410f72fb9f40f99959886d92ef

Observation b4660cd1-02b2-46f5-910e-2475dec1acb2 · outbound

This paper cites Enh ancing molecular design efficiency: Uniting language models and ge nerative networks with genetic algorithms,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Enh ancing molecular design efficiency: Uniting language models and ge nerative networks with genetic algorithms,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.105507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.558119Z digest=sha256:248cb8d45572c8660e3023e8275ec2ec2273744dbf50388de79ad417e6182bab

Observation a873b736-35bf-4859-ae49-63ca1ac51876 · outbound

This paper cites Path-bigbird: An ai-driven transformer appro ach to classi- fication of cancer pathology reports,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Path-bigbird: An ai-driven transformer appro ach to classi- fication of cancer pathology reports,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.091091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.562776Z digest=sha256:3ad933786d1be27fc9f14666694dc93d1bb12b1734672c8d38e3d54840c2ca19

Observation 55288ae4-3069-4aea-bbf5-72de2c81f4fc · outbound

This paper cites Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.076670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.567627Z digest=sha256:c08b9b2632b41f9e17b813541b64afc41d788651bcb551c6374c9424db8c5194

Observation 482daa7e-b360-45fe-83ce-ef2c1586d498 · outbound

This paper cites Attention is all you need,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Attention is all you need,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.061485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.572624Z digest=sha256:37ba105c91139478185d1461acc81bc22e9a7c941cc632b7bd4858f0c74ac629

Observation fe2008fa-86f5-4635-9eb2-17a2a97781aa · outbound

This paper cites Big bird: Transformers for longer sequences,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Big bird: Transformers for longer sequences,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.046896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.577071Z digest=sha256:1be906db0f66bd48bdd38b7b6de3f5ef2420d292e135d3f29536e69a582a0a8f

Observation 111e59df-db05-4aa6-8cc8-4fb5eb861679 · outbound

This paper cites Longnet: Scaling transformers to 1,000,000,00 0 tokens,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Longnet: Scaling transformers to 1,000,000,00 0 tokens,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.032040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.582358Z digest=sha256:afdacc5ff972cb295fc652d30b2cf44d798bf62c78e90c1b6dd66bf229da4086

Observation 519b063f-f8c2-495f-b7b6-70a578983ac2 · outbound

This paper cites Longformer: The Long-Document Transformer.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Longformer: The Long-Document Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.591311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.591311Z digest=sha256:77adb0809165217ba8a9605e003725ecb5bdc6103493f72351433484d2f092b6

Observation a142f2ca-da0e-4d9e-ae0b-fc39cdfd7904 · outbound

This paper cites Scaled Dot Product Attention,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Scaled Dot Product Attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.001995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.596125Z digest=sha256:24d8bcc28e0d7f84612331414024e8a0aaa610ad907c99e9acbbd4e988a57f3c

Observation 0b861f90-77c0-41b4-80be-1ac1764db6ad · outbound

This paper cites xformers: A mod- ular and hackable transformer modelling library,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques xformers: A mod- ular and hackable transformer modelling library,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.987469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.600723Z digest=sha256:3e8f2e218121363aa08ec857894086f5a47b3fcd8e0645619f6a258e9aeb46ce

Observation 07c2b4a2-cc94-449b-b72d-7f07304adabf · outbound

This paper cites Reformer: The Efficient Transformer.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Reformer: The Efficient Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.605321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.605321Z digest=sha256:9b3ec3e8a819493bd7c625cb1f8b24f0877bef53a8f158454381ef7b41aa9063

Observation 31528233-5fe3-4462-bbfb-19c5e332320b · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.609918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.609918Z digest=sha256:d733a9b1ffe1eddbb6482a03a485e43680b419a8dde3dfdf303a7fc1647c8c05

Observation e9c68474-da92-401f-8422-e672b75014f6 · outbound

This paper cites Representing long-range context for graph neural network s with global attention,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Representing long-range context for graph neural network s with global attention,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.971912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.615306Z digest=sha256:3479ee2d27a94f85e0f1f0d4a032534fb6049d7391a964f36e20a896e14f2767

Observation f5b02652-6964-4fe2-9172-75aad3881ef4 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.619889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.619889Z digest=sha256:bfd09ac3b73d66f9f90542004cae551f66c04dd06a78715f29ff9e9c925c31d9

Observation 4f491a2e-409f-4af4-85d1-321a611e4390 · outbound

This paper cites Blockwise parallel transformers for large context models,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Blockwise parallel transformers for large context models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.955929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.624919Z digest=sha256:7a09f6f37c79a86270208a73678ad3fcf542d7020105b167278eab072966120f

Observation d0c6718a-c934-411e-97ea-65eab7e4a81a · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.629276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.629276Z digest=sha256:443dc726bc98b8e913d581172ba9c35181954095e3ec008f1b52f25d71f2e4f5

Observation 3e7a8ee8-796c-4045-a9d0-b86c1b4377c4 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.634125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.634125Z digest=sha256:1103ce8fca2131300d127efa3467d8ba49e2bce37b5e3654920235a9ac0325bd

Observation dd7d6fc3-1fca-4267-9285-c34428813930 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.638827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.638827Z digest=sha256:475be634700aef0f679eb7d8c9de0d37007a80ffe8a030aa31f1e975fb3b6336

Observation b7c3e28b-d0c1-4d88-a2a9-25be97bb4c89 · outbound

This paper cites Flashattention-2: Faster attention with bett er parallelism and work partitioning,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Flashattention-2: Faster attention with bett er parallelism and work partitioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.941376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.643740Z digest=sha256:cb62c457e4c4219a7fc972130dcecd3d31b174066d9d3299528cc7f82e8fb8df

Observation 5094d809-8176-4923-9b10-210a5d1406c7 · outbound

This paper cites Flashattention-3: Fast and accurate attention with async hrony and low-precision,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Flashattention-3: Fast and accurate attention with async hrony and low-precision,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.926557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.648012Z digest=sha256:4f9d982213836051363ae4268a36b2d4029953920cb188e52d67f2cff013addc

Observation 15a5d1cf-8792-4de6-a79b-b2838f604f5c · outbound

This paper cites Faster Causal Attention Over Large Sequences Through Sparse Flash Attention.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Faster Causal Attention Over Large Sequences Through Sparse Flash Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.652472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.652472Z digest=sha256:39d1e8c0dd741596a3b2c5c82f3b9cd1af31f36f3efe5c4a0cf3ef3d4ac59bf6

Observation 708e9999-3208-4809-80aa-f22aaee6f15d · outbound

This paper cites Efficiently Dispatching Flash Attention For Partially Filled Attention Masks.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Efficiently Dispatching Flash Attention For Partially Filled Attention Masks

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-09T19:58:14.748858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.657264Z digest=sha256:58d7c6146b0de9ea529c93ffdc01fe2cfbdd69dcf20e03c54d24bb6dfd679cf1

Observation fb76bb74-b7da-4650-8099-a2379f0770de · outbound

This paper cites Online normalizer calculation for softmax.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Online normalizer calculation for softmax

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.661834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.661834Z digest=sha256:244cbdd0c86f9944e92508f022c93d2f0c4a1f2ac2271883d747dc71601e386b

Observation f45fad95-f6b0-4952-b127-0159a1d532f5 · outbound

This paper cites On the power of some pram models,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques On the power of some pram models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.910917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.666325Z digest=sha256:b7472b7c7050be82b850bf0165fe26f16780c6307708eaccccb203d9ed6abd42

Observation 6b0649cd-c7c1-4705-b0cd-421d6f60bb44 · outbound

This paper cites The Llama 3 Herd of Models.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques The Llama 3 Herd of Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T19:58:14.670570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:58:14.670570Z digest=sha256:47ca38e33c4a8feae7c870fd9a09f3575c34589df8f4b8a97aa18d1d0d0f92eb

Observation 57219432-3d10-46c5-8c68-3d6832828071 · outbound

This paper cites Algorithm 10xx: Suitesparse:graphblas: Graph algorithms in the language of sparse linear algebra,.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Algorithm 10xx: Suitesparse:graphblas: Graph algorithms in the language of sparse linear algebra,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:14.895370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.675856Z digest=sha256:4359ddce2b8975e995e972d0992167d83241ab3b6be9864ba19a9024db009c91

Observation b0074eb4-cd9c-4d93-a543-bcf4e14d4178 · outbound

This paper cites Available: https://arxiv.org/abs/2307.

Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Available: https://arxiv.org/abs/2307

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T19:58:15.016939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-09T19:58:14.586742Z digest=sha256:186e7c5ac42554e518f76b36b860e92b7f351bcceec370498541645ac4a58f81

Pith citing papers

Observation 68793c4a-d8bc-43c8-b927-4c223eb6b933 · inbound

Sparse Fine-Tuning of Transformers for Generative Tasks cites this paper.

Sparse Fine-Tuning of Transformers for Generative Tasks Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:30:15.058609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T17:30:14.955831Z digest=sha256:1134b495325bee381368bc3c83a5848fca3d937ad5dc313e85d0321adb050aea