Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:58:14.675856Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.01659.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T19:58:14.675856Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:30:14.955831Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T17:30:15.054198Z
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ad0472da-339c-4524-80d8-8a306225c2c6 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques A Survey of Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4660cd1-02b2-46f5-910e-2475dec1acb2 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Enh ancing molecular design efficiency: Uniting language models and ge nerative networks with genetic algorithms,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a873b736-35bf-4859-ae49-63ca1ac51876 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Path-bigbird: An ai-driven transformer appro ach to classi- fication of cancer pathology reports,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 55288ae4-3069-4aea-bbf5-72de2c81f4fc · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 482daa7e-b360-45fe-83ce-ef2c1586d498 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Attention is all you need,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe2008fa-86f5-4635-9eb2-17a2a97781aa · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Big bird: Transformers for longer sequences,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 111e59df-db05-4aa6-8cc8-4fb5eb861679 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Longnet: Scaling transformers to 1,000,000,00 0 tokens,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 519b063f-f8c2-495f-b7b6-70a578983ac2 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Longformer: The Long-Document Transformer
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a142f2ca-da0e-4d9e-ae0b-fc39cdfd7904 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Scaled Dot Product Attention,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0b861f90-77c0-41b4-80be-1ac1764db6ad · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques xformers: A mod- ular and hackable transformer modelling library,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 07c2b4a2-cc94-449b-b72d-7f07304adabf · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Reformer: The Efficient Transformer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31528233-5fe3-4462-bbfb-19c5e332320b · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Generating Long Sequences with Sparse Transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c68474-da92-401f-8422-e672b75014f6 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Representing long-range context for graph neural network s with global attention,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f5b02652-6964-4fe2-9172-75aad3881ef4 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f491a2e-409f-4af4-85d1-321a611e4390 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Blockwise parallel transformers for large context models,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0c6718a-c934-411e-97ea-65eab7e4a81a · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7a8ee8-796c-4045-a9d0-b86c1b4377c4 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd7d6fc3-1fca-4267-9285-c34428813930 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c3e28b-d0c1-4d88-a2a9-25be97bb4c89 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Flashattention-2: Faster attention with bett er parallelism and work partitioning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5094d809-8176-4923-9b10-210a5d1406c7 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Flashattention-3: Fast and accurate attention with async hrony and low-precision,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 15a5d1cf-8792-4de6-a79b-b2838f604f5c · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Faster Causal Attention Over Large Sequences Through Sparse Flash Attention
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708e9999-3208-4809-80aa-f22aaee6f15d · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Efficiently Dispatching Flash Attention For Partially Filled Attention Masks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb76bb74-b7da-4650-8099-a2379f0770de · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Online normalizer calculation for softmax
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f45fad95-f6b0-4952-b127-0159a1d532f5 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques On the power of some pram models,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6b0649cd-c7c1-4705-b0cd-421d6f60bb44 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57219432-3d10-46c5-8c68-3d6832828071 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Algorithm 10xx: Suitesparse:graphblas: Graph algorithms in the language of sparse linear algebra,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b0074eb4-cd9c-4d93-a543-bcf4e14d4178 · outbound
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques Available: https://arxiv.org/abs/2307
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 68793c4a-d8bc-43c8-b927-4c223eb6b933 · inbound
Sparse Fine-Tuning of Transformers for Generative Tasks Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.