Pith. sign in

Paper Citation Record · LEDGER

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification

As of 10 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2502.09647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09647 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:46:11.464371Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T13:17:57.627009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T13:21:24.216191Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e8dbd32-4477-423b-9cf2-da50db8eeae2 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.343368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.343368Z digest=sha256:c416555148ec3ef7c3b934bb6290c95193294630c0c2ffd5989bf224c0e4f8ab

Observation a3ae2c5c-deaf-485d-b45a-53cead37c47e · outbound

This paper cites PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.362705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.362705Z digest=sha256:4cabe0e3d2a0f3e280bf0d8c85a65d5c00a8a14ced3402b82f6ff73395f1c4be

Observation 210a51f4-a9b1-4d1c-bb8a-fa88c98b3879 · outbound

This paper cites MagicPIG: LSH Sampling for Efficient LLM Generation.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification MagicPIG: LSH Sampling for Efficient LLM Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.366865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.366865Z digest=sha256:dafc4cd87fc5419f795ba4100918347b51c5771ae9f336eeb57a6af734212870

Observation 7e67fef2-d988-42db-b8e1-78c4b474c4bd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.370692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.370692Z digest=sha256:bac768d9041c099196032e31b7e4ffc466048b44dbb2739e7f85fcb5a4019b03

Observation 4d1a736f-6e03-4e41-a732-64618d141386 · outbound

This paper cites The Llama 3 Herd of Models.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.374583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.374583Z digest=sha256:d01440d0a849042e6f77a1f16f7e3618226286fdd34d89491d0c4e36a79f04b4

Observation 7cdedb49-90d8-40ec-bf61-b7844b04b5ea · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification When Attention Sink Emerges in Language Models: An Empirical View

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.379308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.379308Z digest=sha256:3728eeb865251ea5b39675ac134cc0b160da1f3e5ad8d5f87323390d32325a0b

Observation 1b23ff28-9c93-44f8-b1d1-d7b36a38705d · outbound

This paper cites Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.383126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.383126Z digest=sha256:e2cc445959abf172e5147b33345d8042aaa0bb80a71b630694b9281ed701d480

Observation b0d34772-fb1e-462e-a094-5c7c29496126 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.391132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.391132Z digest=sha256:adbaf0241e62d1c8c91f42f0c6bfc76c8e2145b05b04da93dc7af0f12be55c91

Observation e5a635cf-9922-4b56-b3e5-eca0ad66b825 · outbound

This paper cites Mistral 7B.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Mistral 7B

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.394908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.394908Z digest=sha256:49beaa64dd24d06f10fb175d6175718e003aea157ff7164ee1e88f522904b322

Observation 85ee0a17-706f-4b77-bee4-a04e7c72add4 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SnapKV: LLM Knows What You are Looking for Before Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.398714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.398714Z digest=sha256:416c0e491b9c74b4a76c0d2776d425bbd1f2871612d012e34e7c582da62034a2

Observation 2d344b57-1caa-4697-9de6-70177ce76f36 · outbound

This paper cites DeepSeek-V3 Technical Report.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification DeepSeek-V3 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.402637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.402637Z digest=sha256:95572db67df084ee883959d5f65d0cacbfe793bbd4041e849e50668913513c34

Observation 7085e3eb-6652-4322-a50b-b0d1814b0118 · outbound

This paper cites In-context Learning and Induction Heads.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification In-context Learning and Induction Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.407058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.407058Z digest=sha256:96ad7fb4b830628b4bc55013d0af68a6a8311bd19c56dd8af1b9b0030ead4103

Observation 10e161f8-1e4c-413d-8d14-22e51e458da2 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.415240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.415240Z digest=sha256:055aedcbcb45c89c4a3df23e322c1aef6e71ef2cc073e3e8f87380c02cdcf6d8

Observation e9cacb04-604f-45c6-97e3-80b3e285eaeb · outbound

This paper cites ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.424138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.424138Z digest=sha256:b291d1bb8519cadd9f7827271fdab13b2c4df1fe9615b7af7bc6623c0c4166e7

Observation 2883a270-e414-4807-909d-47b065b6dd6b · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.433213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.433213Z digest=sha256:48c24de2b50f80cfd1f6823211a6df5294977fd20c76688935fcc966d57630df

Observation 139f1823-1281-4720-991a-971bac5260ef · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Efficient Streaming Language Models with Attention Sinks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.437333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.437333Z digest=sha256:d43e47ab874479731656ca78bdb3821751a38afbcebc6e54f62bdbe568182e99

Observation 008df740-affe-4a11-a002-85932160f87c · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.441263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.441263Z digest=sha256:92651f286b26f6e806b1fb093e0132fd6d13b558257e786026ba873acf598aa4

Observation 02ec2ad8-0dd1-4c47-9c1a-0b9dbffcc7da · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Attention Heads of Large Language Models: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.446469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.446469Z digest=sha256:1f5f2594c96d36fda001cb085ecbafa4fde65f7632698dfd1e2de70b4c46affb

Observation 3b08fee2-2923-4f2e-849d-c6c8b4f77bc1 · outbound

This paper cites (2024); Tang et al.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification (2024); Tang et al

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:46:11.773078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:46:11.451445Z digest=sha256:8e21c574efac6b92329f1411a6f919dd27cfcb253bac248eba58f4dc7abb0bcd

Observation f95a754a-db31-4341-b0d2-d7808e626bf5 · outbound

This paper cites qa-1” and “qa-2.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification qa-1” and “qa-2

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:46:11.760733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:46:11.455913Z digest=sha256:a77a84aefa3ed13725352a73a918c83ff6ffc2f6ea34b3296ee2d4a696eb4328

Observation b52d1065-85d1-41a5-b1a0-59cbbdd4747a · outbound

This paper cites an unresolved cited work.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:46:11.749185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:46:11.460284Z digest=sha256:41e2ebb10dd236a2b501943264745e90c579ea00f6d49fc462e17421f2b19afe

Observation 654e4c44-559e-44a3-9442-9f1820fcbcad · outbound

This paper cites observation.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification observation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:46:11.737206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T13:46:11.464371Z digest=sha256:bc2cec416e68a59c1f8a183455b5572352a4e5f52c8b2fddf030e716a5410294

Observation 870ee56a-545b-4041-89b2-ff0acca5654d · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.419712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.419712Z digest=sha256:1b7b92a3e7454e90e2e45353c3be526676c45540ff954159b250ea8cd36e2551

Observation dc37a957-0548-46d9-ab98-bf9802b373e6 · outbound

This paper cites Efficient Large Language Models: A Survey.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Efficient Large Language Models: A Survey

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.428921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.428921Z digest=sha256:abe05f919c411330d382794a4cf9df79950aaad371fa3126da4424b58109330c

Observation decf2b0b-31c7-418f-b8d3-0a81247d7f17 · outbound

This paper cites Qwen Technical Report.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Qwen Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.358715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.358715Z digest=sha256:2ef054371be6fb095830f14100aea651a9a27856f82d09e43f5430ea24f12325

Observation 00232cfb-b8d0-4ffd-92b7-ac3239210597 · outbound

This paper cites Transformers are Multi-State RNNs.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Transformers are Multi-State RNNs

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.411181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.411181Z digest=sha256:f2a79790f287e53bc9336dabd076360fefbc2e969751f4149e0c0f0681ffe884

Observation 9d1b1eb3-2eb3-4236-a4f4-eef2411fa7f7 · outbound

This paper cites ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.348797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.348797Z digest=sha256:c1fe9d96949a370cfbaa8c6cc85c6d1399fa69431923003e7110c0fdc4e98f7c

Observation 8ea75401-2c1b-4358-8536-c1d5d5a5fd8d · outbound

This paper cites Program Synthesis with Large Language Models.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification Program Synthesis with Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.353983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.353983Z digest=sha256:aff41ef16ab35b0a4cba06e29de17f16576442eb56b9863597bf3b935bc15b5d

Observation fc685a8a-5258-401e-b729-4d8d8067d15f · outbound

This paper cites On the token distance modeling ability of higher RoPE attention dimension.

Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification On the token distance modeling ability of higher RoPE attention dimension

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T13:46:11.387098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:46:11.387098Z digest=sha256:431e254b7cc664d3ec84a03567b255922d33ce05c18e8a43fbde0b4452d52d5d

Pith citing papers

Observation 2cb2a063-772e-4b2d-8e67-99c4953ce8be · inbound

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding cites this paper.

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding Unveiling Simplicities of Attention: Adaptive Long-Context Head Identification

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:21:24.218559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T13:17:57.627009Z digest=sha256:c725bf3744ad2becbacbd0215508bc45657c35907a4b4efa866c944022e79e5d