Pith. sign in

Paper Citation Record · LEDGER

Faster Query-Key Learning Sharpens Attention in Self-Attention Models

As of 17 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2608.06776.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06776 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:18.403313Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:48.968402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T04:15:49.290160Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact3
  • verified fuzzy39
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 05db1fb3-1915-4d4d-8c9c-42167c5cf6af · outbound

This paper cites On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.024092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.024092Z digest=sha256:2008b0b1298755519b4b22a81ae2252a5ad245517084dc1f6c0d4cd6b7bc4a9d

Observation 1f04f26b-4471-4ce9-bd8c-9554a47a806a · outbound

This paper cites Self-attention networks localize when qk-eigenspectrum concentrates.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Self-attention networks localize when qk-eigenspectrum concentrates

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.535605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.031138Z digest=sha256:893fc6077d39374a7d3e89b8a4f6d98cb2349e79452ffe3bb84a244ed25e1a23

Observation cd6eb2e9-caea-43e5-ad11-61852b497662 · outbound

This paper cites Birth of a transformer: a memory viewpoint.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Birth of a transformer: a memory viewpoint

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.513920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.047780Z digest=sha256:5d018b8028944da1fb2bd81a440325cab2f98607189740a6b625c613de6b0109

Observation 68441002-f3d1-488f-844c-f3690651deba · outbound

This paper cites Language Models are Few-Shot Learners.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.052596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.052596Z digest=sha256:5b833dd23d624311b19177b015ee666ca07c91db8fef56209517b525e81943ae

Observation c4b2cc58-65ff-4a71-be7b-5feffc94c991 · outbound

This paper cites Universal Transformers.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Universal Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.057149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.057149Z digest=sha256:1e6a6884d6dc345e98d4d8f40cbd540e3d6859729dfd64133d3d04839bf80d93

Observation 067bf91f-cb3a-40d2-ab3e-795be6d9a3c1 · outbound

This paper cites On the optimization and generalization of multi-head attention.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the optimization and generalization of multi-head attention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.494546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.062406Z digest=sha256:25116302e0e9aecb51a58ac58b61e209816d8a8d60662a7b295496c04b89f0c5

Observation e9cbbafc-1209-481a-8049-06b45c4e698f · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Optimization and Generalization of Multi-head Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.066978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.066978Z digest=sha256:780aead4321d8647594818e3e71cc9eddef878fa12a6f59c0fb7a1fccc91408e

Observation 6d32f044-0fca-489c-87a3-948b4f932167 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.077317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.077317Z digest=sha256:5b697ffd38d9d9b192af2084c6234e4265f033372f540f5cc31969a1199d1a0c

Observation e466f636-4ef9-494b-b147-0b4bf080a84d · outbound

This paper cites S., Hu, W., and Lee, J.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models S., Hu, W., and Lee, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.478602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.082324Z digest=sha256:7119b03a142a90663b92709827f6d14aeb6ea653aadb630ec6a1c676e6e612b4

Observation cb061834-8ad3-4064-af60-2b39a56c5cab · outbound

This paper cites A mathematical framework for transformer circuits.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models A mathematical framework for transformer circuits

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.086746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.086746Z digest=sha256:c1008992b85b6d13d5f848b0ecaf9d7c0f0de1f98c548ff95d7d20dfdd3471be

Observation 943218e2-4634-4659-bb85-c76682efc00d · outbound

This paper cites The emergence of clusters in self-attention dynamics.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models The emergence of clusters in self-attention dynamics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.450939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.091642Z digest=sha256:803f9a696063f52d0584b062623ef2406f2e5ed7dc888c0d81a8a5a666523dc5

Observation 2d979015-6b34-4794-b445-e49d6f16c970 · outbound

This paper cites and Wallace, B.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Wallace, B

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.428665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.096096Z digest=sha256:435c57df8758d7c091c967ee7bb53bb5bd4e0855962ea732b15426ad24d27c94

Observation aeb751e1-e6a1-464f-8c3e-9554802a074b · outbound

This paper cites Clustering in causal attention masking.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Clustering in causal attention masking

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.411456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.100541Z digest=sha256:a7c82af492a57c27cfcbd1b9fa95a1d8691dca3d848886ff9a25dcc94168d33d

Observation fbb1ea4c-32ed-400f-b5bf-9c2a5f4e0b06 · outbound

This paper cites Transformers in Speech Processing: A Survey.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Transformers in Speech Processing: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.105188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.105188Z digest=sha256:fbe8b1ca9309aa37c7ed957795557ce9c513ad681a80b15bcc1c335758e662e4

Observation ed667f7c-1045-4aa7-aeb8-e55d0dee41c1 · outbound

This paper cites How do transformers learn topic structure: towards a mechanistic understanding.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models How do transformers learn topic structure: towards a mechanistic understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.394395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.110194Z digest=sha256:7fd2b88aaf8074ecf7c53785811fb0b285341d422cd70e87aed8104be07135bf

Observation 98b45860-5b3b-499c-bdfd-1ab759a7fca4 · outbound

This paper cites Mechanics of Next Token Prediction with Self-Attention.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Mechanics of Next Token Prediction with Self-Attention

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:39:18.616383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.114791Z digest=sha256:bdb5c35638d72109291530dd58a7075b546aa34a0dd4713ec3a1a90e9e1751e4

Observation 3b8c52ce-c7a1-41a4-aff1-6afe9c5d5b1b · outbound

This paper cites On the dynamics of training attention models.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the dynamics of training attention models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.375759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.124678Z digest=sha256:ae512564ea880c8bc1ca2db228f8120a0465804cae7869d1e64d1c836c582c74

Observation 075fc0dd-a114-433e-9230-450eb2083918 · outbound

This paper cites M., Biemann, C., Goyal, P., and Mukherjee, A.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models M., Biemann, C., Goyal, P., and Mukherjee, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.359952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.129625Z digest=sha256:8c5db7194ce8b86a4310b82aa6c7697e0effd808ae228a43a584e6e14a81b8b0

Observation 51cf7cdf-3d68-464b-8078-1685cf2ed3df · outbound

This paper cites In-context learning and induction heads.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models In-context learning and induction heads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.134544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.134544Z digest=sha256:0b7f2f1c821ba3dbc8f7a514c8afb8434fab4a32cafa60dc1e147aeed497f4a2

Observation 52b99241-6a03-4d27-87eb-66ac25ca15bc · outbound

This paper cites N., Vashisht, R., and Ramaswamy, H.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models N., Vashisht, R., and Ramaswamy, H

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.334617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.139452Z digest=sha256:1250467ef74f460f593d424a510582f8735fb142ea172f1e40bf74dee8f1d832

Observation 5fe12347-73ea-4936-9372-1998bc77a2d9 · outbound

This paper cites Attention is turing complete.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is turing complete

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.317498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.144322Z digest=sha256:8983fb57d8ba9ebd0e365f8e4376e1fda568293dea097bb1b8838169119cb720

Observation 1b8e5479-3e76-4b9b-8c91-81e0fd0b96e3 · outbound

This paper cites Improving language understanding by generative pre-training.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Improving language understanding by generative pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.150061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.150061Z digest=sha256:5930d40dc9d24ebd5450f4f0b90d80348d33cca3e4e29b4476b93ba96f311121

Observation 8be5a6b6-9d0b-4cd2-9274-9fba9d662ea9 · outbound

This paper cites Scan and snap: understanding training dynamics and token composition in 1-layer transformer.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Scan and snap: understanding training dynamics and token composition in 1-layer transformer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.287972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.165853Z digest=sha256:a571948d8602276bb617336f5adbb2933590617953bdb8c285618328e6c3c7e3

Observation 1ffb953c-5610-48a1-adee-fdafc300c806 · outbound

This paper cites and Ramaswamy, H.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Ramaswamy, H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.272360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.171140Z digest=sha256:b71da6cb32b664251acce46bf1a9f3f2e660c01954be105c3d7ffa073a4df343

Observation 7f632b6b-0afa-41fb-9f36-6a46896f5417 · outbound

This paper cites N., Kaiser, L.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models N., Kaiser, L

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.176192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.176192Z digest=sha256:89b66618b28b5968887f71df8c152e11d20a3f2278af63f46ca577873abcf8b6

Observation b029877b-2e9f-4876-aa2c-e7bad65ca26b · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.186873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.186873Z digest=sha256:a5dba73c1b2fa5466d76766b37185f81ac75c9f26fb1b1a35374f4701e74e95a

Observation 85017a99-ae25-4307-a77e-cbcdae1831a7 · outbound

This paper cites Attention is All you Need , url =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is All you Need , url =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.247018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.191872Z digest=sha256:6d43f416c7c47a03c8271d73c7c827b19aa82f76c012dfcad8fc9eb00f4efd70

Observation a9962a82-cb5f-4b3b-bfb0-52dc99b649be · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.229176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.197816Z digest=sha256:4b754152763985d4316a43498ae11bd8c7e810259fffc7a510640492b0183ba3

Observation 2ce633c1-dec7-4c60-afb2-4cd12f19af38 · outbound

This paper cites and Kumar, Sanjiv , biburl =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Kumar, Sanjiv , biburl =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.213639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.203365Z digest=sha256:83ab8daf7c435a995dc4e4e551613b9c7deab9109844d424bb736dc17c9c0d0f

Observation 5d45f57c-fe0f-49be-8482-35a94a84243d · outbound

This paper cites On the Computational Power of Transformers and Its Implications in Sequence Modeling.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Computational Power of Transformers and Its Implications in Sequence Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.208510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.208510Z digest=sha256:8ede1ec7484c1e83af7bc5fb6f707ec47b08dc8d1bb196c277de05c4417bdd51

Observation d4cb7314-b594-4f33-8148-cd3c81dac9e1 · outbound

This paper cites On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.213259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.213259Z digest=sha256:35b387bd71d388623c8237d31c67026e92092ad327e3c6fac565eb915fe74e78

Observation 3b8377c4-f28d-453a-acb9-f3fd5d7cd7fd · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.197617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.219271Z digest=sha256:295944e0a11895852320c72d1997212ece7cbc11784515c2c3e9a5b9130e43f4

Observation d77dc4ac-07c0-43b2-ab88-db6a58df865a · outbound

This paper cites Attention is turing complete , year =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is turing complete , year =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.179134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.224200Z digest=sha256:2eb69b27b20ed0cd8fcb56de4b5ea4ecbc294e7adf5d6d8c5deb7e7c47ae0d41

Observation 61fbacbd-5c02-4471-afc2-08d979fb5a98 · outbound

This paper cites and Ba, Jimmy , biburl =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Ba, Jimmy , biburl =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.161972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.231213Z digest=sha256:89d2d20fcaff81534ecc9c590bda94ea494b4c3eb50c4748026d68335beb2c61

Observation c4106af3-60f1-4649-bdb5-0c2563327275 · outbound

This paper cites an unresolved cited work.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.237378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.237378Z digest=sha256:f03ceb881929573edca7b94026c9b6bc528a423cfd7fa898e6c363f7f350913a

Observation b29081aa-6747-4c23-881c-16e2651e7f3f · outbound

This paper cites 2021 , journal=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2021 , journal=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.242650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.242650Z digest=sha256:c55329615a616fd255cdb5ac7706c21855cb4b0795ae50c622fccb6e4a165d4e

Observation a08a058c-d101-4628-a59b-99d93c1190fb · outbound

This paper cites In-context Learning and Induction Heads.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models In-context Learning and Induction Heads

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.122060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.247761Z digest=sha256:cc160342995d4d30a22180d483f80bbcc909f81bae63f03dec48200e2378f2cc

Observation 2eb62838-e40d-4716-ab1a-13ba5b48a26f · outbound

This paper cites Birth of a transformer: a memory viewpoint , year =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Birth of a transformer: a memory viewpoint , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.103887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.253256Z digest=sha256:d650936c5313de170bd1d0711aac924533c603c8c811717c85492f47a46a8fab

Observation 9be6d413-dccf-4002-8194-ad25097c2cab · outbound

This paper cites SQ u AD : 100,000+ Questions for Machine Comprehension of Text.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models SQ u AD : 100,000+ Questions for Machine Comprehension of Text

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.258382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.258382Z digest=sha256:adb357807d778faeaebb52e5ec2bd7f097ab4e121f7d73cfff519b64b48111c0

Observation 068cf1d7-7f45-4414-a4b1-7db80a11e364 · outbound

This paper cites Assessing the Ability of LSTM s to Learn Syntax-Sensitive Dependencies.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Assessing the Ability of LSTM s to Learn Syntax-Sensitive Dependencies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.262761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.262761Z digest=sha256:ed3d2ba21cbc336c3f51e952bfafb33743533397136dec8ab243fca276937d14

Observation 14f7765e-f0c2-4d95-b874-b2dd21a4be76 · outbound

This paper cites AAAI Conference on Artificial Intelligence , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models AAAI Conference on Artificial Intelligence , year=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.086276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.267286Z digest=sha256:21db095b33792ac31094c886197b02679b52a9f06939acd0b5313ff583c9b082

Observation 23c39edd-bddc-4ee3-a11c-98918278029e · outbound

This paper cites Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.068323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.272953Z digest=sha256:21c6cb5c0d86bec5c1d32a0f5ccc6bd1b00b626444e541f54ce3622564605944

Observation 9c8e705f-8870-49fc-9ad2-ebe22dc20016 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , url =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , url =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.277783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.277783Z digest=sha256:aca74bf275b4930cf099d677e14dcba5a21eae14cc1dd41090e3df8c859b8c03

Observation 9a8d2a32-6d35-4dcc-a34c-982f5f57ddba · outbound

This paper cites Language Models are Few-Shot Learners.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Language Models are Few-Shot Learners

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.284319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.284319Z digest=sha256:7567a714b7c562045682914035ccc7a4db36f794b17c5d75b4c882b87a34016b

Observation 2f7b9722-6608-46af-8395-362f5039b842 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.289534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.289534Z digest=sha256:749958d65ff3faa9b9998f7725db5861c315d3c26275e02ea62a7fe58cdc8035

Observation b45d4986-e646-41c3-bde9-6c52bac9508d · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.294174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.294174Z digest=sha256:408d7bfefd1ff2c69e6880d1395d45f083caa6ed136a3a64d579aac7c533f8ff

Observation 5b585f6d-0d2c-489a-8b1f-e72b224ff1f3 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.016311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.299668Z digest=sha256:49198e7e2ef08442a65397b72f5537bd47d8a36c1e0f8cdcfa2cd71e6431efd3

Observation aa5540ea-bd45-4b08-9017-ac9559d068d4 · outbound

This paper cites and Hu, Wei and Lee, Jason D.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Hu, Wei and Lee, Jason D

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.997570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.304807Z digest=sha256:8849af0ede8632e78725c324d7a2d681e893502246ffa973bb4382cb8853307d

Observation f437a744-300c-42be-bb1a-ee39277e3456 · outbound

This paper cites 2022 , journal=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2022 , journal=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.310228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.310228Z digest=sha256:6e9c729150c430e31f182fd28f30f5ce2d1a6d89285b83b0d70b063923a72336

Observation 7163e939-6afe-489e-9b05-27f555586377 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.970434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.316089Z digest=sha256:d181985ec5fc981479af12cc26830c2046e6011af008fe75890d40f6e09b46eb

Observation d1602f52-1aff-4bea-b548-b596476f0b21 · outbound

This paper cites 2025 , eprint=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2025 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.956155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.321322Z digest=sha256:8e92ce2c79e3207401661e1a18ac054932d21f2fb46c94142339c499bad5f528

Observation e2296ce4-2ff9-4269-9729-6c34e0085c92 · outbound

This paper cites Attention is not not Explanation.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is not not Explanation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.327189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.327189Z digest=sha256:2f5c880167d3140611065cac220d765f85088a2bca1f3987dc17965b5cf9aa31

Observation 521285b6-e782-406a-bf6d-d1f40b979d6a · outbound

This paper cites North American Chapter of the Association for Computational Linguistics , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models North American Chapter of the Association for Computational Linguistics , year=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.941436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.332384Z digest=sha256:a25807507abd95f44950db5d92189643e9c1f4b5e12f03aa37f9ff170bfaf5f0

Observation 07c2fcb9-a389-46bd-834e-4e59d86f645a · outbound

This paper cites Is Attention Interpretable?.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Is Attention Interpretable?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.338319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.338319Z digest=sha256:e33ee1b12958ab24d4eba47e59f50a7ef9b294cde8315605634630670cf5ad3f

Observation 300fb36e-0524-4fe8-af54-54b6a7c665d3 · outbound

This paper cites 2024 , eprint=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.925397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.343703Z digest=sha256:d3b1c4ede803646f51cde6389c62541b340d329d26008592053c717478177781

Observation 68feb1c8-a306-497d-a414-4a0f00296363 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.909569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.348421Z digest=sha256:a7d93ba44c45e2950199d790d08451a6ced4d3c0f01aef2eaf170155159dc751

Observation 381831d6-f1f6-428a-a8e6-5cfe10bfe1cf · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 41st International Conference on Machine Learning , articleno =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.893694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.353396Z digest=sha256:7da41d978ba61dbe1817a971a6713fca10ddd532e64cae4ae98de76e60bb6cbe

Observation 18b0ed52-e575-4366-87cd-110ff5be9c6c · outbound

This paper cites Transactions on Machine Learning Research , issn=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Transactions on Machine Learning Research , issn=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.874535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.358371Z digest=sha256:6241b52e082c337fe8c81f83c72c08d1489c67408becd87954bd5c469fba5a56

Observation 15745298-06e3-45ef-afa0-a94c60961d3f · outbound

This paper cites Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence , articleno =

Reference 67

Resolution
verified exact
doi, observed 2026-08-15T14:39:18.468085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.362713Z digest=sha256:50aa10bec6e80fa748edc415f5bf84749f668648c6fdfc113109e1092f76e4dd

Observation 44b4543b-df80-4580-911e-9d83d571ef11 · outbound

This paper cites International Conference on Learning Representations , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models International Conference on Learning Representations , year=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.854925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.367236Z digest=sha256:e5d0ac70476868049d357bdaf4f975aff0fc01f9c3f5ea6c0270ad4a017ad8f8

Observation 9e226e27-f37b-4687-8bf0-40f6d8272bc6 · outbound

This paper cites an unresolved cited work.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Unresolved cited work

Reference 69

Resolution
verified exact
doi, observed 2026-08-15T14:39:18.448554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.371684Z digest=sha256:49043572fccc7f2ec34a9c0e80c8ff4a6b697095acf28a9723ba9756ea8c4a16

Observation c9a9e8f5-7966-4a6f-992f-1ef36ca5cdb3 · outbound

This paper cites Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.835752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.376632Z digest=sha256:576e42c3384164c3d7c28b2ac76027f94cebad4a02d1e153afd94ab1e6afecdc

Observation 3b8a100d-98f6-4ca0-a665-cff24368d6fa · outbound

This paper cites Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.818156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.381279Z digest=sha256:4444d7b8096ceb9b5a4339edcf0a88df5aa5d7788ed5d5d5ae765bd650bdd8bd

Observation af580012-a011-4851-bec7-bd87cb694164 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 40th International Conference on Machine Learning , articleno =

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.796993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.385781Z digest=sha256:86c8977d03d7446d7c0db3eab0542206a972db2980ea38b929be2bc617fab68d

Observation cd807fa5-061e-4dba-b80b-2eb806503025 · outbound

This paper cites Quantifying Attention Flow in Transformers.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Quantifying Attention Flow in Transformers

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.390155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.390155Z digest=sha256:7ad1ebbd891d242abc18fc488f202ea1715a8f1fa29269203a206a0383bd959c

Observation b166557b-8498-4d35-bc1e-b76757058ca0 · outbound

This paper cites ERASER : A Benchmark to Evaluate Rationalized NLP Models.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ERASER : A Benchmark to Evaluate Rationalized NLP Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.394573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.394573Z digest=sha256:79900acd0701a3da1ed553df827ede579a525f33b8100b5192d78c7875a546b2

Observation de0c266e-9a1d-4525-b9c9-abda41873ce3 · outbound

This paper cites Proceedings of The 14th Asian Conference on Machine Learning , pages =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of The 14th Asian Conference on Machine Learning , pages =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.778188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.398617Z digest=sha256:bb9cf73ff8ef896979716fac4a783d2c91339c8be72f5878505586917ebfd2ca

Observation 45bb5563-3142-4a07-b98a-769964221409 · outbound

This paper cites European Conference on Artificial Intelligence , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models European Conference on Artificial Intelligence , year=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.758095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.403313Z digest=sha256:e49a85f6233da0c2d634abfb64c3b5b8bd46a463a9f55ec10eb80a7dbb45c022

Pith citing papers

Observation f6db4ff7-a0e7-480d-a701-6470ca84d468 · inbound

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference cites this paper.

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference Faster Query-Key Learning Sharpens Attention in Self-Attention Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:15:49.294748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T04:15:48.968402Z digest=sha256:1f1424e8c3b6588c3809467fea28fc234a396fdfa0accfb96046467895cd187e