Pith. sign in

Paper Citation Record · LEDGER

Selective Attention: Enhancing Transformer through Principled Context Control

As of 24 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2411.12892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12892 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:12:32.864356Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4c03ae0-6ca2-43c9-8bfd-061929b7d7aa · outbound

This paper cites Generalization on the unseen, logic reasoning and degree curriculum.

Selective Attention: Enhancing Transformer through Principled Context Control Generalization on the unseen, logic reasoning and degree curriculum

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.570684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.615306Z digest=sha256:e9456588aa3f28c5cbb9a5e54fe7f1f76e3b6f6f09cde47263a0fa05b95c416e

Observation d16b9858-1456-49ef-bd25-cb7b6524ae0e · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

Selective Attention: Enhancing Transformer through Principled Context Control Simple linear attention language models balance the recall-throughput tradeoff

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.620286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.620286Z digest=sha256:1c6d0b52950fa7241cfb5c67c0de3976d2b4b4752374bdd3ca3d2c6fd887e4b9

Observation 7e584632-0adf-459b-bcd1-a74b27936164 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Selective Attention: Enhancing Transformer through Principled Context Control Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.626226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.626226Z digest=sha256:195d17e1db060b74a08790d6cd7dfe619ff1bce9d8a80a8c9d8c55234fb927aa

Observation 4e0e51ba-d25e-4ac4-bd08-b1e4c0b4a770 · outbound

This paper cites Piqa: Reasoning about phys- ical commonsense in natural language.

Selective Attention: Enhancing Transformer through Principled Context Control Piqa: Reasoning about phys- ical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.630746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.630746Z digest=sha256:295173346d6ead9337c9534274e52fb77a7b9549c892347fc4687d9d1aa80128

Observation 134d0f6c-7f9b-4447-867b-55b3c34f0d33 · outbound

This paper cites Language models are few-shot learners.

Selective Attention: Enhancing Transformer through Principled Context Control Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.635218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.635218Z digest=sha256:fa785a21250d6bb5456d44d217dc0bf0ea6af3f2a102cf36427dc20e92156628

Observation 465b0d33-9678-4e9e-81d6-1ba9e8434219 · outbound

This paper cites Scatterbrain: Unifying sparse and low-rank attention.

Selective Attention: Enhancing Transformer through Principled Context Control Scatterbrain: Unifying sparse and low-rank attention

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.531266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.639946Z digest=sha256:8e4cb5545157094c1ac081430da8e9aa7d58d02173d4c792ed1d2e68a878e24a

Observation b29b67e7-fe97-4854-ac37-8fb24f3c14bf · outbound

This paper cites Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation.

Selective Attention: Enhancing Transformer through Principled Context Control Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:12:33.252661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.645822Z digest=sha256:420bfc719db833fff633418436f44e14d1b5dd00b78f7a053b237b13bf12c52c

Observation a4f4835b-7aa1-4246-a06d-93fbee0e238a · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Selective Attention: Enhancing Transformer through Principled Context Control Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.650537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.650537Z digest=sha256:149fabda5b03959f3e0661a979e31138762f56922fc0e273223fc36463e2e8dd

Observation 7babb13d-2381-4da3-9191-0bfd022f5961 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Selective Attention: Enhancing Transformer through Principled Context Control Palm: Scaling language modeling with pathways

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.655107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.655107Z digest=sha256:5ae9c89128f295aec5584eaa159bd77a756017d9127a5f8f85c43d7bea70e880

Observation 2a7830f5-5812-4c78-996c-f04f888ab748 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Selective Attention: Enhancing Transformer through Principled Context Control Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.659558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.659558Z digest=sha256:09e9aba31ece782653d9daabdf8fe2a3d5864f1775cbfe9b6acf55cdffc8cbb0

Observation 01245eda-53d6-4d45-ae9e-a5fdfce0243e · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Selective Attention: Enhancing Transformer through Principled Context Control FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.665568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.665568Z digest=sha256:f2913184082db81088a83b3e2060cb11adc8a2e42e389cca1b12c34b40c05d40

Observation 2cf976a9-74fd-43f0-918e-464c5ab7b23d · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Selective Attention: Enhancing Transformer through Principled Context Control Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.670886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.670886Z digest=sha256:f2e9a915d847fd06f5b662357fd4f76a7f88f4d0c7a3379c5b3516d25719b4d8

Observation 0a8f6f18-0558-4bd9-8e71-3efd00f61b86 · outbound

This paper cites Language modeling with gated convolutional networks.

Selective Attention: Enhancing Transformer through Principled Context Control Language modeling with gated convolutional networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.501525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.675400Z digest=sha256:2b74c9c3ad23d4cd41025b96093f91b8300ee105587081d126ece2bbfc98d0fe

Observation 4ccf5561-4705-4e92-ba3c-7ceb6021ba38 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

Selective Attention: Enhancing Transformer through Principled Context Control Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.680612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.680612Z digest=sha256:2768d84dce57504ffbc2c7d5b1b13484fc5da3432bbb92d4ea0be85e97c9b8dc

Observation 6bc3b1c0-9a81-4433-a9e3-bdb308a8a9ea · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Selective Attention: Enhancing Transformer through Principled Context Control Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.685858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.685858Z digest=sha256:d8da3a65a0feba438a914ddc71e730c939f6207ed0981b0496502ff0bdc75611

Observation 0f4c0856-bb0e-4463-bb79-8a24fc5685e7 · outbound

This paper cites The Llama 3 Herd of Models.

Selective Attention: Enhancing Transformer through Principled Context Control The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.691159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.691159Z digest=sha256:6af04b764228ed6d26073684ac97acf36d5f5fc5c42c2222d329a75eebc3ca35

Observation a95a4e96-1b86-4910-ad5e-675d7a23820b · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Selective Attention: Enhancing Transformer through Principled Context Control The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.695780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.695780Z digest=sha256:f72ffb403020e87064d69c6a2df01c1cbc8fba64dea880d5635ef5b809851d60

Observation 6b456d1e-30ec-4684-90d6-45fbc373eb66 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Selective Attention: Enhancing Transformer through Principled Context Control Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.700295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.700295Z digest=sha256:4be3f631086979c60b8db36e0fdd894b1f4577561cca6a97b2beefb5a4c77d06

Observation 8b8be8c8-6bad-486b-b6df-16ee79492774 · outbound

This paper cites Long short-term memory.

Selective Attention: Enhancing Transformer through Principled Context Control Long short-term memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.705197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.705197Z digest=sha256:4477349175977c0432ef6612660e060a557ee0e7899b3156e303d98bcc26b70b

Observation bb862eb3-8d67-4908-8b40-fedb66f1561a · outbound

This paper cites From self-attention to markov models: Unveiling the dynamics of generative transformers.

Selective Attention: Enhancing Transformer through Principled Context Control From self-attention to markov models: Unveiling the dynamics of generative transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.470689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.709320Z digest=sha256:4e4ab37663a67511997783bc8f5f715214ad30f538e3842c9a13238b9b1b8510

Observation 11150208-117f-4ebd-be29-2eb91f3af28d · outbound

This paper cites GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling.

Selective Attention: Enhancing Transformer through Principled Context Control GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.714550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.714550Z digest=sha256:6ccf0b280c9ef9ad0e239e8639e8ca72aaed47b22d9a25871939bd6de93eb2dd

Observation 3e1023c4-1e10-4820-bb86-e58ff4f4a18d · outbound

This paper cites Au- tobalance: Optimized loss functions for imbalanced data.

Selective Attention: Enhancing Transformer through Principled Context Control Au- tobalance: Optimized loss functions for imbalanced data

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.457437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.719357Z digest=sha256:59038f9ad7eab1a7455c6e4cb05afab994ca1765f62879d6d3b5b1e1472cf42f

Observation 84955bea-7ea9-4a27-b52f-3c9bea82cf3f · outbound

This paper cites Exposing attention glitches with flip-flop language modeling.

Selective Attention: Enhancing Transformer through Principled Context Control Exposing attention glitches with flip-flop language modeling

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.442093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.723576Z digest=sha256:bda7109b0b6b467a5e1a76405f487b1231f115d5253ed73bf2ca10a3fb5366c9

Observation 12495330-3e98-4525-bd5d-28429bf2b89a · outbound

This paper cites Decoupled Weight Decay Regularization.

Selective Attention: Enhancing Transformer through Principled Context Control Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.727685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.727685Z digest=sha256:e393e9c21e3ac5b80b319abc8a9fbb481ea506ced279b8cae3d29103d0a5111f

Observation 755132b2-c44f-48e2-8d58-0c342a3ec52f · outbound

This paper cites Long range language modeling via gated state spaces.

Selective Attention: Enhancing Transformer through Principled Context Control Long range language modeling via gated state spaces

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.428298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.731965Z digest=sha256:38c08458a6755f92a0af99c2f46ed97781452f352d32f39115082ab5c75ac9ed

Observation 8af3079e-e0cc-4aa7-b1cf-c95e3b1e17cf · outbound

This paper cites Long-tail learning via logit adjustment.

Selective Attention: Enhancing Transformer through Principled Context Control Long-tail learning via logit adjustment

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.413125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.735839Z digest=sha256:89c116e26e724ccb0a186b074140e8ddbbd9e940b77f9d16311df73cb2a28f32

Observation 01d99d14-a619-4f81-9054-122e6971f143 · outbound

This paper cites Pointer Sentinel Mixture Models.

Selective Attention: Enhancing Transformer through Principled Context Control Pointer Sentinel Mixture Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.739949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.739949Z digest=sha256:33605ffe3c7b4c34eac52d789ddc878170ed18724f7961fe85b72ee118de44d1

Observation 21757244-e032-4d60-bed0-82fe1f80280e · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

Selective Attention: Enhancing Transformer through Principled Context Control Efficient Estimation of Word Representations in Vector Space

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.745042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.745042Z digest=sha256:b4a965c273072b7304b598f2cee587f4a5b7185d7feef2a5f9f84df2f712df4d

Observation c36da27f-e4d3-43f7-87df-58dde1c8838f · outbound

This paper cites Landmark Attention: Random-Access Infinite Context Length for Transformers.

Selective Attention: Enhancing Transformer through Principled Context Control Landmark Attention: Random-Access Infinite Context Length for Transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.749713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.749713Z digest=sha256:da6a0a04597ab72092b8396c6e94293a6fd9a234018e2d6deee053115a36bc33

Observation fe810b72-7f78-4bff-8447-3fd4af5875e2 · outbound

This paper cites In-context Learning and Induction Heads.

Selective Attention: Enhancing Transformer through Principled Context Control In-context Learning and Induction Heads

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.754010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.754010Z digest=sha256:babc6aaeff244274581d42925b3f28c9ff20a60539b8e2cd55f1dce8c1611b57

Observation 80e00af2-9827-4e86-95bc-74c10e5c399e · outbound

This paper cites Fast attention over long sequences with dynamic sparse flash attention.

Selective Attention: Enhancing Transformer through Principled Context Control Fast attention over long sequences with dynamic sparse flash attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.399016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.758299Z digest=sha256:229d4d3831bc47dac4ae015aff7669e42eaf12aa1741bf8320620a1bc4cf1ebe

Observation bb298cb2-55eb-4355-93be-8d9e2befbdb9 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Selective Attention: Enhancing Transformer through Principled Context Control The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.762408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.762408Z digest=sha256:ad28e7868131dfe547e3241340f4d08af07e81fd4332736873ae0952c5e7410b

Observation a2cf47ee-e863-4c1d-a6c3-da4fb6456a47 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

Selective Attention: Enhancing Transformer through Principled Context Control YaRN: Efficient Context Window Extension of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.767257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.767257Z digest=sha256:dcb68f56e6b933a8d6a2c899c55e7e4523490b336e41e75d4c5671f9fc53b715

Observation 7c6e2e70-e19c-44f9-8b7d-b5bc80adeefa · outbound

This paper cites Language models are unsupervised multitask learners.

Selective Attention: Enhancing Transformer through Principled Context Control Language models are unsupervised multitask learners

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.384297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.772241Z digest=sha256:6c5aee9315328bb76aeca1b96056e8c151464d35956356a351dea95aacd45e76

Observation 8eb263be-f9bf-4a60-8f80-1a5e48c70134 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Selective Attention: Enhancing Transformer through Principled Context Control Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.777185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.777185Z digest=sha256:2b3925c81335b30b6cbf1be62fc4bdd393a8a6f12892bae7a431b6af58ef86c0

Observation 4b4ea709-7189-4bf2-b357-90b05c696a41 · outbound

This paper cites Sparse modular activation for efficient sequence modeling.

Selective Attention: Enhancing Transformer through Principled Context Control Sparse modular activation for efficient sequence modeling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.361081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.782068Z digest=sha256:ffe8adcdfc76f10a7ccd141cc096aee0512c057e1654ec5f1af3e2c0bad0e709

Observation 1339a8e0-be48-4ec3-a062-c9bad9871058 · outbound

This paper cites Unraveling attention via convex duality: Analysis and interpretations of vision transformers.

Selective Attention: Enhancing Transformer through Principled Context Control Unraveling attention via convex duality: Analysis and interpretations of vision transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.786327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.786327Z digest=sha256:9b66a81d336c8a4c7be23f871b62a21b11bcdfa27eabff5df2efae60dd4c98eb

Observation ac92af6a-575e-4e2e-90d6-c47f2a62a74f · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Selective Attention: Enhancing Transformer through Principled Context Control Winogrande: An adversarial winograd schema challenge at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.790450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.790450Z digest=sha256:c2a95dbcfaf5a774484219b93af5c2d054e56677569c0ad124e754e342fdc779

Observation a7ee327f-6ecd-42a9-a81f-370b9b71131f · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

Selective Attention: Enhancing Transformer through Principled Context Control Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.794883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.794883Z digest=sha256:a06afeef402ac78e59d04f4506da11ab89a2dd4593113bea0ed0bd6f113131a0

Observation 1a99ab6c-ec1d-4e8f-881c-36bf2aae7cdb · outbound

This paper cites GLU Variants Improve Transformer.

Selective Attention: Enhancing Transformer through Principled Context Control GLU Variants Improve Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.799765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.799765Z digest=sha256:9ab4c098d3331a6644fe0c83115e3b76b2fcd626797d286a9e1bfb2ecfcefcb3

Observation ebe21000-6a5f-45b6-bbfd-3dd12ffc181d · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

Selective Attention: Enhancing Transformer through Principled Context Control SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.805134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.805134Z digest=sha256:a33a2346c87f99b44494d7e422251e237058e6af58fc0bd3046bae92b20fdd04

Observation fe68b948-e9ea-4b01-bf98-0976fd8079f0 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Selective Attention: Enhancing Transformer through Principled Context Control Roformer: Enhanced transformer with rotary position embedding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.810322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.810322Z digest=sha256:83aad12da79614f41bdb4660665bd0620985e74cf811a85623e5aa4f922177ad

Observation dc4a55a4-ae75-4bcf-8855-8eacb7bd399e · outbound

This paper cites Max-margin token selection in attention mechanism.

Selective Attention: Enhancing Transformer through Principled Context Control Max-margin token selection in attention mechanism

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.320118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.815135Z digest=sha256:d2509140459453c675e8fa897311d7f7e2f60904a5ff7ab47dd542f235983006

Observation ced6b819-a208-4a8e-957f-09d6982126fd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Selective Attention: Enhancing Transformer through Principled Context Control LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.819536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.819536Z digest=sha256:e76f3c64f2c2f4c3916bd6f19b9ddaacb93a6670c5d4c793bc9502785de6338d

Observation e41068ad-ec21-4aae-bc13-10b02c123904 · outbound

This paper cites Attention is all you need.

Selective Attention: Enhancing Transformer through Principled Context Control Attention is all you need

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.824430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.824430Z digest=sha256:e0ab4b48326a2cf5d6eb8e9dbdd32c6cce0d3b18c6b4d400a705fc95f86b4ea2

Observation 7d7085b9-10ea-4a38-903b-55c162e0635b · outbound

This paper cites MambaByte: Token-free Selective State Space Model.

Selective Attention: Enhancing Transformer through Principled Context Control MambaByte: Token-free Selective State Space Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.829035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.829035Z digest=sha256:d166018ae78127a8bca2fd662a9c097aea9334491f69582d940265c966bf0707

Observation d199088d-4baa-42e3-8fd6-9d7e7d00faa0 · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

Selective Attention: Enhancing Transformer through Principled Context Control An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.833532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.833532Z digest=sha256:8682765afec5e1bb9fa09540a5274f7a437853339fec06f3b094fd72c1bade05

Observation 4260864f-eceb-4214-8328-86d4a760af9a · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Selective Attention: Enhancing Transformer through Principled Context Control Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.837721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.837721Z digest=sha256:6f3a7e4d4873fe5715d2cc4904f0fd737d7e29c9df0e6f7e96adf1b5608a29f6

Observation dd23a109-9351-4129-bc16-20ae39c55522 · outbound

This paper cites Self-attention networks can process bounded hierarchical languages.

Selective Attention: Enhancing Transformer through Principled Context Control Self-attention networks can process bounded hierarchical languages

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.296537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.841996Z digest=sha256:ffbd84cdb0d0b0bbc499b607e9c283ebc65fc87b63fb7d0c7903ba5a3728b3b6

Observation 072d5fac-31bd-43e9-8efa-494a3927a17b · outbound

This paper cites Differential Transformer.

Selective Attention: Enhancing Transformer through Principled Context Control Differential Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.846216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.846216Z digest=sha256:b9eb6696e4f652106c55032de82960695f0753d3877d26a0ff6535119c4d5394

Observation 066b80dd-0efd-44ee-8478-ff41540eb413 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Selective Attention: Enhancing Transformer through Principled Context Control HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.850896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.850896Z digest=sha256:c9a3db05598c8172b19e08ae8c99cb4c23ce37db298eaf2e048826e8228a2e78

Observation a1b27eb4-7a7c-4d55-8e76-3c194670fcde · outbound

This paper cites Class- attribute priors: Adapting optimization to heterogeneity and fairness objective.

Selective Attention: Enhancing Transformer through Principled Context Control Class- attribute priors: Adapting optimization to heterogeneity and fairness objective

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:12:33.282230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:12:32.855087Z digest=sha256:9c622739d49f3b9d9a7a6befc168b3b6f5059c141a751f91670a931c494c6add

Observation 10b48b63-1475-44e8-ab0d-99c91d95bbcf · outbound

This paper cites What Algorithms can Transformers Learn? A Study in Length Generalization.

Selective Attention: Enhancing Transformer through Principled Context Control What Algorithms can Transformers Learn? A Study in Length Generalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.859190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.859190Z digest=sha256:cc6415b91e6e0aa5e98331592a5d57cdb1abc81d8716c75128d0e478a676caa6

Observation a5a1cb71-073a-400c-a639-2688cb580694 · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

Selective Attention: Enhancing Transformer through Principled Context Control Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.864356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.864356Z digest=sha256:d8f7d21c9648303b3a0baafb0993c4a238e20864dfd1c09739cb95f4cbdf8017

Pith citing papers

No inbound Pith citation observations are available.