Pith. sign in

Paper Citation Record · LEDGER

Transformers Learn Faster with Semantic Focus

As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2506.14095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14095 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:31.193173Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact7
  • verified fuzzy42
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62771866-b7fa-4835-8861-4c85f607bed9 · outbound

This paper cites Attention is all you need.

Transformers Learn Faster with Semantic Focus Attention is all you need

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.918906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.715156Z digest=sha256:1f201ae75d0f3a6ec9f29a4eb2ff2f0f82d6635c922432605ece0f5392c71dee

Observation 22bb8517-1758-44a3-89d5-f49376359346 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a.

Transformers Learn Faster with Semantic Focus Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.902339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.776075Z digest=sha256:af8827c3714c257d91f6dd135934843b61103e2e343ccb4ea8d94333a8cd88d4

Observation 8e9223ce-1562-43b5-a3cb-18c05d591716 · outbound

This paper cites Efficient transformers: A survey.

Transformers Learn Faster with Semantic Focus Efficient transformers: A survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.865774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.865774Z digest=sha256:b1933c53bf815597388338fcb75616c6ce60638143cfd1901c791193df02f5df

Observation 1af2cb21-31bf-4ab7-b069-681e4b50bdc6 · outbound

This paper cites Long range arena: A benchmark for efficient transformers.

Transformers Learn Faster with Semantic Focus Long range arena: A benchmark for efficient transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.887632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.878995Z digest=sha256:a5d971b691470cc20412091ca927c7718a35e1d3cbb1ae298b127d5b9362d2e4

Observation 59ba69a1-cbbe-4543-97ca-d997251586a1 · outbound

This paper cites Cognitive Mechanisms Associated with Auditory Sensory Gating.

Transformers Learn Faster with Semantic Focus Cognitive Mechanisms Associated with Auditory Sensory Gating

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.873539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.884261Z digest=sha256:92273bb19b6c3b6535a59348cec79726b67e1fb5d1d2d84ff1bacb0f0e53ff32

Observation 44f6baf5-da5e-402a-83b1-a523095912f5 · outbound

This paper cites The Senses: A Comprehensive Reference.

Transformers Learn Faster with Semantic Focus The Senses: A Comprehensive Reference

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.859632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.888896Z digest=sha256:fc511dd8ccb208fbd6e7ad5207d7e43e7539c0a18e7fd5d05d4a642ff103ee79

Observation 21ace225-81de-41d0-821b-fd63afaeb096 · outbound

This paper cites Sensory gating deficits in schizophrenia: new results.

Transformers Learn Faster with Semantic Focus Sensory gating deficits in schizophrenia: new results

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:26:31.888917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.894305Z digest=sha256:b18778e0d3c9a9f0998fe345b1d663c326d0ce80f6c21ae5a2f7ab3d43dc182f

Observation f4952284-fd8e-4da6-b4c6-88f49606b297 · outbound

This paper cites The Consciousness Prior.

Transformers Learn Faster with Semantic Focus The Consciousness Prior

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.899429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.899429Z digest=sha256:81202031fb97bd76d55039f9713352baeccd7bb4028754fd261bf05b8216dbc1

Observation b2232304-32ec-42e8-a934-7a304b29b1c2 · outbound

This paper cites URL https://neurosymbolic.github.io/nsss2024.

Transformers Learn Faster with Semantic Focus URL https://neurosymbolic.github.io/nsss2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.845025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.910069Z digest=sha256:3afb6d1d29c25d6f5cd32b4ed2d5a54825cd1f29d8b7e6d70952a6c283dcb4c1

Observation 14e85bea-67d9-43a4-a785-ab7f32ae9204 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Transformers Learn Faster with Semantic Focus Neural Machine Translation by Jointly Learning to Align and Translate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.917132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.917132Z digest=sha256:355eb58d0570b5c6819b6c0823ab7e3da6ef1a915add734f9fd905fbf64e89c1

Observation a9fff3c4-0fbf-4b1b-a1f7-c3f9fc251544 · outbound

This paper cites Neural networks and the chomsky hierarchy.

Transformers Learn Faster with Semantic Focus Neural networks and the chomsky hierarchy

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.830872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.922213Z digest=sha256:c0b93141db9aff40e9e286bb61343b6a843de9f1c58f9ca1115449bf2aadf712

Observation a66a7ba5-854b-43ff-9bbe-e3ff87789bc5 · outbound

This paper cites O(n) connections are expressive enough: Universal approximability of sparse transformers.

Transformers Learn Faster with Semantic Focus O(n) connections are expressive enough: Universal approximability of sparse transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.812327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.926723Z digest=sha256:9fe7f88b8d43e7afa6ad004803f9e5c8da8c836cf8dba9f2191b10212eabd8b7

Observation 61af9724-2b54-4a43-8217-c53f99349493 · outbound

This paper cites Etc: Encoding long and structured inputs in transformers.

Transformers Learn Faster with Semantic Focus Etc: Encoding long and structured inputs in transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.794142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.931698Z digest=sha256:6b48491a5f01ef28a54c344c31f00a203e817ccd93006eeb71e5434719fa14de

Observation 84ce2eca-690d-4bb7-a2aa-9ea0dff21b3f · outbound

This paper cites Big bird: Transformers for longer sequences.

Transformers Learn Faster with Semantic Focus Big bird: Transformers for longer sequences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.778932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.936249Z digest=sha256:37036328f3490f7d9b3129103adec7f2ffb91250ae0538335b0471736ebd6469

Observation 32f6bc68-148f-4bae-a696-780ff6753dc3 · outbound

This paper cites Memory-efficient transformers via top-k attention.

Transformers Learn Faster with Semantic Focus Memory-efficient transformers via top-k attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.940576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.940576Z digest=sha256:0cf5880a5d9536cb0fa32a868c76e472b34c1172e65e4ad7b1efe10e18f0a063

Observation 2f42847e-789a-4545-ad74-dbd1c9a601cb · outbound

This paper cites ZETA : Leveraging z -order curves for efficient top- k attention.

Transformers Learn Faster with Semantic Focus ZETA : Leveraging z -order curves for efficient top- k attention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.764128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.945255Z digest=sha256:fcc9e67a905b1a14bee50deabc1e5760fb91a0c8c57102ebf9a5ce63418ab6ea

Observation 672ff6cb-90c5-4c82-99a3-f10cd42cd3a6 · outbound

This paper cites Algorithmic stability and generalization performance.

Transformers Learn Faster with Semantic Focus Algorithmic stability and generalization performance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.950382Z digest=sha256:4db8867bb91721d0e4bca77e05bd1c0ac559d7f6239cdaf05be98e4f03c16f4e

Observation 8495391a-4fb7-4739-ac0c-aec66433f718 · outbound

This paper cites Train faster, generalize better: Stability of stochastic gradient descent.

Transformers Learn Faster with Semantic Focus Train faster, generalize better: Stability of stochastic gradient descent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.954810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.954810Z digest=sha256:6ad463dd63c9709b1cc81e80ab25b86570b9ad04f030824efd90afce850af422

Observation 9bb6606d-3f55-472b-a431-4d0742abc47c · outbound

This paper cites Formal Algorithms for Transformers.

Transformers Learn Faster with Semantic Focus Formal Algorithms for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.960292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.960292Z digest=sha256:418c794b81e16786433e7de53bf8d734e1d784337f6634009e1692d891e2e47e

Observation dd34a3bd-096a-4936-9e29-698e87b2df1d · outbound

This paper cites A survey of transformers.

Transformers Learn Faster with Semantic Focus A survey of transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.736787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.965609Z digest=sha256:94c5e0af157929ac83b013b69e4e5b5e84c4a3aae70040e46ffd8f979241a519

Observation c1bfa5c9-f0c9-4ec2-9862-5b35f64ee4ca · outbound

This paper cites Image transformer.

Transformers Learn Faster with Semantic Focus Image transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.723660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.969645Z digest=sha256:7916d9f0ee0c44fac5b489668803cf52b317863b853b0f855e9d15d9ca33d201

Observation 79b46b3c-40a9-4820-be5e-130bc2e701cb · outbound

This paper cites Blockwise self-attention for long document understanding.

Transformers Learn Faster with Semantic Focus Blockwise self-attention for long document understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.709229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.974200Z digest=sha256:720a1280b63e0b4a4102249acfb4628f24d7096dd6d519f8620b0e66628ab033

Observation f63968ca-efc8-419d-878a-ddb05d57f944 · outbound

This paper cites Longformer: The Long-Document Transformer.

Transformers Learn Faster with Semantic Focus Longformer: The Long-Document Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.979069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.979069Z digest=sha256:f9bba74a7f8d430d9e7ed2c12c4c0d30eae6296ebaea9211b5e57c0d18957057

Observation 067581ba-98c3-4144-b6bb-576ff06f2877 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Transformers Learn Faster with Semantic Focus Generating Long Sequences with Sparse Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.983693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.983693Z digest=sha256:d7638081a37d43a09fac70fb42513c02b39f4d1083cb191f78df33a63caefc8a

Observation de423e8c-14cf-431b-bb13-0d340ddd288b · outbound

This paper cites Sparse sinkhorn attention.

Transformers Learn Faster with Semantic Focus Sparse sinkhorn attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.693446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.988852Z digest=sha256:7d6d85e2a92c564060fc04119c26e426f0e02917bde4a65a93c41c0cf3c1ed0f

Observation 8cee7eee-810c-4efa-84d8-49e5f85f3321 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Transformers Learn Faster with Semantic Focus Efficient content-based sparse attention with routing transformers

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T00:26:31.355960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.993269Z digest=sha256:befbdd0063fcc341e78b62317e8ac7a638b48f347dda954629a6ab121f320708

Observation 2e7434e7-0ad7-45cb-b59f-e14f2167eae9 · outbound

This paper cites Reformer: The efficient transformer.

Transformers Learn Faster with Semantic Focus Reformer: The efficient transformer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.678445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.998380Z digest=sha256:a84d677026a8a976d35b7a16ca68d057c95e4c718f6838174eb51bde14525f13

Observation d84cf548-c9c4-4cf2-8365-00a069fe0a17 · outbound

This paper cites COGS : A compositional generalization challenge based on semantic interpretation.

Transformers Learn Faster with Semantic Focus COGS : A compositional generalization challenge based on semantic interpretation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.664182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.003042Z digest=sha256:bd682bc66731da5339282b77befe9a3a80b848ce75572cd80dd9c8e8fe19802d

Observation cf06c895-63e9-4ecd-978e-e546c4c82dde · outbound

This paper cites Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks.

Transformers Learn Faster with Semantic Focus Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.649685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.007490Z digest=sha256:bfc0b3ce0848ea59fa324a93125fa8078c972f09daab2b25ca94884392ef354e

Observation 47980f98-7844-48dd-9fd5-0af02175dc94 · outbound

This paper cites When can transformers ground and compose: Insights from compositional generalization benchmarks.

Transformers Learn Faster with Semantic Focus When can transformers ground and compose: Insights from compositional generalization benchmarks

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.339636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.011682Z digest=sha256:9e48142aab0adf54f6719859fdfe08e0b34b6fbd31b2b3e2ea15e05eac820549

Observation 3ceab004-66f3-41e6-84d4-824a19b41134 · outbound

This paper cites The devil is in the detail: Simple tricks improve systematic generalization of transformers.

Transformers Learn Faster with Semantic Focus The devil is in the detail: Simple tricks improve systematic generalization of transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.016326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.016326Z digest=sha256:63cf95949581ba84b17c41f7671e2bb729b1a6e0ed8c653b9bd771ab3b34d5c5

Observation a13d6515-33e0-4e1d-91c8-7e541b43e13d · outbound

This paper cites Making transformers solve compositional tasks.

Transformers Learn Faster with Semantic Focus Making transformers solve compositional tasks

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.312207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.020431Z digest=sha256:a0f302d7429b3eb317166386bf1262817d54a3b9d7038f55337849161be6c816

Observation 2496269a-85e9-44eb-a5db-d03f7139143a · outbound

This paper cites Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks.

Transformers Learn Faster with Semantic Focus Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.295252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.025449Z digest=sha256:f81f40a9bd3a2635e117e1ab64ec6441bf6f8854d52a3ad10b641aa84087c8fc

Observation 84412c7a-508f-461c-a335-db1b31394c91 · outbound

This paper cites a rli, Ekin Aky \.

Transformers Learn Faster with Semantic Focus a rli, Ekin Aky \

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.635646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.029829Z digest=sha256:cc82db4f70f4d6f19bf9ba1291320363018eca835c17bf79c5b873797b3cf21c

Observation f5b2bfd4-f7c9-4872-a2ad-f00fdeacdc96 · outbound

This paper cites What formal languages can transformers express? a survey.

Transformers Learn Faster with Semantic Focus What formal languages can transformers express? a survey

Reference 35

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.279003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.034312Z digest=sha256:208079f57c71d22065a16f0bcdbf10b7db0a490d18792a2ed00e2036be793f8d

Observation 6020b2e6-aa1c-47ce-8d90-8afc314dd586 · outbound

This paper cites On the ability and limitations of transformers to recognize formal languages.

Transformers Learn Faster with Semantic Focus On the ability and limitations of transformers to recognize formal languages

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.620841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.038652Z digest=sha256:c100fb2777760aec1737f16e057182adb902367d60e4980ac28d8289b87d025d

Observation 1604f60a-e988-4af0-8874-2c4f77626b61 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Transformers Learn Faster with Semantic Focus Theoretical limitations of self-attention in neural sequence models

Reference 37

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.263821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.042758Z digest=sha256:4f428657d7fc7c4e75ac7cec361962f002ff346c9220fab49902ac94ad8b4c96

Observation 2ece83c2-9f63-418c-b1d5-c8ae82a60982 · outbound

This paper cites Formal language recognition by hard attention transformers: Perspectives from circuit complexity.

Transformers Learn Faster with Semantic Focus Formal language recognition by hard attention transformers: Perspectives from circuit complexity

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.605490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.046942Z digest=sha256:64682a558f3532eb72876d0f2f39b63ec3e969cf0c5615724a1cb01e56993cc2

Observation bb1ab5b5-9b8c-400a-b686-8b77247e18c6 · outbound

This paper cites Saturated transformers are constant-depth threshold circuits.

Transformers Learn Faster with Semantic Focus Saturated transformers are constant-depth threshold circuits

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.589814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.051326Z digest=sha256:b97e7cb15d071eea9069eca1d08e5ff92f72995dfb8190195df22c49f516e1fc

Observation f3fe64ff-55f2-46e6-80d0-57bf72a7230a · outbound

This paper cites Overcoming a theoretical limitation of self-attention.

Transformers Learn Faster with Semantic Focus Overcoming a theoretical limitation of self-attention

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.573709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.055666Z digest=sha256:e74d137785f7ae7544f74c5b0d2a1e3b534d4b5bf3fd8e0bec1886c5e9c019dc

Observation 5f80651a-9042-41df-8507-c06510ced51d · outbound

This paper cites Tighter bounds on the expressivity of transformer encoders.

Transformers Learn Faster with Semantic Focus Tighter bounds on the expressivity of transformer encoders

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.557973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.059942Z digest=sha256:9bd1818b23775deb02c27151d525b0f7ade90b95bc65cbec3d0fdc46835c4192

Observation e1a1abc0-80ea-4abc-a8b7-530e386ed594 · outbound

This paper cites Transformers as algorithms: Generalization and stability in in-context learning.

Transformers Learn Faster with Semantic Focus Transformers as algorithms: Generalization and stability in in-context learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.543429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.064187Z digest=sha256:d983178bbc6a16c370083284a2557659a3bbacd616abc190cbb7a7da6a790988

Observation af5cac0e-1cef-4924-a631-3d6f7a36699f · outbound

This paper cites Transformers learn in-context by gradient descent.

Transformers Learn Faster with Semantic Focus Transformers learn in-context by gradient descent

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.529452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.068351Z digest=sha256:a29229a82f9292d0de93579517e0cc4682d4abfba41f3dbc0247029907f409a1

Observation bc5a2ec6-5e4e-4209-8bbe-586adb9acbe7 · outbound

This paper cites The emergence of clusters in self-attention dynamics.

Transformers Learn Faster with Semantic Focus The emergence of clusters in self-attention dynamics

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.514809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.072620Z digest=sha256:e816de5d64bff83d4dc088688b9a2d5c831d952b8ad2d291f8f1c665b2142321

Observation 4842f956-d2c0-4cd2-bf07-6a05324354e0 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Transformers Learn Faster with Semantic Focus Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.495826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.077091Z digest=sha256:934eb91fd19a0d1bd3802110d888275315e6181a864043467805a32b706fd993

Observation 0676e91a-49e1-44af-9ffe-d180a49e2d1e · outbound

This paper cites Trained transformers learn linear models in-context.

Transformers Learn Faster with Semantic Focus Trained transformers learn linear models in-context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.480464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.082054Z digest=sha256:72e560d47e1cf34135bd03b8053e506a043db2a19cbdb324e3ffe7d9f1dd0ffb

Observation f1208b8e-f112-4011-bd87-dd3159dc5249 · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020.

Transformers Learn Faster with Semantic Focus Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.464242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.086492Z digest=sha256:cbea46ce05e7a211f9a19ba94bba7d207691fad94946e597256119e790ca67ed

Observation da9fc738-70f8-48b4-8259-7545c05251a4 · outbound

This paper cites Toward understanding why adam converges faster than SGD for transformers.

Transformers Learn Faster with Semantic Focus Toward understanding why adam converges faster than SGD for transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.446017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.091192Z digest=sha256:8cfeb18350dd582976d6c143a75aaf567aa1109bde956d08bf105085659da10d

Observation c9f67186-e199-4d0c-9ad6-15fa02f759f1 · outbound

This paper cites How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023.

Transformers Learn Faster with Semantic Focus How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.431099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.095922Z digest=sha256:fefc923e25d471631825e8c8b00061aecec97464f3bd8a046cd4249eb1fa4935

Observation dbc1b7cc-1d95-47c1-85fa-4b50bf0a234a · outbound

This paper cites Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be.

Transformers Learn Faster with Semantic Focus Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.414833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.101370Z digest=sha256:9a9d6f70d2348c0fcf154bec215de9c74f2a94075d8832033f03f791dc283a61

Observation 884fe00e-92a9-460e-848e-cb3bd47f188e · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Transformers Learn Faster with Semantic Focus Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.399313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.105944Z digest=sha256:73b44e147d9bb6a9f1484a441efb75b3b96ce07f33a4f5cdb00c2e245ac2bc9e

Observation 0b632e26-88f9-40fd-b30c-2f0591d3f489 · outbound

This paper cites On the optimization and generalization of two-layer transformers with sign gradient descent.

Transformers Learn Faster with Semantic Focus On the optimization and generalization of two-layer transformers with sign gradient descent

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.383165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.110777Z digest=sha256:791d594c71519d7e135327936c43cd62e420938778562d14552a4f5a7c46d4c5

Observation 7d1005fb-2019-4208-b0c1-d3163e3f2a97 · outbound

This paper cites Layer Normalization.

Transformers Learn Faster with Semantic Focus Layer Normalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.114796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.114796Z digest=sha256:1a6e38cf88f0ef00beebeaf1b52113f913be3ce982f6ce8b396238a2fbd5392b

Observation 46115e48-be86-45ce-b49a-e9ce4bdda4f9 · outbound

This paper cites Root mean square layer normalization.

Transformers Learn Faster with Semantic Focus Root mean square layer normalization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.119039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.119039Z digest=sha256:630f85de87300e9cf8154073d455ff8d24d7e56b238d5f89b5403bd77167bf22

Observation ee4e4e51-cd62-4925-9254-1f342c52ca72 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Transformers Learn Faster with Semantic Focus BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.123158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.123158Z digest=sha256:a954b1a746716bc64eeb4a9a82345518bb90e338b714a2cff2c766025b5c8bf4

Observation d1ad7295-411a-42bf-9db6-42b3ff685421 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Transformers Learn Faster with Semantic Focus Gaussian Error Linear Units (GELUs)

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.128361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.128361Z digest=sha256:28c0871aac86aad8d8e8897d0ee3d5672e1aa9c340021aca89b6e32f303d0975

Observation d521b3a9-96a1-4475-9a3a-d3942c18b279 · outbound

This paper cites Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs).

Transformers Learn Faster with Semantic Focus Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.133288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.133288Z digest=sha256:2a2d78a4961f4b7c061e4c237d7b89b87fbe1c1b823362e0a1598968ec3dc4d1

Observation 5fa87b52-75d6-4f21-b69c-f76cc01c4f8e · outbound

This paper cites Listops: A diagnostic dataset for latent tree learning.

Transformers Learn Faster with Semantic Focus Listops: A diagnostic dataset for latent tree learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.356077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.141387Z digest=sha256:75480321b21a818e43f2ef3b46a4646b586c588b1f5bfa402f3d58de2f6496ab

Observation bfb46c1d-85a6-4889-b1b2-37873f04faca · outbound

This paper cites Mish: A Self Regularized Non-Monotonic Activation Function.

Transformers Learn Faster with Semantic Focus Mish: A Self Regularized Non-Monotonic Activation Function

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.146060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.146060Z digest=sha256:2d45745f7cc8905c8ffb1a1675dd0c88d4860b2e0e74882808ad8a416e6510d8

Observation 7a944d61-2998-44b7-a34b-676ba1b45bf1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Transformers Learn Faster with Semantic Focus Adam: A Method for Stochastic Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.152690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.152690Z digest=sha256:653210d29ae5a0bcc426a722fa8b67fbf46646651209293318c8392b352ab5bc

Observation 1c0085ca-f463-4103-952b-307c15eb115e · outbound

This paper cites The Lipschitz Constant of Self-Attention.

Transformers Learn Faster with Semantic Focus The Lipschitz Constant of Self-Attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.158224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.158224Z digest=sha256:3c05c8d5c77ccab4ec69755b9ad534dff39fcf2721e39b8df16d603a3a8f0b36

Observation 12797a98-aa4b-437c-b189-96d7d2b884b1 · outbound

This paper cites Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection.

Transformers Learn Faster with Semantic Focus Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.162921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.162921Z digest=sha256:903291824d26573a5287109514caee7c3fb8eed96f38d5369456f5d768bdfbaa

Observation 37334c11-9643-4161-800f-94b4c7a6bbe5 · outbound

This paper cites Visualizing the loss landscape of neural nets.

Transformers Learn Faster with Semantic Focus Visualizing the loss landscape of neural nets

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.275494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.168150Z digest=sha256:a103720d9a21f9d775b7c2f6e35de062d7926345340d93b8cef847e9b166a6a8

Observation 10878189-944f-4b9b-8f6c-9907c722def7 · outbound

This paper cites Never train from scratch: Fair comparison of long-sequence models requires data-driven priors.

Transformers Learn Faster with Semantic Focus Never train from scratch: Fair comparison of long-sequence models requires data-driven priors

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.980495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.172682Z digest=sha256:fdcb329323231a1d29fdec2efa6238438dc9ef06c50570876c5b00faecc23426

Observation c25273f1-6b34-4339-bc1d-998ffb03b3da · outbound

This paper cites Learning overparameterized neural networks via stochastic gradient descent on structured data.

Transformers Learn Faster with Semantic Focus Learning overparameterized neural networks via stochastic gradient descent on structured data

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.851131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.177991Z digest=sha256:84835189f451a92cf1faf19d196e49a2c8e151d1c3bb2db2d4ad832b9c7a055a

Observation 233ed32b-ea5f-4522-944c-d2bb9ef0050c · outbound

This paper cites A convergence theory for deep learning via over-parameterization.

Transformers Learn Faster with Semantic Focus A convergence theory for deep learning via over-parameterization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.582899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.182659Z digest=sha256:9a1554a81272752694085360eafa15c15a1744cc7c56417ce48bf7c591bafd02

Observation ca681d40-ef8b-457c-962c-fffe1a113051 · outbound

This paper cites Gradient descent optimizes over-parameterized deep relu networks.

Transformers Learn Faster with Semantic Focus Gradient descent optimizes over-parameterized deep relu networks

Reference 67

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.234860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.187525Z digest=sha256:e1efc0880401ae75236bb0089ea0cfb764e3139563fe27dcadce25943e415ec1

Observation aa3a87b7-a9a4-47f5-a37e-2fd076f33f5b · outbound

This paper cites Convergence rates for the stochastic gradient descent method for non-convex objective functions.

Transformers Learn Faster with Semantic Focus Convergence rates for the stochastic gradient descent method for non-convex objective functions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.280375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.193173Z digest=sha256:a422ae26232a96022f8ef3537910d2594f8d8564b35e8ac1a9f8aeab948204ed

Pith citing papers

No inbound Pith citation observations are available.