Pith. sign in

Paper Citation Record · LEDGER

Transformers Learn Faster with Semantic Focus

As of 20 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2506.14095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14095 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:31.193173Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact7
  • verified fuzzy42
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62771866-b7fa-4835-8861-4c85f607bed9 · outbound

This paper cites Attention is all you need.

Transformers Learn Faster with Semantic Focus Attention is all you need

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.918906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.715156Z digest=sha256:3655ebe629d3e61ad1dca4e0c3fcdeef29122d4a17fc8786d7da837d718081a2

Observation 22bb8517-1758-44a3-89d5-f49376359346 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a.

Transformers Learn Faster with Semantic Focus Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.902339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.776075Z digest=sha256:5144d2955f0e0857a363fd4b0b6711e939e71b8552a4bcedca354112f37c162d

Observation 8e9223ce-1562-43b5-a3cb-18c05d591716 · outbound

This paper cites Efficient transformers: A survey.

Transformers Learn Faster with Semantic Focus Efficient transformers: A survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.865774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.865774Z digest=sha256:e8b9605f2750705e692c935192dfcc62df006c6f4ec07a1c0f1e290fb6eb6572

Observation 1af2cb21-31bf-4ab7-b069-681e4b50bdc6 · outbound

This paper cites Long range arena: A benchmark for efficient transformers.

Transformers Learn Faster with Semantic Focus Long range arena: A benchmark for efficient transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.887632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.878995Z digest=sha256:e063758819db7ad03c5ab19896b452a02babdc9b237004881b980f413b2bf3ff

Observation 59ba69a1-cbbe-4543-97ca-d997251586a1 · outbound

This paper cites Cognitive Mechanisms Associated with Auditory Sensory Gating.

Transformers Learn Faster with Semantic Focus Cognitive Mechanisms Associated with Auditory Sensory Gating

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.873539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.884261Z digest=sha256:05d7a864af96f02ef7e94762cfcc88c0a37896be7c9f80f0e931cd012594a213

Observation 44f6baf5-da5e-402a-83b1-a523095912f5 · outbound

This paper cites The Senses: A Comprehensive Reference.

Transformers Learn Faster with Semantic Focus The Senses: A Comprehensive Reference

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.859632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.888896Z digest=sha256:70c96546dfbcbc35d7527230af529688c91ab14d911974fa6f77048a8d314939

Observation 21ace225-81de-41d0-821b-fd63afaeb096 · outbound

This paper cites Sensory gating deficits in schizophrenia: new results.

Transformers Learn Faster with Semantic Focus Sensory gating deficits in schizophrenia: new results

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:26:31.888917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.894305Z digest=sha256:6eb535b8bf36a4daad6c59bef8ed78d134e6d9f97c65b522da4f4a39ba6cb410

Observation f4952284-fd8e-4da6-b4c6-88f49606b297 · outbound

This paper cites The Consciousness Prior.

Transformers Learn Faster with Semantic Focus The Consciousness Prior

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.899429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.899429Z digest=sha256:7da063bc97fb6d51c9955827ac9b8aeca1ed7309409879e530956446f42214b5

Observation b2232304-32ec-42e8-a934-7a304b29b1c2 · outbound

This paper cites URL https://neurosymbolic.github.io/nsss2024.

Transformers Learn Faster with Semantic Focus URL https://neurosymbolic.github.io/nsss2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.845025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.910069Z digest=sha256:34114684a9c11d5af83af7fa3429a5d48c74e4c98045c32db96eeb7325899798

Observation 14e85bea-67d9-43a4-a785-ab7f32ae9204 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Transformers Learn Faster with Semantic Focus Neural Machine Translation by Jointly Learning to Align and Translate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.917132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.917132Z digest=sha256:82568ce701ac869da6897b086196747eb2592fe8ec9f51dc8cc0e3608c59c65c

Observation a9fff3c4-0fbf-4b1b-a1f7-c3f9fc251544 · outbound

This paper cites Neural networks and the chomsky hierarchy.

Transformers Learn Faster with Semantic Focus Neural networks and the chomsky hierarchy

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.830872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.922213Z digest=sha256:0d64bbb4c19b3ac714d95688f1515c350bbb9741a179ce3cc1497b1691db7d29

Observation a66a7ba5-854b-43ff-9bbe-e3ff87789bc5 · outbound

This paper cites O(n) connections are expressive enough: Universal approximability of sparse transformers.

Transformers Learn Faster with Semantic Focus O(n) connections are expressive enough: Universal approximability of sparse transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.812327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.926723Z digest=sha256:7e38f9fca0088665aacd90bbf5b79abce4d899bf1e1a3e5cd05ea9715373ae40

Observation 61af9724-2b54-4a43-8217-c53f99349493 · outbound

This paper cites Etc: Encoding long and structured inputs in transformers.

Transformers Learn Faster with Semantic Focus Etc: Encoding long and structured inputs in transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.794142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.931698Z digest=sha256:20a9fdf0cf2d8bc0fe1e835e9c831bb3240805c2daad85fde40f518f93ddac96

Observation 84ce2eca-690d-4bb7-a2aa-9ea0dff21b3f · outbound

This paper cites Big bird: Transformers for longer sequences.

Transformers Learn Faster with Semantic Focus Big bird: Transformers for longer sequences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.778932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.936249Z digest=sha256:74fe7c0608d714d82cc3c16f0d56b73676bb0191f32bdcda8354b9004b867cfa

Observation 32f6bc68-148f-4bae-a696-780ff6753dc3 · outbound

This paper cites Memory-efficient transformers via top-k attention.

Transformers Learn Faster with Semantic Focus Memory-efficient transformers via top-k attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.940576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.940576Z digest=sha256:97b15975eafedefee7af34464c9374d26cbb23ec78cefbbdddc2e7a4f0b5552e

Observation 2f42847e-789a-4545-ad74-dbd1c9a601cb · outbound

This paper cites ZETA : Leveraging z -order curves for efficient top- k attention.

Transformers Learn Faster with Semantic Focus ZETA : Leveraging z -order curves for efficient top- k attention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.764128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.945255Z digest=sha256:5ba30e102852c9a4d2da17f09bee0126b6fd74bc83d1adbdf80158c10f5443cb

Observation 672ff6cb-90c5-4c82-99a3-f10cd42cd3a6 · outbound

This paper cites Algorithmic stability and generalization performance.

Transformers Learn Faster with Semantic Focus Algorithmic stability and generalization performance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.950382Z digest=sha256:3d4a1be881dd8a072cade883bfa932f1fdaa324ad2cd80ad62b4fe328449137f

Observation 8495391a-4fb7-4739-ac0c-aec66433f718 · outbound

This paper cites Train faster, generalize better: Stability of stochastic gradient descent.

Transformers Learn Faster with Semantic Focus Train faster, generalize better: Stability of stochastic gradient descent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.954810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.954810Z digest=sha256:8f8b5bc453dd116c182e8e930bb19d8e4df5515ce49705c9f1a96ab8b52aaef3

Observation 9bb6606d-3f55-472b-a431-4d0742abc47c · outbound

This paper cites Formal Algorithms for Transformers.

Transformers Learn Faster with Semantic Focus Formal Algorithms for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.960292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.960292Z digest=sha256:f3511818599538fefc0a48245e31f3bf95a880f206a05b3798d4a3858d62e335

Observation dd34a3bd-096a-4936-9e29-698e87b2df1d · outbound

This paper cites A survey of transformers.

Transformers Learn Faster with Semantic Focus A survey of transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.736787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.965609Z digest=sha256:0ac64af5b45047ffcb0679d01833a2c396567152dd7c97fdf61d9e71bbbef1c6

Observation c1bfa5c9-f0c9-4ec2-9862-5b35f64ee4ca · outbound

This paper cites Image transformer.

Transformers Learn Faster with Semantic Focus Image transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.723660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.969645Z digest=sha256:856af841dc1154010c143ea3cf1db5fd0d4e0a062e5825dd58b54bda961276f0

Observation 79b46b3c-40a9-4820-be5e-130bc2e701cb · outbound

This paper cites Blockwise self-attention for long document understanding.

Transformers Learn Faster with Semantic Focus Blockwise self-attention for long document understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.709229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.974200Z digest=sha256:c2f4d32df2ceae031868361f9f8f2c23c93b2421081794f40a10ded9ceb8a05a

Observation f63968ca-efc8-419d-878a-ddb05d57f944 · outbound

This paper cites Longformer: The Long-Document Transformer.

Transformers Learn Faster with Semantic Focus Longformer: The Long-Document Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.979069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.979069Z digest=sha256:3ef6f3079ce3bc19061710b8578faeb3b76e2e80b70074ec0238737dda201afc

Observation 067581ba-98c3-4144-b6bb-576ff06f2877 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Transformers Learn Faster with Semantic Focus Generating Long Sequences with Sparse Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.983693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.983693Z digest=sha256:751e5d6c4469a4a50141fa98941dbcf8f535fc4616a9bd8c9c7fe62847eb32e3

Observation de423e8c-14cf-431b-bb13-0d340ddd288b · outbound

This paper cites Sparse sinkhorn attention.

Transformers Learn Faster with Semantic Focus Sparse sinkhorn attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.693446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.988852Z digest=sha256:d3028cd0599a28e2d72108b935239c930bcc5c79838943c4e45f2202f0e650e8

Observation 8cee7eee-810c-4efa-84d8-49e5f85f3321 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Transformers Learn Faster with Semantic Focus Efficient content-based sparse attention with routing transformers

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T00:26:31.355960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.993269Z digest=sha256:490c7982e2d6ec5d16991e457b4140a68dd08422cdf53cc6223600ad76b3b65e

Observation 2e7434e7-0ad7-45cb-b59f-e14f2167eae9 · outbound

This paper cites Reformer: The efficient transformer.

Transformers Learn Faster with Semantic Focus Reformer: The efficient transformer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.678445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.998380Z digest=sha256:a7c558b4984c8426f230d9a35609f811c6e088d579bef1ee20be694d006a9afc

Observation d84cf548-c9c4-4cf2-8365-00a069fe0a17 · outbound

This paper cites COGS : A compositional generalization challenge based on semantic interpretation.

Transformers Learn Faster with Semantic Focus COGS : A compositional generalization challenge based on semantic interpretation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.664182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.003042Z digest=sha256:8999091eff68813be7696f353d7206eda9305740d6d9a57e91833dd62380da02

Observation cf06c895-63e9-4ecd-978e-e546c4c82dde · outbound

This paper cites Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks.

Transformers Learn Faster with Semantic Focus Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.649685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.007490Z digest=sha256:560edf7486915b6e2a0b34637470bb172f69c9196cc4cfcefd0217607bdb5d8e

Observation 47980f98-7844-48dd-9fd5-0af02175dc94 · outbound

This paper cites When can transformers ground and compose: Insights from compositional generalization benchmarks.

Transformers Learn Faster with Semantic Focus When can transformers ground and compose: Insights from compositional generalization benchmarks

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.339636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.011682Z digest=sha256:884ad77ac972880e3f8486c9a91766406781b83c51db1506700272598abe990a

Observation 3ceab004-66f3-41e6-84d4-824a19b41134 · outbound

This paper cites The devil is in the detail: Simple tricks improve systematic generalization of transformers.

Transformers Learn Faster with Semantic Focus The devil is in the detail: Simple tricks improve systematic generalization of transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.016326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.016326Z digest=sha256:81674ccba0b982d077b79181de693ccd2d79b5b291b1222187f8a9d73ed8d7fa

Observation a13d6515-33e0-4e1d-91c8-7e541b43e13d · outbound

This paper cites Making transformers solve compositional tasks.

Transformers Learn Faster with Semantic Focus Making transformers solve compositional tasks

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.312207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.020431Z digest=sha256:400f63e199d4053a665e0b8578093a267db2a6b25683e87b684157d8ff7a0e89

Observation 2496269a-85e9-44eb-a5db-d03f7139143a · outbound

This paper cites Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks.

Transformers Learn Faster with Semantic Focus Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.295252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.025449Z digest=sha256:83fa432999c791f7cb1c60d24cf9f91453e1e54396b401d21c05b01520e978ce

Observation 84412c7a-508f-461c-a335-db1b31394c91 · outbound

This paper cites a rli, Ekin Aky \.

Transformers Learn Faster with Semantic Focus a rli, Ekin Aky \

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.635646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.029829Z digest=sha256:4696113df663eefa1ad680631a776bc8d7ce8c2da672be00a8a63f48868af41f

Observation f5b2bfd4-f7c9-4872-a2ad-f00fdeacdc96 · outbound

This paper cites What formal languages can transformers express? a survey.

Transformers Learn Faster with Semantic Focus What formal languages can transformers express? a survey

Reference 35

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.279003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.034312Z digest=sha256:d708062e352c924542276348a6cffd5049e371ac12ff33c8ccc9ca77c5404a71

Observation 6020b2e6-aa1c-47ce-8d90-8afc314dd586 · outbound

This paper cites On the ability and limitations of transformers to recognize formal languages.

Transformers Learn Faster with Semantic Focus On the ability and limitations of transformers to recognize formal languages

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.620841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.038652Z digest=sha256:9b1437cfcf9f9c97dc9690b64d23c5dc8005ef751341584d719abec78c786492

Observation 1604f60a-e988-4af0-8874-2c4f77626b61 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Transformers Learn Faster with Semantic Focus Theoretical limitations of self-attention in neural sequence models

Reference 37

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.263821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.042758Z digest=sha256:c924dcf63eaa20d670099995c0ad64719d8914e327d251b5d98ffe01ea0e194a

Observation 2ece83c2-9f63-418c-b1d5-c8ae82a60982 · outbound

This paper cites Formal language recognition by hard attention transformers: Perspectives from circuit complexity.

Transformers Learn Faster with Semantic Focus Formal language recognition by hard attention transformers: Perspectives from circuit complexity

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.605490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.046942Z digest=sha256:c6a3470fab8241d98c7da9ec0accfec51d484fb32deae4adb99490cf8008dddc

Observation bb1ab5b5-9b8c-400a-b686-8b77247e18c6 · outbound

This paper cites Saturated transformers are constant-depth threshold circuits.

Transformers Learn Faster with Semantic Focus Saturated transformers are constant-depth threshold circuits

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.589814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.051326Z digest=sha256:aea0dc51e0e613505bc3f5d2f6d4db8884decdc55690691dfb1461cf03decccf

Observation f3fe64ff-55f2-46e6-80d0-57bf72a7230a · outbound

This paper cites Overcoming a theoretical limitation of self-attention.

Transformers Learn Faster with Semantic Focus Overcoming a theoretical limitation of self-attention

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.573709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.055666Z digest=sha256:3bc51c275bcf543f0c2a34bdf6b3dd70589bd88c4600ead8ea89b5f2011cf184

Observation 5f80651a-9042-41df-8507-c06510ced51d · outbound

This paper cites Tighter bounds on the expressivity of transformer encoders.

Transformers Learn Faster with Semantic Focus Tighter bounds on the expressivity of transformer encoders

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.557973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.059942Z digest=sha256:2ed2e4ec35e4663c2a5d804d96520057175a1171153b928a0b6a9a959c4e4fbc

Observation e1a1abc0-80ea-4abc-a8b7-530e386ed594 · outbound

This paper cites Transformers as algorithms: Generalization and stability in in-context learning.

Transformers Learn Faster with Semantic Focus Transformers as algorithms: Generalization and stability in in-context learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.543429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.064187Z digest=sha256:0a19cec6340b69ebe15a6573107d2197db392224d8b3f712b17a5d9edab92255

Observation af5cac0e-1cef-4924-a631-3d6f7a36699f · outbound

This paper cites Transformers learn in-context by gradient descent.

Transformers Learn Faster with Semantic Focus Transformers learn in-context by gradient descent

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.529452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.068351Z digest=sha256:88333a7020829958212a1af7c9809fc87154b5caa3cd14bda630be44ef1b1267

Observation bc5a2ec6-5e4e-4209-8bbe-586adb9acbe7 · outbound

This paper cites The emergence of clusters in self-attention dynamics.

Transformers Learn Faster with Semantic Focus The emergence of clusters in self-attention dynamics

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.514809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.072620Z digest=sha256:a965eb119b2921f004d11f91465ab4ef47e99367865619c37edd79dd18011c09

Observation 4842f956-d2c0-4cd2-bf07-6a05324354e0 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Transformers Learn Faster with Semantic Focus Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.495826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.077091Z digest=sha256:050a0646df69bcd18e8e05deb5c64f1f98d784d1e4d94ed5cd07e341fdb75c57

Observation 0676e91a-49e1-44af-9ffe-d180a49e2d1e · outbound

This paper cites Trained transformers learn linear models in-context.

Transformers Learn Faster with Semantic Focus Trained transformers learn linear models in-context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.480464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.082054Z digest=sha256:9e24938e7f1dee9709e004475b7d24f28d4cb41420bd32339f794e143604e403

Observation f1208b8e-f112-4011-bd87-dd3159dc5249 · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020.

Transformers Learn Faster with Semantic Focus Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.464242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.086492Z digest=sha256:79dbe9b54af869322e040b681ed9f27c26a9bf3a964975baa131cf1c7e8c8523

Observation da9fc738-70f8-48b4-8259-7545c05251a4 · outbound

This paper cites Toward understanding why adam converges faster than SGD for transformers.

Transformers Learn Faster with Semantic Focus Toward understanding why adam converges faster than SGD for transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.446017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.091192Z digest=sha256:5234f2a9134dd5070ba212943df974c6968aa3bf9500c6f589d369681e6124b4

Observation c9f67186-e199-4d0c-9ad6-15fa02f759f1 · outbound

This paper cites How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023.

Transformers Learn Faster with Semantic Focus How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.431099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.095922Z digest=sha256:57e0a4fe95003be8e403d21a1c15ebc71f5c8478d03c3cb9b24b1d1db34f06df

Observation dbc1b7cc-1d95-47c1-85fa-4b50bf0a234a · outbound

This paper cites Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be.

Transformers Learn Faster with Semantic Focus Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.414833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.101370Z digest=sha256:6e9d33ed3c1dae29f2ce936d270438ae0b13e5d9a9af70f96b1982a983a476ea

Observation 884fe00e-92a9-460e-848e-cb3bd47f188e · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Transformers Learn Faster with Semantic Focus Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.399313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.105944Z digest=sha256:569e0073745a0ef7a35fa9553f78c79b61662267c384fca02163bac9db78dc5d

Observation 0b632e26-88f9-40fd-b30c-2f0591d3f489 · outbound

This paper cites On the optimization and generalization of two-layer transformers with sign gradient descent.

Transformers Learn Faster with Semantic Focus On the optimization and generalization of two-layer transformers with sign gradient descent

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.383165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.110777Z digest=sha256:1981bcdd8cec2bb225e1855b2240c242d971060bc1981e9e6a7eec68ba67cc9a

Observation 7d1005fb-2019-4208-b0c1-d3163e3f2a97 · outbound

This paper cites Layer Normalization.

Transformers Learn Faster with Semantic Focus Layer Normalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.114796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.114796Z digest=sha256:26458685d89f3a487d990b23f4dfa3c77922e59302a79f1e7ad260b20a1a2e44

Observation 46115e48-be86-45ce-b49a-e9ce4bdda4f9 · outbound

This paper cites Root mean square layer normalization.

Transformers Learn Faster with Semantic Focus Root mean square layer normalization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.119039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.119039Z digest=sha256:6b7971fc4568371ec2068cf8c024e7cd70ed2b820ee6bc6896ecf801d82c97ff

Observation ee4e4e51-cd62-4925-9254-1f342c52ca72 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Transformers Learn Faster with Semantic Focus BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.123158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.123158Z digest=sha256:53b2c825c65ffba16f37221483941219f5d057a5dd1e742f461cb356fd049bc9

Observation d1ad7295-411a-42bf-9db6-42b3ff685421 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Transformers Learn Faster with Semantic Focus Gaussian Error Linear Units (GELUs)

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.128361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.128361Z digest=sha256:52d9ae33f4a4f35ad7a64b70af7ce806eebe72f6b2110bdb11e2e2864f4d38f5

Observation d521b3a9-96a1-4475-9a3a-d3942c18b279 · outbound

This paper cites Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs).

Transformers Learn Faster with Semantic Focus Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.133288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.133288Z digest=sha256:899f80aeff9055f40716dacb0b6204eeb70466286a6d0704fca9bdffc5851681

Observation 5fa87b52-75d6-4f21-b69c-f76cc01c4f8e · outbound

This paper cites Listops: A diagnostic dataset for latent tree learning.

Transformers Learn Faster with Semantic Focus Listops: A diagnostic dataset for latent tree learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.356077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.141387Z digest=sha256:c38dcf4b07d1b7a62a4c44c22fc85afd9a81a4e4d6f8b18cb566e0d9a9500491

Observation bfb46c1d-85a6-4889-b1b2-37873f04faca · outbound

This paper cites Mish: A Self Regularized Non-Monotonic Activation Function.

Transformers Learn Faster with Semantic Focus Mish: A Self Regularized Non-Monotonic Activation Function

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.146060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.146060Z digest=sha256:b02d07b13752bb50b8431d0b1e21100308e49383a030f952980136a8bd09a14f

Observation 7a944d61-2998-44b7-a34b-676ba1b45bf1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Transformers Learn Faster with Semantic Focus Adam: A Method for Stochastic Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.152690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.152690Z digest=sha256:c701bbeabf17a366c67637a7bf4defafbe1db814e9ef339d13c40ea6fe672131

Observation 1c0085ca-f463-4103-952b-307c15eb115e · outbound

This paper cites The Lipschitz Constant of Self-Attention.

Transformers Learn Faster with Semantic Focus The Lipschitz Constant of Self-Attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.158224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.158224Z digest=sha256:1512a524ce423fdce8310a6f7506ca0c6fa03173af7a08b85fb3fb0286c156db

Observation 12797a98-aa4b-437c-b189-96d7d2b884b1 · outbound

This paper cites Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection.

Transformers Learn Faster with Semantic Focus Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.162921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.162921Z digest=sha256:8876e16a26dac341dd7072c931f85cc8b72a826af4c6426a9e48a9f0c486c9ca

Observation 37334c11-9643-4161-800f-94b4c7a6bbe5 · outbound

This paper cites Visualizing the loss landscape of neural nets.

Transformers Learn Faster with Semantic Focus Visualizing the loss landscape of neural nets

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.275494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.168150Z digest=sha256:3acd90fb3d6b8b0794998d6c5ef1cc10140187ffbf89be3bb3e51030504c93cf

Observation 10878189-944f-4b9b-8f6c-9907c722def7 · outbound

This paper cites Never train from scratch: Fair comparison of long-sequence models requires data-driven priors.

Transformers Learn Faster with Semantic Focus Never train from scratch: Fair comparison of long-sequence models requires data-driven priors

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.980495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.172682Z digest=sha256:d1f823578862659bdc7b9f3857a825ea8a311d8a8f67c1116a5bfc43d17d9c92

Observation c25273f1-6b34-4339-bc1d-998ffb03b3da · outbound

This paper cites Learning overparameterized neural networks via stochastic gradient descent on structured data.

Transformers Learn Faster with Semantic Focus Learning overparameterized neural networks via stochastic gradient descent on structured data

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.851131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.177991Z digest=sha256:26d9ca91447ad28eef18ca84e59ebab28f6f8c66dfcf965ec10bd034a5cda00a

Observation 233ed32b-ea5f-4522-944c-d2bb9ef0050c · outbound

This paper cites A convergence theory for deep learning via over-parameterization.

Transformers Learn Faster with Semantic Focus A convergence theory for deep learning via over-parameterization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.582899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.182659Z digest=sha256:de1f084e4bdb6346ce742c52005b17a9142dd69119f82b72d56fb09dc59f72e5

Observation ca681d40-ef8b-457c-962c-fffe1a113051 · outbound

This paper cites Gradient descent optimizes over-parameterized deep relu networks.

Transformers Learn Faster with Semantic Focus Gradient descent optimizes over-parameterized deep relu networks

Reference 67

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.234860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.187525Z digest=sha256:73525ff673f6a7bdcf9fa73f15c3c18802f2051fc1b9cd58058354580a1d268a

Observation aa3a87b7-a9a4-47f5-a37e-2fd076f33f5b · outbound

This paper cites Convergence rates for the stochastic gradient descent method for non-convex objective functions.

Transformers Learn Faster with Semantic Focus Convergence rates for the stochastic gradient descent method for non-convex objective functions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.280375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.193173Z digest=sha256:553a7716d2e41a051a7e5c8153cbc225e2d6281e02e40c334850e40f77e2fa99

Pith citing papers

No inbound Pith citation observations are available.