Pith. sign in

Paper Citation Record · LEDGER

Theoretical limitations of multi-layer Transformer

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 22 inbound Pith citation observations for arXiv:2412.02975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02975 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:08:41.245145Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:23:02.058863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:06:16.211261Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f07fda5-a328-49ad-9893-3c2ba0ee785a · outbound

This paper cites GPT-4 Technical Report.

Theoretical limitations of multi-layer Transformer GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:39.963350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:39.963350Z digest=sha256:05a66f04ef474d30d471b9a9d72188c82064dc3ed9af4755cc222e9b1691f2ce

Observation f2fc5af2-fdf4-42f8-a517-27dab7d74f10 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Theoretical limitations of multi-layer Transformer Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.164864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.164864Z digest=sha256:e38c7d7f964800d6d46536129423a9973f3163808a90e8c62c0842819b661543

Observation fd83f51b-5990-41df-8240-8ef4e1d69fdb · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Theoretical limitations of multi-layer Transformer Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.255139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.255139Z digest=sha256:a551f3aaf503167c0bcf9859a21ed832ace3b756b0c6124554e1963b61857f58

Observation 10638fbd-74fb-4915-8236-decd6dfc6feb · outbound

This paper cites ENTP: Encoder-only Next Token Prediction.

Theoretical limitations of multi-layer Transformer ENTP: Encoder-only Next Token Prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.404769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.404769Z digest=sha256:8e8652946895f7d0df34d673ea39c178400c2a900e66ff4b598e4b008bde3724

Observation 18246e31-ef52-4c00-a08b-9a152546daf6 · outbound

This paper cites How can s elf-attention networks recog- nize dyck-n languages? In Findings of the Association for Computational Linguistics: EMNLP 2020 , pages 4301–4306,.

Theoretical limitations of multi-layer Transformer How can s elf-attention networks recog- nize dyck-n languages? In Findings of the Association for Computational Linguistics: EMNLP 2020 , pages 4301–4306,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:43.528637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.496367Z digest=sha256:b472ad78dd5da2d6c18fcc0649d94ab0062d3bc5e165972ccca56d3463b05c91

Observation dd51977d-38f1-47e5-be79-8d1b5ef5b678 · outbound

This paper cites Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder.

Theoretical limitations of multi-layer Transformer Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.536263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.536263Z digest=sha256:8daf8ab854ae35b448937d524cb64aaaef26b750a4f9951f9389829ea6ea4d54

Observation cba6c8cc-72de-4319-b2e6-35671cd10e24 · outbound

This paper cites Hoza, Avishay Tal, and Roe i Tell.

Theoretical limitations of multi-layer Transformer Hoza, Avishay Tal, and Roe i Tell

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:43.224843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.587189Z digest=sha256:1a1d47d423c9d9c8f5f9003e52128d8001eeec21667e665f6f05836af367e617

Observation 3afd07c8-b8f0-4bdb-b535-0ba71ce89df0 · outbound

This paper cites Language Models are Few-Shot Learners.

Theoretical limitations of multi-layer Transformer Language Models are Few-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.684312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.684312Z digest=sha256:b2b197591844f42e98d5102f1302144e824a4e975b324edc6c78917bb06ff0f2

Observation 913378a0-8ffb-40a6-acc0-c7a4c305f34a · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

Theoretical limitations of multi-layer Transformer The Expressive Power of Transformers with Chain of Thought

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.697210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.697210Z digest=sha256:1a8a5115ece4e3fde529ffd6040c8404edccb4626b1c763760c2279569600500

Observation 730121a1-99d6-41a2-b410-6e86e4435012 · outbound

This paper cites Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks.

Theoretical limitations of multi-layer Transformer Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.713400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.713400Z digest=sha256:96a17ad9c6a951aee569f5f46a14acde7e8a4b343d1f06e24ab5a34a75747c67

Observation b4da9a69-1fcc-4d76-b413-a5c075946cb7 · outbound

This paper cites The impact of depth on compositional generalizatio n in transformer language models.

Theoretical limitations of multi-layer Transformer The impact of depth on compositional generalizatio n in transformer language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:43.064020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.849997Z digest=sha256:8590a6c5efb04646145af2efc2a9f72b6a568c306ab483c47aa957b9d0dc110e

Observation 769a57df-100f-465e-a3b2-5851eb4c13e3 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Theoretical limitations of multi-layer Transformer Measuring and narrowing the compositionality gap in language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:42.995974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.896185Z digest=sha256:75f73730f47b2971eecd3f3e0e8d08dcfab0b8551016221da1280a7aa5544d5c

Observation de7ad18a-7852-42e4-a10d-7c07edc27c05 · outbound

This paper cites an unresolved cited work.

Theoretical limitations of multi-layer Transformer Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:08:42.894017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.908544Z digest=sha256:8aaf6f3a570b26749bbccf66a7314b37dce01de58ded359afa0a17830f1fef29

Observation b9bb37d7-f136-4928-ba10-c6ccb381f3a5 · outbound

This paper cites One-layer transformers fail to solve the induction heads task.

Theoretical limitations of multi-layer Transformer One-layer transformers fail to solve the induction heads task

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.926187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.926187Z digest=sha256:d55301a56d1d49800d5b1005f19ade68d0b9a5e3d68e233a23d6ecc9c053f980

Observation a0155291-d9dd-460b-9541-e02051813e00 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Theoretical limitations of multi-layer Transformer Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.988082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.988082Z digest=sha256:16560a34e7413ec36125633dc4ff0404c7f395bc326e884ed9aaca28d3677874

Observation a517fd89-f493-4e04-8813-7229279dcb12 · outbound

This paper cites RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval.

Theoretical limitations of multi-layer Transformer RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:41.006808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:41.006808Z digest=sha256:fa2a34b0969a4f9ecce75c3786f745f122ba50b1c52d2f752cbf23ade4e3e633

Observation 1a97bfc1-598d-46af-99a3-a3d58a1cc152 · outbound

This paper cites Emergent Abilities of Large Language Models.

Theoretical limitations of multi-layer Transformer Emergent Abilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:41.082304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:41.082304Z digest=sha256:288f0e6f13bdc824258c863c5f2e905c80f43ee384148803a608fe116d605681

Observation ae0062f6-dad7-4894-af4b-43a6f88b7815 · outbound

This paper cites Do large language models latently perform multi-hop reasoning? In Association for Com- putational Linguistics: ACL-IJCNLP 2024 ,.

Theoretical limitations of multi-layer Transformer Do large language models latently perform multi-hop reasoning? In Association for Com- putational Linguistics: ACL-IJCNLP 2024 ,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:42.688922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:41.206834Z digest=sha256:98ac46216cd4991dde23ab3fa0ee1ee3894130dcd321f40d6935ccfa5c5d089e

Observation 6bd2eb1e-1304-462a-9ca9-02f8863f88f7 · outbound

This paper cites A Theory for Emergence of Complex Skills in Language Models.

Theoretical limitations of multi-layer Transformer A Theory for Emergence of Complex Skills in Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.040192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.040192Z digest=sha256:e8c9fa03350265738d80ff181b127b8c839be683df0e092f6f6ccb7341c50e59

Observation 81157387-3bb9-453e-87e4-67309dafc729 · outbound

This paper cites On medium-unifor mity and circuit lower bounds.

Theoretical limitations of multi-layer Transformer On medium-unifor mity and circuit lower bounds

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:42.762856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.966787Z digest=sha256:61683495814ef05c43a3c2d1806a88193c22edced4b144a60542cd23fdde2533

Observation e5e1eadc-763a-4a0b-8720-b1e6e87826e5 · outbound

This paper cites Bootstrapping results for t hreshold circuits ”just beyond” known lower bounds.

Theoretical limitations of multi-layer Transformer Bootstrapping results for t hreshold circuits ”just beyond” known lower bounds

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:43.599277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.318423Z digest=sha256:e5efc310cd5ca96af9f32fb7ac6906fb58a427c697546655c49587129ad7a7f7

Observation 9e1afbf5-a0a7-487b-9d7a-1552cbb80b11 · outbound

This paper cites Mixture of Parrots: Experts improve memorization more than reasoning.

Theoretical limitations of multi-layer Transformer Mixture of Parrots: Experts improve memorization more than reasoning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.634905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.634905Z digest=sha256:d5e4ed02b94e5bd0340009271545186749417223c646a7b4991e5c55728d2694

Observation ed83ac09-1faf-479a-9d54-1594fc7debb5 · outbound

This paper cites Rnns can generate bounded hierarchical languages with opti mal memory.

Theoretical limitations of multi-layer Transformer Rnns can generate bounded hierarchical languages with opti mal memory

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:43.365175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:40.544760Z digest=sha256:58a71c23656706fdf40f8e2c1d928d7416b3c051eb7577bbaa35774a4860e085

Observation 29394370-e257-4826-bf3f-2f6312204bde · outbound

This paper cites Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process.

Theoretical limitations of multi-layer Transformer Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:41.245145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:41.245145Z digest=sha256:98b596805b560d7989afa3a7239b714df612c55fee307a4025f61ca84fb5e027

Observation 1148aad9-b8ca-4c71-a0b0-68ab6c76928b · outbound

This paper cites Chan, and R.

Theoretical limitations of multi-layer Transformer Chan, and R

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:08:43.760560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T23:08:39.999265Z digest=sha256:fd84a833504a0bee0f9f0a14b6b7a1af6e342cb822e18005c1b4deae9b8455cc

Observation d4209ae2-08da-4405-a76b-c77e42ddd1be · outbound

This paper cites Physics of Language Models: Part 1, Learning Hierarchical Language Structures.

Theoretical limitations of multi-layer Transformer Physics of Language Models: Part 1, Learning Hierarchical Language Structures

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T23:08:40.089817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:08:40.089817Z digest=sha256:d564cee7db7ec8dfaa94042f135cfc74f03adada66e494775bbb2a897f5d3dc0

Pith citing papers

Observation 20fcc26d-e5b9-4c05-9573-cc7ab0ef6084 · inbound

Lower bounds on transformers with infinite precision cites this paper.

Lower bounds on transformers with infinite precision Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:36:27.117603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:36:27.117603Z digest=sha256:8ec5fe58589fcae1e35af94c98577f7ccd0abb1644372c62130e8186ab4f7e62

Observation bb59aa56-0181-4f9f-a88c-00d741943aa3 · inbound

Ehrenfeucht-Haussler Rank and Chain of Thought cites this paper.

Ehrenfeucht-Haussler Rank and Chain of Thought Theoretical limitations of multi-layer Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:38:56.558522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:38:56.558522Z digest=sha256:ebbc03db67fabf99f406e65962b87f440a4b78b88e6ac451ab7fbbe8a98e069b

Observation 2fff638b-a1a9-4bbf-af5d-92983cb89be9 · inbound

Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers cites this paper.

Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers Theoretical limitations of multi-layer Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:57.804636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:23:57.804636Z digest=sha256:acf61191478f6353f61f934b90b1bf1a1c453b5d8372760bb08e64d7805bf900

Observation 54a66fe2-2e84-4fba-a26c-7fc846a45aa3 · inbound

When More is Less: Understanding Chain-of-Thought Length in LLMs cites this paper.

When More is Less: Understanding Chain-of-Thought Length in LLMs Theoretical limitations of multi-layer Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T13:23:22.006914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:23:22.006914Z digest=sha256:2ee254f96de541999e38a76a9fe3637bef09d209e83f891a2dcd124e63254a73

Observation d07a529e-cc28-44a6-85b0-053c7761745c · inbound

Chain-of-Thought Tokens are Computer Program Variables cites this paper.

Chain-of-Thought Tokens are Computer Program Variables Theoretical limitations of multi-layer Transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:23:02.058863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:23:02.058863Z digest=sha256:6d8a50bae3b96c6c011e73dfb6f0863b5f78413a8f916755b41eb46b91d71a3d

Observation b0de0330-ed76-4636-9e78-5e132706a0e8 · inbound

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers cites this paper.

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers Theoretical limitations of multi-layer Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:57.267988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:57.267988Z digest=sha256:db75864ff9a48b49c0ac9fa747f4c36df1de9f3c55741dd5b89c604c8a85695d

Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · inbound

Learning Compositional Functions with Transformers from Easy-to-Hard Data cites this paper.

Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:33.978754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:33.978754Z digest=sha256:0ee0da5f847097246b373fcebe62cf32b54c11a089d02ce88f85b078af496e9c

Observation 0892a673-8ad9-4bb9-8ba8-6a3aa63cfb42 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Theoretical limitations of multi-layer Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:35.409540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:35.409540Z digest=sha256:1c09eaee88fe31b2c2c0550ba19da7b70fb6910199a30f56e7c920d822d0ede6

Observation c7ad6197-2c7d-4432-9d5b-2249cf7b5a9e · inbound

Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity cites this paper.

Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity Theoretical limitations of multi-layer Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T01:18:51.437766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:18:51.437766Z digest=sha256:1c1e334c62b6236b12552991c2a436211de24ae6e5bf7e3b559bd6ea9baec83c

Observation 38a38488-424e-417c-9efa-9dbe9551f1bc · inbound

The Serial Scaling Hypothesis cites this paper.

The Serial Scaling Hypothesis Theoretical limitations of multi-layer Transformer

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.486141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T04:08:11.344622Z digest=sha256:e188012de5747adc859953188a3ca910dfeeb65182120d65237d0d48a285505e

Observation fea037c4-e0f2-44f6-83db-7d76371b6564 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Theoretical limitations of multi-layer Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.879218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.879218Z digest=sha256:47e9c7a20aee570431179a22ece7c5b8d23a1103036cbd9448f103a76a8bb935

Observation 87de76e6-4d0e-4296-a207-0af19beb2fa0 · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Theoretical limitations of multi-layer Transformer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:40:36.194169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:2c35a87a92e0277efb22b9b32a97e93962169897d36e19d2ce6c2fe6548a4761

Observation 06813133-1349-45ba-96d5-f3f7f8f2f588 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Theoretical limitations of multi-layer Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:09.914249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:09.914249Z digest=sha256:976f32047d204740a969b70725cd7ea21fd811b1e78646814906f9500fbce6bb

Observation ef74e4a7-f24d-426e-a99e-21cdeb40f1fb · inbound

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression cites this paper.

When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T13:07:00.928770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:07:00.928770Z digest=sha256:b8d6c7da52691bfe18b66d5165b3ea9ccc5d578660dd1029cca9abf76dd87068

Observation 58cf521f-8294-41cc-9335-9906260b1043 · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:08.240882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T11:49:49.787123Z digest=sha256:0489010f5d8af2173ad3a2dd8f047ddca94dcf4c49f4c991d52405fa41e67c03

Observation 346494fb-6aff-40b1-af21-3c6719f9c40d · inbound

The Power of Power Law: Asymmetry Enables Compositional Reasoning cites this paper.

The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T18:26:05.728364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T18:26:05.728364Z digest=sha256:7a00a502d61ef8b76e64ce455c8b590fb706e2cc78cb0f30a8265499ee9d16c0

Observation 0f6a387a-5972-444b-ad9e-d0e0bd7ea9f1 · inbound

Continuous Latent Contexts Enable Efficient Online Learning in Transformers cites this paper.

Continuous Latent Contexts Enable Efficient Online Learning in Transformers Theoretical limitations of multi-layer Transformer

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:41:23.574147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:06:54.598357Z digest=sha256:2f8b52236a1bae54321e73677ec907b7a0ae96c08f963f19bcf7e3c060bd25cd

Observation cb428005-a4ad-4090-8f1a-d89ede65f0a7 · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning Theoretical limitations of multi-layer Transformer

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:42:49.947749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:62799f010e71b548914c2abefb97b462322ba4476744740594a4e055c203e99c

Observation 4b62c3fd-4b9e-4246-80fe-07705c18a1f7 · inbound

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete cites this paper.

Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete Theoretical limitations of multi-layer Transformer

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:06:16.212640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T15:48:48.046003Z digest=sha256:91d99bba4ee21a7bb9fae48d8bfb35e4b9f2efdd8e5ba88bfad1fcba62f15d06

Observation 8bfa3f70-6a48-42be-9ed7-5e34a0907aa8 · inbound

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D cites this paper.

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D Theoretical limitations of multi-layer Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T21:31:00.903239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T21:31:00.903239Z digest=sha256:511441b16250f4f50af57cf6243624df985acd784abe669177b5a2220fc06454

Observation d52774c1-9cf4-4ffe-bffc-6140e67f8c26 · inbound

Hierarchical Domain Generalization cites this paper.

Hierarchical Domain Generalization Theoretical limitations of multi-layer Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T20:54:05.553053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:54:05.553053Z digest=sha256:15d2472041846767e3387d6fac4b81b26f13ca631e33ad5cf98a9bb41981e7d4

Observation 2b3d7df0-27b4-4bd4-bce4-d824ce19a4e0 · inbound

Attention-based representations for multi-task computation cites this paper.

Attention-based representations for multi-task computation Theoretical limitations of multi-layer Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:21.263783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:18:21.263783Z digest=sha256:aaab4cfe392c9c26e17987a880d2efdbbd9f1d643e682849539b07ae82678a4e