Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:08:41.245145Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 22 inbound Pith citation observations for arXiv:2412.02975.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:08:41.245145Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:23:02.058863Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T22:06:16.211261Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f07fda5-a328-49ad-9893-3c2ba0ee785a · outbound
Theoretical limitations of multi-layer Transformer GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2fc5af2-fdf4-42f8-a517-27dab7d74f10 · outbound
Theoretical limitations of multi-layer Transformer Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd83f51b-5990-41df-8240-8ef4e1d69fdb · outbound
Theoretical limitations of multi-layer Transformer Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10638fbd-74fb-4915-8236-decd6dfc6feb · outbound
Theoretical limitations of multi-layer Transformer ENTP: Encoder-only Next Token Prediction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18246e31-ef52-4c00-a08b-9a152546daf6 · outbound
Theoretical limitations of multi-layer Transformer How can s elf-attention networks recog- nize dyck-n languages? In Findings of the Association for Computational Linguistics: EMNLP 2020 , pages 4301–4306,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd51977d-38f1-47e5-be79-8d1b5ef5b678 · outbound
Theoretical limitations of multi-layer Transformer Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cba6c8cc-72de-4319-b2e6-35671cd10e24 · outbound
Theoretical limitations of multi-layer Transformer Hoza, Avishay Tal, and Roe i Tell
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3afd07c8-b8f0-4bdb-b535-0ba71ce89df0 · outbound
Theoretical limitations of multi-layer Transformer Language Models are Few-Shot Learners
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 913378a0-8ffb-40a6-acc0-c7a4c305f34a · outbound
Theoretical limitations of multi-layer Transformer The Expressive Power of Transformers with Chain of Thought
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730121a1-99d6-41a2-b410-6e86e4435012 · outbound
Theoretical limitations of multi-layer Transformer Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4da9a69-1fcc-4d76-b413-a5c075946cb7 · outbound
Theoretical limitations of multi-layer Transformer The impact of depth on compositional generalizatio n in transformer language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 769a57df-100f-465e-a3b2-5851eb4c13e3 · outbound
Theoretical limitations of multi-layer Transformer Measuring and narrowing the compositionality gap in language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de7ad18a-7852-42e4-a10d-7c07edc27c05 · outbound
Theoretical limitations of multi-layer Transformer Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9bb37d7-f136-4928-ba10-c6ccb381f3a5 · outbound
Theoretical limitations of multi-layer Transformer One-layer transformers fail to solve the induction heads task
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0155291-d9dd-460b-9541-e02051813e00 · outbound
Theoretical limitations of multi-layer Transformer Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a517fd89-f493-4e04-8813-7229279dcb12 · outbound
Theoretical limitations of multi-layer Transformer RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a97bfc1-598d-46af-99a3-a3d58a1cc152 · outbound
Theoretical limitations of multi-layer Transformer Emergent Abilities of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0062f6-dad7-4894-af4b-43a6f88b7815 · outbound
Theoretical limitations of multi-layer Transformer Do large language models latently perform multi-hop reasoning? In Association for Com- putational Linguistics: ACL-IJCNLP 2024 ,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6bd2eb1e-1304-462a-9ca9-02f8863f88f7 · outbound
Theoretical limitations of multi-layer Transformer A Theory for Emergence of Complex Skills in Language Models
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81157387-3bb9-453e-87e4-67309dafc729 · outbound
Theoretical limitations of multi-layer Transformer On medium-unifor mity and circuit lower bounds
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e5e1eadc-763a-4a0b-8720-b1e6e87826e5 · outbound
Theoretical limitations of multi-layer Transformer Bootstrapping results for t hreshold circuits ”just beyond” known lower bounds
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9e1afbf5-a0a7-487b-9d7a-1552cbb80b11 · outbound
Theoretical limitations of multi-layer Transformer Mixture of Parrots: Experts improve memorization more than reasoning
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed83ac09-1faf-479a-9d54-1594fc7debb5 · outbound
Theoretical limitations of multi-layer Transformer Rnns can generate bounded hierarchical languages with opti mal memory
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 29394370-e257-4826-bf3f-2f6312204bde · outbound
Theoretical limitations of multi-layer Transformer Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1148aad9-b8ca-4c71-a0b0-68ab6c76928b · outbound
Theoretical limitations of multi-layer Transformer Chan, and R
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4209ae2-08da-4405-a76b-c77e42ddd1be · outbound
Theoretical limitations of multi-layer Transformer Physics of Language Models: Part 1, Learning Hierarchical Language Structures
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fcc26d-e5b9-4c05-9573-cc7ab0ef6084 · inbound
Lower bounds on transformers with infinite precision Theoretical limitations of multi-layer Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb59aa56-0181-4f9f-a88c-00d741943aa3 · inbound
Ehrenfeucht-Haussler Rank and Chain of Thought Theoretical limitations of multi-layer Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fff638b-a1a9-4bbf-af5d-92983cb89be9 · inbound
Lower Bounds for Chain-of-Thought Reasoning in Hard-Attention Transformers Theoretical limitations of multi-layer Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54a66fe2-2e84-4fba-a26c-7fc846a45aa3 · inbound
When More is Less: Understanding Chain-of-Thought Length in LLMs Theoretical limitations of multi-layer Transformer
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07a529e-cc28-44a6-85b0-053c7761745c · inbound
Chain-of-Thought Tokens are Computer Program Variables Theoretical limitations of multi-layer Transformer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0de0330-ed76-4636-9e78-5e132706a0e8 · inbound
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers Theoretical limitations of multi-layer Transformer
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · inbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0892a673-8ad9-4bb9-8ba8-6a3aa63cfb42 · inbound
Transformers Meet In-Context Learning: A Universal Approximation Theory Theoretical limitations of multi-layer Transformer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ad6197-2c7d-4432-9d5b-2249cf7b5a9e · inbound
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity Theoretical limitations of multi-layer Transformer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a38488-424e-417c-9efa-9dbe9551f1bc · inbound
The Serial Scaling Hypothesis Theoretical limitations of multi-layer Transformer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fea037c4-e0f2-44f6-83db-7d76371b6564 · inbound
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Theoretical limitations of multi-layer Transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87de76e6-4d0e-4296-a207-0af19beb2fa0 · inbound
Deep sequence models tend to memorize geometrically; it is unclear why Theoretical limitations of multi-layer Transformer
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 06813133-1349-45ba-96d5-f3f7f8f2f588 · inbound
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Theoretical limitations of multi-layer Transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef74e4a7-f24d-426e-a99e-21cdeb40f1fb · inbound
When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path Compression Theoretical limitations of multi-layer Transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58cf521f-8294-41cc-9335-9906260b1043 · inbound
The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 346494fb-6aff-40b1-af21-3c6719f9c40d · inbound
The Power of Power Law: Asymmetry Enables Compositional Reasoning Theoretical limitations of multi-layer Transformer
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6a387a-5972-444b-ad9e-d0e0bd7ea9f1 · inbound
Continuous Latent Contexts Enable Efficient Online Learning in Transformers Theoretical limitations of multi-layer Transformer
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb428005-a4ad-4090-8f1a-d89ede65f0a7 · inbound
Agentic Transformers Provably Learn to Search via Reinforcement Learning Theoretical limitations of multi-layer Transformer
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4b62c3fd-4b9e-4246-80fe-07705c18a1f7 · inbound
Rethinking the Role of Positional Encoding: Sliding-Window Transformers without PE Remain Turing Complete Theoretical limitations of multi-layer Transformer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8bfa3f70-6a48-42be-9ed7-5e34a0907aa8 · inbound
Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D Theoretical limitations of multi-layer Transformer
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d52774c1-9cf4-4ffe-bffc-6140e67f8c26 · inbound
Hierarchical Domain Generalization Theoretical limitations of multi-layer Transformer
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3d7df0-27b4-4bd4-bce4-d824ce19a4e0 · inbound
Attention-based representations for multi-task computation Theoretical limitations of multi-layer Transformer
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.