Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:26:36.926253Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2508.12837.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:26:36.926253Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T10:12:29.129465Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
21 of 21 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 6329383b-e0b7-4a9b-9033-79d01fe375e6 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Lemma H.4 (Stationarity of sub-k-tuples)
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9285ac0d-c5c0-4f06-8fec-286356d5fa0e · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers learn through gradual rank increase
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7783ac8e-c731-40b6-a86e-734b189170de · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Understanding Incremental Learning of Gradient Descent: A Fine-grained Analysis of Matrix Sensing
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef88414f-672f-472b-8a34-66c1c885fa2a · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Task Diversity Shortens the ICL Plateau
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd5ad81-8cc8-4652-819d-2f8885f5c2f5 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff6388f-9ea2-47c8-bb47-61d49b1e97cd · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points How Transformers Learn Causal Structure with Gradient Descent
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729f2995-ba6e-4df5-b876-53784724f649 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers on Markov Data: Constant Depth Suffices
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f495f07-29f9-40db-90be-7b519477889b · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Emergent Abilities of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c39ca19-1411-4709-9f16-1479081aa876 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Large Language Models as Markov Chains
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373e80b1-79ea-4309-b859-f5018b7297bb · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Unresolved cited work
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e83276c0-6e19-4f46-ae9a-f62c4c8e15db · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers Can Represent $n$-gram Language Models
Reference 1948
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 364f9266-cd07-4149-9994-3b6002c82579 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Edelman, B
Reference 1956
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e416bfee-006b-4672-92b3-ff68fbf6d459 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Language Models are Few-Shot Learners
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7d66a7-bab5-4c48-8fff-47cedee14160 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization
Reference 1998
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 517fafe7-6b3c-4979-8520-bca981bf77af · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points doi: https://doi.org/10.1016/S0893-6080(00)00009-5
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b261dcc0-6745-4499-8898-aa6ed025e675 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Incremental Learning in Diagonal Linear Networks
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f0c618b3-6213-47c5-b045-289d914670fc · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47303cbe-4018-4c8c-8bad-d3aabd8c4e65 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798ccdd0-48f8-4247-b015-769d65ece607 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aade19d-f781-48df-908c-924a2f0dc9d1 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points In-Context Language Learning: Architectures and Algorithms
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77ef9e6-0216-4253-9b72-dab3de4b96a7 · outbound
Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points A Mechanistic Study of Transformers Training Dynamics
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6175fb5c-89b1-40c8-a995-7fe76b38f892 · inbound
On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 60b03b04-346f-49ce-85f2-3a1766e2527b · inbound
Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.