Pith. sign in

Paper Citation Record · LEDGER

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

As of 18 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2508.12837.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12837 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:26:36.926253Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:12:29.129465Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6329383b-e0b7-4a9b-9033-79d01fe375e6 · outbound

This paper cites Lemma H.4 (Stationarity of sub-k-tuples).

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Lemma H.4 (Stationarity of sub-k-tuples)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:26:37.297607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:26:36.921831Z digest=sha256:07cfef7f581f2ff7edca9a870b5608f98a75eead03de5d5469065787494a0626

Observation 9285ac0d-c5c0-4f06-8fec-286356d5fa0e · outbound

This paper cites Transformers learn through gradual rank increase.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers learn through gradual rank increase

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.850936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.850936Z digest=sha256:b26aa77f5c1cb35a8877063c556705c376765b0bb29202f91ee8bb703eaa0ded

Observation 7783ac8e-c731-40b6-a86e-734b189170de · outbound

This paper cites Understanding Incremental Learning of Gradient Descent: A Fine-grained Analysis of Matrix Sensing.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Understanding Incremental Learning of Gradient Descent: A Fine-grained Analysis of Matrix Sensing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.876555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.876555Z digest=sha256:aaa93ead254748db965d216bca339e85333ecd5a2a0887c10ed5d982b8903f77

Observation ef88414f-672f-472b-8a34-66c1c885fa2a · outbound

This paper cites Task Diversity Shortens the ICL Plateau.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Task Diversity Shortens the ICL Plateau

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.880999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.880999Z digest=sha256:410ba5d66a0cf9caf85a813ea727bbda4880e18fa02637e29a5aa463fb576b4f

Observation 0fd5ad81-8cc8-4652-819d-2f8885f5c2f5 · outbound

This paper cites Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.886042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.886042Z digest=sha256:c9cf2b9143df644abaedbfa4af5a747f5af60350ab2f04b6818c666e93df5b10

Observation 7ff6388f-9ea2-47c8-bb47-61d49b1e97cd · outbound

This paper cites How Transformers Learn Causal Structure with Gradient Descent.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points How Transformers Learn Causal Structure with Gradient Descent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.890264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.890264Z digest=sha256:15aa28aeb869ab72f1db609c976b04cff3daf3520bbe3c8d160a0b94fec8478d

Observation 729f2995-ba6e-4df5-b876-53784724f649 · outbound

This paper cites Transformers on Markov Data: Constant Depth Suffices.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers on Markov Data: Constant Depth Suffices

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.902854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.902854Z digest=sha256:1b7b7fe285c2910d5e1ee80c1d56ac7e9678811057eadb15bb9e9290c1b091dd

Observation 4f495f07-29f9-40db-90be-7b519477889b · outbound

This paper cites Emergent Abilities of Large Language Models.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Emergent Abilities of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.912757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.912757Z digest=sha256:d3e40e4a54d9da666404d96c13179b778fee65dedc25ec5c96fee51a68b2f159

Observation 4c39ca19-1411-4709-9f16-1479081aa876 · outbound

This paper cites Large Language Models as Markov Chains.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Large Language Models as Markov Chains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.917183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.917183Z digest=sha256:2e571450c723ca6601ee92a3dc37e248ef0cd400bdd6887709315279159aeb48

Observation 373e80b1-79ea-4309-b859-f5018b7297bb · outbound

This paper cites an unresolved cited work.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Unresolved cited work

Reference 128

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:26:37.284197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:26:36.926253Z digest=sha256:9d3fe20f3693733d8329eccd320771a5b5ba845754fc773d0ae42f2cddda0327

Observation e83276c0-6e19-4f46-ae9a-f62c4c8e15db · outbound

This paper cites Transformers Can Represent $n$-gram Language Models.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers Can Represent $n$-gram Language Models

Reference 1948

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.907740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.907740Z digest=sha256:b5c11d9d17249180109cda20aafe87930a791410d301ac3016c68f5cf5b90b4e

Observation 364f9266-cd07-4149-9994-3b6002c82579 · outbound

This paper cites Edelman, B.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Edelman, B

Reference 1956

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.859528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.859528Z digest=sha256:4cfd46f3df686caf72355d9e94da0852981453fe54256dfed4c51bffa405e206

Observation e416bfee-006b-4672-92b3-ff68fbf6d459 · outbound

This paper cites Language Models are Few-Shot Learners.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Language Models are Few-Shot Learners

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.855241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.855241Z digest=sha256:e825d14a7385268e629141508ff7b0173509fd03ae110d1d8ba78c00082e94b4

Observation de7d66a7-bab5-4c48-8fff-47cedee14160 · outbound

This paper cites Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization

Reference 1998

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:26:37.110041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:26:36.872360Z digest=sha256:b7d479bee2f603f6c8498d4f476c2c07070d49dcc371ae308433cdcf1fcbf90d

Observation 517fafe7-6b3c-4979-8520-bca981bf77af · outbound

This paper cites doi: https://doi.org/10.1016/S0893-6080(00)00009-5.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points doi: https://doi.org/10.1016/S0893-6080(00)00009-5

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.863641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.863641Z digest=sha256:457871ed215e27799beea9083b058066775c9d3a90a85b707e36ab496804df7d

Observation b261dcc0-6745-4499-8898-aa6ed025e675 · outbound

This paper cites Incremental Learning in Diagonal Linear Networks.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Incremental Learning in Diagonal Linear Networks

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:26:37.256553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:26:36.841860Z digest=sha256:7008637a8153e2960bc8fcb3094a721754c0db31110b33ac19dcf6192a5c104c

Observation f0c618b3-6213-47c5-b045-289d914670fc · outbound

This paper cites Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.867859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.867859Z digest=sha256:4bcad81f10c481133438997aef6fd298a43dc2a87da1e83378cdd1a9db4ab79a

Observation 47303cbe-4018-4c8c-8bad-d3aabd8c4e65 · outbound

This paper cites Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.846150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.846150Z digest=sha256:835bff1b8cf0627c4180e27c746f9f957581f3b04792a4289183d60380a8fbbb

Observation 798ccdd0-48f8-4247-b015-769d65ece607 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.898860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.898860Z digest=sha256:2e05130b676313e8725581f8ef2681e9d04903484ecf9880b2324e0c71bf0d3f

Observation 2aade19d-f781-48df-908c-924a2f0dc9d1 · outbound

This paper cites In-Context Language Learning: Architectures and Algorithms.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points In-Context Language Learning: Architectures and Algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.836988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.836988Z digest=sha256:13e3f754ea6ae6e5b75ee9dc76b9e64644dfc3af9e960850ce3b9a97f07eb78c

Observation d77ef9e6-0216-4253-9b72-dab3de4b96a7 · outbound

This paper cites A Mechanistic Study of Transformers Training Dynamics.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points A Mechanistic Study of Transformers Training Dynamics

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:26:37.039736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:26:36.894671Z digest=sha256:6e8d9cbe457e23dd0f0e4b6cfc52a28c2f842f360e5ca125b7cf853719ee06ff

Pith citing papers

Observation 6175fb5c-89b1-40c8-a995-7fe76b38f892 · inbound

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities cites this paper.

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:51:19.291387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T03:51:08.871267Z digest=sha256:0061274f02f96bdf22dc2269e8507417d64a2fe266708f4687bd4115a796745d

Observation 60b03b04-346f-49ce-85f2-3a1766e2527b · inbound

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models cites this paper.

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T10:12:29.129465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:12:29.129465Z digest=sha256:f0baf81b5ba300f6d2db8dd5bf80e45b44107fb90ea0333b78ab050fef4d7f2e