Pith. sign in

Paper Citation Record · LEDGER

Uncovering mesa-optimization algorithms in Transformers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2309.05858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.05858 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:00:58.100805Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9210d8e8-b104-4e33-9716-b18b348d126a · inbound

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models cites this paper.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Uncovering mesa-optimization algorithms in Transformers

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.172724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:ced82fa66775f10b79c70183526af00fb61eb28c8c5a84c1a2ebd060a7c64097

Observation 1bbfec1d-a87b-449d-bc64-7810b77d1f86 · inbound

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence cites this paper.

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence Uncovering mesa-optimization algorithms in Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:00:58.100805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:00:58.100805Z digest=sha256:695513fa06909c5677686e254ac63d28cb117cfbe160e5951c657dbd06376e63

Observation 998052c0-e19d-47ee-994a-aeab003d2d99 · inbound

The Role of Diversity in In-Context Learning for Large Language Models cites this paper.

The Role of Diversity in In-Context Learning for Large Language Models Uncovering mesa-optimization algorithms in Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:38.439808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:21:38.439808Z digest=sha256:87558ec86ab2fb4ab66c9b25227df509f14f1fd0f1c98a529f2dac74b454c867

Observation ee67979d-b995-4e61-9598-ffee220492d9 · inbound

ATLAS: Learning to Optimally Memorize the Context at Test Time cites this paper.

ATLAS: Learning to Optimally Memorize the Context at Test Time Uncovering mesa-optimization algorithms in Transformers

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:21.592813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:21.592813Z digest=sha256:67e7e6931a40bf71f8bf585f0d60a87cd77c1afbbab426bd36d7790ebc65230b

Observation 6fdb14e6-8e63-428b-9224-93dcb347cd95 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Uncovering mesa-optimization algorithms in Transformers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:42.468990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:42.468990Z digest=sha256:f77275a7d66eef399db3b6f8b3c6f444010cb65a0efe4be3e5420dd61078d9b3

Observation 2d90cb9a-ff6a-4f3a-b45a-132f60b1c6b7 · inbound

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training cites this paper.

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Uncovering mesa-optimization algorithms in Transformers

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:13.976683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:30:13.976683Z digest=sha256:287b690c5f72c77e827b81be677f3f8cdc39e5dd0bbf019feb0dce1bd8f8ae7d

Observation ff999589-2f12-448d-be26-ea9e35402d55 · inbound

Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning cites this paper.

Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning Uncovering mesa-optimization algorithms in Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:11:13.097520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:11:13.097520Z digest=sha256:483f66a8790a128f6183075eb20183d9f41771f77bdf44afcdf48e57cd954c00

Observation 8019fbb2-e19b-4c7a-903c-83027a45460e · inbound

Position: We Need An Algorithmic Understanding of Generative AI cites this paper.

Position: We Need An Algorithmic Understanding of Generative AI Uncovering mesa-optimization algorithms in Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:41:33.871059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:41:33.871059Z digest=sha256:fabdcaab3f623205e30c648d21fa701f10c746a22f9ef414bfa02f34447d4003

Observation dd6a7791-dec0-4451-a94b-9de310699d76 · inbound

Selective Induction Heads: How Transformers Select Causal Structures In Context cites this paper.

Selective Induction Heads: How Transformers Select Causal Structures In Context Uncovering mesa-optimization algorithms in Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:11:45.295838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:11:45.295838Z digest=sha256:a307a8eb3ff6880b309cb47663babe4bc4b3263aad4f33b54337b934c0bb1819

Observation ee68abba-7ec3-481d-a96d-8928d03f439a · inbound

Learning to Remember, Learn, and Forget in Attention-Based Models cites this paper.

Learning to Remember, Learn, and Forget in Attention-Based Models Uncovering mesa-optimization algorithms in Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:07.695368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:07.695368Z digest=sha256:9e77c3c53a512f4070bc7cb28e157e2a6ac3ea9b2927cbf7eecff4db9170013a

Observation 197bb3a4-9682-4dd5-bc3c-dd00706898e0 · inbound

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning cites this paper.

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning Uncovering mesa-optimization algorithms in Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-15T12:09:16.468055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:09:16.468055Z digest=sha256:9d8c4a8ff6872218a292a66adfe22e95a429281c06a0d1f7f1238a73af5f1403

Observation df420f82-5bac-4bfe-85f2-65e78f58701d · inbound

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences cites this paper.

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences Uncovering mesa-optimization algorithms in Transformers

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:19:46.569906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T00:19:24.366922Z digest=sha256:9af89332f35ed2dacce993f80dc256dc68fb008d4e2012c61af183b4dbe36677

Observation 03a68202-8f3a-4669-b546-f82f5756354e · inbound

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention cites this paper.

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Uncovering mesa-optimization algorithms in Transformers

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:11.996270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T15:27:55.566795Z digest=sha256:d2922081975588cd9ba86b6596fdcd9a366461fce3ce7c827ed3586dcfc9f083

Observation 287d35b9-1503-4598-ac53-1920ee193c37 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Uncovering mesa-optimization algorithms in Transformers

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.026408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:04fae8edeebb96868b7f4b4a9926de89ec680dc193577de1df3027b48870243d

Observation a94d68b2-39ea-47a2-b82f-ffad56a48db6 · inbound

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention cites this paper.

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention Uncovering mesa-optimization algorithms in Transformers

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:09:23.215404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:08:18.768344Z digest=sha256:f80dc7e8ea1edf0498fc66a11430f45fcb3ef8f5d75126012906280edd946cb6

Observation 171ea9ba-02c9-4f2f-931a-2a8aba847e8f · inbound

Dynamic Short Convolutions Improve Transformers cites this paper.

Dynamic Short Convolutions Improve Transformers Uncovering mesa-optimization algorithms in Transformers

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:26.970273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:48:50.103004Z digest=sha256:baed9f05b7b3405dc013cc38d6bf9f4d9c07a5d40c41e7f1e110098dd367ae8a

Observation a4675f7f-d2d6-4824-ba9d-83c82dcba78b · inbound

A Systematic Study of Behavioral Cloning for Scientific Data Annotation cites this paper.

A Systematic Study of Behavioral Cloning for Scientific Data Annotation Uncovering mesa-optimization algorithms in Transformers

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:38.998915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T16:23:08.402194Z digest=sha256:d077894d77e0b6573aa2b66492a228e94cea7b8671d07c9f5c664a9d13f1f175

Observation 37c49333-b27e-4e33-9633-e93a09d270ae · inbound

Blurry Window Attention cites this paper.

Blurry Window Attention Uncovering mesa-optimization algorithms in Transformers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.996400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:43:34.429061Z digest=sha256:a66c2d7713eba1c54882e018e4fc237dccae728e7575ef13559737d989d329d3

Observation 5da9a9d6-6222-42e3-93d6-db9447c0916a · inbound

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence cites this paper.

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence Uncovering mesa-optimization algorithms in Transformers

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:13.105960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:13:40.546738Z digest=sha256:e909e32a4a2b5dabdf22f3151fa579fd0271e5b49887fb7b19ec084fe67cd4be

Observation 105d4032-d5cf-4cb6-9aad-7d3dac356f7a · inbound

Understanding Large Language Models cites this paper.

Understanding Large Language Models Uncovering mesa-optimization algorithms in Transformers

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:56.433366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T12:53:00.254754Z digest=sha256:f93cdfdc38fecdb00f3a5e0e44ea5efe562c3a4026e287856088216ab221868b

Observation 18d52398-c2af-4d83-a126-10ba7f2d07c4 · inbound

Induction Heads Interpolate N-Grams cites this paper.

Induction Heads Interpolate N-Grams Uncovering mesa-optimization algorithms in Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T06:58:39.213932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:58:39.213932Z digest=sha256:346bf15b0729bf255821ec00c0bfce0697d831f76ad4d3d0da3b04d7a7e292a2

Observation 11b9ee44-bb4f-4e4b-a8a4-ba0b103da30e · inbound

The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory cites this paper.

The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory Uncovering mesa-optimization algorithms in Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T09:06:10.979909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:06:10.979909Z digest=sha256:58eb2f908d342c2c295f2ecd3b664d6d9cb1f95ed2de8dc858dfddb232d20c18

Observation 536edb80-fd75-4937-b225-6627b360b02f · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Uncovering mesa-optimization algorithms in Transformers

Reference 192

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:06.172263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:06.172263Z digest=sha256:47e0e4f7db342e0fd0ab46a53271eea0c833f10fe52b34ebc6692e1361bdbbfb

Observation 9cbabda1-561f-4c03-8b92-d818308ce118 · inbound

Grounding latent algorithm routing in transformer reasoning cites this paper.

Grounding latent algorithm routing in transformer reasoning Uncovering mesa-optimization algorithms in Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T14:04:25.225241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:04:25.225241Z digest=sha256:718eb269f65a7af3a77829ac96ca6fdf8177dc75aa87c1b2d2972af35f7ed405