Pith. sign in

Paper Citation Record · LEDGER

Transformers learn in-context by gradient descent

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2212.07677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.07677 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:06:05.127999Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:50.967150Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32247463-6fcb-4843-910b-dbbee5c8c1cb · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks Transformers learn in-context by gradient descent

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:17:26.845013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:80b36f30d8637ff329f2ae42395b670141c50272bb265d6d108e78c9830ebb07

Observation 6f4b0280-b613-4a58-8e59-968a275442a1 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Transformers learn in-context by gradient descent

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.266138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:88db9d47762799780c55030382947570700b488d5e1fe256937052445ded7528

Observation cb185528-a245-4a6d-874a-9e9490c651b1 · inbound

Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective cites this paper.

Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective Transformers learn in-context by gradient descent

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T14:20:23.787636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:20:23.787636Z digest=sha256:028be76928ef122b564312f04691afe4823676abee7fc1e6ff1bdf35152ffc13

Observation 8d5df9e5-15db-415a-b534-f6ba0ec63fa0 · inbound

In-context learning for medical image segmentation cites this paper.

In-context learning for medical image segmentation Transformers learn in-context by gradient descent

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:18:09.773508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:18:09.773508Z digest=sha256:17e72470900e7ae06ac2062b9e19faa3ab3da04123558692859685a10aff5402

Observation 25c9cb9a-76ff-478a-af77-ef1deba04bb8 · inbound

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit cites this paper.

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit Transformers learn in-context by gradient descent

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.417098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.417098Z digest=sha256:3d412c5153985e189fb0452f74dc3f63d15d866882cbebcbcfca3107b6b6e810

Observation 40568d91-16b8-44a8-973a-598430fd89e7 · inbound

Scaling sparse feature circuit finding for in-context learning cites this paper.

Scaling sparse feature circuit finding for in-context learning Transformers learn in-context by gradient descent

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:06:05.127999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:06:05.127999Z digest=sha256:98907f0269245c6e2d19e299aa0fa61e0962bf8c02ba901d220e755e683bbf64

Observation 3bbc278e-c836-42ff-85b2-191fca430072 · inbound

ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers cites this paper.

ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers Transformers learn in-context by gradient descent

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:34.073164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:58:34.073164Z digest=sha256:3b6feb654f61c85957e24d0d5f41548d838abb56eb087ad6af567b1c2b460e32

Observation 609e8753-af36-4bf4-9a0c-47b8338a16eb · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Transformers learn in-context by gradient descent

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:38.610485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:38.610485Z digest=sha256:686abff658d39fc9cf3b4675156fc49971d3d94667721f151417f6724b5986b1

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:26d239cb2a6f860087a9aafeececcf12f61b118dc7728fdefc5c626f87746ed8

Observation 3a58dcd2-8f83-4ea3-a2c8-9e9d3f0984ed · inbound

In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly cites this paper.

In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly Transformers learn in-context by gradient descent

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:42:34.508285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:42:34.508285Z digest=sha256:bcd4b530abd19c54a70871c5923c9565b6c80e21c98a4830638fe63ba54c0292

Observation f01fd6a8-5cee-4206-817b-028894d7fec8 · inbound

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models cites this paper.

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models Transformers learn in-context by gradient descent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:16.278676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:16.278676Z digest=sha256:b6f4bb9c823065b2c99d57c624bb40308b6bb249a11c4487d4bdbc854268e248

Observation 2bed3420-686e-41eb-bf14-6f62c9bffe54 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Transformers learn in-context by gradient descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:12.669528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:12.669528Z digest=sha256:09481e77ba71a736b8ae326e89c79dd0e2b32c65c21c3223a5540eaf52666884

Observation 5f09dc21-10d2-4352-9f20-797acd8af406 · inbound

When Context Sticks: Studying Interference in In-Context Learning cites this paper.

When Context Sticks: Studying Interference in In-Context Learning Transformers learn in-context by gradient descent

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:09.553762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T08:31:14.231710Z digest=sha256:1f93ce3bb19d37d6f0446851b6f8176d934a1f1f002466d3536eb49d5c24bac6

Observation 585e48d1-f584-4eee-b9bb-fcb131af08d1 · inbound

SMolLM: Small Language Models Learn Small Molecular Grammar cites this paper.

SMolLM: Small Language Models Learn Small Molecular Grammar Transformers learn in-context by gradient descent

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:16.906551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T12:57:47.721361Z digest=sha256:ccdcff6e1950804f5103e393b8969769dde8380807ac2098b644c1a101546a5d

Observation 64e90ae5-66f6-45ca-ab58-f592057636a1 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Transformers learn in-context by gradient descent

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.950075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:40b39fc18f5b72dbdc7270ace214813fa1db8e17bfd0429105be3d6ab33f5922

Observation 99d9c37f-1c2f-4824-b922-a66497b167a9 · inbound

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models cites this paper.

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models Transformers learn in-context by gradient descent

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.778045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T19:07:54.236182Z digest=sha256:4031abc57ab588261760538ab3fe5d5f69ae4f4149942d952aa81d2daa3b4593

Observation 74199b95-59dc-4b8c-90ed-120a5b46db56 · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction Transformers learn in-context by gradient descent

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:50.968747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:512d926a75b6788fde5008f7e6a9382779c5a2231ef2ac7121d323f27ff5b945

Observation 91c15dc8-0d55-435c-86b5-a4519d575197 · inbound

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models cites this paper.

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Transformers learn in-context by gradient descent

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:11.269139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:11.269139Z digest=sha256:000e6806c11c495b9041daad16963573fab1613057f6100f3f477a12f1601476

Observation 937cce38-7c21-47f1-8935-1769e2e8f247 · inbound

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention cites this paper.

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention Transformers learn in-context by gradient descent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:19:54.206559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:19:54.206559Z digest=sha256:b18899a4eb5faefe9698478ad75ea9514ce5d72831413f072b4d608b2af5f98a

Observation b6f33ba5-fc60-4e30-8f39-4d21678e2008 · inbound

Bayesian Wind Tunnels for Model Selection cites this paper.

Bayesian Wind Tunnels for Model Selection Transformers learn in-context by gradient descent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:02.240344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:19:02.240344Z digest=sha256:46fe8e41bca910c4626625238c716a5e595e0e8255dd51ad16b8e8ec17324e56

Observation fb9d9422-1de6-444c-adaa-25b1bd307e80 · inbound

Context-Adaptive Inference: A Unified Statistical and Foundation-Model View cites this paper.

Context-Adaptive Inference: A Unified Statistical and Foundation-Model View Transformers learn in-context by gradient descent

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T23:53:01.264186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:53:01.264186Z digest=sha256:fcea485ce08c5326eddc44f2a5327cf8b4dc24f7814e73b1301be792e3db085a

Observation 90560fd6-493c-48fc-92a2-bc098899d474 · inbound

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners cites this paper.

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners Transformers learn in-context by gradient descent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:16:27.017180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:16:27.017180Z digest=sha256:17be1844ed299ae2eb444aa404e7452398d5461b729d334b41f61a3f87a07a9d

Observation daf18c9b-1bbb-450d-88d6-8fd8d53d52c1 · inbound

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL cites this paper.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Transformers learn in-context by gradient descent

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.524771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.524771Z digest=sha256:d9ca07cb4d2398e0adf5bc71ad7f505510bbe5243069a481018a948541ca34f5