Pith. sign in

Paper Citation Record · LEDGER

Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2212.10559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.10559 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:21:58.250300Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

25
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 50dbbee2-35d4-437d-9146-bec7b25b55bf · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:17:26.697618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:4c8c1cfee00e1e491e6cf58b03faa1403d928b733f96dc4d8b52ab8314fb84f6

Observation 63fd8123-6f64-4735-bd67-d11d758eb59e · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:46:40.660977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:7a17d51b3ac6cd834098ddcef7ab2a914515b78701d177dbdcd58b192de86f88

Observation 15ae8f89-7997-4cf9-9e3c-183d48dbacd2 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.829878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:e5d68bb3bb6ab9453592dd20f3e2501e707dd2dea979341a54114c028ddf626e

Observation 5a6b0803-177a-4485-a7ac-de0eda355bb2 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.422912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:fc8087747858432489879e657e6d29893fc1f21c98ee04a8cce1f23624c3a394

Observation c806e01e-2fba-43c1-bed1-771b9d5298b4 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:26:06.309911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:90315e9f87164f4ff69ced688e43bc39acc53827a81868bc33f3d381d79c6759

Observation 3da142e9-65e9-47e2-99e6-72d879ce859f · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.250300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.250300Z digest=sha256:fed42c3cd40f8e4b7ae45ce8c7d84a7bd81055994e23827a96e0cd083300a244

Observation 1b2c741d-5e3f-4b31-8fc0-a82672e51db4 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:24.005948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:5800bf2abac07e5fdf3b84e319f5c056ac2fc40aa4b63153a07d1b6e377a62e8

Observation f2a18f1c-ddf2-4360-96b5-16cbb6868900 · inbound

The Prompt is Mightier than the Example cites this paper.

The Prompt is Mightier than the Example Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:10.101025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:10.101025Z digest=sha256:285717b10efab27c614dd327d7a63eb9bece6293209c521bb12ff4ac7e748017

Observation 3b5f9650-bca0-4ed6-b0b5-4582b762fb54 · inbound

Optimization-Inspired Few-Shot Adaptation for Large Language Models cites this paper.

Optimization-Inspired Few-Shot Adaptation for Large Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:28.780094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:28.780094Z digest=sha256:e0cd685c9dfaf3441e07743887f80c205eb76dba409d41778874fbba6f613e3e

Observation 37303a8d-8c66-443b-b845-98f7ab01a44b · inbound

Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models cites this paper.

Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.806691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:40:35.806691Z digest=sha256:4e0acd4bc93ecf4473aee4be020be884fa0d8e3c20c270fe6c12746ecdb2965e

Observation a3cea4c2-1328-4005-bd43-bd54952659da · inbound

Adaptive Task Vectors for Large Language Models cites this paper.

Adaptive Task Vectors for Large Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:00.986510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:00.986510Z digest=sha256:ac5e0e3201953f901dd65c61fe6a472f8f2db7ad417c133e7c347b08ec93be16

Observation bc3b31fb-3efc-45a2-9ef0-166de5dcd8aa · inbound

ConText: Driving In-context Learning for Text Removal and Segmentation cites this paper.

ConText: Driving In-context Learning for Text Removal and Segmentation Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:08.177043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:01:08.177043Z digest=sha256:b2cfefa74351e8a0cbd202abc5f0585c33f0d9d8a04945bfbf220e96c4cbedbd

Observation f530a241-cb5d-45f9-98f0-db8dbcbee486 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:36.244991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:36.244991Z digest=sha256:ef0c794f07b52c975067c09130fe290be97fc08e37732ff919f6417c394d4ef2

Observation 189287f6-70ad-4f4b-a374-df3b7489b3b9 · inbound

Prompting Wireless Networks: Reinforced In-Context Learning for Power Control cites this paper.

Prompting Wireless Networks: Reinforced In-Context Learning for Power Control Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:02.033815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:02.033815Z digest=sha256:df7b9c396e83d31eeb8b605e49e834875724461d5734eb069126c3cd8d4ab5b2

Observation 85552c65-25d6-4d33-ba2c-693c03187b7c · inbound

Can Gradient Descent Simulate Prompting? cites this paper.

Can Gradient Descent Simulate Prompting? Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:50.199758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:41:50.199758Z digest=sha256:1d9230f9956ee4455eb359aaebd2467e30025d4e73c8dc10934521ea7911d8d0

Observation 9e863460-63fb-42d3-b1ed-6a727aab1fb1 · inbound

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models cites this paper.

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:16.429781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:16.429781Z digest=sha256:352e8d712eb5c8ca6244bbb01c252d2dea5df284f8a8d59034d436ef3699b017

Observation f33f08b3-2cb9-41e1-ab57-7d174a2ed617 · inbound

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning cites this paper.

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:35.454020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:35.454020Z digest=sha256:8a2a37508e914d5c9502c0d49ca6b460c20386a4f393c99a239906cae318ee18

Observation deca8069-da0c-437e-813b-8acbc46c6ab6 · inbound

Provable Low-Frequency Bias of In-Context Learning of Representations cites this paper.

Provable Low-Frequency Bias of In-Context Learning of Representations Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:02.712441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:02.712441Z digest=sha256:93123309500bb0f5eeede0aaaf98c9ff932e8c2257a969c69033387d6de4d6f5

Observation 4416e9c6-7f19-4c47-a553-1d75d4f02152 · inbound

FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design cites this paper.

FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:56.605376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:49:56.605376Z digest=sha256:592bea821c65c2373b1cec08369d46036a09b3f428440a855cbb1235b93e0589

Observation ba35a10a-481a-42e2-9508-fdf5e60b80a3 · inbound

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention cites this paper.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.436932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.436932Z digest=sha256:ec2fde018ae8f0634fc76605d5ed0e6bcee8e28d8a9677d70a9d23a17de6cbeb

Observation 035df581-eae6-4b96-a93a-b08e5d629744 · inbound

Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding cites this paper.

Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.798403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:42:33.975433Z digest=sha256:06ca3ed4f59cc036037d0090268cb0ce5872231a7718ba5f3bc303edb65df279

Observation 454ea40e-9d3f-4977-ae76-bf7186f1293e · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.215723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:0012bb8fd7f83761fcd86cd1dac4498658188572537a0bc154e71647d37eff44

Observation 09dcf600-7f5c-4d2d-979b-a4de53e955d8 · inbound

When Context Sticks: Studying Interference in In-Context Learning cites this paper.

When Context Sticks: Studying Interference in In-Context Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:09.472739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:31:14.231710Z digest=sha256:1c2f050e3baa09c41ec1f4bc37af2c048f6bec30c993ecc03be8d00c95964616

Observation 9ea03faa-0a0d-4735-bbe3-5912714f6a17 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:32.066177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:730602d8ceae97b4553458b6f8360e84a3beaba8a2de8b6bc2dbad6660e2810f

Observation 5d61b067-8c82-4d16-8790-406bdc1b0a09 · inbound

Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm cites this paper.

Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:25.555016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:47:54.466097Z digest=sha256:57eebecae3e7a08006673ce3db5c8817715782297f1c30135e85cae85c678f43

Observation 884600e5-edd9-4d87-99db-7326b94408fd · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.289708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:8ab7fdd7d9ed526483480b139e74c235a0f3384dc6b0eb94eb2a2890a0cafddf

Observation 52b36e7f-c9cd-4886-878b-1e1dd404afde · inbound

A Human-in-the-Loop Framework for Efficient Prompt Selection in Microscopy Vision-Language Models cites this paper.

A Human-in-the-Loop Framework for Efficient Prompt Selection in Microscopy Vision-Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.436740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:05:19.381117Z digest=sha256:41181b5dcb5401cf23d861801864417a055ffe3ebfc9a17529455906b08aeead

Observation e2aac111-2d4b-43a6-b390-6a6c03f2c5a9 · inbound

Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis cites this paper.

Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:43.288286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T17:34:37.552706Z digest=sha256:9120397ca6716b28f63271c436f5505fb9922991019bd8795ce958d6a32b7051

Observation 9826500b-d439-411f-a5ba-babb80ea6390 · inbound

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention cites this paper.

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:19:53.717359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:19:53.717359Z digest=sha256:e275abcb5d16d21a95693a4565a85b0cc9a073bfc32b9db11bb71a07cb3e042c

Observation c62be06b-bd3c-4f73-b347-59328c7af37d · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:33.862109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:33.862109Z digest=sha256:e256162de2971ee4cedab712e4095ecdff0865a83436faca39b4aab8ffb86afe