Pith. sign in

Paper Citation Record · LEDGER

A Survey of In-Context Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 14 inbound Pith citation observations for arXiv:2502.07978.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07978 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:17:16.825411Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:52:47.188471Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T13:44:40.651347Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy66
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5142e2b-79b4-4023-835a-04a2f51217fd · outbound

This paper cites an unresolved cited work.

A Survey of In-Context Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:17:17.608599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.582744Z digest=sha256:94be155167fe7546520ca9b0238227aa744ac1c0dd02e2b4f07161c66fcd4ee4

Observation 156d4b8b-1571-40dd-97cf-3f46d689df32 · outbound

This paper cites In- Context Language Learning : Architectures and Algorithms.

A Survey of In-Context Reinforcement Learning In- Context Language Learning : Architectures and Algorithms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.598129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.586650Z digest=sha256:0073d229564baa18c3117a844c168050a8493cd14ced822797fde825a01de182

Observation 799bf4c8-e3c5-4242-b7c4-0dd23a3b8e02 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

A Survey of In-Context Reinforcement Learning Minimax regret bounds for reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.588089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.589773Z digest=sha256:f36391d4495337315f9f016e0cd683f4c2dc652e32f274bcae5c8d08cdb719f1

Observation 2e310bd5-15f1-4967-9000-4bb11b482cc2 · outbound

This paper cites Provable self-play algorithms for competitive reinforcement learning.

A Survey of In-Context Reinforcement Learning Provable self-play algorithms for competitive reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.577408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.593242Z digest=sha256:cdfbe74b3b033d6aa034edaea6802ada2939eb952a292561da1c738f5bc10690

Observation d9e53b31-2c45-4294-83a2-4b26e5035c4e · outbound

This paper cites an unresolved cited work.

A Survey of In-Context Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:17:17.567437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.596014Z digest=sha256:b406f49861f695428b97758d766c08c703b20b6d22dfde5328fa6871881fa4de

Observation 03e198c8-6fb2-4d68-a910-2bdc8b0bff22 · outbound

This paper cites A Survey of Meta-Reinforcement Learning.

A Survey of In-Context Reinforcement Learning A Survey of Meta-Reinforcement Learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.557401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.598876Z digest=sha256:9b69f967962600e04e7b57735113fe35126f33d596efb48edbe0c088ce10ba65

Observation c1b9a8b9-08cf-4526-976c-15467d361668 · outbound

This paper cites o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \.

A Survey of In-Context Reinforcement Learning o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.547479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.602466Z digest=sha256:5ad113a2a0e544ebda4a609e3e2cea5cc99722096318c312a13d5e5471fc71b3

Observation cb6f6179-b361-4a24-aca9-f6f244d6064c · outbound

This paper cites When does return-conditioned supervised learning work for offline reinforcement learning? In Advances in Neural Information Processing Systems , 2022.

A Survey of In-Context Reinforcement Learning When does return-conditioned supervised learning work for offline reinforcement learning? In Advances in Neural Information Processing Systems , 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.538507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.606335Z digest=sha256:83787bcc5833fbd0e360dd7ccaee5fa5929f3a79227033bd27071fd14927ffd5

Observation 3805d8f6-5833-4cb5-a6c9-7073e8d28922 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

A Survey of In-Context Reinforcement Learning RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T11:17:16.609545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:17:16.609545Z digest=sha256:df48116489b3fcc812f8430163d903ef91787083403d6584bba2feae18a2c0a7

Observation f15eb66e-cd5c-4f10-afe8-9f622dedc18d · outbound

This paper cites Learning to Cooperate with Unseen Agent via Meta-Reinforcement Learning.

A Survey of In-Context Reinforcement Learning Learning to Cooperate with Unseen Agent via Meta-Reinforcement Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-08T11:17:16.855399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.615009Z digest=sha256:d314b66b7113d0d596738f9b495191a235dc655580b6bf1ac1af85679657554a

Observation 0db95420-7735-4e59-a0b9-05e828335e27 · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

A Survey of In-Context Reinforcement Learning Decision transformer: Reinforcement learning via sequence modeling

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.529671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.619692Z digest=sha256:56d660799e2c3e67af1fe57396e47588d3c2437caf65dab75fceac6a40432ca2

Observation 1cbe1c46-8225-420c-ad68-d665a0b170b0 · outbound

This paper cites Contextual bandits with linear payoff functions.

A Survey of In-Context Reinforcement Learning Contextual bandits with linear payoff functions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.520891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.623037Z digest=sha256:637039cb6000aa25fa20f1262d275a7abe9068c0e6dee35ac33b407e2f0c74ab

Observation d187e1c2-61b5-4c60-b65a-6554c4e286b7 · outbound

This paper cites Leveraging procedural generation to benchmark reinforcement learning.

A Survey of In-Context Reinforcement Learning Leveraging procedural generation to benchmark reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.510282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.626980Z digest=sha256:2383b50be6007ad6510f135b366df145dda42e949ff83084a4f140b574c1bd67

Observation 30da6640-49e4-4136-beae-e4035620846f · outbound

This paper cites Leibo, and Jakob Nicolaus Foerster.

A Survey of In-Context Reinforcement Learning Leibo, and Jakob Nicolaus Foerster

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.501022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.630776Z digest=sha256:e2538f24a55bcf71e14ac7fc1e02ffcac9eb4f75d690e4b288d7a109fc16f1ee

Observation 9a0788b5-0b3c-4ee5-a2e2-89bb2bbdac68 · outbound

This paper cites In-context Exploration-Exploitation for Reinforcement Learning.

A Survey of In-Context Reinforcement Learning In-context Exploration-Exploitation for Reinforcement Learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.491430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.634649Z digest=sha256:20c7a14abf6878e8578790c07221ba4799ff1e2dd64dac434ad58ca43e3fbfbb

Observation f7aa4d54-b25f-470d-9158-cddeb14186d8 · outbound

This paper cites A Survey on In-context Learning.

A Survey of In-Context Reinforcement Learning A Survey on In-context Learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.480905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.638723Z digest=sha256:26c3358d0130b191bfc3207567527f391dfef423608156a15176eb1dca0a65d5

Observation 3390d03f-44d1-4113-a98e-6dd500d0d328 · outbound

This paper cites Fang, Zhuoran Yang, and Vahid Tarokh.

A Survey of In-Context Reinforcement Learning Fang, Zhuoran Yang, and Vahid Tarokh

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.471641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.642896Z digest=sha256:7f16bd9a51b5bfdcf421fe14ee88ecd9f9241dfd7b1db91e79412df14463d30f

Observation 229b1152-d29f-44f1-80b2-b77206c00638 · outbound

This paper cites Bartlett, Ilya Sutskever, and Pieter Abbeel.

A Survey of In-Context Reinforcement Learning Bartlett, Ilya Sutskever, and Pieter Abbeel

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.461704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.647211Z digest=sha256:3001ca1454ab8ae02748da782757f44cb1b812eec50b4e6710761b4f2147c174

Observation 1dfd90c9-a79e-411f-9e1a-586778f4dd26 · outbound

This paper cites ReLIC : A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI.

A Survey of In-Context Reinforcement Learning ReLIC : A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.450960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.651122Z digest=sha256:8a10ca961c7dee6d16e26d6bcb98ff4c0cdca11f276189217f715f9b708213f4

Observation 3362b539-4485-4c12-a88b-d3e5046b1432 · outbound

This paper cites Rvs: What is essential for offline RL via supervised learning? In Proceedings of the International Conference on Learning Representations , 2022.

A Survey of In-Context Reinforcement Learning Rvs: What is essential for offline RL via supervised learning? In Proceedings of the International Conference on Learning Representations , 2022

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.440145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.655750Z digest=sha256:cb9b330f967320af5c8505aee2df314ce4a27d692c12e0b5239283e73df88968

Observation 818dc4f3-8438-478f-8349-4d4e360166ea · outbound

This paper cites Generalized decision transformer for offline hindsight information matching.

A Survey of In-Context Reinforcement Learning Generalized decision transformer for offline hindsight information matching

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.429363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.659120Z digest=sha256:aa27faafc49752958d91171b717a8f2b0b2d4f71b45ba845f838568db9d0253a

Observation 0e0ff42a-1014-402f-b26e-4fd4d17a99fa · outbound

This paper cites Meta-rl for multi-agent rl: Learning to adapt to evolving agents.

A Survey of In-Context Reinforcement Learning Meta-rl for multi-agent rl: Learning to adapt to evolving agents

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.419316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.662842Z digest=sha256:7bb1d843d6db73906f86a69b8222817670188dab62af3f613ea053a151ebfbee

Observation 59227d69-e035-421a-9cb4-0240c35415b5 · outbound

This paper cites AMAGO : Scalable In-Context Reinforcement Learning for Adaptive Agents.

A Survey of In-Context Reinforcement Learning AMAGO : Scalable In-Context Reinforcement Learning for Adaptive Agents

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.408938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.667117Z digest=sha256:dc67744fb60766d65ded278f8954b12b0814a8c04fe7233ad99bf3c3d4a7c8d0

Observation 296fcb9c-ade6-4fe9-830b-6a3743b85c22 · outbound

This paper cites AMAGO-2 : Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers.

A Survey of In-Context Reinforcement Learning AMAGO-2 : Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.399788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.670757Z digest=sha256:1d46396f69f17ced9a689e1901f32c8a7e4c322a150b7a35303ba7d72b851a13

Observation 3b40f51c-d553-4d8f-843f-703ce0f32da5 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

A Survey of In-Context Reinforcement Learning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.390380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.674363Z digest=sha256:0f35213e1582f5d926de13a2328e9278f93a6b9806f0c6afab5f9a46de9474f3

Observation 8a9be31d-5da3-4e15-93b3-18e3be91ee68 · outbound

This paper cites Muesli: Combining improvements in policy optimization.

A Survey of In-Context Reinforcement Learning Muesli: Combining improvements in policy optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.381078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.677706Z digest=sha256:238c1d16c1a66b67d3a65ae7bdfb7808b89bb5277e94e64d97d0eb4faa932221

Observation 74cedd12-82e6-4c2f-a1fd-88accccdba19 · outbound

This paper cites In- Context Decision Transformer : Reinforcement Learning via Hierarchical Chain-of-Thought.

A Survey of In-Context Reinforcement Learning In- Context Decision Transformer : Reinforcement Learning via Hierarchical Chain-of-Thought

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.371898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.680398Z digest=sha256:732717b469567fd30381befa8f0e302ee00d2a677ced886477b895ca9e38344f

Observation 029964cc-6c59-48b2-8703-99d1423ef9f6 · outbound

This paper cites Decision Mamba : Reinforcement Learning via Hybrid Selective Sequence Modeling.

A Survey of In-Context Reinforcement Learning Decision Mamba : Reinforcement Learning via Hybrid Selective Sequence Modeling

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.362073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.683046Z digest=sha256:a644d3d41d606f5a4e41cf1907cd5552855f900c840a10f9837c52a47150d8ae

Observation 31b4c2a6-a1f0-433b-b38b-99357e9a7a5e · outbound

This paper cites V-learning—a simple, efficient, decentralized algorithm for multiagent reinforcement learning.

A Survey of In-Context Reinforcement Learning V-learning—a simple, efficient, decentralized algorithm for multiagent reinforcement learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.350883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.685609Z digest=sha256:c51bcddbb8fc107cd2647f11af8fb03259c5ebae7f36da906764fa015055f1cf

Observation 8cbbe734-8232-489a-a856-c9f1b301a60d · outbound

This paper cites Kingma and Jimmy Ba.

A Survey of In-Context Reinforcement Learning Kingma and Jimmy Ba

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.339791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.688085Z digest=sha256:e4e4a32e2052b1a943c80d048b573319b79885c375d898b0a3ac02928b599778

Observation f9528815-93e5-4853-9ea4-cc79fee72be4 · outbound

This paper cites A Survey of Zero-shot Generalisation in Deep Reinforcement Learning.

A Survey of In-Context Reinforcement Learning A Survey of Zero-shot Generalisation in Deep Reinforcement Learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.330386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.690670Z digest=sha256:44fdd01334ecf6dddbab9b7a66a1cff3047976b216bb8fd45f133a45ba7c8ec5

Observation 3f9400ac-8e29-4650-a793-ce5c449addfb · outbound

This paper cites Daniel Freeman, Jascha Sohl-Dickstein , and J \"u rgen Schmidhuber.

A Survey of In-Context Reinforcement Learning Daniel Freeman, Jascha Sohl-Dickstein , and J \"u rgen Schmidhuber

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.319455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.693165Z digest=sha256:a9bb656e4f7b523e0840153fd4efd6b209934b11fb1d8d0e429aa3af1745667f

Observation 4b8fee5b-b0fb-4e4e-9ad3-c385c518b69c · outbound

This paper cites Foster, Cyril Zhang, and Aleksandrs Slivkins.

A Survey of In-Context Reinforcement Learning Foster, Cyril Zhang, and Aleksandrs Slivkins

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.308851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.696034Z digest=sha256:04b484c8aea78089f5f5bec2b88aa95401759913755a264ed865294fd93f1959

Observation 31167b1f-b57d-4920-8f47-6d85d28b2cc3 · outbound

This paper cites Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh, and Volodymyr Mnih.

A Survey of In-Context Reinforcement Learning Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh, and Volodymyr Mnih

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.298193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.698567Z digest=sha256:b7c6ee050cc223fec232dec2761ca5ef22ff0fd8de81282d16e8783bb135bb1a

Observation feca94a5-1738-4888-8dd1-aa806819c85a · outbound

This paper cites Supervised pretraining can learn in-context reinforcement learning.

A Survey of In-Context Reinforcement Learning Supervised pretraining can learn in-context reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.287849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.701150Z digest=sha256:ad04cc99076577a49831182fa539bee2d1f658d516ec84df138b1aee91db56de

Observation d4ba52b1-8f3e-43cf-be8e-3d2a83d33354 · outbound

This paper cites Lillicrap, Jonathan J.

A Survey of In-Context Reinforcement Learning Lillicrap, Jonathan J

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.278283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.704863Z digest=sha256:cb55ef9ee8f685d3cc9de767e942b7ca7ca063bab5657af64092a94fee5a2000

Observation b06f1201-dcba-45c8-9488-df751d906ab0 · outbound

This paper cites Transformers as Decision Makers : Provable In-Context Reinforcement Learning via Supervised Pretraining.

A Survey of In-Context Reinforcement Learning Transformers as Decision Makers : Provable In-Context Reinforcement Learning via Supervised Pretraining

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.269470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.707316Z digest=sha256:77cb8eed57537b8840c3b931475a723b310084a42c8d936dc13af87086cb4fe3

Observation 93520946-1586-4a17-888d-5d13126c56a9 · outbound

This paper cites Emergent agentic transformer from chain of hindsight experience.

A Survey of In-Context Reinforcement Learning Emergent agentic transformer from chain of hindsight experience

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.259666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.709906Z digest=sha256:a4e9e4bd536d6435212f7a40c95c20dc75223275002e1b8ecc21395ae5eb1fbc

Observation 18d587dc-5e56-412f-a3d6-715ff8f8c9a7 · outbound

This paper cites Goal-conditioned reinforcement learning: Problems and solutions.

A Survey of In-Context Reinforcement Learning Goal-conditioned reinforcement learning: Problems and solutions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.250369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.712913Z digest=sha256:91176ad6e8588a9d7eb92431a885cb82b022cfdf7cb26ff67d5c0430897b4c79

Observation 78e7fbde-bb21-4cb1-b0b9-7567d6ee983b · outbound

This paper cites Foerster, Satinder Singh, and Feryal M.

A Survey of In-Context Reinforcement Learning Foerster, Satinder Singh, and Feryal M

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.240733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.716408Z digest=sha256:d6ca4cca9f8aafc7bb2028427b9b6d6ec68be6a0aa649eba019a02e231c8ca71

Observation a9a13e95-e135-49f6-bcd2-b81221697025 · outbound

This paper cites an unresolved cited work.

A Survey of In-Context Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:17:17.231286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.718958Z digest=sha256:3cc6a8b854fdcfa76131e9fd61e98c6e90c72c5af8cddd9117c583e7f5fb5b5d

Observation f8ab1639-37e8-4c6f-a68b-5031b7938460 · outbound

This paper cites Offline pre-trained multi-agent decision transformer.

A Survey of In-Context Reinforcement Learning Offline pre-trained multi-agent decision transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.221828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.721675Z digest=sha256:b3eb64fffe6fc90e2e74beeb8239f548dd098cc732b4e2ba2726ee4d464d8e1b

Observation 1c5ae4ba-8210-4c11-a67c-80a5194594b2 · outbound

This paper cites A simple neural attentive meta-learner.

A Survey of In-Context Reinforcement Learning A simple neural attentive meta-learner

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.212032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.724595Z digest=sha256:252210ee7c8b4114ff858cf519752faf47e8dee9a8b7c37524334a0d35c4e92c

Observation 842631a6-c72c-4a2e-9085-2aaa58fed02c · outbound

This paper cites Human-level control through deep reinforcement learning.

A Survey of In-Context Reinforcement Learning Human-level control through deep reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.201207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.727671Z digest=sha256:bd6129cad457c4723a22ad4ec122ce37c0255ec56b9e123516b76d51b003a00a

Observation 6e5072ed-4b60-4e8a-838d-21b06a543f30 · outbound

This paper cites First- Explore , then Exploit : Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs.

A Survey of In-Context Reinforcement Learning First- Explore , then Exploit : Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.191020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.730174Z digest=sha256:8fa3f9ce970928cb486d1a1f2cd5b47e0cde426b0cacb8f5a5f2289a6f36f3e9

Observation f37dda9c-f985-4f68-ad9d-2940a7640f1d · outbound

This paper cites Do LLM Agents Have Regret ? A Case Study in Online Learning and Games.

A Survey of In-Context Reinforcement Learning Do LLM Agents Have Regret ? A Case Study in Online Learning and Games

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.180707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.733241Z digest=sha256:888656b47be89a8f247da402cf1b60af62a3ee84ea4f8e0964c35f3a28a00ee8

Observation e25b7145-2082-4cef-b396-baddc3af9ecb · outbound

This paper cites Efficient off-policy meta-reinforcement learning via probabilistic context variables.

A Survey of In-Context Reinforcement Learning Efficient off-policy meta-reinforcement learning via probabilistic context variables

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.169543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.736892Z digest=sha256:8b49d58c5ae70d332ff31347d124c53f9bf41a66208f6360ab839d758c1aea7e

Observation 060223cc-bdbd-4bc5-9e8e-6455752ee667 · outbound

This paper cites Generalization to New Sequential Decision Making Tasks with In-Context Learning.

A Survey of In-Context Reinforcement Learning Generalization to New Sequential Decision Making Tasks with In-Context Learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.159843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.740536Z digest=sha256:d060f47c846d0b34815c23631c436ef50426b174bdbd76b7d1d3401dacda4abc

Observation 7ffc8df1-8f57-44da-a1bf-dae1aed80783 · outbound

This paper cites Wang, Zeb Kurth - Nelson, Siddhant M.

A Survey of In-Context Reinforcement Learning Wang, Zeb Kurth - Nelson, Siddhant M

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.149786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.743854Z digest=sha256:aa1dbd212fe823a90ff7830f48c0be31c790fe4f5b027dbec78895e7b38c910e

Observation e4234989-ec69-46d4-b00d-b9659e3dcc9f · outbound

This paper cites A tutorial on thompson sampling.

A Survey of In-Context Reinforcement Learning A tutorial on thompson sampling

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.139470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.747584Z digest=sha256:b590a821d3322a30ebbb9d396a527cc08ef4e33731684e7bf1ef2a153059c6d0

Observation a3985274-a346-4458-b189-339491a18d86 · outbound

This paper cites o ppel, Johannes Brandstetter, G \.

A Survey of In-Context Reinforcement Learning o ppel, Johannes Brandstetter, G \

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.130760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.750608Z digest=sha256:c86f66bb82b3365020fd9136b2bc4fd5bf7bac3bf3f78873f398ccea2061db1b

Observation 52d16628-b98b-46d3-9213-cc8d6a66baca · outbound

This paper cites Proximal policy optimization algorithms.

A Survey of In-Context Reinforcement Learning Proximal policy optimization algorithms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.122114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.754602Z digest=sha256:b2efcd74e72cb5b8b19bed1f052628d43ba987a00e8a907f811f7a43772c4e0a

Observation d99a9082-00de-4b49-b11f-bc53519165df · outbound

This paper cites A primal-dual perspective of online learning algorithms.

A Survey of In-Context Reinforcement Learning A primal-dual perspective of online learning algorithms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.111531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.758103Z digest=sha256:7a868ea8ac7e53cd14a13e0b2ac1c172fac68701d8eb61cf6370710f814763a3

Observation c1063b9e-4bde-4ae6-82f8-a06108fd5dd8 · outbound

This paper cites Cross-episodic curriculum for transformer agents.

A Survey of In-Context Reinforcement Learning Cross-episodic curriculum for transformer agents

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.101169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.761240Z digest=sha256:e9ac206d7ceddaaf6d5fc11c2fa0074c7f8aaaecd438ec3cf895ddf6300de4da

Observation 235deb61-b564-464a-aafa-16bdcb1e25e8 · outbound

This paper cites Transformers as Game Players : Provable In-context Game-playing Capabilities of Pre-trained Models.

A Survey of In-Context Reinforcement Learning Transformers as Game Players : Provable In-context Game-playing Capabilities of Pre-trained Models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.089859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.764507Z digest=sha256:dbf9138a769697cb563595f85d3c2c75ac5d407a66b1684b42e56c64401671cb

Observation 95a5558b-d3e7-4744-be73-6156a36606ca · outbound

This paper cites In- Context Reinforcement Learning for Variable Action Spaces.

A Survey of In-Context Reinforcement Learning In- Context Reinforcement Learning for Variable Action Spaces

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.078794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.767770Z digest=sha256:7a3c0a35a7ef8849bcf5f6935cd39cfe7d0906bf4f401c40d451cb8170fae736

Observation 77b5e56c-d43d-4687-8b2a-af2164a672c6 · outbound

This paper cites an unresolved cited work.

A Survey of In-Context Reinforcement Learning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:17:17.068149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.771229Z digest=sha256:9067b46d1439910c8c3172a026122762fb82a9e438ceb09cff7a743b784d77d0

Observation 09d39497-c5fa-40ba-aec0-2f3cc28a251c · outbound

This paper cites Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever.

A Survey of In-Context Reinforcement Learning Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.055910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.774526Z digest=sha256:dfe7431ac7428b8f733f5b0e53d848e585abcea9c904484d6a0f3fb867ead77a

Observation 5ceb6185-b8ab-4e50-a7cc-365651f8b86d · outbound

This paper cites Reinforcement Learning: An Introduction (2nd Edition).

A Survey of In-Context Reinforcement Learning Reinforcement Learning: An Introduction (2nd Edition)

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.042916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.777661Z digest=sha256:e62767d2130e918de9e53bc619d8736330b105094b08a441380ec23f1eec56ff

Observation 226e5757-ad40-4cc5-894c-288d4c33ad5d · outbound

This paper cites an unresolved cited work.

A Survey of In-Context Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:17:17.031817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.780792Z digest=sha256:8dabecc4acdb51091586ab44123d5acb63d0be11ca61ad654369d17906395b2c

Observation 2891a65e-027e-4964-8b55-05cd128c3343 · outbound

This paper cites Mujoco: A physics engine for model-based control.

A Survey of In-Context Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.021079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.783930Z digest=sha256:54d16d3e40f37eb204a7f5d2ddc9ace83d1ca4cf138bb28225abca0b11d63ae0

Observation 82c2755a-032d-4538-a535-19a161be12e9 · outbound

This paper cites Tsitsiklis and Benjamin Van Roy.

A Survey of In-Context Reinforcement Learning Tsitsiklis and Benjamin Van Roy

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:17.011052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.787602Z digest=sha256:514173c79c8a8cdf6fdfaf0789eb74dcc99abdadc08d412d995c2e897c2810a4

Observation 84a2644d-7168-4393-bcfa-ccf442a12bba · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

A Survey of In-Context Reinforcement Learning Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T11:17:16.790910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:17:16.790910Z digest=sha256:7553c710090e1247d43d652df2633a72858a06a7675230747342ba6f7f4e12cb

Observation 4794ef81-e0bb-4c3e-add4-c925b9ef8773 · outbound

This paper cites Wang, Zeb Kurth-Nelson , Dhruva Tirumala, Hubert Soyer, Joel Z.

A Survey of In-Context Reinforcement Learning Wang, Zeb Kurth-Nelson , Dhruva Tirumala, Hubert Soyer, Joel Z

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.993524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.793962Z digest=sha256:361c12c7a5cfe8e14ab79e420201a164899166151a84ef06a33566f15c420219

Observation 281bf302-ecac-481b-9828-ff0df44a7a72 · outbound

This paper cites Transformers Learn Temporal Difference Methods for In-Context Reinforcement Learning.

A Survey of In-Context Reinforcement Learning Transformers Learn Temporal Difference Methods for In-Context Reinforcement Learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.983690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.796930Z digest=sha256:e96675bcdc4020643720c6471dfbdf8cbb9e8a13a78b3ac762a16cbd6e92af0d

Observation 195591af-9331-4fb6-9745-9fabfa569286 · outbound

This paper cites Hierarchical Prompt Decision Transformer : Improving Few-Shot Policy Generalization with Global and Adaptive Guidance.

A Survey of In-Context Reinforcement Learning Hierarchical Prompt Decision Transformer : Improving Few-Shot Policy Generalization with Global and Adaptive Guidance

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.973320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.799992Z digest=sha256:ff0e8b4842ebc41fe3dbf746fe8aa87d8fe3721e86d8dab17c2ab806e9be09b0

Observation bfec3a4b-656f-4071-9126-7509ec8a0c10 · outbound

This paper cites Large Sequence Models for Sequential Decision-Making : A Survey.

A Survey of In-Context Reinforcement Learning Large Sequence Models for Sequential Decision-Making : A Survey

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.963309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.803045Z digest=sha256:07d0bacf6420c73da156263abe82caf8888664902c20227ce2a86abade349bf1

Observation 435ea215-b39f-41f1-853c-6112f6434591 · outbound

This paper cites Tenenbaum, and Chuang Gan.

A Survey of In-Context Reinforcement Learning Tenenbaum, and Chuang Gan

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.952484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.806279Z digest=sha256:0c095281bb586879a46bc0b153138fe6bb49dfe7cd0909f0a1758a5832bab7f8

Observation b94d2d2c-ef44-4cd2-8d88-c960a499e755 · outbound

This paper cites Meta- Reinforcement Learning Robust to Distributional Shift Via Performing Lifelong In-Context Learning.

A Survey of In-Context Reinforcement Learning Meta- Reinforcement Learning Robust to Distributional Shift Via Performing Lifelong In-Context Learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.939818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.808846Z digest=sha256:177b733d4e658a9ede0c5a9d3b1ab303de6c4ce08026eeaa7ffe4f4a60f258c3

Observation 3e2a8c8a-dad3-4c8d-8787-446bf36d3ccf · outbound

This paper cites Meta- World : A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

A Survey of In-Context Reinforcement Learning Meta- World : A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.928109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.811685Z digest=sha256:adc69806ed318cec9bcd87bfbb1dd38ad72ce351900beca23048065078bfaae2

Observation 97ca88a2-3481-4a08-b4c5-bd4b4ffbaa1e · outbound

This paper cites A survey on model compression for large language models.

A Survey of In-Context Reinforcement Learning A survey on model compression for large language models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.916455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.814691Z digest=sha256:761cd8079ab73cf6453dfa1d478d4020a3ce1205612d4ba8d72e7283ca049b93

Observation 5af23720-b35c-4605-bb0c-59d9d09147c3 · outbound

This paper cites Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson.

A Survey of In-Context Reinforcement Learning Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.904377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.817728Z digest=sha256:48ce2d144a3bf8134a905a258db86a64be261c0ecaf7118a8911ab43e27a3e99

Observation dfa01cda-48a1-4e2a-872c-1bf8f700b99f · outbound

This paper cites Emergence of In-Context Reinforcement Learning from Noise Distillation.

A Survey of In-Context Reinforcement Learning Emergence of In-Context Reinforcement Learning from Noise Distillation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.893835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.820301Z digest=sha256:c3f19e1b55a8c27ba89779fe19e48d84b6eddabbc60e8ba76ab69b69d03f5733

Observation 924b8fcc-6729-4f80-acab-4814a72f0e5d · outbound

This paper cites N- Gram Induction Heads for In-Context RL : Improving Stability and Reducing Data Needs.

A Survey of In-Context Reinforcement Learning N- Gram Induction Heads for In-Context RL : Improving Stability and Reducing Data Needs

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:17:16.883528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T11:17:16.822902Z digest=sha256:4de87a65a8dc27a0943623a71dcfa3a3caa4e06f3a3b4c3b3b97d6283d2e043b

Observation 5e73efff-96b8-4f8a-a685-b580af5a54d5 · outbound

This paper cites write newline.

A Survey of In-Context Reinforcement Learning write newline

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T11:17:16.825411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:17:16.825411Z digest=sha256:0e7769cb27a3b15443556f43f9a14176efc192f118f4394b47d61bf8def4c55c

Pith citing papers

Observation 1746c1c3-08e4-4fab-af78-76e71a11315e · inbound

Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks cites this paper.

Robust In-Context Reinforcement Learning Under Reward Poisoning Attacks A Survey of In-Context Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:52:47.188471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:52:47.188471Z digest=sha256:1f86781e3cc454c4e287308e31241e94be07714db8aeebdbe6085257ce6b11ae

Observation 981418f6-b0d5-49b8-8672-c37bdd9ff3ec · inbound

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra cites this paper.

LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra A Survey of In-Context Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:28:45.819093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:28:45.819093Z digest=sha256:0cf554706960dcb5c2ead88b977ae20968694bf7ba6e78cae6b2625cca8a254a

Observation d519dff5-6433-4b44-819f-3357bb3c41ec · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence A Survey of In-Context Reinforcement Learning

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.883101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:064d3da2f5778b004772d4f5eeecfd0816cba091c5a7ce1a266f8a8eaf82d863

Observation 6de6714d-47c2-47d4-8036-ac7410f5e057 · inbound

In-Context Reinforcement Learning via Communicative World Models cites this paper.

In-Context Reinforcement Learning via Communicative World Models A Survey of In-Context Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T22:39:48.540944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:39:48.540944Z digest=sha256:fb4ac1b6c146bd0493d13d25f441370675a918e75c0a2dcb070e9596108f74dd

Observation c61b9786-540f-4b8a-b864-1cc67f788d85 · inbound

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean cites this paper.

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean A Survey of In-Context Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:11:35.898118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T16:09:39.975762Z digest=sha256:cc93f0c4b73bb67433a262ff4e1fd0197eb6b41ceb205508a35336d517458bd7

Observation 2b23d148-4e36-4552-818e-6c57ecbdc7db · inbound

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean cites this paper.

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean A Survey of In-Context Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T16:41:00.086178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:41:00.086178Z digest=sha256:91dbdcb0aca1c63dddbf63b6fe5920481bf9574db83052dbe3483043457fb6e0

Observation 35e8ed47-e3c7-408c-8410-1b327db3676d · inbound

How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off cites this paper.

How Does the Pretraining Distribution Shape In-Context Learning? A Fundamental Trade-Off A Survey of In-Context Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:26.025046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:26.025046Z digest=sha256:863240c72dababdfaf913a113e31da87dac4f5b3284396270216fad16a374ceb

Observation b97a3c38-4cf1-4a41-b230-4b7b78dd812b · inbound

Bridging Natural Language and Microgrid Dynamics: A Context-Aware Simulator and Dataset cites this paper.

Bridging Natural Language and Microgrid Dynamics: A Context-Aware Simulator and Dataset A Survey of In-Context Reinforcement Learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:45:49.834480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:35:57.142293Z digest=sha256:a7195de3cfd0746674ba6df7ae75f078426185c8771cb9520d286548ba9f2507

Observation deed3684-0eec-4925-a1b1-8b3f95ece377 · inbound

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits cites this paper.

When Do We Need LLMs? A Diagnostic for Language-Driven Bandits A Survey of In-Context Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:25:51.227419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T18:32:12.396568Z digest=sha256:e1dab6d6ca7dfd2831b2f58192b2339f23f2d1196bcbd3d67b37daf6e522486b

Observation 31992104-4253-486b-842e-a7b024dd0b6d · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning A Survey of In-Context Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:32.009372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:02f6919335b3cf22dcd64d3d0c76dfb4746fbd692b8e1c1e2ef2e497a3d727a3

Observation 7a82df82-f228-4b6c-bd3a-7d0df59583ca · inbound

MATE: Solving Contextual Markov Decision Processes with Memory of Accumulated Transition Embeddings cites this paper.

MATE: Solving Contextual Markov Decision Processes with Memory of Accumulated Transition Embeddings A Survey of In-Context Reinforcement Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:43:19.568432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:40:19.541614Z digest=sha256:bcd943793e0a70c917481e318da1a7bb16508cd547dadad61ee1036e633ad9ce

Observation 9d35b37b-59e1-43eb-9bb7-aa9eedff9f76 · inbound

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork cites this paper.

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork A Survey of In-Context Reinforcement Learning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:44:40.652995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:41:33.048760Z digest=sha256:b0835fbab1deb4686a06c15be98091ff6cd75d7f1b9ac07e04c4c1668da355ee

Observation b81742fe-e83e-4f47-9e00-b5a7fc03714a · inbound

LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis cites this paper.

LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis A Survey of In-Context Reinforcement Learning

Reference 20

Resolution
malformed identifier
no resolver link, observed 2026-07-12T13:53:50.979339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:53:50.979339Z digest=sha256:77ae343df48e7eb3e9ba7e914cc94f719ac9ec81ed65d4fc8afccfd8cc64217d

Observation 7074d225-71c3-4598-a3c1-0f7418f83691 · inbound

ReBRAC-v2: The Return of the King cites this paper.

ReBRAC-v2: The Return of the King A Survey of In-Context Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:43.025125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:32:43.025125Z digest=sha256:0357d8066211b7eaa55203e987c6d960b8970ba6a80b158dd385b8611ffb1a9a