Pith. sign in

Paper Citation Record · LEDGER

Supervised Pretraining Can Learn In-Context Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2306.14892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14892 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:58:47.657835Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f4e02161-42d8-43f8-9165-6ef40fa7e2bd · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:35.080068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:14437a5e6973b3b22e5142346e95b5603d70c1868448204c68ba76b9789318c6

Observation c42bf182-47db-44f6-9c24-324797390aa0 · inbound

A Research Agenda for Usability and Generalisation in Reinforcement Learning cites this paper.

A Research Agenda for Usability and Generalisation in Reinforcement Learning Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T05:58:47.657835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:58:47.657835Z digest=sha256:6f99ee1d7b318793fea393eead1e92794a4a81e1cc5dee44438fafc10997bc6e

Observation 30666b6b-eab3-4d8e-a4c7-c9fb3254bc55 · inbound

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit cites this paper.

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.324021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.324021Z digest=sha256:274937122bda55db72df6a3792a57c661ebff13db2adc259d731c5d52de5570f

Observation b808f807-603a-42fb-bcf0-4b2f612a05ff · inbound

Interaction as Intelligence: Deep Research With Human-AI Partnership cites this paper.

Interaction as Intelligence: Deep Research With Human-AI Partnership Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:30:40.975135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:30:40.975135Z digest=sha256:d3685fd42a4ee37ead110ccb3874ba5baa93261330925ca6285772fa1367d7b0

Observation 51a9fedf-71d2-4fba-a256-aea2cab9baa1 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.935629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:10e182a9eeba630c31c5dcd0d031a700985a90f770b00fb2dc8f8b67f1476cbb

Observation 0d394af4-8626-4273-aa97-f8663abe2f41 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.371685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:cbbf4214fe8e433fcfe8af91a8173a452ab6c726fd2a9b2a2773e982de4a9606

Observation c55a68a0-d79c-4972-838b-c8e9b2fef293 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.497953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:67cad859d4eb882ee6b2acb58df5b78becbad9d75bb09d9f2d5eef84d8f07d67

Observation 3a18d0c7-fd70-4d9a-9dd7-2e0936c7274d · inbound

Reinforcement Learning Foundation Models Should Already Be A Thing cites this paper.

Reinforcement Learning Foundation Models Should Already Be A Thing Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:05.055344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T21:50:00.974590Z digest=sha256:714a4c196c7a4cd33a4ea07e4c9c2a13e51166105802dea83ac8a0a0d921042f

Observation e7d950d4-a278-414c-ae1e-03b695a73b5c · inbound

Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models cites this paper.

Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.124700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T00:38:48.912290Z digest=sha256:b3e828e9e5ea917ee20ab70e01db15bf860b06fdd7fd68b84492a49f94dc2ce1