Pith. sign in

Paper Citation Record · LEDGER

Supervised Pretraining Can Learn In-Context Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2306.14892.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14892 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:30:40.975135Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f4e02161-42d8-43f8-9165-6ef40fa7e2bd · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:35.080068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:77ecead08c5c5169e0326d72205bce2169a8b5076492ebf97fe8c801b72cc074

Observation b808f807-603a-42fb-bcf0-4b2f612a05ff · inbound

Interaction as Intelligence: Deep Research With Human-AI Partnership cites this paper.

Interaction as Intelligence: Deep Research With Human-AI Partnership Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:30:40.975135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:30:40.975135Z digest=sha256:cf24b783f3aeb2d63547c372f98b9b4c1638230f307bb2c563fdb1695eac91d9

Observation 51a9fedf-71d2-4fba-a256-aea2cab9baa1 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.935629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:3d5d4e5107e98d0181c718bffd2cdb9f28518c350dc205c20f0804b30b3d1baf

Observation 0d394af4-8626-4273-aa97-f8663abe2f41 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.371685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:3606a95e2a3e6cf602aa5a4bd74bdc916098feeb7421a4de12090a3619551201

Observation c55a68a0-d79c-4972-838b-c8e9b2fef293 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.497953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:f89250427a11746404ade418dc0cfd237f8d8f51883a9c907791d96f38e4f0ee

Observation 3a18d0c7-fd70-4d9a-9dd7-2e0936c7274d · inbound

Reinforcement Learning Foundation Models Should Already Be A Thing cites this paper.

Reinforcement Learning Foundation Models Should Already Be A Thing Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:39:05.055344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T21:50:00.974590Z digest=sha256:e805f41933c39ae8074ec2210b86e5e09df538a98bb28b8a1cb94a6e9e7bffcc

Observation e7d950d4-a278-414c-ae1e-03b695a73b5c · inbound

Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models cites this paper.

Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models Supervised Pretraining Can Learn In-Context Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.124700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T00:38:48.912290Z digest=sha256:25cac230b8ec5a2bdeed33d55f144ca2ddc6cf9ac9af784231c1802a4caa57e8