Pith. sign in

Paper Citation Record · LEDGER

Training Agents using Upside-Down Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:1912.02877.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.02877 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:45:57.537078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:57:25.915173Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 80bf680b-5b57-4d18-8841-ae1b5076cacc · inbound

Decision Transformer: Reinforcement Learning via Sequence Modeling cites this paper.

Decision Transformer: Reinforcement Learning via Sequence Modeling Training Agents using Upside-Down Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:11.117837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T15:11:11.056013Z digest=sha256:a6f001b418ec99dbee1d4a4c40d8048ba9d161111e92b97861e6877327e6092a

Observation 1baf92be-02ec-4bbb-be46-cfa2500d5d2e · inbound

Is Conditional Generative Modeling all you need for Decision-Making? cites this paper.

Is Conditional Generative Modeling all you need for Decision-Making? Training Agents using Upside-Down Reinforcement Learning

Reference 209

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:35:10.858704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-15T15:35:10.593969Z digest=sha256:2df034f1d8e1b567ace9258dbc9b05c06c0a3ffc65a5e0050e999f6312e2fc62

Observation 38f45699-690a-442d-9483-3fb582ba6590 · inbound

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework cites this paper.

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework Training Agents using Upside-Down Reinforcement Learning

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:43:19.125664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T03:43:18.632292Z digest=sha256:f8645f968a99647022a988c640aa681ec1cd9299cb14124b4b7a22cf4482e319

Observation 0ff21285-d901-4b47-901b-4138ce934c81 · inbound

Upside-Down Reinforcement Learning for More Interpretable Optimal Control cites this paper.

Upside-Down Reinforcement Learning for More Interpretable Optimal Control Training Agents using Upside-Down Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:35:59.778062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:35:59.778062Z digest=sha256:78b04ab5995f236005b432cec3fb7f0a5cd09c4ce13e874326f4fb51429511b4

Observation 9bd1bbdf-79b5-440e-9209-b86158fa4055 · inbound

Heuristically Adaptive Diffusion-Model Evolutionary Strategy cites this paper.

Heuristically Adaptive Diffusion-Model Evolutionary Strategy Training Agents using Upside-Down Reinforcement Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T16:32:34.761175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:32:34.761175Z digest=sha256:7e34952311b5661c1b26db9db3517ac032699d475b3ddf544177aebbe5ca7826

Observation 30b68d5f-2cf3-4bf6-9ce7-1411877a6585 · inbound

The broader spectrum of in-context learning cites this paper.

The broader spectrum of in-context learning Training Agents using Upside-Down Reinforcement Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T22:10:22.971826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:10:22.971826Z digest=sha256:ac5c2b997e27274b24c48bb9f871f0f87c7135f4e7efb723b912455beac65ec8

Observation bb27b557-e7ef-4f33-b183-0d538c203ee5 · inbound

MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning cites this paper.

MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning Training Agents using Upside-Down Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:44.262016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:02:44.262016Z digest=sha256:99a3913a638d2a7495f7b5adbc541f0e2825580828d1b38f889edfcb4871aac2

Observation 788f03c6-d461-45cc-a387-7cfd7518c8fc · inbound

Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer cites this paper.

Beyond the Known: Decision Making with Counterfactual Reasoning Decision Transformer Training Agents using Upside-Down Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:57.537078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:57.537078Z digest=sha256:8b82e6c40810bb5b7e4091f7596e8a042607b84e08e34d8e564244ce6a26a24d

Observation fc5c38cc-235b-47fc-a44a-eabe153027c2 · inbound

A Provable Approach for End-to-End Safe Reinforcement Learning cites this paper.

A Provable Approach for End-to-End Safe Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.270417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.270417Z digest=sha256:bedd098b62c28e0d5eacbe9b4e7e95339ae9c9827fb801a4736489d668a04a4d

Observation 1e69c1d7-7926-44ab-b3ac-bb44ef679d6c · inbound

How to Provably Improve Return Conditioned Supervised Learning? cites this paper.

How to Provably Improve Return Conditioned Supervised Learning? Training Agents using Upside-Down Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:08.014844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:08.014844Z digest=sha256:cd7c9f047eae457a3b9666f7fa7472fc69163f44f62093bf7dff1bd401e08f5c

Observation 325a3a7e-8201-4ff1-971f-797eabf7a0e8 · inbound

Behavioral Exploration: Learning to Explore via In-Context Adaptation cites this paper.

Behavioral Exploration: Learning to Explore via In-Context Adaptation Training Agents using Upside-Down Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:46.817188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:15:46.817188Z digest=sha256:98e4e801adfe43130cc4d910273614aa387a684b06ed123b89ff39549919f3f9

Observation 715459ac-bfe6-476e-b793-aea24c7a4984 · inbound

Equivariant Goal Conditioned Contrastive Reinforcement Learning cites this paper.

Equivariant Goal Conditioned Contrastive Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.543536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.543536Z digest=sha256:1608fe37826c7ca69e71cd37e64ea56c31abdb39c6f1667a0f2a169b799ee953

Observation 28fa6eef-fb21-4a98-897d-7278f3b6ea36 · inbound

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration cites this paper.

GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration Training Agents using Upside-Down Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:14.448302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:14.448302Z digest=sha256:1384f8935e3dfef28e4029b711c7e7d457d3960ea863ca5b3e9145b39bea4750

Observation 310d9c01-bbdb-4460-8592-6e9a51c52323 · inbound

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success cites this paper.

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Training Agents using Upside-Down Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:02.544642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:02.544642Z digest=sha256:abaf0132b2103db6ee9380624c5999ba70335ab543c1fa4e07257cbb909cf281

Observation 07d1701b-aded-4818-9ad6-1aa0f6fc72eb · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Training Agents using Upside-Down Reinforcement Learning

Reference 232

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.178293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:010e382203a8dc9ce6c4dfe2fbc6bc750c16824120732e2def599bcf8d571344

Observation b33bb334-208a-4f7b-b6bd-9a6d9280bba5 · inbound

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies cites this paper.

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies Training Agents using Upside-Down Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.916771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T19:23:32.966503Z digest=sha256:efe8b2f04a9cf8e1a6bf9318bfc3c93f27ccfcf38f64b2e1a87d0dbd2305b560

Observation fc21b51e-44c7-4011-86bc-1d869c0540fc · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models Training Agents using Upside-Down Reinforcement Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:14.105428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:14.105428Z digest=sha256:28cfda145416363e4c9b61bf5e8fb9a1a8deb524999297e3b539664bcc100a43

Observation 008ffa89-a4da-4946-90dd-a2c1f29c159c · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Training Agents using Upside-Down Reinforcement Learning

Reference 269

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.197052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.197052Z digest=sha256:372ef46283232378f835f749f80e81090ef016ddc1397d58be3eb7b5a32038e3