Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:1912.02875.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:29.406729Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:58.082706Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c848b185-a977-43b6-96b0-762c8ebfcd70 · inbound
Is Conditional Generative Modeling all you need for Decision-Making? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 935963c6-4d05-4f6a-9792-79c1ea9c2621 · inbound
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 90cddf7c-eeb2-4a7d-bc5b-4a7371aa3876 · inbound
A Provable Approach for End-to-End Safe Reinforcement Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8788ce0a-e4a4-4fde-8217-1fade5dd987c · inbound
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5717991e-91ec-4bfa-ba3a-df3a467f88cf · inbound
How to Provably Improve Return Conditioned Supervised Learning? Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69a88021-63dc-4679-9ef6-aef353c8e3b9 · inbound
Self-Predictive Representations for Combinatorial Generalization in Behavioral Cloning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28f66d78-9c97-40d7-8194-3a772848df4f · inbound
Single-pass Adaptive Image Tokenization for Minimum Program Search Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159f7b58-4c8a-4f70-9822-cd365b443697 · inbound
Behavioral Exploration: Learning to Explore via In-Context Adaptation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e036314-dc6f-45b0-adea-608e157c7c2c · inbound
GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7542ae43-ba7b-45cf-b733-1f13a4991a15 · inbound
$\pi^{*}_{0.6}$: a VLA That Learns From Experience Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 221a88ec-2bdf-4443-ac4e-ba250853f5da · inbound
Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5bc580-a521-4071-ab1e-75788d2bc868 · inbound
QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc5ca21d-ace3-47b1-8197-7cc9f2486d35 · inbound
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1748dc48-5e81-4e46-8503-0c24aed948ee · inbound
Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f190c8be-e006-41af-abc7-8f35fd549430 · inbound
Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa1c7e24-98bb-4437-b66c-88cf4272c7f4 · inbound
Freeform Preference Learning for Robotic Manipulation Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b479fce-53c0-4401-a849-44e8ea6307e2 · inbound
Reinforcement Learning: From Algorithms To Foundation Models Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 175
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d15564-9343-4d9e-9703-70d5d4ee1517 · inbound
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
Reference 268
Source-reported events for the cited work
Unavailable: canonical work link unavailable.