Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:23:50.399402Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.21782.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:23:50.399402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8b63a5da-96b5-493f-8e37-08a7a46ee441 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Proximal Policy Optimization Algorithms
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09574043-4761-413b-90cb-947ed4e59639 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb8cb75a-7d18-4af7-a56e-c25d5df41659 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Mastering Atari with Discrete World Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0310f045-6613-4153-9c81-e1a73e7a9ee7 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Mastering Diverse Domains through World Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1ce2b2-c1a1-4059-8f5d-e8a477180d76 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Trust Region Policy Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53c20a3b-706c-4f6d-9404-16c0b31728ef · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Playing Atari with Deep Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaec0511-2f74-4194-8720-facda86ca18a · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Continuous control with deep reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ec3561-8e4e-4b0b-89fa-e7d6918ec85a · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5d5d53-fcec-4e23-9fe9-741c0f8baf33 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Policy Optimization with Model-based Explorations
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d1937661-5cd9-4174-87f5-1da27bb68ce8 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization dmcontrol: Software and tasks for continuous control,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81a6273c-6ca9-480b-9cf9-1d8e263842e3 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f525bea-36dc-4468-8bc8-62f2c11cc64c · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Deepmind lab,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 758f57f7-b5a3-47c9-8015-bdaf72ee23d4 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Temporal Difference Learning for Model Predictive Control
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ac20c4-3aa6-452d-a54d-0a8400623856 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization When to trust your model: Model-based policy optimization,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28891731-d7b9-4553-ad66-23096209da70 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Model-based Policy Optimization using Symbolic World Model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4d6e0d5e-b2b8-4d9f-a08c-a9838100d3c3 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Deep reinforcement learning and the deadly triad,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eccfcef3-33cc-44a8-ab83-e31c69edb686 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Distributed prioritized experience replay,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fbce273d-6627-43f7-a564-fb5ab9e96ec9 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Recurrent experience replay in distributed reinforcement learning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 24d7b1a2-8640-4d15-aa2a-d20528885b22 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 527e8fa6-5e32-4a07-aef6-2ff75013b1c1 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Distributed Prioritized Experience Replay
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7278c1d5-47f6-445d-be5f-400836e15977 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Model predictive path integral control using covariance variable importance sampling,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13d28f2e-6497-452f-bc00-5a7cdb735774 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization A markovian decision process,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b079b1f-c613-4002-a894-41ca4f39e383 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Model Predictive Path Integral Control using Covariance Variable Importance Sampling
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e43fe24-537d-4370-b356-6cbde2b73c48 · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization DeepMind Lab
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f728dbd0-ed32-4b5a-8764-ceb866476c1e · outbound
M3PO: Massively Multi-Task Model-Based Policy Optimization Deep Reinforcement Learning and the Deadly Triad
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.