Pith. sign in

Paper Citation Record · LEDGER

M3PO: Massively Multi-Task Model-Based Policy Optimization

As of 14 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2506.21782.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21782 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:23:50.399402Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact3
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b63a5da-96b5-493f-8e37-08a7a46ee441 · outbound

This paper cites Proximal Policy Optimization Algorithms.

M3PO: Massively Multi-Task Model-Based Policy Optimization Proximal Policy Optimization Algorithms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:47.273749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:47.273749Z digest=sha256:986addfc24d821525fdc8ad26311dfe53d187783013e291b8b4df0e038b23fb7

Observation 09574043-4761-413b-90cb-947ed4e59639 · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

M3PO: Massively Multi-Task Model-Based Policy Optimization Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:52.843527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:47.418901Z digest=sha256:4c59a845daf69d15f1c7e6901c07377fccec2ff5d759f8f82d9600023d1113d9

Observation cb8cb75a-7d18-4af7-a56e-c25d5df41659 · outbound

This paper cites Mastering Atari with Discrete World Models.

M3PO: Massively Multi-Task Model-Based Policy Optimization Mastering Atari with Discrete World Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:47.586569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:47.586569Z digest=sha256:e5970f8b8a81f301cd7aa32414f213793980c8f5b84ec12ea127e3e4406f4a89

Observation 0310f045-6613-4153-9c81-e1a73e7a9ee7 · outbound

This paper cites Mastering Diverse Domains through World Models.

M3PO: Massively Multi-Task Model-Based Policy Optimization Mastering Diverse Domains through World Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:47.743937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:47.743937Z digest=sha256:fa9042ac4f342d355d9fa6f368b97a9c0b0f2dcd4b433b1bfbb39f43eea3b7c2

Observation da1ce2b2-c1a1-4059-8f5d-e8a477180d76 · outbound

This paper cites Trust Region Policy Optimization.

M3PO: Massively Multi-Task Model-Based Policy Optimization Trust Region Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:47.886009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:47.886009Z digest=sha256:9063557c67488d5bb4b5c4a7d931dd07442f665391bbe0a5fe304e82fcf6cb9d

Observation 53c20a3b-706c-4f6d-9404-16c0b31728ef · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

M3PO: Massively Multi-Task Model-Based Policy Optimization Playing Atari with Deep Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.017663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.017663Z digest=sha256:bacb10b8d02907319279597c2b53997cdbd82cd8acff89c0f2054e980cc4bda1

Observation aaec0511-2f74-4194-8720-facda86ca18a · outbound

This paper cites Continuous control with deep reinforcement learning.

M3PO: Massively Multi-Task Model-Based Policy Optimization Continuous control with deep reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.141617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.141617Z digest=sha256:e7bd10ecbf3bf4fe7ff4e9fe0082386cd9f02efeec14e3de2907f17d2a0bdfd8

Observation a6ec3561-8e4e-4b0b-89fa-e7d6918ec85a · outbound

This paper cites TD-MPC2: Scalable, Robust World Models for Continuous Control.

M3PO: Massively Multi-Task Model-Based Policy Optimization TD-MPC2: Scalable, Robust World Models for Continuous Control

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.240430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.240430Z digest=sha256:b56f7cb9cb3b09b70a1c359b27c42005443be479b5e916448580ea3397cd8751

Observation ee5d5d53-fcec-4e23-9fe9-741c0f8baf33 · outbound

This paper cites Policy Optimization with Model-based Explorations.

M3PO: Massively Multi-Task Model-Based Policy Optimization Policy Optimization with Model-based Explorations

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:23:51.379798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:48.340474Z digest=sha256:97c3575cd364d08f0526b3524f693d5d4b4287ab4fece0626248a0521b6ad762

Observation d1937661-5cd9-4174-87f5-1da27bb68ce8 · outbound

This paper cites dmcontrol: Software and tasks for continuous control,.

M3PO: Massively Multi-Task Model-Based Policy Optimization dmcontrol: Software and tasks for continuous control,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.436730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.436730Z digest=sha256:f0bd45c20a86f4c7257654a12d72b8fdb351b2de5a8b73f3c917e8a416cf6840

Observation 81a6273c-6ca9-480b-9cf9-1d8e263842e3 · outbound

This paper cites Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning.

M3PO: Massively Multi-Task Model-Based Policy Optimization Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.547874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.547874Z digest=sha256:2c35bf0f8e59187cabeb726d1e3eb1b320bcf9398d6ed91d0ffccebb08f306ba

Observation 6f525bea-36dc-4468-8bc8-62f2c11cc64c · outbound

This paper cites Deepmind lab,.

M3PO: Massively Multi-Task Model-Based Policy Optimization Deepmind lab,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:52.622746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:48.692072Z digest=sha256:96aa4784f1cfceb3de42dc191a929cf00b2a250a16ca796221deb61c5787e504

Observation 758f57f7-b5a3-47c9-8015-bdaf72ee23d4 · outbound

This paper cites Temporal Difference Learning for Model Predictive Control.

M3PO: Massively Multi-Task Model-Based Policy Optimization Temporal Difference Learning for Model Predictive Control

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.993986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.993986Z digest=sha256:d925d43f185228d04f37ed7893ec3b3535cb227c2cf1b7b3321a86cedb3723bc

Observation e6ac20c4-3aa6-452d-a54d-0a8400623856 · outbound

This paper cites When to trust your model: Model-based policy optimization,.

M3PO: Massively Multi-Task Model-Based Policy Optimization When to trust your model: Model-based policy optimization,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:49.076812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:49.076812Z digest=sha256:7bb0745d0730144ac677027d9bd1b3f1b0de7dbe98bff0993c4adb66c1f2a217

Observation 28891731-d7b9-4553-ad66-23096209da70 · outbound

This paper cites Model-based Policy Optimization using Symbolic World Model.

M3PO: Massively Multi-Task Model-Based Policy Optimization Model-based Policy Optimization using Symbolic World Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:23:51.036310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:49.148300Z digest=sha256:c03a10fb3719dec7437ec686d52a5f6a8442e7c3432adf6ee7602383446dfa16

Observation 4d6e0d5e-b2b8-4d9f-a08c-a9838100d3c3 · outbound

This paper cites Deep reinforcement learning and the deadly triad,.

M3PO: Massively Multi-Task Model-Based Policy Optimization Deep reinforcement learning and the deadly triad,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:52.427891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:49.234846Z digest=sha256:55366043f8e558ae8f27024b365d7abe24d20b9990ebd55b112726048443270d

Observation eccfcef3-33cc-44a8-ab83-e31c69edb686 · outbound

This paper cites Distributed prioritized experience replay,.

M3PO: Massively Multi-Task Model-Based Policy Optimization Distributed prioritized experience replay,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:52.189114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:49.515614Z digest=sha256:5a1a4344367148dd47f740e24dbde86fc2beee6a092e402b3aacaffbba1e7455

Observation fbce273d-6627-43f7-a564-fb5ab9e96ec9 · outbound

This paper cites Recurrent experience replay in distributed reinforcement learning,.

M3PO: Massively Multi-Task Model-Based Policy Optimization Recurrent experience replay in distributed reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:51.915891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:49.854733Z digest=sha256:c6e5201f3cc236f8dfa8f2fb91ff64bbdeef7a4b70cc62f78fa198c73538a1c1

Observation 24d7b1a2-8640-4d15-aa2a-d20528885b22 · outbound

This paper cites IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.

M3PO: Massively Multi-Task Model-Based Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:50.021754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:50.021754Z digest=sha256:4d01407ecdd74bdbe5f322bffe07622655209bc7281fede8b7418a56645ba85a

Observation 527e8fa6-5e32-4a07-aef6-2ff75013b1c1 · outbound

This paper cites Distributed Prioritized Experience Replay.

M3PO: Massively Multi-Task Model-Based Policy Optimization Distributed Prioritized Experience Replay

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:49.630943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:49.630943Z digest=sha256:0dfdea9a2f3a36d4f9dfd94493daa9fa6ff6ec9d9d35a9629728c53e315c2cea

Observation 7278c1d5-47f6-445d-be5f-400836e15977 · outbound

This paper cites Model predictive path integral control using covariance variable importance sampling,.

M3PO: Massively Multi-Task Model-Based Policy Optimization Model predictive path integral control using covariance variable importance sampling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:23:51.703998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:50.268097Z digest=sha256:860218bbbe8ea157e3d96a13de2f42e09876f247f85f80b25bf1d0884e5c5197

Observation 13d28f2e-6497-452f-bc00-5a7cdb735774 · outbound

This paper cites A markovian decision process,.

M3PO: Massively Multi-Task Model-Based Policy Optimization A markovian decision process,

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T22:23:50.728011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T22:23:50.170608Z digest=sha256:84d2331ad454c8eacca44f3bb6d39e167e2478ae19ba40f1b1a981bf6b5ad580

Observation 3b079b1f-c613-4002-a894-41ca4f39e383 · outbound

This paper cites Model Predictive Path Integral Control using Covariance Variable Importance Sampling.

M3PO: Massively Multi-Task Model-Based Policy Optimization Model Predictive Path Integral Control using Covariance Variable Importance Sampling

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:50.399402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:50.399402Z digest=sha256:3b75660b5a90ba1a14d338e0c0cf2b184691a6b91c9b5f3c9fa46dbc3d8a2d69

Observation 4e43fe24-537d-4370-b356-6cbde2b73c48 · outbound

This paper cites DeepMind Lab.

M3PO: Massively Multi-Task Model-Based Policy Optimization DeepMind Lab

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:48.857752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:48.857752Z digest=sha256:4f2ad958a4b05918f698c984afdfe86bc78f3fcb5d4e1eacdd94ae8f225d6ac7

Observation f728dbd0-ed32-4b5a-8764-ceb866476c1e · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

M3PO: Massively Multi-Task Model-Based Policy Optimization Deep Reinforcement Learning and the Deadly Triad

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:49.384353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:49.384353Z digest=sha256:4830179154f89d2b71ffcb2a4b1889cede48660a1bffa32c2a148c67f3a808d5

Pith citing papers

No inbound Pith citation observations are available.