Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:10.092408Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2501.02774.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:13:10.092408Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f3ce5166-102e-49cb-b77e-86b7d2e3b02e · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Human-level control through deep reinforcement learning,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16c667f4-1af0-449b-9d6a-6317cbd823f6 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Continuous control with deep reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a6db6c-117f-423a-9601-4f742d222a3a · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Trust Region Policy Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf5726b-06b7-42e8-8f35-4871cf29760a · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Reinforcement Learning framework for Autonomous Driving
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc5a620d-d7d0-4110-9afc-d63e310814f9 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Continuous mdp homomorphisms and homomorphic policy gradient,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48462379-f791-4650-bc6b-fe83478d06c5 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes TD-MPC2: Scalable, Robust World Models for Continuous Control
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773dc85a-275e-4afa-a622-a962d8bcf345 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Movie: Visual model-based policy adaptation for view generalization,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c90270ed-d52e-4f73-a5af-0d61e1f651db · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Making better decision by directly planning in continuous control,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation efd40ed2-e94b-4ee3-8e81-250cc28bf14c · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Hierarchical advantage for reinforcement learning in parameterized action space,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dc94fb86-bcb7-49ae-af24-eb4ae35ee822 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Reinforcement Learning in Parameterized Action Space
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fbc369-6f81-4cd0-ac1f-861e330543f7 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Deep Multi-Agent Reinforcement Learning with Discrete-Continuous Hybrid Action Spaces
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e967c4e-ef86-4707-83df-62f7a5cb5816 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Parametrized Deep Q-Networks Learning: Reinforcement Learning with Discrete-Continuous Hybrid Action Space
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 214518cb-6f71-4875-bb3e-b8c8c8a3b67a · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1030316-f089-4206-8969-67613adac404 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023c83db-31dc-44f9-adae-bef1638bc15b · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Reinforcement learning with parameterized actions,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35cbf09d-f9b8-4de6-8456-c4e8ee3bdedf · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Dream to Control: Learning Behaviors by Latent Imagination
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f76ae8c-4ccf-4d0f-bd9a-2335b5b9b408 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Mastering Atari with Discrete World Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9e0995-93ab-4e25-bbb2-dc5d3442620a · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Privileged Sensing Scaffolds Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34d92f07-a7cd-49b7-9bb0-2ca85dc350f9 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based Reinforcement Learning for Parameterized Action Spaces
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ef3cd29-9f52-46e3-9eaa-a75d1a33c9a7 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The Benefits of Model-Based Generalization in Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a393001f-e5fd-4011-9e24-a7db96cc6e9f · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Diminishing return of value expansion methods in model-based reinforcement learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43c3c435-2a7c-4e9e-a36b-c1bef140e093 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Models, Pixels, and Rewards: Evaluating Design Trade-offs in Visual Model-Based Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55f67307-798f-4012-bfd5-64cbd6767d41 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes DreamSmooth: Improving Model-based Reinforcement Learning via Reward Smoothing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 57fc4af9-3968-4d15-aff1-64006f9c2c8b · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Multi-Pass Q-Networks for Deep Reinforcement Learning with Parameterised Action Spaces
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a66bf1-3717-430b-b6d1-4807988fa085 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Learning insertion primitives with discrete-continuous hybrid action space for robotic assembly tasks,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99753433-1d7b-4a4f-955a-f554b0454e9b · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Lipschitz continuity in model- based reinforcement learning,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e95acb9-b3ea-4e4b-b779-c6c713777e18 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Markov processes over denumerable products of spaces, describing large systems of automata,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 055f35d8-f7a8-4de8-9abc-70969ac87f0d · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Explaining and Harnessing Adversarial Examples
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e37482-ebdb-46f1-afea-f32b426a8185 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes A mathematical theory of communication,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa9f94b-ea28-4d56-957a-ae76c286218d · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The im algorithm: a variational approach to information maximization,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4c1b5286-a103-4fe4-ac63-cfef51d5a6bb · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Learning-based model predictive control for markov decision processes,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55fd597e-cee8-4b0d-92d6-e2bed94a51f7 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Optimization of computer simulation models with rare events,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bbf4ba56-0ff9-4a24-9d8a-20309191a49d · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Plan to predict: Learning an uncertainty-foreseeing model for model-based reinforce- ment learning,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2ba40424-c73c-4155-b3b7-e6a2ca0c21d5 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Choreographer: Learning and Adapting Skills in Imagination
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df786e93-e2c2-480a-97e1-3820ec7f11a7 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Mismatched no more: Joint model-policy optimization for model- based rl,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c677bc8f-2160-41dc-b0dd-0eb365f470d0 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Differen- tiable mpc for end-to-end planning and control,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70895f64-e819-42f1-b1ac-c051685bf061 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes On information and sufficiency,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e934efb5-2c2d-432e-a0a3-ba4114c01bd4 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Measures of distance between probability distributions,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a1cab5a-896b-4d5f-ba88-fbc9d853d983 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Dr. Strategy: Model-Based Generalist Agents with Strategic Dreaming
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ce6fd07d-e6c5-4bb2-a43a-1b59251b0e2e · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes When to trust your model: Model-based policy optimization,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e6b81818-ff75-43bf-b7aa-7adf24c03920 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based reinforcement learning via meta-policy optimization,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 24d689a0-cebb-460a-ba81-f0f1b7f8df2c · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes The virtues of laziness in model-based rl: A unified objective and algo- rithms,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c3bcd5d5-fc70-48c4-ab46-475e1479b127 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Iterative value-aware model learning,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7eb6f6a4-bf98-48ea-8459-17304275825f · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Model-based value expansion for efficient model-free rein- forcement learning,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d7bce9d-9972-4c4f-8e45-2228a24eb222 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Villani et al
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a423560-c737-4b7b-9eb0-ba5fc99a2d78 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Generative adversarial nets,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21fad6e0-ba0c-4e0c-94a5-bf23d2e010f7 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Towards Principled Methods for Training Generative Adversarial Networks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1629a347-0aa7-448f-853a-47d4f3ec0bae · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Wasserstein generative ad- versarial networks,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a9e0545-bac5-4c2e-9959-ff8fe7f60c67 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Harpy, a connected speech recognition system,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 44c6830e-c21a-4e80-bd0a-d0cdc6c7c663 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f74544f5-fdee-48fc-bab3-6d7f128b19fd · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Walk Wisely on Graph: Knowledge Graph Reasoning with Dual Agents via Efficient Guidance-Exploration
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ae52c77-d93e-43f9-b08d-39c637e62b6b · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Spectral Norm Regularization for Improving the Generalizability of Deep Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a214a2d9-fd2c-4359-be98-2a65c8aed1b3 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Categorical Reparameterization with Gumbel-Softmax
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a172a2a-29bd-4124-916a-e77d37a0d58b · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Introduction to online convex optimization,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d0d3d94-3f37-47e5-841f-80391b92d272 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Is Model Ensemble Necessary? Model-based RL via a Single Model with Lipschitz Regularized Value Function
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd4be036-c677-46ec-81b8-6c3558dcd1b7 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Visualizing data using t-sne.,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f6f2d4-73a8-4e7e-8c07-df67b7a7e86f · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes General boundary conditions for denumberable markov processes,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51b87c37-82ee-431a-aae4-056eb7bf6398 · outbound
Learn A Flexible Exploration Model for Parameterized Action Markov Decision Processes Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.