Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:33:04.223413Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2602.07764.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T03:33:04.223413Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 367208e2-3346-453d-b789-ce28b1603cea · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77951548-8b8f-4916-baf2-c4370b614546 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1102797b-f85c-4156-b26c-aaf30457b684 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7944cded-417b-4130-93cf-2a99628f429d · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization 16 Decomposed, Diversity-Driven Policy Optimization Then gradient ascent converges to the set of stationary points ofJ(π)
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37102d1d-272a-45dd-9e15-0eaa11fbe060 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589e2b91-2307-430d-98ee-d1fce62815c7 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788c2421-dd9f-468f-9697-e0c7dfa12f2e · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e06cd141-afc1-441e-b972-60b8ec1386f9 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Then lim t→∞ ∥∇J(θ t)∥= 0almost surely
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4e88ab7-453b-47b8-94d0-02fa5f27cc80 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization This environment showcases D3PO’s ability to reliably improve both reward quality and the structure of Pareto-optimal solutions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ee189ac-1bb1-40cc-9b7f-d738107d7a79 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad509f00-5405-418c-bcc5-b4a4c19d15a3 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization These results highlight D 3PO’s robustness in high-dimensional, unstable regimes where conventional MORL baselines often struggle
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b854ae7-ad8b-4bd3-9997-df649c7a83ab · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Importantly, in nearly all such cases, D 3PO still attains better mean performance, but the tests are dominated by large variance, typically from C-MORL
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d017fcf7-0138-492d-9d84-c2e46d78296a · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Pareto Set Learning for Multi-Objective Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d761e5c-c4fe-4972-a8b5-e03497ac9220 · outbound
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization 10 Decomposed, Diversity-Driven Policy Optimization A
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.