Pith. sign in

Paper Citation Record · LEDGER

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

As of 16 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2602.07764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.07764 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:33:04.223413Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 367208e2-3346-453d-b789-ce28b1603cea · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.187857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.187857Z digest=sha256:a8c3d971a723f8b61dc7f77af865ef1337211191cacd95435ad775a5ba7a9c8c

Observation 77951548-8b8f-4916-baf2-c4370b614546 · outbound

This paper cites A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.179060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.179060Z digest=sha256:d5a9219dbe0eb1c2aac002eae23a7a8f2a282d567f6d98ff68066c58e19fbbfb

Observation 1102797b-f85c-4156-b26c-aaf30457b684 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.194989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.194989Z digest=sha256:98616514b7523509efae54c9700eca398169ec9d39db6efce0bf22d91dc9f8f6

Observation 7944cded-417b-4130-93cf-2a99628f429d · outbound

This paper cites 16 Decomposed, Diversity-Driven Policy Optimization Then gradient ascent converges to the set of stationary points ofJ(π).

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization 16 Decomposed, Diversity-Driven Policy Optimization Then gradient ascent converges to the set of stationary points ofJ(π)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.198801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.198801Z digest=sha256:6623a2cf1d469f85bda69c7a5912718dc0d274ae18f3ab8ee0a432865f4d5f9d

Observation 37102d1d-272a-45dd-9e15-0eaa11fbe060 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.191639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.191639Z digest=sha256:7e37c9e4f5e0dad65aeed7412e1351af6c40dac4cdf2d286e8ab412638c27afd

Observation 589e2b91-2307-430d-98ee-d1fce62815c7 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.202501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.202501Z digest=sha256:43e4d214ab3f24772d1b19c9808d6654fff80ed1243505371853987becbc9cdc

Observation 788c2421-dd9f-468f-9697-e0c7dfa12f2e · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.205697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.205697Z digest=sha256:1ef8a42bfd1265e732e30d8162e021fad510c2d4b2ea6dbd86be70279794c159

Observation e06cd141-afc1-441e-b972-60b8ec1386f9 · outbound

This paper cites Then lim t→∞ ∥∇J(θ t)∥= 0almost surely.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Then lim t→∞ ∥∇J(θ t)∥= 0almost surely

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.209291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.209291Z digest=sha256:bb569fb9801dac4621c23442c303a094b8c2d53e40b85c144a057687a5260fe1

Observation d4e88ab7-453b-47b8-94d0-02fa5f27cc80 · outbound

This paper cites This environment showcases D3PO’s ability to reliably improve both reward quality and the structure of Pareto-optimal solutions.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization This environment showcases D3PO’s ability to reliably improve both reward quality and the structure of Pareto-optimal solutions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.213020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.213020Z digest=sha256:7d507eadd8d58cf642d1aafb79becd52afba482886733c2dbe17967257739541

Observation 3ee189ac-1bb1-40cc-9b7f-d738107d7a79 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.216280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.216280Z digest=sha256:9c0148d46e2a0ff886361b1bb415f961cce920b818a45780495a6caaf70de7f8

Observation ad509f00-5405-418c-bcc5-b4a4c19d15a3 · outbound

This paper cites These results highlight D 3PO’s robustness in high-dimensional, unstable regimes where conventional MORL baselines often struggle.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization These results highlight D 3PO’s robustness in high-dimensional, unstable regimes where conventional MORL baselines often struggle

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.219958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.219958Z digest=sha256:7b366249f47bb34048e85be2a33fa281c321ae5f7c9356748ab0f21dd1e27e57

Observation 0b854ae7-ad8b-4bd3-9997-df649c7a83ab · outbound

This paper cites Importantly, in nearly all such cases, D 3PO still attains better mean performance, but the tests are dominated by large variance, typically from C-MORL.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Importantly, in nearly all such cases, D 3PO still attains better mean performance, but the tests are dominated by large variance, typically from C-MORL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.223413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.223413Z digest=sha256:459e3b774c19b0e9d9a0fbc29ea91d3d68f8d5c11811ef13edc39cc9791a53aa

Observation d017fcf7-0138-492d-9d84-c2e46d78296a · outbound

This paper cites Pareto Set Learning for Multi-Objective Reinforcement Learning.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Pareto Set Learning for Multi-Objective Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.174258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.174258Z digest=sha256:ef7b773a0515e38814260342345befb4b653e38c788b1149a20d0a676b8963eb

Observation 5d761e5c-c4fe-4972-a8b5-e03497ac9220 · outbound

This paper cites 10 Decomposed, Diversity-Driven Policy Optimization A.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization 10 Decomposed, Diversity-Driven Policy Optimization A

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.183177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.183177Z digest=sha256:ba2c2db4e82e5df6e12a5830da03f7e0f988296acaeb1a80ebfc143f06ba5a4d

Pith citing papers

No inbound Pith citation observations are available.