Pith. sign in

Paper Citation Record · LEDGER

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

As of 15 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2602.07764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.07764 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:33:04.223413Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 367208e2-3346-453d-b789-ce28b1603cea · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.187857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.187857Z digest=sha256:a5eaf736165f913d8fe2ab4a57d680b56d2f19b43130409eda4085fb7d98dda8

Observation 77951548-8b8f-4916-baf2-c4370b614546 · outbound

This paper cites A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.179060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.179060Z digest=sha256:e7c8505053817e65c5fdb2bc54a3c8b1782b8c472e2453d9ab02d8ebff34b08d

Observation 1102797b-f85c-4156-b26c-aaf30457b684 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.194989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.194989Z digest=sha256:6c12c60064da07d835739c58bffe2db7dca19a784b1031ecd5e3069780602cb5

Observation 7944cded-417b-4130-93cf-2a99628f429d · outbound

This paper cites 16 Decomposed, Diversity-Driven Policy Optimization Then gradient ascent converges to the set of stationary points ofJ(π).

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization 16 Decomposed, Diversity-Driven Policy Optimization Then gradient ascent converges to the set of stationary points ofJ(π)

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.198801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.198801Z digest=sha256:b74e0994487bbf00e5e66b8a2dfafe68728e36e7378939692764eb24886e7155

Observation 37102d1d-272a-45dd-9e15-0eaa11fbe060 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.191639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.191639Z digest=sha256:cad83b2f97d2ff50bfcb5852257fcbc66ba7337b83db0a9106c777a9db582f1c

Observation 589e2b91-2307-430d-98ee-d1fce62815c7 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.202501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.202501Z digest=sha256:46ca2d2215f7c13359aab646aae1f9f576827f30a93045ed0cc379c7e7f2d0a1

Observation 788c2421-dd9f-468f-9697-e0c7dfa12f2e · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.205697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.205697Z digest=sha256:4c2fd204f908aa778605e94f9d29dcf364eef1b1146eaa9a0f29fa00fd33d067

Observation e06cd141-afc1-441e-b972-60b8ec1386f9 · outbound

This paper cites Then lim t→∞ ∥∇J(θ t)∥= 0almost surely.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Then lim t→∞ ∥∇J(θ t)∥= 0almost surely

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.209291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.209291Z digest=sha256:5295e3af3678f9c0402faa153f1ae30b80b1752da09b009422bca79dab58099b

Observation d4e88ab7-453b-47b8-94d0-02fa5f27cc80 · outbound

This paper cites This environment showcases D3PO’s ability to reliably improve both reward quality and the structure of Pareto-optimal solutions.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization This environment showcases D3PO’s ability to reliably improve both reward quality and the structure of Pareto-optimal solutions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.213020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.213020Z digest=sha256:ffc8da82164f8638666eb6f3c537d2b4fea2d2a0e48af8669b447f65bb75d63f

Observation 3ee189ac-1bb1-40cc-9b7f-d738107d7a79 · outbound

This paper cites an unresolved cited work.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.216280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.216280Z digest=sha256:da34b436b29815ba795ffe2154ed394ccd5b47f14b316628702d7af0d55055d4

Observation ad509f00-5405-418c-bcc5-b4a4c19d15a3 · outbound

This paper cites These results highlight D 3PO’s robustness in high-dimensional, unstable regimes where conventional MORL baselines often struggle.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization These results highlight D 3PO’s robustness in high-dimensional, unstable regimes where conventional MORL baselines often struggle

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.219958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.219958Z digest=sha256:9c1667faea13c3aba8934d43713542d79100e5f628fd133bb8479d301a369214

Observation 0b854ae7-ad8b-4bd3-9997-df649c7a83ab · outbound

This paper cites Importantly, in nearly all such cases, D 3PO still attains better mean performance, but the tests are dominated by large variance, typically from C-MORL.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Importantly, in nearly all such cases, D 3PO still attains better mean performance, but the tests are dominated by large variance, typically from C-MORL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.223413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.223413Z digest=sha256:d769dee29f1887e053b3482fdcdfda302ff5dbbdda4d11ddd8b6b90fd2fc22c1

Observation d017fcf7-0138-492d-9d84-c2e46d78296a · outbound

This paper cites Pareto Set Learning for Multi-Objective Reinforcement Learning.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization Pareto Set Learning for Multi-Objective Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.174258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.174258Z digest=sha256:f6ec96fb9df14347e710601383444f0e1f3ff79f79147ba2ebd28b0fd76a640a

Observation 5d761e5c-c4fe-4972-a8b5-e03497ac9220 · outbound

This paper cites 10 Decomposed, Diversity-Driven Policy Optimization A.

Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization 10 Decomposed, Diversity-Driven Policy Optimization A

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:04.183177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:04.183177Z digest=sha256:b26fc71271605c132c08fd0c6c6c27f03be8bafafcbc306b242e245658eebe54

Pith citing papers

No inbound Pith citation observations are available.