Pith. sign in

Paper Citation Record · LEDGER

Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.17466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.17466 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:53:09.027483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T06:44:00.901151Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b41119ee-d208-4a11-b610-82002cf7e145 · inbound

Multi-objective Large Language Model Alignment with Hierarchical Experts cites this paper.

Multi-objective Large Language Model Alignment with Hierarchical Experts Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:12.434324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:12.434324Z digest=sha256:6c8226ad922a5015b2f8a767d98d592d116b7ebe1227244352c41e97ca0b55d8

Observation 51ab15c2-aa1a-4d51-9507-df65c499ee88 · inbound

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach cites this paper.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.106357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.106357Z digest=sha256:7c871d560abbb47ebeca4cc7647522b3ea9ec2bfeee71b0c62392b0ca719c978

Observation bf288d27-4fa1-4a5c-8e0b-ee6f1ca3725e · inbound

A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning cites this paper.

A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:51:29.689165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:04:19.891421Z digest=sha256:f1a201fc1086a1d05cfe16fa29be60c5bec21c92f06d82a45662bec2004edcac

Observation 7d4ad250-cbd8-4c5f-9b50-c20acec4f67b · inbound

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems cites this paper.

Contextual Multi-Objective Optimization: Rethinking Objectives in Frontier AI Systems Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:18.942894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T16:26:04.152045Z digest=sha256:96b0b75b542b2af989f1e810a03841db35b40e13f16685e714ce4b7d79f751b1

Observation ce8807aa-1619-4c9c-816a-47c858360871 · inbound

Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization cites this paper.

Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:51.671927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:32:35.197431Z digest=sha256:6426c36c4f253fe512c22e470ee09f5bd3db6b88224dc012eb30961f84ddb65f

Observation fdf2ba28-614f-43f0-b7f6-e57dda7c0738 · inbound

SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front cites this paper.

SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:44:00.902598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T06:42:15.135148Z digest=sha256:1e87662adaf65c7c610b40ccc199ec1756fd3cf6b92c736bd24be5c25f44467f

Observation 573a5b80-221b-4ab5-b35d-a7b27625e9ea · inbound

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability cites this paper.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:5aa731c4d1a7ba4cf08b07397e3d5c70ce0c96690f26fa0a970f94b49dbd4e83

Observation 1c10cf1e-909c-499b-9ac3-035e593912f3 · inbound

Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits cites this paper.

Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:53:09.027483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:53:09.027483Z digest=sha256:4d38116724685755c75b844b9dd477b732330b77dc160e973999fab8e255a3ed

Observation 0972ff51-7be1-487c-8a16-2764c24ce67a · inbound

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation cites this paper.

Cost-Aware Multi-Objective Bandits: Theory and Application to Budgeted LLM Configuration Evaluation Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T19:48:08.805698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T19:48:08.805698Z digest=sha256:c5ac1bc3dcad9e70df00f2552c03af84e7a377351f9c789012ce0b3a72567f4f