Pith. sign in

Paper Citation Record · LEDGER

Preference-based Multi-Objective Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.14066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14066 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:17:35.918503Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fc1a82f-f4fb-4255-be81-4d7da518b296 · outbound

This paper cites A survey on modeling and optimizing multi-objective systems,.

Preference-based Multi-Objective Reinforcement Learning A survey on modeling and optimizing multi-objective systems,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.328027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:32.960135Z digest=sha256:1bad659cc6a349535c13f2d51bf23180fbe00e7b4a42ab380c65318412bd7213

Observation ae06a308-6755-455b-a074-e2e0c3686134 · outbound

This paper cites Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,.

Preference-based Multi-Objective Reinforcement Learning Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.320638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:33.086379Z digest=sha256:476dca4accb1d24adac8e02a035f9f4d5f003c6fd0541257b1648bd3abb54bcf

Observation 0c8692a9-07b7-45d4-b321-303a35428fbd · outbound

This paper cites Constrained ordinal opti- mization—a feasibility model based approach,.

Preference-based Multi-Objective Reinforcement Learning Constrained ordinal opti- mization—a feasibility model based approach,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.313585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:33.158951Z digest=sha256:4a1060a611bea65ca8a6bb03bf4fee29b46f5795a02255fd72eda4e8e9f8ae43

Observation c696ac33-1444-4b15-8189-5e10c27456d2 · outbound

This paper cites Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,.

Preference-based Multi-Objective Reinforcement Learning Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.306194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:33.319524Z digest=sha256:6b2e7c5777afc185fd9d3de52fb620627cc3d4e6215855986d41b6979a104245

Observation 2d48f134-df5b-413c-a3aa-0c3ee75c410e · outbound

This paper cites Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,.

Preference-based Multi-Objective Reinforcement Learning Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.298042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:33.469179Z digest=sha256:9efa434f77451c97fb7cb3e5d99aa883c6045d6325c11dc0eb4bc9f75f645a85

Observation 716eb782-3d6c-4a6c-8ed9-1be6e67e00ff · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning B-pref: Benchmarking preference-based reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.290058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:33.685131Z digest=sha256:960cbdaac2e5d710f901e70470d45536dd9b295923d92d800b9588e5fb231e47

Observation 9e9b1cfd-c2d6-4a14-9fc6-cbd0d5bc2d51 · outbound

This paper cites Self- supervised online reward shaping in sparse-reward environments,.

Preference-based Multi-Objective Reinforcement Learning Self- supervised online reward shaping in sparse-reward environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.282576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:33.851165Z digest=sha256:2764263dbcb90959dca275c56fff105214c51959746e22796a437ea6a4a4902d

Observation 13ceb81f-37c5-41c2-83c9-72311c454326 · outbound

This paper cites Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.275241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:34.081232Z digest=sha256:a14bb97b0021b9ca6f27e91065f9475138c52c8a30aa1d995fdda50585346bb9

Observation a2583ef5-ca33-4e10-82ce-64ef063c2802 · outbound

This paper cites Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,.

Preference-based Multi-Objective Reinforcement Learning Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.267833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:34.340535Z digest=sha256:4e0053a1d89682ea111d3cebfed54745c4347d6dd67c9755feaae557f2b3f919

Observation 4537b1db-9c38-4270-ae4b-71a51d79da40 · outbound

This paper cites Reinforcement learning and the reward engineering princi- ple,.

Preference-based Multi-Objective Reinforcement Learning Reinforcement learning and the reward engineering princi- ple,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.260407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:34.526104Z digest=sha256:646119fd34157394b52380a8a0215354bee61a300dd12665549faa22a0e19da8

Observation 24ca2cd7-5064-40f5-99ea-12b31bbaef82 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Preference-based Multi-Objective Reinforcement Learning Deep reinforcement learning from human preferences,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.529527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.529527Z digest=sha256:3432fe208c4819f28c03ed379290575a9b01afa7153f927f9dc35dd533dc51a4

Observation b220a6de-e529-44a7-a136-6f85cbaf7a7f · outbound

This paper cites A bayesian approach for policy learning from trajectory preference queries,.

Preference-based Multi-Objective Reinforcement Learning A bayesian approach for policy learning from trajectory preference queries,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.247729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:34.569067Z digest=sha256:11cf8d53648d8c45e49c24d730025683546a6941671ec68050350fbecea9d049

Observation 5e73303c-0400-4634-b2d6-680945ac21e7 · outbound

This paper cites an unresolved cited work.

Preference-based Multi-Objective Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.641984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.641984Z digest=sha256:1652d906fe2496be561cb515ef010b4a440df2856ec41c8cd9baf4efd60519b0

Observation 59e50602-f0f1-4bc3-8a65-75e1f0d8e058 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.803156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.803156Z digest=sha256:19546d346093334560f56c70dd9ed800abc623f6fb4fa226c6e4916dd0359633

Observation 45c39390-16cf-419b-8bc3-0cf91b5b9b7f · outbound

This paper cites E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,.

Preference-based Multi-Objective Reinforcement Learning E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.235748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:34.908704Z digest=sha256:5776c4a51679abd34ade22a08ab9ff9dd08ee8358fb41709a27bf32c0b38d67d

Observation 188d3ba9-859c-4067-93ac-7c801054fb7e · outbound

This paper cites Mastering the game of go without human knowledge,.

Preference-based Multi-Objective Reinforcement Learning Mastering the game of go without human knowledge,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.990348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.990348Z digest=sha256:72b9ee4a803598c8a293441a53b7a49797b03700a554eee9fae1962dd3f75b65

Observation 191fe72d-0e3f-4e18-92d0-ca2e6acdd8a8 · outbound

This paper cites Simplify twin crane scheduling in railway yard by spatial task assignment,.

Preference-based Multi-Objective Reinforcement Learning Simplify twin crane scheduling in railway yard by spatial task assignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.223250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.073240Z digest=sha256:c342c94d4cbfd4641d784e53bf8c508934de10de6ee5794757ce826b1fa64d1e

Observation b5bb24a4-fef6-4148-a288-c2e0654dcf78 · outbound

This paper cites Large-scale data center cooling control via sample-efficient reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Large-scale data center cooling control via sample-efficient reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.215576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.159987Z digest=sha256:392543c7a6310433e13624cf2a533ec7c88e060d8e4b2708abfc905c79cb9744

Observation a82b67cb-516f-4951-92b1-4de16f88d1de · outbound

This paper cites An efficient real- time railway container yard management method based on partial de- coupling,.

Preference-based Multi-Objective Reinforcement Learning An efficient real- time railway container yard management method based on partial de- coupling,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.208046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.240063Z digest=sha256:37499204fbaa71a4f0fb753f56c27f08a8ecfc1f42f2bda96d2738a25a1ccab3

Observation 97ff2a81-6704-405c-af47-15839da3abfc · outbound

This paper cites Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,.

Preference-based Multi-Objective Reinforcement Learning Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.200746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.322733Z digest=sha256:43f4d6a64cad259ca69b2a0a2b519d1552847315f9ed3901897a0f448e360065

Observation 1dffd4ba-39e9-4fea-b1e1-816f02af9399 · outbound

This paper cites Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.192649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.382374Z digest=sha256:263f56f3964d3a08e73e62ff82d71cdfa055a786448c89c9e16261e91a7a45c3

Observation 7786b8cc-d84f-4ec2-a684-b758ac9015f7 · outbound

This paper cites OpenAI Gym.

Preference-based Multi-Objective Reinforcement Learning OpenAI Gym

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.450783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.450783Z digest=sha256:f5e5a0d36daa58ab737ae789cccf47c5d21f5d206d95ae419c5b6ef9e1903c39

Observation 7be9d931-503a-43e0-9f1a-065a9f2119ae · outbound

This paper cites Exploration by random network distillation,.

Preference-based Multi-Objective Reinforcement Learning Exploration by random network distillation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.184804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.511221Z digest=sha256:5ebf1a55343b8bb7e119a442551de9d035f3744097c81821137a55b588e9927d

Observation 88d8419b-af3a-4b5b-8697-f17d167b40be · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Preference-based Multi-Objective Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.584119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.584119Z digest=sha256:e990d4383f571ae08da23afd9dc0802715c6afa955fd42631de0151b7a7d3fc3

Observation 2af1f740-84f9-45f4-8b03-a71c0e055626 · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

Preference-based Multi-Objective Reinforcement Learning Reward learning from human preferences and demonstrations in atari,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.173015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.668501Z digest=sha256:3081733aa22956feef753eb8f77ea4835b48d78e1edbc957840a86b64b9207d2

Observation 86d955f6-956d-453c-b71f-bf247c652a92 · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:17:35.967349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.749954Z digest=sha256:b89f54f30765f1da6ef56902f84d5047b05923a5c9070ccfce8ed71c6ef5e4c7

Observation 8e1175aa-1ece-41da-a025-3a324564b00f · outbound

This paper cites Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.165275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.810695Z digest=sha256:35f941f39be21c2a8c709c4e080dfd097704fa898c75a58fd1c9ece8d3583f82

Observation b70ba05f-268c-45fb-9cf8-5e797f7c0d36 · outbound

This paper cites Few-shot preference learning for human- in-the-loop rl,.

Preference-based Multi-Objective Reinforcement Learning Few-shot preference learning for human- in-the-loop rl,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.156940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.840299Z digest=sha256:1699da5d23dc9c3dbbe0b6d91890e817ce59cdb19f6c6d87f8f8f78a6cd6e308

Observation ad4b84f9-c650-4b4c-a6a4-c2e3eba99b98 · outbound

This paper cites Learning to summarize with human feedback,.

Preference-based Multi-Objective Reinforcement Learning Learning to summarize with human feedback,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.844045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.844045Z digest=sha256:bfc42d4f55eaaddab46d6bfe87e20b7141280c0444665f4c1519d2348c7e4511

Observation 468c6574-e772-4935-be5d-7550d1d20647 · outbound

This paper cites A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.145259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.848219Z digest=sha256:cf8e3704fde40bb08f46e6fe1d9a321e578516f55389aff29c3275e2851834d5

Observation 112be415-f344-4a58-af0f-03915d8a1d6e · outbound

This paper cites A practical guide to multi-objective reinforcement learning and planning,.

Preference-based Multi-Objective Reinforcement Learning A practical guide to multi-objective reinforcement learning and planning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.137836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.851254Z digest=sha256:901615736b9ed2adad408b47fe812e0d4eab1ad32aeb2cf5d6eedf0cd97bdcdf

Observation 420d363e-a9f4-47b2-bd3e-6ef4f759e91c · outbound

This paper cites Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:17:35.956243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.855636Z digest=sha256:91fd8c3656b819b103b60f952094e932e46c58e37de9ac021305afbcc1bdf23b

Observation 2712168f-f393-4094-ab35-cbfe5a78983a · outbound

This paper cites A generalized algorithm for multi-objective reinforcement learning and policy adaptation,.

Preference-based Multi-Objective Reinforcement Learning A generalized algorithm for multi-objective reinforcement learning and policy adaptation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.130297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.859598Z digest=sha256:80731a840ebeadaab61e19fdec62226e6716c93b98ae09e3f427cca9c2394be9

Observation 9eb1a771-cbb8-4150-a54b-76732f7c49cf · outbound

This paper cites Multi-objective rein- forcement learning for the expected utility of the return,.

Preference-based Multi-Objective Reinforcement Learning Multi-objective rein- forcement learning for the expected utility of the return,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.122647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.863547Z digest=sha256:1418ed18ab8651d3a728bbd2007348742bed730532c2365fb5748f9c1e02b4e2

Observation e893cdc1-7bfb-4641-8956-8dd0749bdc23 · outbound

This paper cites Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,.

Preference-based Multi-Objective Reinforcement Learning Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.865902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.865902Z digest=sha256:a47a82a385b3fdf595e0d2f84acebbef7989908cd88dcc8814c7c918a0a1ed6b

Observation a2dca11c-79cc-4b92-a95e-a824b1f95d7a · outbound

This paper cites Pareto conditioned net- works,.

Preference-based Multi-Objective Reinforcement Learning Pareto conditioned net- works,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.109162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.869150Z digest=sha256:04f516cf5fa1d7e5df4fff8b5dc88b85510f3ab8716cd3e873635c11976d1bb4

Observation 1386df91-dd9e-4e2f-b578-ea283fac0c47 · outbound

This paper cites Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,.

Preference-based Multi-Objective Reinforcement Learning Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.101150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.871852Z digest=sha256:f9cd652d2f6b1b3d50729dcb357638a06688286b50503ff52d0574ab53cfe775

Observation 08479a28-5f7d-426b-b471-a1c561ec2fc0 · outbound

This paper cites Q-learning,.

Preference-based Multi-Objective Reinforcement Learning Q-learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.874650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.874650Z digest=sha256:f018710e9dc24f8cc7e63c8aa836b2c65a60da726d3ce7d12daaec49b47a7885

Observation 8c28b02e-16f2-441c-8153-7ae6a43bdfd5 · outbound

This paper cites an unresolved cited work.

Preference-based Multi-Objective Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:17:36.089262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.876997Z digest=sha256:d4c84af4146c1e807440bef85e22ac8bb3ced165fbcbba7c10821be87e9234ae

Observation cdb7ec45-fc02-495e-9814-7e508caa2089 · outbound

This paper cites Convergence of q-learning: A simple proof,.

Preference-based Multi-Objective Reinforcement Learning Convergence of q-learning: A simple proof,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.879279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.879279Z digest=sha256:8000e18fc3e9262993889060fd0e7d5cc1bb1b64bda54283a1c85b885a951844

Observation 4dfa9a84-855f-4b7c-b202-e87470ba748c · outbound

This paper cites Decentralized multi-agent reinforcement learning: An off-policy method,.

Preference-based Multi-Objective Reinforcement Learning Decentralized multi-agent reinforcement learning: An off-policy method,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.076420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.881717Z digest=sha256:93ed6db78b9b304727ece31c6c14e728f1302109fc8ec0b0d7969c638717207f

Observation 7a7e2417-c332-43fb-bf14-0778f084600e · outbound

This paper cites An ocba-based method for efficient sample collection in reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning An ocba-based method for efficient sample collection in reinforcement learning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.068792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.884101Z digest=sha256:abfdf2bc33b8ad57e8a69434ba795caa08aa96d5f0587a0d3f090828bb7a11eb

Observation 3ddde7a4-ef4e-4b4b-beb2-67bb2f31f995 · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Preference-based Multi-Objective Reinforcement Learning Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.886959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.886959Z digest=sha256:cc56c3029b93202f1a1353d9709fdb8cdbdaae6adfbeacd38aab3f8fb2cd3ad7

Observation efa1052b-c7ae-4221-9dfe-a7afba883828 · outbound

This paper cites Preference-based multi-objective reinforcement learning with explicit reward modeling,.

Preference-based Multi-Objective Reinforcement Learning Preference-based multi-objective reinforcement learning with explicit reward modeling,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.057027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.889314Z digest=sha256:a4a75d86d32b0cbc38870b3b719a958cac9f3e8a985440ac84b11375aeeca71a

Observation fb1ad97a-90c7-48d6-a16e-e32ffce1f627 · outbound

This paper cites Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,.

Preference-based Multi-Objective Reinforcement Learning Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.048623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.891737Z digest=sha256:d3a085c4caf8243780ba2c47ef208941dd33907652dcb1a9ceb113c29b57e490

Observation 01e14327-69ca-4bd8-9cc6-d519c38dbc89 · outbound

This paper cites Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications.

Preference-based Multi-Objective Reinforcement Learning Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.040417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.894131Z digest=sha256:e57d858e79dc7e912786e1ce25161edad46132ee7d0a5efb1326011a2f9fb35b

Observation 5470b709-8978-47df-bc91-e5981e8a03b3 · outbound

This paper cites Query-policy mis- alignment in preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Query-policy mis- alignment in preference-based reinforcement learning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.032810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.896952Z digest=sha256:c651a690a8e17c88bca8f161b8a85e81c2357cc2c2ba2172f3b0221c1c94a5dc

Observation 04d09981-ffca-4a2a-bd60-afbd07ba336f · outbound

This paper cites Empirical evaluation methods for multiobjective reinforcement learning algorithms,.

Preference-based Multi-Objective Reinforcement Learning Empirical evaluation methods for multiobjective reinforcement learning algorithms,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.024987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.899502Z digest=sha256:b68f8cd4539556f7462d4bb33e0ad3c6ed452cf6d268d13428135baddfaac3ad

Observation 5c6d56b8-ce60-4bb6-bd21-65d694dd9faf · outbound

This paper cites Learning all optimal policies with multiple criteria,.

Preference-based Multi-Objective Reinforcement Learning Learning all optimal policies with multiple criteria,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.017157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.902317Z digest=sha256:71379c43d319d8710103012b22d0e4c67bd3e1c85dd1862ce7c04c41d45a5edc

Observation c6af2e34-a117-40a6-89e9-3164030e5849 · outbound

This paper cites An environment for autonomous driving decision-making,.

Preference-based Multi-Objective Reinforcement Learning An environment for autonomous driving decision-making,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.905119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.905119Z digest=sha256:a72c22039291e9671a4eaa025251f8f5d8b485761d32dd61170463b898b0afee

Observation 99c5bb15-f318-4c62-bf1a-764392ecd9d8 · outbound

This paper cites Congested traffic states in empirical observations and microscopic simulations,.

Preference-based Multi-Objective Reinforcement Learning Congested traffic states in empirical observations and microscopic simulations,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.907272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.907272Z digest=sha256:91e4933b477ba75817b876e17b108e65bacf6cd116e765c73a4ed13e01e64496

Observation dc9de4e9-cb9f-4b63-b0de-a2ba264dd6e8 · outbound

This paper cites General lane-changing model mobil for car-following models,.

Preference-based Multi-Objective Reinforcement Learning General lane-changing model mobil for car-following models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.909933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.909933Z digest=sha256:cdf0be87e10f2ceb4ee3600d86f8ec1a2f89b5791cb412791dcd6975fac5d701

Observation c6181032-1c83-4bd9-8cd5-c8e95428db13 · outbound

This paper cites Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,.

Preference-based Multi-Objective Reinforcement Learning Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:35.996133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.912851Z digest=sha256:581902ae401ff71a694339806b120b303c93c680172fa624f0f1fc55e33cb5df

Observation c10ce3c5-a198-4c91-8f6b-450aeb1c1485 · outbound

This paper cites Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation.

Preference-based Multi-Objective Reinforcement Learning Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.915292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.915292Z digest=sha256:037b2beb93a7130c4ab51fec1c4d893799a7b5cc92465867b7c1cf82ba3ab8fd

Observation 2ab81b5d-f8d5-4169-aca5-ff1f79acf333 · outbound

This paper cites Listwise reward estimation for offline preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Listwise reward estimation for offline preference-based reinforcement learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:35.988790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T16:17:35.918503Z digest=sha256:6aeced1a37d9471c928c13348596875a19ee4d244debf333834f9286b2eca43e

Pith citing papers

No inbound Pith citation observations are available.