Pith. sign in

Paper Citation Record · LEDGER

Preference-based Multi-Objective Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.14066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14066 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:17:35.918503Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fc1a82f-f4fb-4255-be81-4d7da518b296 · outbound

This paper cites A survey on modeling and optimizing multi-objective systems,.

Preference-based Multi-Objective Reinforcement Learning A survey on modeling and optimizing multi-objective systems,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.328027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:32.960135Z digest=sha256:a68b6f6cc18446f090e90c4fcb5d947733075456b563f86cb36e4986b5fc1bac

Observation ae06a308-6755-455b-a074-e2e0c3686134 · outbound

This paper cites Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,.

Preference-based Multi-Objective Reinforcement Learning Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.320638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:33.086379Z digest=sha256:a299d1883d7a6bb1f8fb89b61486f6990c880e8f115cc55fd914e163d58be35a

Observation 0c8692a9-07b7-45d4-b321-303a35428fbd · outbound

This paper cites Constrained ordinal opti- mization—a feasibility model based approach,.

Preference-based Multi-Objective Reinforcement Learning Constrained ordinal opti- mization—a feasibility model based approach,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.313585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:33.158951Z digest=sha256:a1466ac763b2bba64f5ee33328721edc3df8645191e9f7b785c1bd251a4e0add

Observation c696ac33-1444-4b15-8189-5e10c27456d2 · outbound

This paper cites Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,.

Preference-based Multi-Objective Reinforcement Learning Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.306194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:33.319524Z digest=sha256:d1e0e6fd43f338785eedc546a9c95dad904b1ec841b88e998fbcc504621dfdc7

Observation 2d48f134-df5b-413c-a3aa-0c3ee75c410e · outbound

This paper cites Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,.

Preference-based Multi-Objective Reinforcement Learning Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.298042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:33.469179Z digest=sha256:b9cf371341b770d39ce5426c519fe157a6c4046873e1392ca3c1ac5daaaf65db

Observation 716eb782-3d6c-4a6c-8ed9-1be6e67e00ff · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning B-pref: Benchmarking preference-based reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.290058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:33.685131Z digest=sha256:a56e0d7f68b02ccd71587f11a14bd52a7edee2d8d9e2165415de1775ca4d7bfe

Observation 9e9b1cfd-c2d6-4a14-9fc6-cbd0d5bc2d51 · outbound

This paper cites Self- supervised online reward shaping in sparse-reward environments,.

Preference-based Multi-Objective Reinforcement Learning Self- supervised online reward shaping in sparse-reward environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.282576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:33.851165Z digest=sha256:ded22f615c3009d6e4cc08b03ad74863a51cfbc78d5fc42c7e2b2d5196cd9797

Observation 13ceb81f-37c5-41c2-83c9-72311c454326 · outbound

This paper cites Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.275241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:34.081232Z digest=sha256:b227128c84f7d51131fbe21708c0203480b5f07a34c27dd028378422c511e172

Observation a2583ef5-ca33-4e10-82ce-64ef063c2802 · outbound

This paper cites Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,.

Preference-based Multi-Objective Reinforcement Learning Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.267833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:34.340535Z digest=sha256:30801ab43d3516ff3466d4e2623a9a5e3d685b7ed9d27cdf6966686d482b8887

Observation 4537b1db-9c38-4270-ae4b-71a51d79da40 · outbound

This paper cites Reinforcement learning and the reward engineering princi- ple,.

Preference-based Multi-Objective Reinforcement Learning Reinforcement learning and the reward engineering princi- ple,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.260407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:34.526104Z digest=sha256:4ef062c26a481ee55355ccb91ea1ec23b72fa6333eb20a23daa3afb0bc3348cd

Observation 24ca2cd7-5064-40f5-99ea-12b31bbaef82 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Preference-based Multi-Objective Reinforcement Learning Deep reinforcement learning from human preferences,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.529527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.529527Z digest=sha256:3432fe208c4819f28c03ed379290575a9b01afa7153f927f9dc35dd533dc51a4

Observation b220a6de-e529-44a7-a136-6f85cbaf7a7f · outbound

This paper cites A bayesian approach for policy learning from trajectory preference queries,.

Preference-based Multi-Objective Reinforcement Learning A bayesian approach for policy learning from trajectory preference queries,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.247729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:34.569067Z digest=sha256:b2f9fb2b05f1d2757ac403e7b06e28297dba4b667c4908e25ca1107c206af6ef

Observation 5e73303c-0400-4634-b2d6-680945ac21e7 · outbound

This paper cites an unresolved cited work.

Preference-based Multi-Objective Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.641984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.641984Z digest=sha256:1652d906fe2496be561cb515ef010b4a440df2856ec41c8cd9baf4efd60519b0

Observation 59e50602-f0f1-4bc3-8a65-75e1f0d8e058 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.803156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.803156Z digest=sha256:19546d346093334560f56c70dd9ed800abc623f6fb4fa226c6e4916dd0359633

Observation 45c39390-16cf-419b-8bc3-0cf91b5b9b7f · outbound

This paper cites E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,.

Preference-based Multi-Objective Reinforcement Learning E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.235748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:34.908704Z digest=sha256:31f04cc8341e079c619a5b8af7c44a3b7ca9b75265b2017db5ecb4b0064817c4

Observation 188d3ba9-859c-4067-93ac-7c801054fb7e · outbound

This paper cites Mastering the game of go without human knowledge,.

Preference-based Multi-Objective Reinforcement Learning Mastering the game of go without human knowledge,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.990348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.990348Z digest=sha256:72b9ee4a803598c8a293441a53b7a49797b03700a554eee9fae1962dd3f75b65

Observation 191fe72d-0e3f-4e18-92d0-ca2e6acdd8a8 · outbound

This paper cites Simplify twin crane scheduling in railway yard by spatial task assignment,.

Preference-based Multi-Objective Reinforcement Learning Simplify twin crane scheduling in railway yard by spatial task assignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.223250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.073240Z digest=sha256:3375783af9686a1700527ada24d646d7d7a2d6b5a6e631db3f4dae8f9edb512a

Observation b5bb24a4-fef6-4148-a288-c2e0654dcf78 · outbound

This paper cites Large-scale data center cooling control via sample-efficient reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Large-scale data center cooling control via sample-efficient reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.215576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.159987Z digest=sha256:cb4c7128369e418959f88d322a25adba510653d16f309048a1b5eb6a766312c9

Observation a82b67cb-516f-4951-92b1-4de16f88d1de · outbound

This paper cites An efficient real- time railway container yard management method based on partial de- coupling,.

Preference-based Multi-Objective Reinforcement Learning An efficient real- time railway container yard management method based on partial de- coupling,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.208046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.240063Z digest=sha256:32b5706092fc0e0d2416b998a173512d42c25340e6d57792e82de5bd938f2dbe

Observation 97ff2a81-6704-405c-af47-15839da3abfc · outbound

This paper cites Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,.

Preference-based Multi-Objective Reinforcement Learning Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.200746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.322733Z digest=sha256:ccb7f87f92e43cabc82ab9f1a4cbf78333e2901b901ad34dca4ebce6ffd3a27b

Observation 1dffd4ba-39e9-4fea-b1e1-816f02af9399 · outbound

This paper cites Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.192649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.382374Z digest=sha256:e6a2db4958962b3a160edcc78df8faf5d4045b5892ac72ad66d4c0343353ec35

Observation 7786b8cc-d84f-4ec2-a684-b758ac9015f7 · outbound

This paper cites OpenAI Gym.

Preference-based Multi-Objective Reinforcement Learning OpenAI Gym

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.450783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.450783Z digest=sha256:f5e5a0d36daa58ab737ae789cccf47c5d21f5d206d95ae419c5b6ef9e1903c39

Observation 7be9d931-503a-43e0-9f1a-065a9f2119ae · outbound

This paper cites Exploration by random network distillation,.

Preference-based Multi-Objective Reinforcement Learning Exploration by random network distillation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.184804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.511221Z digest=sha256:8bfdf3996f3b9718ca57a3e02d687610856f95f45c1a025e747e2406d665a807

Observation 88d8419b-af3a-4b5b-8697-f17d167b40be · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Preference-based Multi-Objective Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.584119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.584119Z digest=sha256:e990d4383f571ae08da23afd9dc0802715c6afa955fd42631de0151b7a7d3fc3

Observation 2af1f740-84f9-45f4-8b03-a71c0e055626 · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

Preference-based Multi-Objective Reinforcement Learning Reward learning from human preferences and demonstrations in atari,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.173015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.668501Z digest=sha256:76bdcd0b0f372f8c36c5d329798bedbf8542279341495656e6d935655c9b2cad

Observation 86d955f6-956d-453c-b71f-bf247c652a92 · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:17:35.967349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.749954Z digest=sha256:2dc19cad0ea403adca905bba0c21380bdf976a5b5e98a6a87faf27b83784c07b

Observation 8e1175aa-1ece-41da-a025-3a324564b00f · outbound

This paper cites Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.165275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.810695Z digest=sha256:5812bfd197d4d78db66f4727f8d345f4f51d4f8627692bfd09ead671303075ed

Observation b70ba05f-268c-45fb-9cf8-5e797f7c0d36 · outbound

This paper cites Few-shot preference learning for human- in-the-loop rl,.

Preference-based Multi-Objective Reinforcement Learning Few-shot preference learning for human- in-the-loop rl,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.156940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.840299Z digest=sha256:5aa142f3e6ff476f9a19882ff8b7869e18ceda221d7100d495f378c92d64b091

Observation ad4b84f9-c650-4b4c-a6a4-c2e3eba99b98 · outbound

This paper cites Learning to summarize with human feedback,.

Preference-based Multi-Objective Reinforcement Learning Learning to summarize with human feedback,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.844045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.844045Z digest=sha256:bfc42d4f55eaaddab46d6bfe87e20b7141280c0444665f4c1519d2348c7e4511

Observation 468c6574-e772-4935-be5d-7550d1d20647 · outbound

This paper cites A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.145259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.848219Z digest=sha256:c386d90acbb294f2f0b5c5ec092e6c14d030c387897bb58353591d7559556ebe

Observation 112be415-f344-4a58-af0f-03915d8a1d6e · outbound

This paper cites A practical guide to multi-objective reinforcement learning and planning,.

Preference-based Multi-Objective Reinforcement Learning A practical guide to multi-objective reinforcement learning and planning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.137836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.851254Z digest=sha256:72d5d65e3176f7a7a6870de72b729aa491fcbf81369240805bdb7e2e4dda549c

Observation 420d363e-a9f4-47b2-bd3e-6ef4f759e91c · outbound

This paper cites Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:17:35.956243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.855636Z digest=sha256:9af0a6bf5576f6fae0b5d3e50864c942e1832c8b34192af635ee1868e3808759

Observation 2712168f-f393-4094-ab35-cbfe5a78983a · outbound

This paper cites A generalized algorithm for multi-objective reinforcement learning and policy adaptation,.

Preference-based Multi-Objective Reinforcement Learning A generalized algorithm for multi-objective reinforcement learning and policy adaptation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.130297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.859598Z digest=sha256:419c65d224434e5eb761663b0e2a15a2c7a86252354ad7530d6d7128d9e015b4

Observation 9eb1a771-cbb8-4150-a54b-76732f7c49cf · outbound

This paper cites Multi-objective rein- forcement learning for the expected utility of the return,.

Preference-based Multi-Objective Reinforcement Learning Multi-objective rein- forcement learning for the expected utility of the return,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.122647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.863547Z digest=sha256:9aad8f56443e364b8dfbfd9fe9533ac53d1e167a789d5d2dcc6255d3c0914227

Observation e893cdc1-7bfb-4641-8956-8dd0749bdc23 · outbound

This paper cites Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,.

Preference-based Multi-Objective Reinforcement Learning Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.865902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.865902Z digest=sha256:a47a82a385b3fdf595e0d2f84acebbef7989908cd88dcc8814c7c918a0a1ed6b

Observation a2dca11c-79cc-4b92-a95e-a824b1f95d7a · outbound

This paper cites Pareto conditioned net- works,.

Preference-based Multi-Objective Reinforcement Learning Pareto conditioned net- works,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.109162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.869150Z digest=sha256:d11a1609bff5802ab0bb2e8361785393ee95a7bdadc7322ae2862766ca5a1cd8

Observation 1386df91-dd9e-4e2f-b578-ea283fac0c47 · outbound

This paper cites Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,.

Preference-based Multi-Objective Reinforcement Learning Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.101150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.871852Z digest=sha256:802212a8eb5e26164d6a89598c3fae23af4dd2b4a735a06556447b5314e695f8

Observation 08479a28-5f7d-426b-b471-a1c561ec2fc0 · outbound

This paper cites Q-learning,.

Preference-based Multi-Objective Reinforcement Learning Q-learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.874650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.874650Z digest=sha256:f018710e9dc24f8cc7e63c8aa836b2c65a60da726d3ce7d12daaec49b47a7885

Observation 8c28b02e-16f2-441c-8153-7ae6a43bdfd5 · outbound

This paper cites an unresolved cited work.

Preference-based Multi-Objective Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:17:36.089262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.876997Z digest=sha256:68f398dd549479046e9ac14d52f454ff78b5e0adfbbbe488caaf9462c9c264d5

Observation cdb7ec45-fc02-495e-9814-7e508caa2089 · outbound

This paper cites Convergence of q-learning: A simple proof,.

Preference-based Multi-Objective Reinforcement Learning Convergence of q-learning: A simple proof,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.879279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.879279Z digest=sha256:8000e18fc3e9262993889060fd0e7d5cc1bb1b64bda54283a1c85b885a951844

Observation 4dfa9a84-855f-4b7c-b202-e87470ba748c · outbound

This paper cites Decentralized multi-agent reinforcement learning: An off-policy method,.

Preference-based Multi-Objective Reinforcement Learning Decentralized multi-agent reinforcement learning: An off-policy method,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.076420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.881717Z digest=sha256:803ba69f759bd1e0d7ea858a2455d054894eb974262b6c4f0c470bee8ab1a265

Observation 7a7e2417-c332-43fb-bf14-0778f084600e · outbound

This paper cites An ocba-based method for efficient sample collection in reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning An ocba-based method for efficient sample collection in reinforcement learning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.068792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.884101Z digest=sha256:c2dd4860149bed5019047bb6a8b9c04b1ae57c7e54ab2cc6779b143328abadd0

Observation 3ddde7a4-ef4e-4b4b-beb2-67bb2f31f995 · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Preference-based Multi-Objective Reinforcement Learning Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.886959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.886959Z digest=sha256:cc56c3029b93202f1a1353d9709fdb8cdbdaae6adfbeacd38aab3f8fb2cd3ad7

Observation efa1052b-c7ae-4221-9dfe-a7afba883828 · outbound

This paper cites Preference-based multi-objective reinforcement learning with explicit reward modeling,.

Preference-based Multi-Objective Reinforcement Learning Preference-based multi-objective reinforcement learning with explicit reward modeling,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.057027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.889314Z digest=sha256:1e22e44fe26d895a9f43b4bb642bcae54ccd82e23b307e68294a1f6a51286915

Observation fb1ad97a-90c7-48d6-a16e-e32ffce1f627 · outbound

This paper cites Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,.

Preference-based Multi-Objective Reinforcement Learning Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.048623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.891737Z digest=sha256:d2b7c540dd1cff1dc01e82117eabbc986cb45c2ef308dc408514612b7de6118b

Observation 01e14327-69ca-4bd8-9cc6-d519c38dbc89 · outbound

This paper cites Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications.

Preference-based Multi-Objective Reinforcement Learning Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.040417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.894131Z digest=sha256:06537e8e2102e8e51451bb929bffd0b429412d3d73faa360cd619fa99822a981

Observation 5470b709-8978-47df-bc91-e5981e8a03b3 · outbound

This paper cites Query-policy mis- alignment in preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Query-policy mis- alignment in preference-based reinforcement learning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.032810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.896952Z digest=sha256:5e073ea78d53d178ee23f753f8be4d3b22a290cf64165ea31fa5fc56c88b0a42

Observation 04d09981-ffca-4a2a-bd60-afbd07ba336f · outbound

This paper cites Empirical evaluation methods for multiobjective reinforcement learning algorithms,.

Preference-based Multi-Objective Reinforcement Learning Empirical evaluation methods for multiobjective reinforcement learning algorithms,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.024987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.899502Z digest=sha256:b640c69baf79932990ece0fbb1cab3e4af00049fd74787ec2d9c595969d53fba

Observation 5c6d56b8-ce60-4bb6-bd21-65d694dd9faf · outbound

This paper cites Learning all optimal policies with multiple criteria,.

Preference-based Multi-Objective Reinforcement Learning Learning all optimal policies with multiple criteria,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.017157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.902317Z digest=sha256:6a5647ac94f7530cd29cfe1d0dc2cb882582614a04b1f6f139da2eb88b47d7b8

Observation c6af2e34-a117-40a6-89e9-3164030e5849 · outbound

This paper cites An environment for autonomous driving decision-making,.

Preference-based Multi-Objective Reinforcement Learning An environment for autonomous driving decision-making,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.905119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.905119Z digest=sha256:a72c22039291e9671a4eaa025251f8f5d8b485761d32dd61170463b898b0afee

Observation 99c5bb15-f318-4c62-bf1a-764392ecd9d8 · outbound

This paper cites Congested traffic states in empirical observations and microscopic simulations,.

Preference-based Multi-Objective Reinforcement Learning Congested traffic states in empirical observations and microscopic simulations,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.907272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.907272Z digest=sha256:91e4933b477ba75817b876e17b108e65bacf6cd116e765c73a4ed13e01e64496

Observation dc9de4e9-cb9f-4b63-b0de-a2ba264dd6e8 · outbound

This paper cites General lane-changing model mobil for car-following models,.

Preference-based Multi-Objective Reinforcement Learning General lane-changing model mobil for car-following models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.909933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.909933Z digest=sha256:cdf0be87e10f2ceb4ee3600d86f8ec1a2f89b5791cb412791dcd6975fac5d701

Observation c6181032-1c83-4bd9-8cd5-c8e95428db13 · outbound

This paper cites Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,.

Preference-based Multi-Objective Reinforcement Learning Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:35.996133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.912851Z digest=sha256:2bafc435f423e4c34ad3a94d9e163a128546fc9558737bb13a678ad5990f1beb

Observation c10ce3c5-a198-4c91-8f6b-450aeb1c1485 · outbound

This paper cites Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation.

Preference-based Multi-Objective Reinforcement Learning Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.915292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.915292Z digest=sha256:037b2beb93a7130c4ab51fec1c4d893799a7b5cc92465867b7c1cf82ba3ab8fd

Observation 2ab81b5d-f8d5-4169-aca5-ff1f79acf333 · outbound

This paper cites Listwise reward estimation for offline preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Listwise reward estimation for offline preference-based reinforcement learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:35.988790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T16:17:35.918503Z digest=sha256:928feaed780b105a4f236940e62cb3679a2779d5d539d4a0801c06e91cafeb44

Pith citing papers

No inbound Pith citation observations are available.