Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T05:38:58.035337Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.11432.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T05:38:58.035337Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 554e12db-c1dc-4230-9f60-07df4766662e · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Concrete Problems in AI Safety
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f231be-1acc-4612-8e83-f3a3a9ba6fef · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Active Preference-Based Gaussian Process Regression for Reward Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c80310a-3406-43f2-b43a-25b1b79ce8e8 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 573a5b80-221b-4ab5-b35d-a7b27625e9ea · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f42a0cd-1d04-476e-80ce-b0d5d087b732 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Proximal Policy Optimization Algorithms
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 641d77a2-c8d0-49a0-9ccf-8864d3194904 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Multi-Objective Reinforcement Learning.MORL addresses the challenge of optimizing multiple, often conflicting, criteria to discover a set of Pareto-optimal policies
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77ecc75e-c5a9-4ce1-92cb-e8c05c63fd82 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability The CbRL setting has the same learning goal, i.e., recovering the Pareto frontier
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9038cc68-5e3c-4179-aa88-30ed3e1f79d5 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f577a497-ff57-44a1-a624-95479a8935e1 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability weight” to refer to the scalarization weight of the multiple objectives rather than the term “preference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14d1e812-55a5-4c12-979f-51514a55e527 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability In recent years, novel rationality models have been proposed to tackle specific limitations of BT
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd800939-999e-49a1-8751-35540713434e · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Finally, thegeneral preference optimization (Zhang et al.,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12fb97f-66ef-41bf-b6db-5624aba5c736 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0526321-cdfe-4ba5-bd8a-30c18673c5b2 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability threshold of the sensory perception of the judge
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39fd50f1-3fbe-4eb1-9687-85274b0d49ed · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability The standard approach in the presence of incomparability in PbRL is to discard the query and consider the sample as erroneous (Christiano et al., 2017)
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7746d8fb-8d28-41de-9673-0c30bb03e7a6 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability With respect toα, we have: ››∇αhąpτ, τ 1|θq ›› 2 “0, ››∇αhăpτ, τ 1|θq ›› 2 “0,››∇αh—pτ, τ 1|θq ›› 2 “1, ››∇αh∥pτ, τ 1|θq ›› 2 “0
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db5559e-cbee-458d-9dc1-aff1040a8a0e · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d90144-a987-4d32-b713-d28d1f8cfcb2 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability First, we discuss the methodology of our experiments in Appendix D.1, covering the procedures for generating trajectories and comparison labels
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6386e3-b5ae-4c7c-95bb-acde4d51fd6c · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability MO-Hopper.We use the mo-hopper-2obj-v5 environment from the MO-Gymnasium suite (Fel- ten et al., 2023), the multi-objective extension of Hopper-v5 (Towers et al., 2025)
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c06ebbf-2c68-4d63-9ceb-e050687b2a5d · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability The Pareto frontier admits a closed-form solution under the LQR formalism, providing an exact reference for policy-level evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd652548-bd16-4374-b0b2-2d7c0d388737 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability One common practical refinement to improve the overall performance is to employ an ensemble of predictors trained from randomized initializations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 714c95c1-523b-485a-b87e-0a979c81b9e3 · outbound
Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.