Pith. sign in

Paper Citation Record · LEDGER

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

As of 11 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.11432.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.11432 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T05:38:58.035337Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 554e12db-c1dc-4230-9f60-07df4766662e · outbound

This paper cites Concrete Problems in AI Safety.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:11b9daaac6fce11346aa1c11a0cc9922aad8e3f0edb4625a4e42c081cac73dee

Observation f0f231be-1acc-4612-8e83-f3a3a9ba6fef · outbound

This paper cites Active Preference-Based Gaussian Process Regression for Reward Learning.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Active Preference-Based Gaussian Process Regression for Reward Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:955eaf613b566e9d4f42e883f4bfe909ce78c529bbb6f855c954b8f59694dd07

Observation 4c80310a-3406-43f2-b43a-25b1b79ce8e8 · outbound

This paper cites an unresolved cited work.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:19a214cba09742be352b9c6d87f56b9c47f4f3e806d420beb24cbbf68b524342

Observation 573a5b80-221b-4ab5-b35d-a7b27625e9ea · outbound

This paper cites Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:5aa731c4d1a7ba4cf08b07397e3d5c70ce0c96690f26fa0a970f94b49dbd4e83

Observation 9f42a0cd-1d04-476e-80ce-b0d5d087b732 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:0a666941dbf6e206c574edc731fcd3adc85d8aa0cdc95a7cdb11b655a30f31e2

Observation 641d77a2-c8d0-49a0-9ccf-8864d3194904 · outbound

This paper cites Multi-Objective Reinforcement Learning.MORL addresses the challenge of optimizing multiple, often conflicting, criteria to discover a set of Pareto-optimal policies.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Multi-Objective Reinforcement Learning.MORL addresses the challenge of optimizing multiple, often conflicting, criteria to discover a set of Pareto-optimal policies

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:cc20029418b9eb4f9f6c152a311d0ed839608c885116e9494d3c35aba929e8cb

Observation 77ecc75e-c5a9-4ce1-92cb-e8c05c63fd82 · outbound

This paper cites The CbRL setting has the same learning goal, i.e., recovering the Pareto frontier.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability The CbRL setting has the same learning goal, i.e., recovering the Pareto frontier

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:689cf2509047d392800cd6aebe59dbd07d3bb027c81609efc2ab8e3422a71409

Observation 9038cc68-5e3c-4179-aa88-30ed3e1f79d5 · outbound

This paper cites an unresolved cited work.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:10ffa1edf081f1d21cbd208a4f81747fc13e2bba3d3663df66d7b9ebcda3ab01

Observation f577a497-ff57-44a1-a624-95479a8935e1 · outbound

This paper cites weight” to refer to the scalarization weight of the multiple objectives rather than the term “preference.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability weight” to refer to the scalarization weight of the multiple objectives rather than the term “preference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:73af76b7970d243747875f343fc95733e43e77985188cd0267534d08ac85c48d

Observation 14d1e812-55a5-4c12-979f-51514a55e527 · outbound

This paper cites In recent years, novel rationality models have been proposed to tackle specific limitations of BT.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability In recent years, novel rationality models have been proposed to tackle specific limitations of BT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:b28d6d44048c2e70a53cd5db45702ffe725bcc17252687536f1ab47e9052a5fe

Observation dd800939-999e-49a1-8751-35540713434e · outbound

This paper cites Finally, thegeneral preference optimization (Zhang et al.,.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Finally, thegeneral preference optimization (Zhang et al.,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:c86b58d4b405bd87ea65ebb3ef21bc6fb0948805882240e44824bd8338d5778e

Observation d12fb97f-66ef-41bf-b6db-5624aba5c736 · outbound

This paper cites an unresolved cited work.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:4876fc0cc0716077740d7ef95e718347c10ac8540049298d15516d8361622262

Observation c0526321-cdfe-4ba5-bd8a-30c18673c5b2 · outbound

This paper cites threshold of the sensory perception of the judge.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability threshold of the sensory perception of the judge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:5e09c61bbd977fa07bfd7c784c2cc41b645f5251e7e42d35bc25f3af6563c530

Observation 39fd50f1-3fbe-4eb1-9687-85274b0d49ed · outbound

This paper cites The standard approach in the presence of incomparability in PbRL is to discard the query and consider the sample as erroneous (Christiano et al., 2017).

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability The standard approach in the presence of incomparability in PbRL is to discard the query and consider the sample as erroneous (Christiano et al., 2017)

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:d31191134c23308a12f442599e5e4ca39403b87c04b90bb9aee3b48320612ffd

Observation 7746d8fb-8d28-41de-9673-0c30bb03e7a6 · outbound

This paper cites With respect toα, we have: ››∇αhąpτ, τ 1|θq ›› 2 “0, ››∇αhăpτ, τ 1|θq ›› 2 “0,››∇αh—pτ, τ 1|θq ›› 2 “1, ››∇αh∥pτ, τ 1|θq ›› 2 “0.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability With respect toα, we have: ››∇αhąpτ, τ 1|θq ›› 2 “0, ››∇αhăpτ, τ 1|θq ›› 2 “0,››∇αh—pτ, τ 1|θq ›› 2 “1, ››∇αh∥pτ, τ 1|θq ›› 2 “0

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:3e0038132383d94ae85435af0253a6817e590be6edce822946bc8eb11c0e4579

Observation 8db5559e-cbee-458d-9dc1-aff1040a8a0e · outbound

This paper cites an unresolved cited work.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:9f2584dc5c1cd9570f4980edb3919102771bc4fc013cc12cdfe2e086defe0601

Observation b6d90144-a987-4d32-b713-d28d1f8cfcb2 · outbound

This paper cites First, we discuss the methodology of our experiments in Appendix D.1, covering the procedures for generating trajectories and comparison labels.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability First, we discuss the methodology of our experiments in Appendix D.1, covering the procedures for generating trajectories and comparison labels

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:35bc130a48afdcef8ac28e8ab108b3db47cf951a98b9dfba2302b15877d6de06

Observation 0f6386e3-b5ae-4c7c-95bb-acde4d51fd6c · outbound

This paper cites MO-Hopper.We use the mo-hopper-2obj-v5 environment from the MO-Gymnasium suite (Fel- ten et al., 2023), the multi-objective extension of Hopper-v5 (Towers et al., 2025).

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability MO-Hopper.We use the mo-hopper-2obj-v5 environment from the MO-Gymnasium suite (Fel- ten et al., 2023), the multi-objective extension of Hopper-v5 (Towers et al., 2025)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:18c35b451569b48602ac7c857373bea649bfdf1e78c56ca6da3d288a47f42620

Observation 7c06ebbf-2c68-4d63-9ceb-e050687b2a5d · outbound

This paper cites The Pareto frontier admits a closed-form solution under the LQR formalism, providing an exact reference for policy-level evaluation.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability The Pareto frontier admits a closed-form solution under the LQR formalism, providing an exact reference for policy-level evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:224c8c81e98b4f0b191289e1b87001519ec31695cf31816da9ada6fd0c434563

Observation dd652548-bd16-4374-b0b2-2d7c0d388737 · outbound

This paper cites One common practical refinement to improve the overall performance is to employ an ensemble of predictors trained from randomized initializations.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability One common practical refinement to improve the overall performance is to employ an ensemble of predictors trained from randomized initializations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:135e0774285c443fb7285f1381adfc4a99a80a9df6777447153158c725b4b354

Observation 714c95c1-523b-485a-b87e-0a979c81b9e3 · outbound

This paper cites an unresolved cited work.

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T05:38:58.035337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:38:58.035337Z digest=sha256:62d6d0c02d4b4f35aab780608cc09fd8866231ade95703adedf8e96728741230

Pith citing papers

No inbound Pith citation observations are available.