Pith. sign in

Paper Citation Record · LEDGER

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.12529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12529 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:13.730009Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a838ea7f-1367-4cab-ade4-ff13cad9272b · outbound

This paper cites Sutton and Andrew G.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Sutton and Andrew G

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.636568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.574686Z digest=sha256:0e6f31f32d306d6f09046f49062286997a51626e75a855f032e3bba3eda484e5

Observation e41adb7b-fc05-47b2-8087-650acdee650b · outbound

This paper cites Reinforcement learning can be more efficient with multiple rewards.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reinforcement learning can be more efficient with multiple rewards

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.629294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.577558Z digest=sha256:2e3ae472699e7e6e85984ea3304d4517e9c7e3106c653fd908db5bfcf5689a25

Observation 31528ee8-6aa9-4fcc-830b-25e50396a312 · outbound

This paper cites Learning Agile Robotic Locomotion Skills by Imitating Animals.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning Agile Robotic Locomotion Skills by Imitating Animals

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.580172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.580172Z digest=sha256:02c3b8bd1abb41c14f70aae4ef0f288567f3ca09ab66e3ef71ec742f73948516

Observation 668de56f-1921-46d4-b2ec-7ffbe1be87e2 · outbound

This paper cites Deep object-centric represen- tations for generalizable robot learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep object-centric represen- tations for generalizable robot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.583522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.583522Z digest=sha256:43de0facec4a041dbe4d37f17d8af3eef41a1321c00f12ca17f42e4785819e2f

Observation b89ad36b-5273-462d-9fa7-3068b8804939 · outbound

This paper cites The ingredients of real world robotic reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning The ingredients of real world robotic reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.622631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.586670Z digest=sha256:b7005f9d4f01c0673437474967e65c649db61b42ed78d6fb067007f67b0e4c70

Observation 1b95cad8-1d9e-4eea-b430-9a45b1b06e2e · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.589096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.589096Z digest=sha256:0eb51cb9fa925b3fa7b04a017d12574d466bc921c290ba9b23495a2dc018e9e7

Observation 1211a9fc-d637-440a-8732-934998ae1c9a · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.591993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.591993Z digest=sha256:86c8483d33f60da7de1b98f9e59a907b391d0720903493a7cf7dc927c7a91dc5

Observation bcdaafa4-37b0-4089-8eca-e979692aa58c · outbound

This paper cites GPT-4 Technical Report.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.594860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.594860Z digest=sha256:ddf16d9a5196aa8755a8eb778cc95786c5dde9da1a363d338b729cc03e8fe418

Observation 69641a7f-1abc-4504-9b55-3c54a6ceec2f · outbound

This paper cites Training language models to follow instructions with human feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Training language models to follow instructions with human feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.615042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.597295Z digest=sha256:c56e006987fffe4a7666831b495f45a7ff839b50406eee18054f5a65feb1aa3d

Observation 19b879c1-f247-4979-b6df-efeccc565ea6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.599621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.599621Z digest=sha256:28cbb5bd7204d44d465e7404ec182a883069a0668e8c04b9c62fde1ead0efa05

Observation a2f8fbe3-ce5d-4f95-a555-30270e9eab9b · outbound

This paper cites Active Preference-Based Learning of Reward Functions.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Active Preference-Based Learning of Reward Functions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.602307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.602307Z digest=sha256:605e5b07ed3cfb2369590bfb5713597a41c0150f2917808795f7db811bea6741

Observation ca135057-2688-40fe-a535-3bc88ad74e26 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep Reinforcement Learning from Human Preferences

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.608347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.604747Z digest=sha256:698bafeead17dc860c219ebf2730f9c9789b25585e64563a38d1cd52d99afb6b

Observation f60f215e-05b2-4d97-b993-3dbe38972c4d · outbound

This paper cites A Survey of Preference-Based Reinforcement Learning Methods.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Survey of Preference-Based Reinforcement Learning Methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.601692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.607066Z digest=sha256:94a4123e57e48ba87834ca9d98a7003a98df06aa1cf81467ae2bfabca30b9c38

Observation dcc01f40-2cf0-4b44-9e41-5b13da365c11 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.609596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.609596Z digest=sha256:3e619e3378e6db59d9558beae0d17f99881c335295e969fe16078f244fc4165b

Observation bbed6be8-6b7e-4898-b8a2-4b2515e6542b · outbound

This paper cites Smith, and Pieter Abbeel.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Smith, and Pieter Abbeel

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.594628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.611610Z digest=sha256:996c6a0753cde3e914a1c9eafdd6a95b5179eb91b6fd58144a519216da3ba672

Observation 5a1f93fd-7264-4d23-8023-83386e640017 · outbound

This paper cites Few-shot preference learning for human-in-the-loop RL.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Few-shot preference learning for human-in-the-loop RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.579789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.616579Z digest=sha256:aa6661e8dba8af45c64828b9724a1eb738a3f83aaeef3ebf6d16d0fb0c03d763

Observation 886092de-c884-4349-aee7-dec74b9d0023 · outbound

This paper cites Bradley Knox, and Dorsa Sadigh.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Bradley Knox, and Dorsa Sadigh

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.572904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.618745Z digest=sha256:7613bb857a9b62c51b83ee518dbb96f906db71af5d150b6ee56b67fe1e7649e6

Observation aa775772-5afe-4056-bb72-43b3df819895 · outbound

This paper cites Inverse Preference Learning: Preference-based RL without a Reward Function.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Inverse Preference Learning: Preference-based RL without a Reward Function

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.559555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.623886Z digest=sha256:8a489b511af280df453d9ef98c4db93c88728e830ad6ece2a73e671ebc0d6595

Observation 327d9590-55cc-4c53-9ddb-7df41ee30b9b · outbound

This paper cites Direct Preference-based Policy Optimization without Reward Modeling.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference-based Policy Optimization without Reward Modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.552744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.626123Z digest=sha256:8c344101260d4eafd9b79df41795bddbabce6376f3e257bbbbc429bbe6868deb

Observation 364218c0-d7c7-44f6-a6e0-26c1e2dd405c · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.545970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.628391Z digest=sha256:23f2019dccad0662d3b6d88659fa48ddf28ab3d0b90539be0e9b4060c7583f78

Observation cba5540a-8275-4414-afb8-118e0aa90b58 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.538656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.631200Z digest=sha256:c65bff8c1aa435fd64c14aaa88608851d67cef212c7ec4faa67099335853c926

Observation 1cc52139-f6aa-46bc-8f50-89faafdb063a · outbound

This paper cites Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.531895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.633432Z digest=sha256:84ffcadb027f62903972cc7baf176f5c1840722d2f09c3118c22990e002051bb

Observation fb61f357-0b37-4a7c-bee7-13dff66eaae2 · outbound

This paper cites Rethinking reward modeling in preference-based large language model alignment.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking reward modeling in preference-based large language model alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.524528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.635700Z digest=sha256:feef7feb34d0f28a43d10a6bc9f4b9efbe1d5db919692974228090c81a4a0cba

Observation 4533caba-3b52-4501-8e24-c7b60990f44a · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Generalized preference optimization: A unified approach to offline alignment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.517863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.638150Z digest=sha256:c8143421cb60761dae9d67b13701c6adbae2d5a406815f4a79dacfe32f0a2975

Observation 8bad5c92-683e-4ad4-a002-cc55117e83cb · outbound

This paper cites Nash learning from human feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Nash learning from human feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.511138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.640321Z digest=sha256:9032a83b73ca3dee9f18609907829787594b4ae7eb00750fdf11eb4d1745962b

Observation 13bdd45a-96b2-43aa-865f-65b560504c57 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A general theoretical paradigm to understand learning from human preferences

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.504228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.643844Z digest=sha256:3c575ffce1bc16ebc87cc027f5c2ed54a15abcfe497db243144d2bbb3e5950ca

Observation 70cd97ea-e593-49fa-9ef1-b44d30a25af8 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.497602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.646342Z digest=sha256:dc087e19e0549cfaa930e776b971c7b07c2181b049443186eb798cc0116ae27b

Observation 681f80ec-a3bf-4e3e-96a8-0c1c4f32eda0 · outbound

This paper cites Intransitivity of preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Intransitivity of preferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.648723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.648723Z digest=sha256:e9ef75400dab47d56ee8b336aa7ccd6c45f6436353036de2e6fe44af8c6536d2

Observation a5f4cb82-180f-41ad-bc80-26a2ee1c3fe5 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:14.491103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.651354Z digest=sha256:00703dce9b6e81a78f9a262c1fa78db44d0a4faf12ebb267a40eb158db705ac5

Observation fa32ebb1-b6de-484f-865b-45bde68c1df8 · outbound

This paper cites RIME: robust preference-based reinforcement learning with noisy preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning RIME: robust preference-based reinforcement learning with noisy preferences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.484703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.653713Z digest=sha256:c9f1ab554fe679afad6b00491dc841a574bb6e9f80a3eb9ec176e534cfa60e41

Observation e1b4705a-362a-4749-a5dc-5b177b88445a · outbound

This paper cites AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.477470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.656095Z digest=sha256:b3b0976865854b6f8ae72f02959fcfae84e2d05763752c5ef69e414104a432bc

Observation 08f7bed2-0bca-414a-bf50-9ebb75e284de · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reward model ensembles help mitigate overoptimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.470636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.658610Z digest=sha256:d73b2210dd48186fafc9ce224240db10ffc5408d97ee3296df13a00005bffa59

Observation 49ca57c0-ffb5-4bfc-bdde-debf0a6536f2 · outbound

This paper cites B-pref: Benchmarking preference- based reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning B-pref: Benchmarking preference- based reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.462842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.662379Z digest=sha256:ad79d77b7a65af553c89d689e1dd6485e52d67e321f58b79f90028294ac988f1

Observation f35ae298-7b63-4b7c-b111-39ab708b06b1 · outbound

This paper cites Rethinking decision transformer via hierarchical reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking decision transformer via hierarchical reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.456224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.665757Z digest=sha256:300b240d48b8494f5b2d45e1af79a7d6898468d0b270584d0691169842e7821b

Observation 93960999-efcb-4eb0-8a44-5b54d0d724f8 · outbound

This paper cites WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.449568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.668108Z digest=sha256:335c54d9248eafc7a656b6c65101ad5c9dc7d55cece8e78071f30b482ffca762

Observation f7b64e6a-3356-4f98-b0e9-4dda53306843 · outbound

This paper cites OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.442807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.670493Z digest=sha256:128c4897f60bcefa92204ccd9ff14fc554617db12b315fda177e50eb82663537

Observation fd154f9e-c94f-4d55-adf5-b1d0e5a085e7 · outbound

This paper cites Preference Alignment with Flow Matching.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Alignment with Flow Matching

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.435365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.672687Z digest=sha256:e3763a6014a2180ba375f4a80d02f4956769c6615d957d8c3a158a8ab7a38247

Observation 275949b1-55c3-4527-a651-7e1c23fb4b66 · outbound

This paper cites Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.428600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.675423Z digest=sha256:783bdbb55735218a3c21d7782051fa7a626998b224fa0c42a82894e19999bd4b

Observation cd767918-e406-4296-b1ad-44c438625715 · outbound

This paper cites Denoising Implicit Feedback for Recommendation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Denoising Implicit Feedback for Recommendation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.677623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.677623Z digest=sha256:80ea1287982131a12381930e6dfc449ab18c247d467aefa8097874538bba2a08

Observation 356483ba-d9a8-4434-8c00-6fd4d6b272cb · outbound

This paper cites mixup: BEYOND EMPIRICAL RISK MINIMIZATION.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning mixup: BEYOND EMPIRICAL RISK MINIMIZATION

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.420313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.679660Z digest=sha256:ff54a3b7673d4011dce23ede2360b0edb3cdb90b5c287b3bf1c94b45c7145ed0

Observation be9b2e01-c247-47f0-8baa-332bbc521971 · outbound

This paper cites Co-teaching: Robust training of deep neural networks with extremely noisy labels.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Co-teaching: Robust training of deep neural networks with extremely noisy labels

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.413068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.681877Z digest=sha256:78131a745f4ae844abd8b4cd184e770107ed44eec7edd34b8c46a594679131a5

Observation 29d4218a-fc9c-49e8-aa23-01dd9802a5bf · outbound

This paper cites Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.405932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.684352Z digest=sha256:6201b7d536dd419ca664182578ec1042f756dda9dafa8fd5293440dbeb1dd74e

Observation 5bb06fb2-59c1-4f04-81d9-f9e6de1d2d0b · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning from noisy labels with deep neural networks: A survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.686484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.686484Z digest=sha256:abc4a494cf76c515f82ccf1fa0eefc6df6a2ca1a03521e5a54770d6425487371

Observation 340f57f2-d046-4e68-93fb-8f4a829786b0 · outbound

This paper cites Attention is All you Need.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Attention is All you Need

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.399011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.688693Z digest=sha256:02730de8933df4fedb80fd45ed4126e9b89459d3a4e91113f27eeb8d78aaa233

Observation 07b25ea8-2a8e-4604-810c-a4532a8e2b03 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Simple Framework for Contrastive Learning of Visual Representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.392260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.691043Z digest=sha256:2f5432e9d0305730d1cf1a8c9007fa62abe4fa58f521065824f7970f5d1424bb

Observation d4dcb6c3-5bc3-4569-a471-ca88d75bbcca · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.693302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.693302Z digest=sha256:95aa60f90c8e73a8ca401333beb88837c8620eed0c999580b9e8a22bbe962584

Observation ebefa7eb-9474-40c3-84b9-61dc1a853f1b · outbound

This paper cites D4RL: Building Better Benchmarks for Offline Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Building Better Benchmarks for Offline Reinforcement Learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.385421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.695735Z digest=sha256:15b4bcf319b76a842d7837f8802ad10daf72971781aa191832db012742ca8a14

Observation 0c06b20e-a5dc-42b0-86c5-cf226630a8c7 · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/index.html.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/index.html

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.378249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.698093Z digest=sha256:db9763dd5c5b080b20ea66f1785fb586b4152d0ea382652e9ced5bd9e53b6659

Observation 5883a3ed-b7f8-4bac-8410-542612030a70 · outbound

This paper cites Offlinerl-kit: An elegant pytorch offline reinforcement learning library.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Offlinerl-kit: An elegant pytorch offline reinforcement learning library

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.370970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.700406Z digest=sha256:20d9d515f3515aa8a893ccfdb206c8b645a83b685cd235e2ba0cf2d38d750b8a

Observation acc20467-b73a-4e61-aafb-0e89c97011d2 · outbound

This paper cites Hierarchically decoupled imitation for morpho- logical transfer.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Hierarchically decoupled imitation for morpho- logical transfer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.363018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.703454Z digest=sha256:94f7c442397a6e56bedbcb97be87ef2c59b924634685b19ad2b69f91eaa81cbd

Observation b8b6faf0-90c2-4efb-a873-f36946864a93 · outbound

This paper cites PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.354877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.705831Z digest=sha256:145a80c33e811e234a00f534b88e3fecfc5ecaac3bb1e424126d88097984f569

Observation faa778ae-004a-4dde-8ba4-0f0d7a677f57 · outbound

This paper cites Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.708152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.708152Z digest=sha256:508cca99e430a055f5f83897a6e5d1c2b472c4e679119975f7c92b1b0f1e893b

Observation 69bc5927-4172-4956-8c45-9083497c7694 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.710490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.710490Z digest=sha256:32bfc6cf51cd7d42c29be0e8b4e58bd80ba7a745976f603bf5df4394fb237a6f

Observation ab1552f6-f733-4b52-9d09-1b90227dd761 · outbound

This paper cites dm_control: Software and tasks for continuous control.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning dm_control: Software and tasks for continuous control

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.713036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.713036Z digest=sha256:feaaf7c949384f1c38a2d32ec884da6556ea4f0dd126416853e0809baa5fec92

Observation 8e0ccee2-5793-45ba-b7b4-5e6d7602df08 · outbound

This paper cites URLB: Unsupervised reinforcement learning benchmark.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URLB: Unsupervised reinforcement learning benchmark

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.346513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.715496Z digest=sha256:8c963ff076b2c933d10feea0510a3fa3fdea7aa29acf51740cf37f221d72b0d9

Observation 103bccc0-2245-448a-b121-76970064a4ac · outbound

This paper cites Continuous control with deep reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Continuous control with deep reinforcement learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.717937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.717937Z digest=sha256:4ce934fa80a676312ef8336cfb65aae5f7b508afa846e2157ef47b60bb9a60d7

Observation ce58efbb-6332-4e03-81a4-6bd802df36c9 · outbound

This paper cites URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.338969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.721084Z digest=sha256:3d6785bb9f09a9c1471de94649dd7ed626e24a77e9a0c870a6b8cef619654055

Observation 9c3c1e64-16a1-4f96-92c9-bfcd397821ca · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/hopper/.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/hopper/

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.331172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.724721Z digest=sha256:fa15fb3e858bd0a6c3c33eb921c974e92fd693b96532a259572a4478dce87caa

Observation 26b71533-3a9e-45e8-96ae-1e3ad80676d8 · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.323213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.727226Z digest=sha256:033be182dca7d1cceb4990b7161e2beaa94823a0b5b77be2b168f46bc78bebbd

Observation ab7ec7ee-36e9-419b-9380-ac83076703d1 · outbound

This paper cites For the Franka Kitchen tasks, we use the preference datasets from An et al.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning For the Franka Kitchen tasks, we use the preference datasets from An et al

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.314920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.730009Z digest=sha256:520fbb69d726859cf50c2bdba95ed863e2264e2cf3faaa84b87e09d45d790edf

Observation 41a941b2-d2f4-404f-8b60-2b89f0b89425 · outbound

This paper cites ISSN: 2640-3498.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning ISSN: 2640-3498

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.588090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.613923Z digest=sha256:708b1e117edd1ed523e8f2f050d2beea84bea3ed28e828df8b832e26ac89216f

Observation 0347b73b-0a39-4a05-b278-31a057847404 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:14.566095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T00:53:13.621171Z digest=sha256:7ded9b20648b7cc9f7fd989217507999d663ca6fba7fa8bd8a560da8b36cab40

Pith citing papers

No inbound Pith citation observations are available.