Pith. sign in

Paper Citation Record · LEDGER

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2506.12529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12529 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:53:13.730009Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a838ea7f-1367-4cab-ade4-ff13cad9272b · outbound

This paper cites Sutton and Andrew G.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Sutton and Andrew G

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.636568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.574686Z digest=sha256:ccfd814bf050f0917693fe0e6af69195494f0dcf4ff9cd086b15d7c820c34514

Observation e41adb7b-fc05-47b2-8087-650acdee650b · outbound

This paper cites Reinforcement learning can be more efficient with multiple rewards.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reinforcement learning can be more efficient with multiple rewards

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.629294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.577558Z digest=sha256:24be14b05766b8de0d0fde9df74249a788480cdeceeed40cd21949c294731d81

Observation 31528ee8-6aa9-4fcc-830b-25e50396a312 · outbound

This paper cites Learning Agile Robotic Locomotion Skills by Imitating Animals.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning Agile Robotic Locomotion Skills by Imitating Animals

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.580172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.580172Z digest=sha256:9673d45381ab88a84d1a83562bd5677d0db459d7a3d7c547b11b240fdea95da3

Observation 668de56f-1921-46d4-b2ec-7ffbe1be87e2 · outbound

This paper cites Deep object-centric represen- tations for generalizable robot learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep object-centric represen- tations for generalizable robot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.583522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.583522Z digest=sha256:c6d164a768607bd507f75606250c93edf46936336f5c0cf423a1ef808507e436

Observation b89ad36b-5273-462d-9fa7-3068b8804939 · outbound

This paper cites The ingredients of real world robotic reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning The ingredients of real world robotic reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.622631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.586670Z digest=sha256:bd9db8eee357e4d27c8f93436bd7fc6dda0f2f0407f772c1d6bc3a922c20ce12

Observation 1b95cad8-1d9e-4eea-b430-9a45b1b06e2e · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.589096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.589096Z digest=sha256:4b46482384d26bfd445d544f163f46147ebede90604fdfd7a95474e645e77a0d

Observation 1211a9fc-d637-440a-8732-934998ae1c9a · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.591993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.591993Z digest=sha256:8d64ab2c4e813bd6cd6a172b3f66bd86ee927ae6afadd2b35933dbe35d7e0423

Observation bcdaafa4-37b0-4089-8eca-e979692aa58c · outbound

This paper cites GPT-4 Technical Report.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning GPT-4 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.594860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.594860Z digest=sha256:3205f63cf36318eb74e3e53943d11c4e0f79bf3c9f44fe2a7cf3a7dbc8d1c942

Observation 69641a7f-1abc-4504-9b55-3c54a6ceec2f · outbound

This paper cites Training language models to follow instructions with human feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Training language models to follow instructions with human feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.615042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.597295Z digest=sha256:355d851d18cff7d1fe3c48efa7088e932a51cb8cb326be35d9cb25ef9717112f

Observation 19b879c1-f247-4979-b6df-efeccc565ea6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.599621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.599621Z digest=sha256:d561387830cfa7b2daa900f65aca46ed1aed17881519d54f571adad006019c50

Observation a2f8fbe3-ce5d-4f95-a555-30270e9eab9b · outbound

This paper cites Active Preference-Based Learning of Reward Functions.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Active Preference-Based Learning of Reward Functions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.602307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.602307Z digest=sha256:450b75fc49a2c10d83c239f07367e46ac493c2e80bb940356b73706c6345222d

Observation ca135057-2688-40fe-a535-3bc88ad74e26 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Deep Reinforcement Learning from Human Preferences

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.608347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.604747Z digest=sha256:71d95d118d794de43bcbda07c829476b945901cb55969b1fa3af8a6400360940

Observation f60f215e-05b2-4d97-b993-3dbe38972c4d · outbound

This paper cites A Survey of Preference-Based Reinforcement Learning Methods.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Survey of Preference-Based Reinforcement Learning Methods

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.601692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.607066Z digest=sha256:03ac7619939f5a85a1076832ed9d11e097897fd87f9cf39ba7da70100250fde0

Observation dcc01f40-2cf0-4b44-9e41-5b13da365c11 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.609596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.609596Z digest=sha256:564ab790656617ce2d7f0d34bb8e753c466ee93ffb5d5fd0f3a54288eb47ccbb

Observation bbed6be8-6b7e-4898-b8a2-4b2515e6542b · outbound

This paper cites Smith, and Pieter Abbeel.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Smith, and Pieter Abbeel

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.594628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.611610Z digest=sha256:03f10648db2c2789c2c2de2676be0b2fd5095ea6995f3d437ea87098de09ef08

Observation 5a1f93fd-7264-4d23-8023-83386e640017 · outbound

This paper cites Few-shot preference learning for human-in-the-loop RL.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Few-shot preference learning for human-in-the-loop RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.579789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.616579Z digest=sha256:385a151613fa1acc5071be91ef679d1096efc1f9f521346270886a00461bc6cb

Observation 886092de-c884-4349-aee7-dec74b9d0023 · outbound

This paper cites Bradley Knox, and Dorsa Sadigh.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Bradley Knox, and Dorsa Sadigh

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.572904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.618745Z digest=sha256:58fbfa8335766a31c91c4741775f9fec470908bcc96421994ec69b5ed764ec9b

Observation aa775772-5afe-4056-bb72-43b3df819895 · outbound

This paper cites Inverse Preference Learning: Preference-based RL without a Reward Function.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Inverse Preference Learning: Preference-based RL without a Reward Function

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.559555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.623886Z digest=sha256:8ea1c525ca9efb36025535b3327c18d17c550332826279339bf1029cebd68246

Observation 327d9590-55cc-4c53-9ddb-7df41ee30b9b · outbound

This paper cites Direct Preference-based Policy Optimization without Reward Modeling.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference-based Policy Optimization without Reward Modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.552744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.626123Z digest=sha256:975014deed86893d7013c0a15b5b5682730356a5da3e2beece6fef9ac18d3617

Observation 364218c0-d7c7-44f6-a6e0-26c1e2dd405c · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Beyond Reward: Offline Preference-guided Policy Optimization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.545970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.628391Z digest=sha256:791ac24cfd882d2d86f3328c193da915072961d2b05d78dbe75faecbd966bb0e

Observation cba5540a-8275-4414-afb8-118e0aa90b58 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.538656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.631200Z digest=sha256:bd5cc68a2e61da89cf3144c70241a8447c91e14e410ed9f8dd8e68b2ab050a5a

Observation 1cc52139-f6aa-46bc-8f50-89faafdb063a · outbound

This paper cites Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning to Discern: Imitating Heterogeneous Human Demonstrations with Preference and Representation Learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.531895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.633432Z digest=sha256:ff5ec9133498c69eeebefb29c48fe5231e0a9f66ba63c738b269ca6dc8d5ccc5

Observation fb61f357-0b37-4a7c-bee7-13dff66eaae2 · outbound

This paper cites Rethinking reward modeling in preference-based large language model alignment.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking reward modeling in preference-based large language model alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.524528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.635700Z digest=sha256:df8a30242c2fdf18d85edb2a314234e8932e3bb988a8c144d6b8d9125d9ac2c7

Observation 4533caba-3b52-4501-8e24-c7b60990f44a · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Generalized preference optimization: A unified approach to offline alignment

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.517863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.638150Z digest=sha256:b8299c90c89b85facf0c78a391bc81efc5f5dcb3fe35c5e8067551fbf979fa62

Observation 8bad5c92-683e-4ad4-a002-cc55117e83cb · outbound

This paper cites Nash learning from human feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Nash learning from human feedback

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.511138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.640321Z digest=sha256:4216dd32ee49b0ac799de65ae2924651c3fe9ff9742a528caec1dc6d38020169

Observation 13bdd45a-96b2-43aa-865f-65b560504c57 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A general theoretical paradigm to understand learning from human preferences

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.504228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.643844Z digest=sha256:4a58fa6c66038046ccf2b421ac8058de2aa535bbd3d1ce5955632fa907a0e7ef

Observation 70cd97ea-e593-49fa-9ef1-b44d30a25af8 · outbound

This paper cites Online Iterative Reinforcement Learning from Human Feedback with General Preference Model.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Online Iterative Reinforcement Learning from Human Feedback with General Preference Model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.497602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.646342Z digest=sha256:056cbf57e026672b1c50b4e0cb28b2ccfe96dd9d8c0e13740b91e98534b0b125

Observation 681f80ec-a3bf-4e3e-96a8-0c1c4f32eda0 · outbound

This paper cites Intransitivity of preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Intransitivity of preferences

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.648723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.648723Z digest=sha256:fe1a397b6028ad77f898ddcdd39dd092bb16a3555d35425a5d8eb0d3f392eb15

Observation a5f4cb82-180f-41ad-bc80-26a2ee1c3fe5 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:14.491103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.651354Z digest=sha256:be921644620a476d7e17375b564cb30e555b24bad6a92f54ed44239655366ba0

Observation fa32ebb1-b6de-484f-865b-45bde68c1df8 · outbound

This paper cites RIME: robust preference-based reinforcement learning with noisy preferences.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning RIME: robust preference-based reinforcement learning with noisy preferences

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.484703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.653713Z digest=sha256:bf6c461170bb572d90b024cd9a193c39814af8f3d5429165e3f5bc2a1b200c25

Observation e1b4705a-362a-4749-a5dc-5b177b88445a · outbound

This paper cites AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning AlpacaFarm: A Sim- ulation Framework for Methods that Learn from Human Feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.477470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.656095Z digest=sha256:faed0d4eda3e95366e8cb0d638664ea58ac8a37f705ee1dcace982a9d0dab439

Observation 08f7bed2-0bca-414a-bf50-9ebb75e284de · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Reward model ensembles help mitigate overoptimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.470636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.658610Z digest=sha256:dc6bdc10bcff38f867fba17e17e04c89989e26b6c0b2fc44407bfcabb9bdb767

Observation 49ca57c0-ffb5-4bfc-bdde-debf0a6536f2 · outbound

This paper cites B-pref: Benchmarking preference- based reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning B-pref: Benchmarking preference- based reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.462842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.662379Z digest=sha256:5ab772df0f1c74e21a539fd650c84efaa57dacd14385758e3ba923aa2af12a0c

Observation f35ae298-7b63-4b7c-b111-39ab708b06b1 · outbound

This paper cites Rethinking decision transformer via hierarchical reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Rethinking decision transformer via hierarchical reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.456224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.665757Z digest=sha256:5696d46b66f2b36814c5d10410bdd044972ef996baeac7da77d1742f3ef51194

Observation 93960999-efcb-4eb0-8a44-5b54d0d724f8 · outbound

This paper cites WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning WHEN SHOULD WE PREFER DECISION TRANSFORMERS FOR OFFLINE REINFORCE- MENT LEARNING? 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.449568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.668108Z digest=sha256:126a3acc8268d358ad8b130434ff76f0b7237fd0cca660ad0fdbacad2545a97a

Observation f7b64e6a-3356-4f98-b0e9-4dda53306843 · outbound

This paper cites OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning OFFLINE REINFORCEMENT LEARNING WITH IMPLICIT Q-LEARNING

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.442807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.670493Z digest=sha256:170d71f00eb8ca95024ee0e0b21769cd92a72f6d84d7e5d57eb86ddb8964a93f

Observation fd154f9e-c94f-4d55-adf5-b1d0e5a085e7 · outbound

This paper cites Preference Alignment with Flow Matching.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Preference Alignment with Flow Matching

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.435365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.672687Z digest=sha256:0c79d5207dfe5d14e17066b41dad5c12dfa088f8af9f5ae03a61b64d245d997f

Observation 275949b1-55c3-4527-a651-7e1c23fb4b66 · outbound

This paper cites Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Flow to better: Offline preference-based reinforcement learning via preferred trajectory generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.428600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.675423Z digest=sha256:c9e3088f99a8791afa1dddad1914e926a56f6107a70daa08b32881c0c7bb98d8

Observation cd767918-e406-4296-b1ad-44c438625715 · outbound

This paper cites Denoising Implicit Feedback for Recommendation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Denoising Implicit Feedback for Recommendation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.677623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.677623Z digest=sha256:f078a5c3dd53dd513195662355b4de8186b72e400fe3f0c158c22ae245667d78

Observation 356483ba-d9a8-4434-8c00-6fd4d6b272cb · outbound

This paper cites mixup: BEYOND EMPIRICAL RISK MINIMIZATION.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning mixup: BEYOND EMPIRICAL RISK MINIMIZATION

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.420313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.679660Z digest=sha256:131546e2e1f8380f0840680c38d1537c7e27498936d471e0e1edc38bac722e13

Observation be9b2e01-c247-47f0-8baa-332bbc521971 · outbound

This paper cites Co-teaching: Robust training of deep neural networks with extremely noisy labels.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Co-teaching: Robust training of deep neural networks with extremely noisy labels

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.413068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.681877Z digest=sha256:4a50b985bf88d070c25e3a8ee29425542d16c55bdbaa845c6a4ccf9f690ee9ce

Observation 29d4218a-fc9c-49e8-aa23-01dd9802a5bf · outbound

This paper cites Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Does label smoothing mitigate label noise? In Proceedings of the 37th International Conference on Machine Learning, volume 119 of ICML’20, pages 6448–6458

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.405932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.684352Z digest=sha256:b97bed8a74dd81ced8a73f73d2c13edfb0a75a5e4a5d581836d285f15f3d15ea

Observation 5bb06fb2-59c1-4f04-81d9-f9e6de1d2d0b · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning from noisy labels with deep neural networks: A survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.686484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.686484Z digest=sha256:7ae4dcd20020a77ec07c478010b8940213f6d4a5b45bc9ea527023c4037f8df4

Observation 340f57f2-d046-4e68-93fb-8f4a829786b0 · outbound

This paper cites Attention is All you Need.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Attention is All you Need

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.399011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.688693Z digest=sha256:9b18232c6c8efe40088403a8de786edb1fcb5f54974f7c92d1739031fe9f4aad

Observation 07b25ea8-2a8e-4604-810c-a4532a8e2b03 · outbound

This paper cites A Simple Framework for Contrastive Learning of Visual Representations.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning A Simple Framework for Contrastive Learning of Visual Representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.392260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.691043Z digest=sha256:909064d31b79a00cb3d3b43992f8b89e8188d629853d9645bfb6fcdae630cc08

Observation d4dcb6c3-5bc3-4569-a471-ca88d75bbcca · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.693302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.693302Z digest=sha256:f9e1525e222534ec051ac256bd9a797932bacbd714146af1a47b3a5b9d3814ed

Observation ebefa7eb-9474-40c3-84b9-61dc1a853f1b · outbound

This paper cites D4RL: Building Better Benchmarks for Offline Reinforcement Learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning D4RL: Building Better Benchmarks for Offline Reinforcement Learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.385421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.695735Z digest=sha256:1cc0235ac0c8298276b952e19cbd725b4ff75c2c1b27808557a1c168859f8881

Observation 0c06b20e-a5dc-42b0-86c5-cf226630a8c7 · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/index.html.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/index.html

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.378249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.698093Z digest=sha256:a60bbca7e4f1f59cee24d6b33be3543239db5645a278fa4945af3d94586633ae

Observation 5883a3ed-b7f8-4bac-8410-542612030a70 · outbound

This paper cites Offlinerl-kit: An elegant pytorch offline reinforcement learning library.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Offlinerl-kit: An elegant pytorch offline reinforcement learning library

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.370970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.700406Z digest=sha256:3d5320a0ac857047179c732ad3c5cd42e0b8e3afc3ccdb8495e6f0930be9371e

Observation acc20467-b73a-4e61-aafb-0e89c97011d2 · outbound

This paper cites Hierarchically decoupled imitation for morpho- logical transfer.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Hierarchically decoupled imitation for morpho- logical transfer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.363018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.703454Z digest=sha256:b34f677308199ca805c7d7f21f5a9730de15cefab9ce09b197c886ee0239d33a

Observation b8b6faf0-90c2-4efb-a873-f36946864a93 · outbound

This paper cites PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning PEARL: Zero-shot cross-task preference alignment and robust reward learning for robotic manipulation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.354877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.705831Z digest=sha256:6c5cd921adc4fe95b0cb05bbe28ec5bb1c625bd4f22dc7e0a6adbd719091f4a7

Observation faa778ae-004a-4dde-8ba4-0f0d7a677f57 · outbound

This paper cites Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Adversarial Motion Priors Make Good Substitutes for Complex Reward Functions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.708152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.708152Z digest=sha256:e8f6f0dde927c39a7020026d6daf67d6b5c6e82356ac87ab17a2b45b66772aca

Observation 69bc5927-4172-4956-8c45-9083497c7694 · outbound

This paper cites Learning robust perceptive locomotion for quadrupedal robots in the wild.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Learning robust perceptive locomotion for quadrupedal robots in the wild

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.710490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.710490Z digest=sha256:78a1b4cd3909ece0e173317ff30d2e0f209883ce89b8863f3f28631edbcda495

Observation ab1552f6-f733-4b52-9d09-1b90227dd761 · outbound

This paper cites dm_control: Software and tasks for continuous control.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning dm_control: Software and tasks for continuous control

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.713036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.713036Z digest=sha256:1369ac3bd0cca67df85b3ba40572861776ad20ae5cf843e50551d15417316a06

Observation 8e0ccee2-5793-45ba-b7b4-5e6d7602df08 · outbound

This paper cites URLB: Unsupervised reinforcement learning benchmark.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URLB: Unsupervised reinforcement learning benchmark

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.346513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.715496Z digest=sha256:094ba27d264180aeab6a5ef5ebe38024716a2569d54a45019139d633052a4f68

Observation 103bccc0-2245-448a-b121-76970064a4ac · outbound

This paper cites Continuous control with deep reinforcement learning.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Continuous control with deep reinforcement learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:13.717937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:53:13.717937Z digest=sha256:f67f7320f7a948c127ee70532ddc91de8ceb5c0ea6e3f50de41e8e2143b3a99e

Observation ce58efbb-6332-4e03-81a4-6bd802df36c9 · outbound

This paper cites URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://docs.pytorch.org/ docs/stable/generated/torch.nn.TransformerEncoder.html

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.338969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.721084Z digest=sha256:558dc9cc762a0ded255b0f0b5037f7eb1daed83145f07ccdf8aa145bb75a87d2

Observation 9c3c1e64-16a1-4f96-92c9-bfcd397821ca · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/hopper/.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/hopper/

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.331172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.724721Z digest=sha256:7ec44bf22d38738a68722601f05f10b0d6426b74080165e78d2d9b1c0a1e43f3

Observation 26b71533-3a9e-45e8-96ae-1e3ad80676d8 · outbound

This paper cites URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning URL https://www.gymlibrary.dev/environments/ mujoco/walker2d/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.323213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.727226Z digest=sha256:20145bc2ae8676291cfe8a73629ba7d2dd5fcc2752018d074ead088b27d71d5a

Observation ab7ec7ee-36e9-419b-9380-ac83076703d1 · outbound

This paper cites For the Franka Kitchen tasks, we use the preference datasets from An et al.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning For the Franka Kitchen tasks, we use the preference datasets from An et al

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.314920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.730009Z digest=sha256:d9b785ebd28e186f8e2ac023e779dd2ae48238d5565baaa73bff7d694921f82d

Observation 41a941b2-d2f4-404f-8b60-2b89f0b89425 · outbound

This paper cites ISSN: 2640-3498.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning ISSN: 2640-3498

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:53:14.588090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.613923Z digest=sha256:781dc14255889d82c695bf66a7298c36a1d2778adbfe49df951ae50a5509f1c0

Observation 0347b73b-0a39-4a05-b278-31a057847404 · outbound

This paper cites an unresolved cited work.

Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:53:14.566095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:53:13.621171Z digest=sha256:6138a6efa54b62e4579eaa0288528952530fb6bb2819222a48e8905e361e8995

Pith citing papers

No inbound Pith citation observations are available.