Pith. sign in

Paper Citation Record · LEDGER

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.00388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00388 v3

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:01.213190Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 250a93c1-e861-440f-b914-2164a8bd24a1 · outbound

This paper cites G., Dabney, W., and Munos, R.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries G., Dabney, W., and Munos, R

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.074379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.074379Z digest=sha256:a0b069dbe254167adde9cb98b800a06a47e863af427b2bdb3f00dd538656658f

Observation 54d7213e-edb2-437b-95bc-fa5e10e26b29 · outbound

This paper cites G., Candido, S., Castro, P.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries G., Candido, S., Castro, P

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.592044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.147263Z digest=sha256:52c6f45a3058d3f8de0a434fed900e59fc1fef080d81ae436546976bc8da42b3

Observation e04203ac-c923-4752-a3f6-a23e54f4bb58 · outbound

This paper cites Learning to distinguish: shared perceptual features and discrimination practice tune behavioural pattern separation.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning to distinguish: shared perceptual features and discrimination practice tune behavioural pattern separation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.418448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.238918Z digest=sha256:0eaa2b0b97263859daac91b4828809635655302ae07cdeec4cf10838ea14744e

Observation 1a505b6a-ca0d-44e2-aba9-e49c090cce6d · outbound

This paper cites Active preference-based gaussian process regression for reward learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Active preference-based gaussian process regression for reward learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.268428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.330118Z digest=sha256:015dd0d657e8c0ddc46919633c0a013032849b417fba716ac88843e9f9f3f4d9

Observation e479e1df-6574-48d8-a2e8-b097915e16e3 · outbound

This paper cites an unresolved cited work.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.354573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.354573Z digest=sha256:2f6d38048d2c3ea2b7886f0b0ade107ce25fdb2e50898d5c829f5e9b56822533

Observation 1bd4b567-0992-476d-aa06-2b0964c20d66 · outbound

This paper cites and He, K.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and He, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.439208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.439208Z digest=sha256:29055123f5ac2a6ca21311d382a9b79a8766409084edaecb5332aa3519ad6668

Observation d7ec4a76-b6b0-47f6-ae4f-ccec4cdd5be6 · outbound

This paper cites RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.502734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.502734Z digest=sha256:6418ec2569477601ca64302932e711ca885d981f61f02274b8cf445260b24b54

Observation c9248b55-c092-436d-9845-69455364ba71 · outbound

This paper cites Listwise Reward Estimation for Offline Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Listwise Reward Estimation for Offline Preference-based Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:01.753321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.547010Z digest=sha256:8ec4de720e971645d0014bac83862ea2bcde4437502515553c805d28171ed555

Observation 89a3358b-236d-486f-80f2-80e68e37d0a8 · outbound

This paper cites Learning a similarity metric discriminatively, with application to face verification.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning a similarity metric discriminatively, with application to face verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.626849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.626849Z digest=sha256:92ae6d96664eac0712608e4a12c55636118a6515bea54d95463f217b1bc85d11

Observation 43919007-ec05-4297-9da6-36c5f86c37bd · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.702927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.702927Z digest=sha256:fbd5f55ef86ead021d2f0854b1d564daf4c5915960b3f820eb57a738f01ea3d6

Observation 1a34d678-3436-4a3a-a4fe-65d1b7b3e432 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.766405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.766405Z digest=sha256:3e08bd1552fd32e930b97c48501e94c99ecd5203c2a82cb7a3b1cc4c8a6442bb

Observation abc8b36e-fc44-4deb-ad0b-d09b727f5b1d · outbound

This paper cites Generalized Decision Transformer for Offline Hindsight Information Matching.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Generalized Decision Transformer for Offline Hindsight Information Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.859163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.859163Z digest=sha256:5d5db59cde798cf5b2507aab3725c88f4b7e78009288bba420a4e5734589c7b5

Observation b45fc730-8cc9-4621-9f44-ee2ae67fa062 · outbound

This paper cites an unresolved cited work.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:03.000325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.997015Z digest=sha256:faa9afda4c0b642470cd021bcd8971d72e1ac79d1b73c49d24a4743515484664

Observation 1179e094-1d9b-4423-bdc4-75bde49d1d28 · outbound

This paper cites Hindsight Preference Learning for Offline Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Hindsight Preference Learning for Offline Preference-based Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.048180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.048180Z digest=sha256:d73661f081d05c98f7dff709645ab2e91c0ccdc9020523b3a79d2864f27ecc94

Observation b71a3d6c-8600-487c-8108-c3687755c31a · outbound

This paper cites and Sadigh, D.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and Sadigh, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.802723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.112002Z digest=sha256:f74fae37745794a3dd2052a74a4ab091cecf5912457ca9db0412e445103f425a

Observation 144323eb-6a86-404c-afb8-5cdfbfdcd49f · outbound

This paper cites Reward learning from human preferences and demonstrations in atari.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reward learning from human preferences and demonstrations in atari

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.174843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.174843Z digest=sha256:1ebe08142e4f8f8486276428539a925a28c2f19ecedd124684f9b8d8cb372e0b

Observation 0378777a-f6aa-451c-8e86-d440816a9297 · outbound

This paper cites Tuning in to how neurons distinguish between stimuli.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Tuning in to how neurons distinguish between stimuli

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.486917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.256962Z digest=sha256:8123123dfd3aff5f2db35631d01df8f183faeeb71257c5ce98e369d84ee5b943

Observation 99d4ff69-fea4-443c-83dc-edd97301a97c · outbound

This paper cites Episodic novelty through temporal distance.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Episodic novelty through temporal distance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.287093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.309735Z digest=sha256:67e03396ad71ce30c6af10632c38ac09b5da65ed2a7eeee09cbbd2544c2c8d43

Observation 769b3f9d-210d-4324-9761-b8f37d1b5367 · outbound

This paper cites Transferring policy of deep reinforcement learning from simulation to reality for robotics.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Transferring policy of deep reinforcement learning from simulation to reality for robotics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.157844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.354488Z digest=sha256:13c554f41b091c128834d577eec7d1fb7253f23db63fc146e58f01d443eed7a5

Observation e33ead37-5bc7-4215-84bf-5ac5a8506a86 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.463092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.463092Z digest=sha256:9f0cab5edcbaf136a6f7ca14794e33b1fd1b39ff5d4d25b3b55f5ef57ce57abd

Observation 39893ad9-687e-45ad-949f-f43fe432dbf1 · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Beyond Reward: Offline Preference-guided Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.566244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.566244Z digest=sha256:b7350f4b411f131efb6c8108f33f310f9e094f549b4e0003eced4e63b666cf7a

Observation f1aa0140-78bb-4265-974a-256780365eb6 · outbound

This paper cites Preference transformer: Modeling human preferences using transformers for rl.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Preference transformer: Modeling human preferences using transformers for rl

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.066230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.705934Z digest=sha256:f00bb482732974252be7cb6e719254e896f1a5487189198b6e389bf8562632d4

Observation 573e15d2-3710-4a67-9edd-851b0dc63bef · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Offline Reinforcement Learning with Implicit Q-Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.782816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.782816Z digest=sha256:fdd0e2c3dc6ae011f24185760c9f719590a59008b30ca757d93e4c5d915fee50

Observation 90389d52-fd13-4079-84f8-66984933191a · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Curl: Contrastive unsupervised representations for reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.924071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.924071Z digest=sha256:e555044e9b81176df7980204df14fe97273b05e54dbea1419e71c2250e9afee2

Observation 8190da58-22ab-43de-a3f7-162654f53e83 · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.030007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.030007Z digest=sha256:0568d029028d6d8fe87da6aaefb8ba23fb258c1a9241d0e414a0703cfe3d07b9

Observation 6c8e3ae6-9185-403b-869a-96c023284f05 · outbound

This paper cites B-Pref: Benchmarking Preference-Based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries B-Pref: Benchmarking Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.151835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.151835Z digest=sha256:bc5aa79b16efe856798180d894d79da41563feeb9375907193185178caff60e7

Observation f3ab3b8f-f2e9-4a75-93fe-7b11dd15946e · outbound

This paper cites Survival instinct in offline reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Survival instinct in offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:01.970434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:12:59.288060Z digest=sha256:1f23c66f9ceb49d5e53087f7bcf493fc37762bab90cfc222f964bc655cf05af4

Observation 1387e94c-3fdf-408f-b021-b70d45fbb5b9 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.440374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.440374Z digest=sha256:06c3fc414f5d7d3b1fb823f89f5020ddf784408619e3ab5822b90856c2a15fd2

Observation 8e0ac5bf-e130-47d6-adab-5a92f2968f88 · outbound

This paper cites A., Veness, J., Bellemare, M.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries A., Veness, J., Bellemare, M

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.633355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.633355Z digest=sha256:d949a56d0c0b148ada827841c8a0813e2dd0579bf2aa124c9077ee5b5f0a441f

Observation f89b50c1-8f96-4fc7-b1c9-87a7b865bfcf · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.815667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.815667Z digest=sha256:da25eeff7b56aee77eeaa79ef241533813a0989a99557169086957d4769ab4f7

Observation faf88960-dd49-4059-a571-3472806c415b · outbound

This paper cites Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.969023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.969023Z digest=sha256:6edd1560533c57d3c858fb82c7a9ada3de3e72fe2174768c1d12dd151ea8018e

Observation c95aa10b-79be-4370-a553-f3c490942db1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.091060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.091060Z digest=sha256:d526d426c082d8e2c14401c146d8be9ad212c2fb71cb5895b6ef8384f96dc73f

Observation 4c46c2c8-2b1f-4eb0-a9a4-b39f94a9886a · outbound

This paper cites Training language models to follow instructions with human feedback.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.201111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.201111Z digest=sha256:bb3a196022575b4df1ce27dad867cf313b375ade4d7a98d0b973a6cd6401982b

Observation cba3728a-2163-4aaa-8eee-1e95afce2404 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.332041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.332041Z digest=sha256:3303c843c6ea868621c7958ac02258e8834a1a3b02e7668f3fd23fbdfc393d85

Observation 2d234ffd-ef7e-4178-98fd-518cc1164d1e · outbound

This paper cites Trust Region Policy Optimization.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Trust Region Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.476431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.476431Z digest=sha256:87d5c691c3627b53cce16f0db5a4eb8344c7f892f7adb6ccb315c073ffdb32a8

Observation 3931ac7b-a868-4eec-9af1-bf714c641b6c · outbound

This paper cites Benchmarks and Algorithms for Offline Preference-Based Reward Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.672841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.672841Z digest=sha256:110b6202b0f7f2f1cb4e035bfc6ed994cf54dae9302c699b386e2edeb2d0652d

Observation 17e588d6-466d-4205-938a-469714d91c97 · outbound

This paper cites DeepMind Control Suite.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries DeepMind Control Suite

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.774609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.774609Z digest=sha256:1fbbe3e4051eaf877890ec8f3079bce40c75e4a0fcf07e7af6706878917aea41

Observation 38bed9a6-7783-490b-a08f-063e354bd693 · outbound

This paper cites Reinforcement Learning from Diverse Human Preferences.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reinforcement Learning from Diverse Human Preferences

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:01.391953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T12:13:00.884525Z digest=sha256:8ef5dc849caca0356e131d07143d41656d58be4c137e253d1135d236451196ef

Observation 6cb82dbc-6008-47ce-8d91-09beca49acdc · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:01.053360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:01.053360Z digest=sha256:419b188cc2e7ba4f99b96e414298cd59e1f809e921f11167d11fee5353138f14

Observation f5f1a679-ecbf-42ec-8793-0f69ebc1ae7a · outbound

This paper cites and Lu, Z.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and Lu, Z

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:01.213190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:01.213190Z digest=sha256:e2a18ad624f1920c71edba5c9f489945b4dd01c7a234ea18f3da01628ee140c8

Pith citing papers

No inbound Pith citation observations are available.