Pith. sign in

Paper Citation Record · LEDGER

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries

As of 14 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.00388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00388 v3

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:01.213190Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 250a93c1-e861-440f-b914-2164a8bd24a1 · outbound

This paper cites G., Dabney, W., and Munos, R.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries G., Dabney, W., and Munos, R

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.074379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.074379Z digest=sha256:81a2a7556c7fac6ae907f5c3fbd2359f92518395e7c01dcc296559a3366524c7

Observation 54d7213e-edb2-437b-95bc-fa5e10e26b29 · outbound

This paper cites G., Candido, S., Castro, P.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries G., Candido, S., Castro, P

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.592044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.147263Z digest=sha256:2740c87d06c59a4c0da94e06bf50f996f6bd264b28802a43c1ecafd0abb29076

Observation e04203ac-c923-4752-a3f6-a23e54f4bb58 · outbound

This paper cites Learning to distinguish: shared perceptual features and discrimination practice tune behavioural pattern separation.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning to distinguish: shared perceptual features and discrimination practice tune behavioural pattern separation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.418448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.238918Z digest=sha256:453a73bcefd68490419415ba7e997b7f38af83349682ef875b7d1fb75597c68b

Observation 1a505b6a-ca0d-44e2-aba9-e49c090cce6d · outbound

This paper cites Active preference-based gaussian process regression for reward learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Active preference-based gaussian process regression for reward learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.268428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.330118Z digest=sha256:bd45e72ac98905d25ad4fabca6bf3e27f4fae94d511fee93fac4b90094dabaae

Observation e479e1df-6574-48d8-a2e8-b097915e16e3 · outbound

This paper cites an unresolved cited work.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.354573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.354573Z digest=sha256:98a834a1d0013d029cfc1f29c4fbba6d6c27a597c7c5a9c144bd6ae5d0c645db

Observation 1bd4b567-0992-476d-aa06-2b0964c20d66 · outbound

This paper cites and He, K.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and He, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.439208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.439208Z digest=sha256:db6fc57723fd0b643ce3e750a9c0846660596e7ad3c106d9b6e7373fa9448077

Observation d7ec4a76-b6b0-47f6-ae4f-ccec4cdd5be6 · outbound

This paper cites RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.502734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.502734Z digest=sha256:d95a3ebd491e27b25557ea1c80a8eebef6774fc268de9600a55a0ead72f8bf5f

Observation c9248b55-c092-436d-9845-69455364ba71 · outbound

This paper cites Listwise Reward Estimation for Offline Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Listwise Reward Estimation for Offline Preference-based Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:01.753321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.547010Z digest=sha256:f712d6edcb4994d19e8462214702ff6c72e204efc92ecc4941992f5d388443dd

Observation 89a3358b-236d-486f-80f2-80e68e37d0a8 · outbound

This paper cites Learning a similarity metric discriminatively, with application to face verification.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning a similarity metric discriminatively, with application to face verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.626849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.626849Z digest=sha256:4f47cad3c4f152b43678c47cd593ed41ded5c22a907b43f79e1aa318ec1f3830

Observation 43919007-ec05-4297-9da6-36c5f86c37bd · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.702927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.702927Z digest=sha256:36fe82003b6b38cb3cc6c14e28d99a0f4b3e4d10d45025d31b90bc128204f9f3

Observation 1a34d678-3436-4a3a-a4fe-65d1b7b3e432 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.766405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.766405Z digest=sha256:98fea521678396d5461c08ff004d81f5cbc7cf78b5373d7eb6fdf1a26d1c4f8f

Observation abc8b36e-fc44-4deb-ad0b-d09b727f5b1d · outbound

This paper cites Generalized Decision Transformer for Offline Hindsight Information Matching.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Generalized Decision Transformer for Offline Hindsight Information Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.859163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.859163Z digest=sha256:5749016576e4dbfb391105c5f95560ad5dd92eea97a871ec30289d9e7be31248

Observation b45fc730-8cc9-4621-9f44-ee2ae67fa062 · outbound

This paper cites an unresolved cited work.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:03.000325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.997015Z digest=sha256:365ec1779f63a60b731adfca18aed9bed4bb28d886dad0b63a666e9193152e29

Observation 1179e094-1d9b-4423-bdc4-75bde49d1d28 · outbound

This paper cites Hindsight Preference Learning for Offline Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Hindsight Preference Learning for Offline Preference-based Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.048180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.048180Z digest=sha256:c105d86c9d1daccdd7d5a0921a3664ca96ececa2279564a180774adec6f3d7cb

Observation b71a3d6c-8600-487c-8108-c3687755c31a · outbound

This paper cites and Sadigh, D.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and Sadigh, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.802723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.112002Z digest=sha256:91e0199d30853f6dccdb6681dc73d9ac220e66733aaf60d417048d956596b50b

Observation 144323eb-6a86-404c-afb8-5cdfbfdcd49f · outbound

This paper cites Reward learning from human preferences and demonstrations in atari.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reward learning from human preferences and demonstrations in atari

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.174843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.174843Z digest=sha256:c05a589907744e2b54c6a4d8f4d8a141699d8927a69754eaf941770ba0ca8b85

Observation 0378777a-f6aa-451c-8e86-d440816a9297 · outbound

This paper cites Tuning in to how neurons distinguish between stimuli.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Tuning in to how neurons distinguish between stimuli

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.486917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.256962Z digest=sha256:222ce82e132fc1ddc985ec5e462a001e6b45c1fb145fa060595afe66bdcd8545

Observation 99d4ff69-fea4-443c-83dc-edd97301a97c · outbound

This paper cites Episodic novelty through temporal distance.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Episodic novelty through temporal distance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.287093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.309735Z digest=sha256:22b8151f3dcb1f335d833f00d35e60359f16975d6c40718f19ffddeb7edc149d

Observation 769b3f9d-210d-4324-9761-b8f37d1b5367 · outbound

This paper cites Transferring policy of deep reinforcement learning from simulation to reality for robotics.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Transferring policy of deep reinforcement learning from simulation to reality for robotics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.157844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.354488Z digest=sha256:66c44fe90748ddcf0ad3acafd446ee688fa91cc89c53d7ce5a1d3c81d46d321b

Observation e33ead37-5bc7-4215-84bf-5ac5a8506a86 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.463092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.463092Z digest=sha256:ffeae0fed7da8f911f927599641ad28c37962f6530c2b2c1ae36bc53bdc2c003

Observation 39893ad9-687e-45ad-949f-f43fe432dbf1 · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Beyond Reward: Offline Preference-guided Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.566244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.566244Z digest=sha256:c368926b0511555791471fe3dff6682bcac7c12526c4a24feb0d60acc129e876

Observation f1aa0140-78bb-4265-974a-256780365eb6 · outbound

This paper cites Preference transformer: Modeling human preferences using transformers for rl.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Preference transformer: Modeling human preferences using transformers for rl

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.066230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.705934Z digest=sha256:55ad5f4ba2a650f15aef382c6e0a3a49a7eb18f76f6b3d1475aa0b7d278680d5

Observation 573e15d2-3710-4a67-9edd-851b0dc63bef · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Offline Reinforcement Learning with Implicit Q-Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.782816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.782816Z digest=sha256:53de251e319a53795b1fe9eb461d36aec6268946b3dcd2a8d3b332a37cb6a1ec

Observation 90389d52-fd13-4079-84f8-66984933191a · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Curl: Contrastive unsupervised representations for reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.924071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.924071Z digest=sha256:81a15cac9f0037d0e1805b4f36fc7adcd8ddab10f6256fa55b2054d71dfe6d3b

Observation 8190da58-22ab-43de-a3f7-162654f53e83 · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.030007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.030007Z digest=sha256:8dda03494c634d0587c3be134541bb33dc907ff6d56287b3bf17eaf96d9d4d98

Observation 6c8e3ae6-9185-403b-869a-96c023284f05 · outbound

This paper cites B-Pref: Benchmarking Preference-Based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries B-Pref: Benchmarking Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.151835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.151835Z digest=sha256:83d195434e63f670ee2e610b95d6e95189040d0625a3542a5538c34d0a2fdb79

Observation f3ab3b8f-f2e9-4a75-93fe-7b11dd15946e · outbound

This paper cites Survival instinct in offline reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Survival instinct in offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:01.970434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:12:59.288060Z digest=sha256:48adf66f89f210c8f2a6ba48034ef8c2751987e92b6162a439c47813d1d1f147

Observation 1387e94c-3fdf-408f-b021-b70d45fbb5b9 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.440374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.440374Z digest=sha256:1f9aa846779d7fb9a40cb114e01d2eea087e99cecc4f63cb72a587010e8e5986

Observation 8e0ac5bf-e130-47d6-adab-5a92f2968f88 · outbound

This paper cites A., Veness, J., Bellemare, M.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries A., Veness, J., Bellemare, M

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.633355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.633355Z digest=sha256:0d6ac83fb7f9410bdd970b91f93ba1250412e572922aea40b8de02b1d7bd752a

Observation f89b50c1-8f96-4fc7-b1c9-87a7b865bfcf · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.815667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.815667Z digest=sha256:d6a8b2a332f0440ec436b37dc71669b0d1e47ba95569ffed39155445878e1046

Observation faf88960-dd49-4059-a571-3472806c415b · outbound

This paper cites Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.969023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.969023Z digest=sha256:69cc5a4eddb3fa0dda708ed3d5dc80fd5ef2a45988f3426c2a0181d3afb94033

Observation c95aa10b-79be-4370-a553-f3c490942db1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.091060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.091060Z digest=sha256:debce08915e1d6281828813431be648ebe476d4943a3c5bfe643ea25e60cab04

Observation 4c46c2c8-2b1f-4eb0-a9a4-b39f94a9886a · outbound

This paper cites Training language models to follow instructions with human feedback.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.201111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.201111Z digest=sha256:b54dee051833394a8c11773921c509b9e197345588321774a6fbe1b91440725d

Observation cba3728a-2163-4aaa-8eee-1e95afce2404 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.332041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.332041Z digest=sha256:0608e67de2827d8b96aa9c2647b4a269e320bae915d62eb66d5e7bb786c14abf

Observation 2d234ffd-ef7e-4178-98fd-518cc1164d1e · outbound

This paper cites Trust Region Policy Optimization.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Trust Region Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.476431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.476431Z digest=sha256:25300830e9ecad8843e3749f80107a1b1791a2afb0d01eaac2769ab0192bcd44

Observation 3931ac7b-a868-4eec-9af1-bf714c641b6c · outbound

This paper cites Benchmarks and Algorithms for Offline Preference-Based Reward Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.672841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.672841Z digest=sha256:5d6abd32438eb1ca95bf959b0b6139289ccf6a8b33eece2de266c67099d18225

Observation 17e588d6-466d-4205-938a-469714d91c97 · outbound

This paper cites DeepMind Control Suite.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries DeepMind Control Suite

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.774609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.774609Z digest=sha256:ccff6d046d3cc7f591c9c65c732cbfd80bbe9ded40abc060fafc6b57e0a2e147

Observation 38bed9a6-7783-490b-a08f-063e354bd693 · outbound

This paper cites Reinforcement Learning from Diverse Human Preferences.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reinforcement Learning from Diverse Human Preferences

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:01.391953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T12:13:00.884525Z digest=sha256:c6f28d60143d58a0711063d87edb9eb672ae56def36dbe8b23d74dde55a5fe73

Observation 6cb82dbc-6008-47ce-8d91-09beca49acdc · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:01.053360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:01.053360Z digest=sha256:0932de9c08d09d2e52b0f0a93dd9f80d32bdecc6fbfd3bd585b057906cd0d835

Observation f5f1a679-ecbf-42ec-8793-0f69ebc1ae7a · outbound

This paper cites and Lu, Z.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and Lu, Z

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:01.213190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:01.213190Z digest=sha256:2f8f7b03e6314b90e465273ea43a945acdb10eb501359ac416929b0cee8b3a2e

Pith citing papers

No inbound Pith citation observations are available.