Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T10:43:29.936815Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2502.02921.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T10:43:29.936815Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c90ea59d-776a-4812-be9d-b6bbc4d740a8 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f72cd497-3718-4577-b12b-5f3fd559a746 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting April: Active preference learning-based reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6caed3f8-8e48-47b1-b2dc-50f7a5c7e7ad · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Programming by feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47dc248a-b856-4cd0-95b0-e0b1eaa7db73 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting K., Anil, R., and Koren, T
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46256e2b-c53c-4186-97b0-34bce3b35b05 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Fine-tuning language models to find agreement among humans with diverse preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d415488-0680-4e94-b953-c9f1da5d1dbb · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting L., Harvey, N., Liaw, C., and Mehrabian, A
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baa68838-deda-4f74-a280-a177ebdb6271 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting and Sadigh, D
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aea9c93c-9e6c-447a-b4f4-e2a9186800b8 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Batch active learning of reward functions from human preferences
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7cf28edb-17f1-4b13-a035-419a8b36a5b0 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting J., and Sadigh, D
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79b9f3ad-7e0b-4948-80c6-501a781abbf9 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7768408-fd99-484b-8b2d-e2ac559d44e5 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75b5cb6f-ee2e-4780-97a8-8b2224c690d8 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Rime: Robust preference-based reinforcement learning with noisy preferences
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 207fe773-bd59-4b53-bfcd-ff9bbe4e72fd · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90cc916-4e2f-4847-b594-6e2d1da110cf · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Active reward learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be4b8444-a54b-4481-9b09-616db01f7354 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3556975a-9c5d-4e36-a74d-37e12791b809 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa62631-b756-4657-887b-66b3bdf78413 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Dextreme: Transfer of agile in-hand manipulation from simulation to reality
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90768d1-0497-4b55-9d14-9fc12726d149 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting A bound on the label complexity of agnostic active learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5adef6a4-abde-478a-9fde-e2396446d4aa · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Contrastive preference learning: Learning from human feedback without rl
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b157bd0-0fbf-4204-ab79-7901b94de3d3 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c450e8-3c25-4f76-acbf-a1854f0dce3c · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting J., Kim, J., Kwak, M
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fd883f38-0857-4f0e-8e3e-3df5712aebd3 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Anymal parkour: Learning agile navigation for quadrupedal robots
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 105dff5c-cca8-4fcb-98a4-0cc41fa7cc85 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Bayesian active learning for classification and preference learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d47eb4f-7368-43bb-9402-13169fc760d9 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c590db-5767-421d-818f-3ec92b8b7e6d · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Reward learning from human preferences and demonstrations in atari
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51bf2fa3-9ec5-44f8-ad7b-1c626319c2e4 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ba41ef-b407-40d3-aecb-2c43d8d25f0a · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting D., Lu, Z., and Mou, S
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f11811a7-a8b8-4459-aecf-fc6097ce0a29 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Adam: A Method for Stochastic Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95374d00-e105-4bc7-ab63-08b4cc1eb262 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Crafting papers on machine learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0cd69ee-c286-4026-8d6c-d99fc2107da2 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Robust inference via generative classifiers for handling noisy labels
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d2c12e43-d28d-4bc1-83a9-02cf4b1bb1b7 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting B-pref: Benchmarking preference-based reinforcement learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6475070-01e0-4310-81f0-89150137b3d4 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting M., and Abbeel, P
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation de70b1c5-8dd4-4ecf-a9d0-38348615a996 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting CANDERE-COACH: Reinforcement Learning from Noisy Feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1633ce63-34da-4a49-ba0a-f9916dd8ea40 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Reward uncertainty for exploration in preference-based reinforcement learning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b241ae0-1068-414b-835b-ce3093103431 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a36d83c-1c8c-489b-8898-e4c8d06a4d8d · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp.\ 6448--6458
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 600468a7-b4e1-4554-96fe-91d8ec0893ec · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Normalized loss functions for deep learning with noisy labels
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e173d207-2496-40bd-83f0-81aaf3a58a02 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Learning multimodal rewards from rankings
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 329a7e06-693f-481f-92dd-b8ffa2be7630 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Active reward learning from online preferences
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d982223b-d224-4711-af23-c842965d9327 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b803d692-1b8f-4421-96ca-b7c9ab70b817 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting In-hand object rotation via rapid motor adaptation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 493d9cf1-8e79-4aeb-9717-1b5d1dee2469 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Real-world humanoid locomotion with reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fa30657b-5875-447a-b850-6f022d688130 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting D., Sastry, S
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 98783067-ef83-449d-8754-5ff5bceaee6c · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Active Learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 84607223-87fc-4817-aced-ef6d4cecf125 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting and Joachims, T
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65af65c1-3e78-4ed1-8d39-e699357bd6dd · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Mujoco: A physics engine for model-based control
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c461b5e-e3bc-4537-af70-750e50c71fb1 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Deepgait: Planning and control of quadrupedal gaits using deep reinforcement learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7b5061c5-85a9-4756-8663-a7dbaf852d99 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Statistical learning theory
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9eaba56-96bd-4b1e-b4da-63c707eb2255 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Model Predictive Path Integral Control using Covariance Variable Importance Sampling
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d612c389-17e6-488a-ad68-bcbbcbb6b901 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting A survey of preference-based reinforcement learning methods
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c84d1ad-ae80-4237-bf06-93a289dd362a · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting J., and Jin, W
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 98541519-e0c2-4f32-bf4e-d08c7828c5eb · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Reinforcement learning from diverse human preferences
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 222d7e08-795e-4030-a439-79181c413d65 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting W., Zhang, Y., Sun, J., Zhang, C., and Zhang, R
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 612708e9-1ebf-4201-a4a4-50eba0878a15 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Rotating without seeing: Towards in-hand dexterity through touch
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b58ab11-be9c-4095-98f7-ebb51bfb6c50 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf8037f1-49f0-45ea-9177-7b698dfb0fcc · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting N., and Lopez-Paz, D
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f602393-683b-4a46-b1c6-a2c40e78eaca · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting Robust curriculum learning: from clean label detection to noisy label self-correction
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 306be71d-1d9d-4666-9187-3a028e1bf514 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting A., Atkeson, C
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cfd5d2b2-3fdc-41af-b6ed-da0903ac5905 · outbound
Robust Reward Alignment via Hypothesis Space Batch Cutting write newline
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.