Pith. sign in

Paper Citation Record · LEDGER

Robust Reward Alignment via Hypothesis Space Batch Cutting

As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2502.02921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02921 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:43:29.936815Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c90ea59d-776a-4812-be9d-b6bbc4d740a8 · outbound

This paper cites GPT-4 Technical Report.

Robust Reward Alignment via Hypothesis Space Batch Cutting GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.811463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.811463Z digest=sha256:19fe74c23658b5a99c37c1601231688c76ad8f195562e90f6e72325b6bfdc840

Observation f72cd497-3718-4577-b12b-5f3fd559a746 · outbound

This paper cites April: Active preference learning-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting April: Active preference learning-based reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.814983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.814983Z digest=sha256:3d6d7478a420889e0b84cc0dca6c82393b04f68d31aaeffc2bdebe4d427078e6

Observation 6caed3f8-8e48-47b1-b2dc-50f7a5c7e7ad · outbound

This paper cites Programming by feedback.

Robust Reward Alignment via Hypothesis Space Batch Cutting Programming by feedback

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:31.054467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.817953Z digest=sha256:d6ee2fe02ca9595e308c0305b9232a0d6beea7d00a709a50b4e4f892850db860

Observation 47dc248a-b856-4cd0-95b0-e0b1eaa7db73 · outbound

This paper cites K., Anil, R., and Koren, T.

Robust Reward Alignment via Hypothesis Space Batch Cutting K., Anil, R., and Koren, T

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.956444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.820995Z digest=sha256:b04924a30cb399ff4901e8da88b0caf79b582d63fbb3f2a013888734444b9ded

Observation 46256e2b-c53c-4186-97b0-34bce3b35b05 · outbound

This paper cites Fine-tuning language models to find agreement among humans with diverse preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Fine-tuning language models to find agreement among humans with diverse preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.823812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.823812Z digest=sha256:3d921460ddf6eddd4ce250afdffbec61a7a150f311dad35c493eea0fe5f265fe

Observation 8d415488-0680-4e94-b953-c9f1da5d1dbb · outbound

This paper cites L., Harvey, N., Liaw, C., and Mehrabian, A.

Robust Reward Alignment via Hypothesis Space Batch Cutting L., Harvey, N., Liaw, C., and Mehrabian, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.826189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.826189Z digest=sha256:3a8cada32f7343181e45c4d895597df0d8aa9d17e0e347d41da369da474c443f

Observation baa68838-deda-4f74-a280-a177ebdb6271 · outbound

This paper cites and Sadigh, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting and Sadigh, D

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.834634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.828918Z digest=sha256:4a41462ca1fd5da89149affe023a7ad9a294930e851c1f57869ee91775c00095

Observation aea9c93c-9e6c-447a-b4f4-e2a9186800b8 · outbound

This paper cites Batch active learning of reward functions from human preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Batch active learning of reward functions from human preferences

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.788939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.831254Z digest=sha256:4da07b236dab8361b4d3e885b369e3c60b050b081fcdba2c0839b7dbb6d783ca

Observation 7cf28edb-17f1-4b13-a035-419a8b36a5b0 · outbound

This paper cites J., and Sadigh, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting J., and Sadigh, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.781660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.833947Z digest=sha256:0e4db305a54ba372159b8aebb72d5c7ee117eb8aa7131b9407dcd111c518acfd

Observation 79b9f3ad-7e0b-4948-80c6-501a781abbf9 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.836824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.836824Z digest=sha256:89df505ec6170d697f02071e3b38c885d9a6103bfde4fc68fc431de2f8a0294e

Observation a7768408-fd99-484b-8b2d-e2ac559d44e5 · outbound

This paper cites K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T.

Robust Reward Alignment via Hypothesis Space Batch Cutting K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.770751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.868127Z digest=sha256:1edc3f2c652ac13559206126071de22746695f9ac5e515bb1c424b29e85fedd7

Observation 75b5cb6f-ee2e-4780-97a8-8b2224c690d8 · outbound

This paper cites Rime: Robust preference-based reinforcement learning with noisy preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Rime: Robust preference-based reinforcement learning with noisy preferences

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.763353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.898076Z digest=sha256:7329a53ed3380d84bf2bce4a8df7f7cc08e8f16e9396b9bc1febfb7261988766

Observation 207fe773-bd59-4b53-bfcd-ff9bbe4e72fd · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.923060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.923060Z digest=sha256:a62829b842ca2509b1da4174c08a7d5f24de02187a02fb4b00a19f79f57f96b0

Observation a90cc916-4e2f-4847-b594-6e2d1da110cf · outbound

This paper cites Active reward learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Active reward learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.751851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.969360Z digest=sha256:1ef3b75ba392ce3781d893d8cfb0971cd9b79e2962fd6d460329a5ef4c5f8ac7

Observation be4b8444-a54b-4481-9b09-616db01f7354 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T10:43:30.744853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.014153Z digest=sha256:922959caf4e2ab695be1b7402d7a0c3b8cd7f20aee1e8f2d441a1078c2920c8f

Observation 3556975a-9c5d-4e36-a74d-37e12791b809 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.040157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.040157Z digest=sha256:18a168341fdb13d30a2b50a3cac0fba5d16b5e3b589ec3318e975bcd853db305

Observation 2aa62631-b756-4657-887b-66b3bdf78413 · outbound

This paper cites Dextreme: Transfer of agile in-hand manipulation from simulation to reality.

Robust Reward Alignment via Hypothesis Space Batch Cutting Dextreme: Transfer of agile in-hand manipulation from simulation to reality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.087356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.087356Z digest=sha256:e9eded8a1decf6a3260636ec6e208dc586a33ff462e76a213ce4b7c7a17bd258

Observation c90768d1-0497-4b55-9d14-9fc12726d149 · outbound

This paper cites A bound on the label complexity of agnostic active learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting A bound on the label complexity of agnostic active learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.728172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.136280Z digest=sha256:a4b437e065f9e6f3e18860f797170e919e6c41b9104e42490efc72d2000c6499

Observation 5adef6a4-abde-478a-9fde-e2396446d4aa · outbound

This paper cites Contrastive preference learning: Learning from human feedback without rl.

Robust Reward Alignment via Hypothesis Space Batch Cutting Contrastive preference learning: Learning from human feedback without rl

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.720704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.195602Z digest=sha256:903964878d7f7e29cfd8543ab5a36bd425381457545422b298a1a6bdd091eeca

Observation 6b157bd0-0fbf-4204-ab79-7901b94de3d3 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.260706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.260706Z digest=sha256:40ebd9bea4ca6e92d9f8558d83058762029c2e8f62b6a9f6bb0f8305f7bb630e

Observation e2c450e8-3c25-4f76-acbf-a1854f0dce3c · outbound

This paper cites J., Kim, J., Kwak, M.

Robust Reward Alignment via Hypothesis Space Batch Cutting J., Kim, J., Kwak, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.708222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.288699Z digest=sha256:eba33254f1c5f943d6847a5fd9df13949412948575eed78f415d8a4c8a5ffd85

Observation fd883f38-0857-4f0e-8e3e-3df5712aebd3 · outbound

This paper cites Anymal parkour: Learning agile navigation for quadrupedal robots.

Robust Reward Alignment via Hypothesis Space Batch Cutting Anymal parkour: Learning agile navigation for quadrupedal robots

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.700971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.318392Z digest=sha256:d8eb54c109b5e9e504cd7bb95a25ca69787cf7790fc637f53579e5c0e8bf8473

Observation 105dff5c-cca8-4fcb-98a4-0cc41fa7cc85 · outbound

This paper cites Bayesian active learning for classification and preference learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Bayesian active learning for classification and preference learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.623589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.342740Z digest=sha256:fdf6e95b82f4650c5b671ba5748263f00882ec8d5d2ff838cc5f98bdcf64e714

Observation 5d47eb4f-7368-43bb-9402-13169fc760d9 · outbound

This paper cites Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo.

Robust Reward Alignment via Hypothesis Space Batch Cutting Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.355325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.355325Z digest=sha256:273da67ee2c44f509ec8c77c7e135342109dca72c2aa602f705f8edb63ef8e95

Observation e7c590db-5767-421d-818f-3ec92b8b7e6d · outbound

This paper cites Reward learning from human preferences and demonstrations in atari.

Robust Reward Alignment via Hypothesis Space Batch Cutting Reward learning from human preferences and demonstrations in atari

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.371725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.371725Z digest=sha256:cc0aaaed02e495e28fbbd51b5295dba7f86410bb1c1cb1c522e3d568ce344fe2

Observation 51bf2fa3-9ec5-44f8-ad7b-1c626319c2e4 · outbound

This paper cites Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels.

Robust Reward Alignment via Hypothesis Space Batch Cutting Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.381374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.381374Z digest=sha256:60828bc7970fe8a07198dc00f39b6b0f3d0b2f7424e703daf22ba9e95aed6562

Observation f1ba41ef-b407-40d3-aecb-2c43d8d25f0a · outbound

This paper cites D., Lu, Z., and Mou, S.

Robust Reward Alignment via Hypothesis Space Batch Cutting D., Lu, Z., and Mou, S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.561042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.400413Z digest=sha256:72f04fe97c5b2ae39f39f910b33e0717935149302396973531940656cfb9f1bd

Observation f11811a7-a8b8-4459-aecf-fc6097ce0a29 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Robust Reward Alignment via Hypothesis Space Batch Cutting Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.419649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.419649Z digest=sha256:bfdee7dccada23a54facdab9f6dd718110ffc152c4f926210fe9084a7c76d87c

Observation 95374d00-e105-4bc7-ab63-08b4cc1eb262 · outbound

This paper cites Crafting papers on machine learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Crafting papers on machine learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.430378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.430378Z digest=sha256:8d2fd570df2f167be5dd74452b147c9b0425e94645192a4298a1de60b8059cc4

Observation c0cd69ee-c286-4026-8d6c-d99fc2107da2 · outbound

This paper cites Robust inference via generative classifiers for handling noisy labels.

Robust Reward Alignment via Hypothesis Space Batch Cutting Robust inference via generative classifiers for handling noisy labels

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.539682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.433242Z digest=sha256:4553549e4775ddcdc35d2b1a8943d66729134ea57b487f07f7d5f446432f69b0

Observation d2c12e43-d28d-4bc1-83a9-02cf4b1bb1b7 · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting B-pref: Benchmarking preference-based reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.532911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.435489Z digest=sha256:ca84d71b66c06f39c8d943b38eec2a432a6eb29713248906b34e66ba5aa2d53f

Observation a6475070-01e0-4310-81f0-89150137b3d4 · outbound

This paper cites M., and Abbeel, P.

Robust Reward Alignment via Hypothesis Space Batch Cutting M., and Abbeel, P

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.526717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.438238Z digest=sha256:d8e8cdee8bc0c8081637fb67dbd933ea4d6eec31bb519683249262beb34ccbe8

Observation de70b1c5-8dd4-4ecf-a9d0-38348615a996 · outbound

This paper cites CANDERE-COACH: Reinforcement Learning from Noisy Feedback.

Robust Reward Alignment via Hypothesis Space Batch Cutting CANDERE-COACH: Reinforcement Learning from Noisy Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.440352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.440352Z digest=sha256:f3c5fa7668ddd590ca7b3f89f21d8d4da12f3a8dface9f21662921e6487fc152

Observation 1633ce63-34da-4a49-ba0a-f9916dd8ea40 · outbound

This paper cites Reward uncertainty for exploration in preference-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Reward uncertainty for exploration in preference-based reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.520002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.443382Z digest=sha256:c5f7af204486e8dce8633d3c0ebb8cdeec17277d7b5a25ded7890131b40485de

Observation 0b241ae0-1068-414b-835b-ce3093103431 · outbound

This paper cites Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.513205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.446214Z digest=sha256:fb63d9f19e3689c6220b55c917015bba43b8fc4075bb7a1178a93c5365457ef7

Observation 2a36d83c-1c8c-489b-8898-e4c8d06a4d8d · outbound

This paper cites Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp.\ 6448--6458.

Robust Reward Alignment via Hypothesis Space Batch Cutting Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp.\ 6448--6458

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.505542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.448711Z digest=sha256:1e2599589c4caae183f5cdb165c5b1ce399b811163ba83caa51163a7fe576242

Observation 600468a7-b4e1-4554-96fe-91d8ec0893ec · outbound

This paper cites Normalized loss functions for deep learning with noisy labels.

Robust Reward Alignment via Hypothesis Space Batch Cutting Normalized loss functions for deep learning with noisy labels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.497933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.451264Z digest=sha256:415bdbe4753bf9696f4e65df63c75878750782e517d0c34b7656e8eb391c3fa3

Observation e173d207-2496-40bd-83f0-81aaf3a58a02 · outbound

This paper cites Learning multimodal rewards from rankings.

Robust Reward Alignment via Hypothesis Space Batch Cutting Learning multimodal rewards from rankings

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.490607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.453805Z digest=sha256:149bbf36ca540ece0722d02afb394e4a1836c0cc30518786828cc49c6fe9220b

Observation 329a7e06-693f-481f-92dd-b8ffa2be7630 · outbound

This paper cites Active reward learning from online preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Active reward learning from online preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.456719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.456719Z digest=sha256:328f3410f6a5ef2072c7221be9b5aa9aa4d78a438c752f11486a01de16cecf4f

Observation d982223b-d224-4711-af23-c842965d9327 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.459491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.459491Z digest=sha256:c6b4fcb5d20b1edb220a461ae707c17ec7a79d212e950ef374521da662f1d857

Observation b803d692-1b8f-4421-96ca-b7c9ab70b817 · outbound

This paper cites In-hand object rotation via rapid motor adaptation.

Robust Reward Alignment via Hypothesis Space Batch Cutting In-hand object rotation via rapid motor adaptation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.454283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.462261Z digest=sha256:84e624d5664d00b759c93012958a0b761d04d6161c1d283406735c0c97a0d223

Observation 493d9cf1-8e79-4aeb-9717-1b5d1dee2469 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Real-world humanoid locomotion with reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.379298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.464334Z digest=sha256:4a9f6e8c8892ed356acf4ea52f463ab5c9caa1c3630f20be6204b05bb452ab2b

Observation fa30657b-5875-447a-b850-6f022d688130 · outbound

This paper cites D., Sastry, S.

Robust Reward Alignment via Hypothesis Space Batch Cutting D., Sastry, S

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.372186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.467115Z digest=sha256:68cab373c5956fd44a91fb2130a57c98a2124797c0dac9a4204bfa96847446c6

Observation 98783067-ef83-449d-8754-5ff5bceaee6c · outbound

This paper cites Active Learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Active Learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.365180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.469869Z digest=sha256:a7d554c191b4ca83569e58b2ca0d4697518fe6361bf71b93355dc18d594fa9d6

Observation 84607223-87fc-4817-aced-ef6d4cecf125 · outbound

This paper cites and Joachims, T.

Robust Reward Alignment via Hypothesis Space Batch Cutting and Joachims, T

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.358134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.472562Z digest=sha256:1bf22198bf1a4ad73f0f2a93d48a8585213559284cfacb4b3908955573debd03

Observation 65af65c1-3e78-4ed1-8d39-e699357bd6dd · outbound

This paper cites Mujoco: A physics engine for model-based control.

Robust Reward Alignment via Hypothesis Space Batch Cutting Mujoco: A physics engine for model-based control

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.505481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.505481Z digest=sha256:5b6774e6927e48d5666f4821b2bc7e735fa34e62da4144e716b8ea540c16f79a

Observation 8c461b5e-e3bc-4537-af70-750e50c71fb1 · outbound

This paper cites Deepgait: Planning and control of quadrupedal gaits using deep reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Deepgait: Planning and control of quadrupedal gaits using deep reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.350777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.531962Z digest=sha256:b393d281579f4e11785eed0a90e8b887a99e4ac94d72e244ad882a5c07f90327

Observation 7b5061c5-85a9-4756-8663-a7dbaf852d99 · outbound

This paper cites Statistical learning theory.

Robust Reward Alignment via Hypothesis Space Batch Cutting Statistical learning theory

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.342935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.563927Z digest=sha256:3226abfbcf42f12e21d5afdb8dd4db092a6d4d7af9efb3e8e5059c03b0bc3f67

Observation d9eaba56-96bd-4b1e-b4da-63c707eb2255 · outbound

This paper cites Model Predictive Path Integral Control using Covariance Variable Importance Sampling.

Robust Reward Alignment via Hypothesis Space Batch Cutting Model Predictive Path Integral Control using Covariance Variable Importance Sampling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.626321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.626321Z digest=sha256:08807d62c8810759b01d815513a7f0d419cec692d811b084452d9efc38f2ab98

Observation d612c389-17e6-488a-ad68-bcbbcbb6b901 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Robust Reward Alignment via Hypothesis Space Batch Cutting A survey of preference-based reinforcement learning methods

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.688311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.688311Z digest=sha256:9679ca667c5608c726fd17f1b9d9e300c29a9cbaa4e0ed843a2ad62b4e320ca0

Observation 6c84d1ad-ae80-4237-bf06-93a289dd362a · outbound

This paper cites J., and Jin, W.

Robust Reward Alignment via Hypothesis Space Batch Cutting J., and Jin, W

Reference 51

Resolution
verified exact
raw_fallback, observed 2026-08-09T10:43:30.044746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.751931Z digest=sha256:2dd424d487f453055b89e8b1fbb296b6c94bceebeeaa46d0b741cd0a9a3a664f

Observation 98541519-e0c2-4f32-bf4e-d08c7828c5eb · outbound

This paper cites Reinforcement learning from diverse human preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Reinforcement learning from diverse human preferences

Reference 52

Resolution
verified exact
doi, observed 2026-08-09T10:43:29.960746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.803585Z digest=sha256:957a790d8455c6c222a7ac907b713d6d6c44ff93e63154014198347939f6d454

Observation 222d7e08-795e-4030-a439-79181c413d65 · outbound

This paper cites W., Zhang, Y., Sun, J., Zhang, C., and Zhang, R.

Robust Reward Alignment via Hypothesis Space Batch Cutting W., Zhang, Y., Sun, J., Zhang, C., and Zhang, R

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.331409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.861494Z digest=sha256:cfafff0e1104b69c7d78d49d97544bba0157decbe7acb434b40159efff252ac6

Observation 612708e9-1ebf-4201-a4a4-50eba0878a15 · outbound

This paper cites Rotating without seeing: Towards in-hand dexterity through touch.

Robust Reward Alignment via Hypothesis Space Batch Cutting Rotating without seeing: Towards in-hand dexterity through touch

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.323771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.906482Z digest=sha256:a4bdfdcc5095ffa72654068fb729467995c960868f5565ead0923b03e41c3476

Observation 4b58ab11-be9c-4095-98f7-ebb51bfb6c50 · outbound

This paper cites Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback.

Robust Reward Alignment via Hypothesis Space Batch Cutting Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.316013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.924836Z digest=sha256:170589c49062225355d187fc2ab616464a0d0d60d13ec17a3c0fd212eae809bf

Observation cf8037f1-49f0-45ea-9177-7b698dfb0fcc · outbound

This paper cites N., and Lopez-Paz, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting N., and Lopez-Paz, D

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.308089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.929348Z digest=sha256:3c7f7c1164aaf3bb97d8701749c471404b9a93de2ae05295eb14af00d10033a0

Observation 3f602393-683b-4a46-b1c6-a2c40e78eaca · outbound

This paper cites Robust curriculum learning: from clean label detection to noisy label self-correction.

Robust Reward Alignment via Hypothesis Space Batch Cutting Robust curriculum learning: from clean label detection to noisy label self-correction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.931811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.931811Z digest=sha256:3d31ad5dc11501f6233ee8a589dff45f7c1b5f259ae7dab7b2e6dc08be27f477

Observation 306be71d-1d9d-4666-9187-3a028e1bf514 · outbound

This paper cites A., Atkeson, C.

Robust Reward Alignment via Hypothesis Space Batch Cutting A., Atkeson, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.257401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.933967Z digest=sha256:307f2e8afb442ad436b0a90577dba0c30d23855d4c815fda9a1ee72aa1f76a2e

Observation cfd5d2b2-3fdc-41af-b6ed-da0903ac5905 · outbound

This paper cites write newline.

Robust Reward Alignment via Hypothesis Space Batch Cutting write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.936815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.936815Z digest=sha256:38fef4e5a7deb9e454a993fb050f7a2bafaa1bae2bbba605b539ae1531492808

Pith citing papers

No inbound Pith citation observations are available.