Pith. sign in

Paper Citation Record · LEDGER

Robust Reward Alignment via Hypothesis Space Batch Cutting

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2502.02921.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02921 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:43:29.936815Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c90ea59d-776a-4812-be9d-b6bbc4d740a8 · outbound

This paper cites GPT-4 Technical Report.

Robust Reward Alignment via Hypothesis Space Batch Cutting GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.811463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.811463Z digest=sha256:dc38effaf8d0a988397c6f65f6264c80d23a3f1c3e0a835de5ce89eb94a57288

Observation f72cd497-3718-4577-b12b-5f3fd559a746 · outbound

This paper cites April: Active preference learning-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting April: Active preference learning-based reinforcement learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.814983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.814983Z digest=sha256:595edd571c3903c3920b7329626d899feddc83990b4e72deb190715b1bcbe8dc

Observation 6caed3f8-8e48-47b1-b2dc-50f7a5c7e7ad · outbound

This paper cites Programming by feedback.

Robust Reward Alignment via Hypothesis Space Batch Cutting Programming by feedback

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:31.054467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.817953Z digest=sha256:2705a4f52975928825a1592b4d43f67ce1e3683f3d589bcd4f686037ac32d6af

Observation 47dc248a-b856-4cd0-95b0-e0b1eaa7db73 · outbound

This paper cites K., Anil, R., and Koren, T.

Robust Reward Alignment via Hypothesis Space Batch Cutting K., Anil, R., and Koren, T

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.956444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.820995Z digest=sha256:15f669dd596b89c33fe628be73354c8df4b95e1fe2ba9d9233407c3a5d75132b

Observation 46256e2b-c53c-4186-97b0-34bce3b35b05 · outbound

This paper cites Fine-tuning language models to find agreement among humans with diverse preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Fine-tuning language models to find agreement among humans with diverse preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.823812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.823812Z digest=sha256:58f7a20ab954281b7206903a557fe8b49ebea939ad10dc39157ee34671a166b7

Observation 8d415488-0680-4e94-b953-c9f1da5d1dbb · outbound

This paper cites L., Harvey, N., Liaw, C., and Mehrabian, A.

Robust Reward Alignment via Hypothesis Space Batch Cutting L., Harvey, N., Liaw, C., and Mehrabian, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.826189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.826189Z digest=sha256:2a9db03e73df74c09665bfe547f066b5eda6ca0cba7a65ec76a943631ac22bc6

Observation baa68838-deda-4f74-a280-a177ebdb6271 · outbound

This paper cites and Sadigh, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting and Sadigh, D

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.834634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.828918Z digest=sha256:47ff133f809d790c8357c7b2949d67b211269c6c6c12afad844af478ebb4fcf2

Observation aea9c93c-9e6c-447a-b4f4-e2a9186800b8 · outbound

This paper cites Batch active learning of reward functions from human preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Batch active learning of reward functions from human preferences

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.788939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.831254Z digest=sha256:849bd05974e024ebd87e4c9ee59674b206bf6832e57584c69ea66ae81689d8aa

Observation 7cf28edb-17f1-4b13-a035-419a8b36a5b0 · outbound

This paper cites J., and Sadigh, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting J., and Sadigh, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.781660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.833947Z digest=sha256:803cd8388f1b9dcf144c4e42200d15a3a9979bf48192d0c9093c9f98b9e95f46

Observation 79b9f3ad-7e0b-4948-80c6-501a781abbf9 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.836824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.836824Z digest=sha256:a2b2f8458188076e9bb5dcb9d0c45a0d6a69c153865872e27e5fc3c23ec9f61c

Observation a7768408-fd99-484b-8b2d-e2ac559d44e5 · outbound

This paper cites K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T.

Robust Reward Alignment via Hypothesis Space Batch Cutting K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., Wang, T

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.770751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.868127Z digest=sha256:6cf13521c5c9c874b0aff04e95b2bd41b83bc2636e74f16bc0258069ae8492cc

Observation 75b5cb6f-ee2e-4780-97a8-8b2224c690d8 · outbound

This paper cites Rime: Robust preference-based reinforcement learning with noisy preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Rime: Robust preference-based reinforcement learning with noisy preferences

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.763353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.898076Z digest=sha256:ab418cfdeabbb144a84c051db8a56c0abc12d0fc7a93452397cec1208f9d266d

Observation 207fe773-bd59-4b53-bfcd-ff9bbe4e72fd · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:28.923060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:28.923060Z digest=sha256:b383213440132ce0c568412cb62e198ce34bac01ef2a5e3312da8e5c30b648c3

Observation a90cc916-4e2f-4847-b594-6e2d1da110cf · outbound

This paper cites Active reward learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Active reward learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.751851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:28.969360Z digest=sha256:e6df32920d2d939edb53dc4dd37c60f65c3bd6871b1bbb3410f07597ad4214c8

Observation be4b8444-a54b-4481-9b09-616db01f7354 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T10:43:30.744853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.014153Z digest=sha256:91ef0ded802da213ed11aae826f3ca281d87ce17b052318a785c2daae03948ad

Observation 3556975a-9c5d-4e36-a74d-37e12791b809 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.040157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.040157Z digest=sha256:a14a210cdcfb0c3d338ab665f87c5bbad598931c87fa8b84ec9b12b6d417bb99

Observation 2aa62631-b756-4657-887b-66b3bdf78413 · outbound

This paper cites Dextreme: Transfer of agile in-hand manipulation from simulation to reality.

Robust Reward Alignment via Hypothesis Space Batch Cutting Dextreme: Transfer of agile in-hand manipulation from simulation to reality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.087356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.087356Z digest=sha256:56ee66a8d239235a97d3c41cac0fc3fd21263bcd03a646cf4c2efeb643998f81

Observation c90768d1-0497-4b55-9d14-9fc12726d149 · outbound

This paper cites A bound on the label complexity of agnostic active learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting A bound on the label complexity of agnostic active learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.728172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.136280Z digest=sha256:861f1f7f241865f73adc4070e5c18479a8891138de8b80633d747ce0aa54452a

Observation 5adef6a4-abde-478a-9fde-e2396446d4aa · outbound

This paper cites Contrastive preference learning: Learning from human feedback without rl.

Robust Reward Alignment via Hypothesis Space Batch Cutting Contrastive preference learning: Learning from human feedback without rl

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.720704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.195602Z digest=sha256:c4fee5e43178feaacf8e12cdeb9da8c236a6c0fd045819a77f458087cbf2a087

Observation 6b157bd0-0fbf-4204-ab79-7901b94de3d3 · outbound

This paper cites an unresolved cited work.

Robust Reward Alignment via Hypothesis Space Batch Cutting Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.260706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.260706Z digest=sha256:09c1f9739597ac3ee68162d2227e1b489bd7c56a45f6084b9eae4b1a633cf27e

Observation e2c450e8-3c25-4f76-acbf-a1854f0dce3c · outbound

This paper cites J., Kim, J., Kwak, M.

Robust Reward Alignment via Hypothesis Space Batch Cutting J., Kim, J., Kwak, M

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.708222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.288699Z digest=sha256:4f2ae70299cf4ff14d090272f71e739c55b6b6589c71de1c1a928a070524c1a1

Observation fd883f38-0857-4f0e-8e3e-3df5712aebd3 · outbound

This paper cites Anymal parkour: Learning agile navigation for quadrupedal robots.

Robust Reward Alignment via Hypothesis Space Batch Cutting Anymal parkour: Learning agile navigation for quadrupedal robots

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.700971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.318392Z digest=sha256:48a58a8a48a76c5b8cf84f0fa787e8a2cac89d9459e4168f27da50091b5b4677

Observation 105dff5c-cca8-4fcb-98a4-0cc41fa7cc85 · outbound

This paper cites Bayesian active learning for classification and preference learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Bayesian active learning for classification and preference learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.623589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.342740Z digest=sha256:1749398b9c47058ed89183ed9b0c1882886176805250d2149a91e9b0ee3248d5

Observation 5d47eb4f-7368-43bb-9402-13169fc760d9 · outbound

This paper cites Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo.

Robust Reward Alignment via Hypothesis Space Batch Cutting Predictive Sampling: Real-time Behaviour Synthesis with MuJoCo

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.355325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.355325Z digest=sha256:f8fc6a42715a7ab5c249809201a1d05bc15dffd7e0d7140f5966cccdf80e5494

Observation e7c590db-5767-421d-818f-3ec92b8b7e6d · outbound

This paper cites Reward learning from human preferences and demonstrations in atari.

Robust Reward Alignment via Hypothesis Space Batch Cutting Reward learning from human preferences and demonstrations in atari

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.371725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.371725Z digest=sha256:9eff9c37ce5aa3f7e0aaa654deef06b3ea46de31cac7dfad815da8fcd14b2efb

Observation 51bf2fa3-9ec5-44f8-ad7b-1c626319c2e4 · outbound

This paper cites Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels.

Robust Reward Alignment via Hypothesis Space Batch Cutting Mentornet: Learning data-driven curriculum for very deep neural networks on corrupted labels

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.381374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.381374Z digest=sha256:1894b79befc586cd9800ba8f0ab74fca83e83aa1bdacb9ab99a1513e15ed7e1a

Observation f1ba41ef-b407-40d3-aecb-2c43d8d25f0a · outbound

This paper cites D., Lu, Z., and Mou, S.

Robust Reward Alignment via Hypothesis Space Batch Cutting D., Lu, Z., and Mou, S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.561042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.400413Z digest=sha256:b95dd337fc63208a3a574717a633e7e723f5929a777bcaf7190c151544b9d64e

Observation f11811a7-a8b8-4459-aecf-fc6097ce0a29 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Robust Reward Alignment via Hypothesis Space Batch Cutting Adam: A Method for Stochastic Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.419649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.419649Z digest=sha256:23dfe55a3663eb3908f0475ac50bf9243ac2cac234a4fc6cbd57819de11df51f

Observation 95374d00-e105-4bc7-ab63-08b4cc1eb262 · outbound

This paper cites Crafting papers on machine learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Crafting papers on machine learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.430378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.430378Z digest=sha256:f0ec4ed1187610f03ec6517674e6bd00009427cccc02aa57181d9fb4faee241d

Observation c0cd69ee-c286-4026-8d6c-d99fc2107da2 · outbound

This paper cites Robust inference via generative classifiers for handling noisy labels.

Robust Reward Alignment via Hypothesis Space Batch Cutting Robust inference via generative classifiers for handling noisy labels

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.539682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.433242Z digest=sha256:25cd331082e686d857fb581bad102a981d356aa04146ff0eba926f183754d408

Observation d2c12e43-d28d-4bc1-83a9-02cf4b1bb1b7 · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting B-pref: Benchmarking preference-based reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.532911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.435489Z digest=sha256:ae4c617b6ebf83aa637b94d15e9d9af699a77bc334aa9664c6f056296d0bcc2c

Observation a6475070-01e0-4310-81f0-89150137b3d4 · outbound

This paper cites M., and Abbeel, P.

Robust Reward Alignment via Hypothesis Space Batch Cutting M., and Abbeel, P

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.526717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.438238Z digest=sha256:c6fae50c2cff9702b3f4641484c56510ed67218f74f260cbe800c9320819b7a6

Observation de70b1c5-8dd4-4ecf-a9d0-38348615a996 · outbound

This paper cites CANDERE-COACH: Reinforcement Learning from Noisy Feedback.

Robust Reward Alignment via Hypothesis Space Batch Cutting CANDERE-COACH: Reinforcement Learning from Noisy Feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.440352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.440352Z digest=sha256:cb214fe743d7e4d9709d8e182416606821b34b2dba752803774df44d3b3102da

Observation 1633ce63-34da-4a49-ba0a-f9916dd8ea40 · outbound

This paper cites Reward uncertainty for exploration in preference-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Reward uncertainty for exploration in preference-based reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.520002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.443382Z digest=sha256:e4dd5da2c047e07c39296b39e1726a1bbf184edf985f4ae5d88581161ad3500d

Observation 0b241ae0-1068-414b-835b-ce3093103431 · outbound

This paper cites Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Meta-reward-net: Implicitly differentiable reward learning for preference-based reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.513205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.446214Z digest=sha256:02ac27b450379ef53445d6315601896819c7b1d2c99c4a0f984ad05a0a9b1fc6

Observation 2a36d83c-1c8c-489b-8898-e4c8d06a4d8d · outbound

This paper cites Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp.\ 6448--6458.

Robust Reward Alignment via Hypothesis Space Batch Cutting Does label smoothing mitigate label noise? In International Conference on Machine Learning, pp.\ 6448--6458

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.505542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.448711Z digest=sha256:f28c1d1ecd6c7a4a56ccc4bf5a9150b4eb835ed30f8c39cfac4045d33214c84d

Observation 600468a7-b4e1-4554-96fe-91d8ec0893ec · outbound

This paper cites Normalized loss functions for deep learning with noisy labels.

Robust Reward Alignment via Hypothesis Space Batch Cutting Normalized loss functions for deep learning with noisy labels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.497933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.451264Z digest=sha256:c6b202b6ea3585ad953ecdeed77cb772ff22112aba8b8917136ef710c3b313b5

Observation e173d207-2496-40bd-83f0-81aaf3a58a02 · outbound

This paper cites Learning multimodal rewards from rankings.

Robust Reward Alignment via Hypothesis Space Batch Cutting Learning multimodal rewards from rankings

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.490607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.453805Z digest=sha256:1f8ad4b2a789b7b52316355121270766d3502eec751210ca56286bba89d07c47

Observation 329a7e06-693f-481f-92dd-b8ffa2be7630 · outbound

This paper cites Active reward learning from online preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Active reward learning from online preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.456719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.456719Z digest=sha256:fb72e844577e9257e3e376287ec6a3654a8c9177551e7d9c1f46c75af9491236

Observation d982223b-d224-4711-af23-c842965d9327 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.459491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.459491Z digest=sha256:313c41c4f65586b1d408a7d35be84440222a7fb10902264d01e27480b427dfc5

Observation b803d692-1b8f-4421-96ca-b7c9ab70b817 · outbound

This paper cites In-hand object rotation via rapid motor adaptation.

Robust Reward Alignment via Hypothesis Space Batch Cutting In-hand object rotation via rapid motor adaptation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.454283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.462261Z digest=sha256:c7ba5938fb187848f2b7540ee090aa731b5bb44c64d18e9755b4dc53434c2ed3

Observation 493d9cf1-8e79-4aeb-9717-1b5d1dee2469 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Real-world humanoid locomotion with reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.379298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.464334Z digest=sha256:c8f9a274b69006e18480d145dea2c478b5af00945965326c8b2430b5731b87c4

Observation fa30657b-5875-447a-b850-6f022d688130 · outbound

This paper cites D., Sastry, S.

Robust Reward Alignment via Hypothesis Space Batch Cutting D., Sastry, S

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.372186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.467115Z digest=sha256:e9bfccc247891280bcb20a4c9cfecb79ff8c11b34f6b23f465ca628be514794c

Observation 98783067-ef83-449d-8754-5ff5bceaee6c · outbound

This paper cites Active Learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Active Learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.365180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.469869Z digest=sha256:3e1d626012427a7874a7f5a8570b813bbec97c18d4d3a7727934fe47793bd1a6

Observation 84607223-87fc-4817-aced-ef6d4cecf125 · outbound

This paper cites and Joachims, T.

Robust Reward Alignment via Hypothesis Space Batch Cutting and Joachims, T

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.358134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.472562Z digest=sha256:e1dec1b40ff059085527e73afd8de0da9dfb8a69833f139503982827d658de35

Observation 65af65c1-3e78-4ed1-8d39-e699357bd6dd · outbound

This paper cites Mujoco: A physics engine for model-based control.

Robust Reward Alignment via Hypothesis Space Batch Cutting Mujoco: A physics engine for model-based control

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.505481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.505481Z digest=sha256:f09e46367c485aff6dc2d634d2eb5baf20d4afbed75ea0d6958d90bb258e2a9e

Observation 8c461b5e-e3bc-4537-af70-750e50c71fb1 · outbound

This paper cites Deepgait: Planning and control of quadrupedal gaits using deep reinforcement learning.

Robust Reward Alignment via Hypothesis Space Batch Cutting Deepgait: Planning and control of quadrupedal gaits using deep reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.350777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.531962Z digest=sha256:be85bb0a15cb80b340ba5eea877dbcf5a8d7fcdfb1c0087418b06d4e2e02eb5b

Observation 7b5061c5-85a9-4756-8663-a7dbaf852d99 · outbound

This paper cites Statistical learning theory.

Robust Reward Alignment via Hypothesis Space Batch Cutting Statistical learning theory

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.342935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.563927Z digest=sha256:58bb63c192a748f4a897bb034aa2731447374e4ae8489791a0af5b066bc591d8

Observation d9eaba56-96bd-4b1e-b4da-63c707eb2255 · outbound

This paper cites Model Predictive Path Integral Control using Covariance Variable Importance Sampling.

Robust Reward Alignment via Hypothesis Space Batch Cutting Model Predictive Path Integral Control using Covariance Variable Importance Sampling

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.626321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.626321Z digest=sha256:9d54257b598bf64e33bfee7f016f0eb08bbfeff9ae30581c08d9f27e51eaa27c

Observation d612c389-17e6-488a-ad68-bcbbcbb6b901 · outbound

This paper cites A survey of preference-based reinforcement learning methods.

Robust Reward Alignment via Hypothesis Space Batch Cutting A survey of preference-based reinforcement learning methods

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.688311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.688311Z digest=sha256:f234293f887d4ddad7c03f269d90bd3f3bad7dc721892dfacebc2e16d204000d

Observation 6c84d1ad-ae80-4237-bf06-93a289dd362a · outbound

This paper cites J., and Jin, W.

Robust Reward Alignment via Hypothesis Space Batch Cutting J., and Jin, W

Reference 51

Resolution
verified exact
raw_fallback, observed 2026-08-09T10:43:30.044746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.751931Z digest=sha256:74f3feb6415f6f178ef51a5f93fd49f7e5e8648613dd28a6ac83e419f6fed42b

Observation 98541519-e0c2-4f32-bf4e-d08c7828c5eb · outbound

This paper cites Reinforcement learning from diverse human preferences.

Robust Reward Alignment via Hypothesis Space Batch Cutting Reinforcement learning from diverse human preferences

Reference 52

Resolution
verified exact
doi, observed 2026-08-09T10:43:29.960746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.803585Z digest=sha256:c9d0a1454d7b600fcc6020edbd1dc2cb35b94ac7ff2db7a7bad27cb415550126

Observation 222d7e08-795e-4030-a439-79181c413d65 · outbound

This paper cites W., Zhang, Y., Sun, J., Zhang, C., and Zhang, R.

Robust Reward Alignment via Hypothesis Space Batch Cutting W., Zhang, Y., Sun, J., Zhang, C., and Zhang, R

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.331409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.861494Z digest=sha256:c1e9d43d3a7ef205e7b1c493ff25dd6205f6300c7cdeb13dd92d37e464de3921

Observation 612708e9-1ebf-4201-a4a4-50eba0878a15 · outbound

This paper cites Rotating without seeing: Towards in-hand dexterity through touch.

Robust Reward Alignment via Hypothesis Space Batch Cutting Rotating without seeing: Towards in-hand dexterity through touch

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.323771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.906482Z digest=sha256:e2666ddb14e7914cefd61b4a08a4461aa6666fee29a965225c2398b98619b9d3

Observation 4b58ab11-be9c-4095-98f7-ebb51bfb6c50 · outbound

This paper cites Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback.

Robust Reward Alignment via Hypothesis Space Batch Cutting Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.316013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.924836Z digest=sha256:99716af90a50680b00cfde2145444824cf9abc79c9ee89e257a4955faf1d0e3d

Observation cf8037f1-49f0-45ea-9177-7b698dfb0fcc · outbound

This paper cites N., and Lopez-Paz, D.

Robust Reward Alignment via Hypothesis Space Batch Cutting N., and Lopez-Paz, D

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.308089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.929348Z digest=sha256:15e40eadc3a1fc597cbaec4bf3fdfea5451ea218dd689232d6ad090303479d9f

Observation 3f602393-683b-4a46-b1c6-a2c40e78eaca · outbound

This paper cites Robust curriculum learning: from clean label detection to noisy label self-correction.

Robust Reward Alignment via Hypothesis Space Batch Cutting Robust curriculum learning: from clean label detection to noisy label self-correction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.931811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.931811Z digest=sha256:fa407412cd357a1ebaa4c6a030e63d14c61df7afa59c6d1e0ad232b05d84f525

Observation 306be71d-1d9d-4666-9187-3a028e1bf514 · outbound

This paper cites A., Atkeson, C.

Robust Reward Alignment via Hypothesis Space Batch Cutting A., Atkeson, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:43:30.257401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T10:43:29.933967Z digest=sha256:20b8495d04c9c5d8b71f9436193f7ca2751e7dcff49de8752e798bcd2ba3c21c

Observation cfd5d2b2-3fdc-41af-b6ed-da0903ac5905 · outbound

This paper cites write newline.

Robust Reward Alignment via Hypothesis Space Batch Cutting write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T10:43:29.936815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:43:29.936815Z digest=sha256:a5b9f4446237fde403449341ae4d36f2227eb2cbcfd3da5b4bca552fad81aa94

Pith citing papers

No inbound Pith citation observations are available.