Pith. sign in

Paper Citation Record · LEDGER

Multi-Task Reward Learning from Human Ratings

As of 19 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.09183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09183 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:51.021706Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f6bb597-f374-4dc0-9f5c-80702e8d0973 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Multi-Task Reward Learning from Human Ratings F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.760680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.760680Z digest=sha256:bb371846bf425ec67f6ad06b3d4fdbe8ae0fae536214a8dcd2b01a23efc84c28

Observation 32e67037-1214-456a-be17-292daea69630 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Multi-Task Reward Learning from Human Ratings Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.832394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.832394Z digest=sha256:f4e6856e4eb9ae78bbb752fbdaffa8eba1376fe07ad9e99ce188caf8dace2dcb

Observation e1ddc9f2-0cb7-47ac-9dbe-9e248d529878 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics.

Multi-Task Reward Learning from Human Ratings Multi-task learning using uncertainty to weigh losses for scene geometry and semantics

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.806826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:49.916874Z digest=sha256:a0736f974f1d2649ea959788a444c1667b1fa0384c19340cecb24644d386fd15

Observation a3f0b867-8090-403b-99ff-95b8e16ce518 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Multi-Task Reward Learning from Human Ratings Playing Atari with Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.996007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.996007Z digest=sha256:92adb96cab9b83afc954aca93c6d5db862244fbc5448532826e9cbe4d240c7c4

Observation 198b248d-6de3-42a4-93d3-e0d195550c7f · outbound

This paper cites Performance Optimization of Ratings-Based Reinforcement Learning.

Multi-Task Reward Learning from Human Ratings Performance Optimization of Ratings-Based Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:59:51.596559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.092283Z digest=sha256:72806726c8dffb2b5d4397b45a20105fba74db0abdeeeb3b7be1ab52c9ad01ad

Observation 193da9e2-238b-4da1-b164-5096c803f4a9 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.709502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.156703Z digest=sha256:981299487377e1627118d8731f340e54d1ca494cfd4fa953280d129ec7ac21ca

Observation d213a956-cdbe-4e94-ad33-4dadd1eb6b94 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multi-Task Reward Learning from Human Ratings Proximal Policy Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.248231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.248231Z digest=sha256:c446bcb7720280955ad24c26a6b8f8175c8d2817df8697b50dbefe8740755497

Observation 79f6a1e6-9583-4904-9999-8c30e3110ad4 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.556661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.328564Z digest=sha256:5c2966566f9670066ad48851e122a90a11928994710cdcf1a03616cdaef81088

Observation 5dc880e9-1868-4bfc-82c9-3e2c3b0c3851 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.281107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.424254Z digest=sha256:16a775176d81c8ea5e2d03d37aabe0d0a62c831c7c4664a4c6301ad3bd83552b

Observation 8a9c1ad1-640c-4584-826a-93eef9aba65b · outbound

This paper cites DeepMind Control Suite.

Multi-Task Reward Learning from Human Ratings DeepMind Control Suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.587696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.587696Z digest=sha256:18449b1457bd874671a945edc922587aa6c33a89658c26315b8a56cb2649fc56

Observation 860d5170-a297-4dbd-a06f-a95d0eb4ea47 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Multi-Task Reward Learning from Human Ratings Mujoco: A physics engine for model-based control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.678056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.678056Z digest=sha256:bb2a03d272392f300f3ac649565f0511ee83b18972efee521266472a5e2bbc26

Observation 92368bed-3358-4f79-be95-3e301767955c · outbound

This paper cites J., Waytowich, N., and Cao, Y.

Multi-Task Reward Learning from Human Ratings J., Waytowich, N., and Cao, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.122919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.777524Z digest=sha256:73590c27def3be015a391048b4338777f3d68cba40b37d87fb84ed6e44954a73

Observation 48c9184b-f062-4a8c-afa4-d0394dbd217d · outbound

This paper cites Value of potential field in reward specification for robotic control via deep reinforcement learning.

Multi-Task Reward Learning from Human Ratings Value of potential field in reward specification for robotic control via deep reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.004313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.847757Z digest=sha256:4de344113fb4cd29c7848084ac00f304ccb582e7235c7db860f2081514fd1671

Observation 2bcbd690-7010-4138-9b8d-6227b1256353 · outbound

This paper cites Offline reinforcement learning with failure under sparse reward environments.

Multi-Task Reward Learning from Human Ratings Offline reinforcement learning with failure under sparse reward environments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:51.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.891547Z digest=sha256:76d4a3bd3eddbec9be524d84cfdb2a50ed4f1bcfe3f20cb7ea53fe0baf5c9a37

Observation b5ba111c-b936-4da8-af73-df95b2eae750 · outbound

This paper cites R., and Cao, Y.

Multi-Task Reward Learning from Human Ratings R., and Cao, Y

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:59:51.343300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.949341Z digest=sha256:3d0eaa47bcab37c814e7d50d73648a96edd1dd76dc52650d2fb3eddb1ddf4493

Observation 93cd9e43-9014-42cc-9656-347a30e9c078 · outbound

This paper cites write newline.

Multi-Task Reward Learning from Human Ratings write newline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:51.021706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:51.021706Z digest=sha256:c938d0c49d4bf84b8ce2ffebd95dbbfaef64ba55f95fd9e3b47bd152e16a7c4d

Pith citing papers

No inbound Pith citation observations are available.