Pith. sign in

Paper Citation Record · LEDGER

Multi-Task Reward Learning from Human Ratings

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.09183.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09183 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:51.021706Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f6bb597-f374-4dc0-9f5c-80702e8d0973 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Multi-Task Reward Learning from Human Ratings F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.760680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.760680Z digest=sha256:d5b687ddd0a668fc14bf29e7df01e6e51877319b4068ec90d8079fa6993ee38e

Observation 32e67037-1214-456a-be17-292daea69630 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Multi-Task Reward Learning from Human Ratings Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.832394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.832394Z digest=sha256:a660c0f17075e198fce76dfc5346ec07e2b2df2940c3357c53b7ee6759cb864b

Observation e1ddc9f2-0cb7-47ac-9dbe-9e248d529878 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics.

Multi-Task Reward Learning from Human Ratings Multi-task learning using uncertainty to weigh losses for scene geometry and semantics

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.806826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:49.916874Z digest=sha256:58f2715bfa3239e5219fd1ff3b6a2d43c73255fac814c6208d7d2a1463143a04

Observation a3f0b867-8090-403b-99ff-95b8e16ce518 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Multi-Task Reward Learning from Human Ratings Playing Atari with Deep Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:49.996007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:49.996007Z digest=sha256:5639af1709b3378c224660f7ea9f49b0358b0acd01b84335b6c7a6eb231ec9de

Observation 198b248d-6de3-42a4-93d3-e0d195550c7f · outbound

This paper cites Performance Optimization of Ratings-Based Reinforcement Learning.

Multi-Task Reward Learning from Human Ratings Performance Optimization of Ratings-Based Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:59:51.596559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.092283Z digest=sha256:3457be18efc8759562231c1ccbcde2cfed7e00b5748883f7728515a839c98f77

Observation 193da9e2-238b-4da1-b164-5096c803f4a9 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.709502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.156703Z digest=sha256:d753df59ad86630a62493943cf0d886ee58a63c5300e2506393200881a2e43d3

Observation d213a956-cdbe-4e94-ad33-4dadd1eb6b94 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Multi-Task Reward Learning from Human Ratings Proximal Policy Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.248231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.248231Z digest=sha256:0621ad8717bf8926c2276f10a5ba629601a847f149c658e435fc5360ecc7d186

Observation 79f6a1e6-9583-4904-9999-8c30e3110ad4 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.556661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.328564Z digest=sha256:f50e3a2f0110fb211ef19e1cd9ffc5d54d1a672e8cd1ad6320ea94469fd5228b

Observation 5dc880e9-1868-4bfc-82c9-3e2c3b0c3851 · outbound

This paper cites an unresolved cited work.

Multi-Task Reward Learning from Human Ratings Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:59:52.281107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.424254Z digest=sha256:046ff74c6f0158e21a0d8d4d77499c226a6d0cb38075caa7bae4d4a9f33f3886

Observation 8a9c1ad1-640c-4584-826a-93eef9aba65b · outbound

This paper cites DeepMind Control Suite.

Multi-Task Reward Learning from Human Ratings DeepMind Control Suite

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.587696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.587696Z digest=sha256:f8ae2c39082075f4f63a2bbcd52f6eb60430c6b2d5169f76abb8c58c043d67df

Observation 860d5170-a297-4dbd-a06f-a95d0eb4ea47 · outbound

This paper cites Mujoco: A physics engine for model-based control.

Multi-Task Reward Learning from Human Ratings Mujoco: A physics engine for model-based control

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:50.678056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:50.678056Z digest=sha256:1cb0683cb555c334092e69c66c1a20516e6f942bceb9796ee87ebd48c97932fc

Observation 92368bed-3358-4f79-be95-3e301767955c · outbound

This paper cites J., Waytowich, N., and Cao, Y.

Multi-Task Reward Learning from Human Ratings J., Waytowich, N., and Cao, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.122919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.777524Z digest=sha256:c2c1c4cd1fd8c51e64216c5eab8d198176e7f6f1b5b2389d7b3dbd9dd3169389

Observation 48c9184b-f062-4a8c-afa4-d0394dbd217d · outbound

This paper cites Value of potential field in reward specification for robotic control via deep reinforcement learning.

Multi-Task Reward Learning from Human Ratings Value of potential field in reward specification for robotic control via deep reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:52.004313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.847757Z digest=sha256:495dd25792f700acd2467786537cf4c64d2ab7b51d2924097ccd9bb7ca859245

Observation 2bcbd690-7010-4138-9b8d-6227b1256353 · outbound

This paper cites Offline reinforcement learning with failure under sparse reward environments.

Multi-Task Reward Learning from Human Ratings Offline reinforcement learning with failure under sparse reward environments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:59:51.876339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.891547Z digest=sha256:68b677de4399c9d10163c61f098b16d1c5a842d01970ed4eea37123f44ccc5f0

Observation b5ba111c-b936-4da8-af73-df95b2eae750 · outbound

This paper cites R., and Cao, Y.

Multi-Task Reward Learning from Human Ratings R., and Cao, Y

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:59:51.343300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.949341Z digest=sha256:20552df2cda1455fcec9f35c343cc5ac40e4dfc40f5068932bad925cacb10dad

Observation 93cd9e43-9014-42cc-9656-347a30e9c078 · outbound

This paper cites write newline.

Multi-Task Reward Learning from Human Ratings write newline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:51.021706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:59:51.021706Z digest=sha256:454f0145adbf0fb2c9c5ff7cb2d02165f661883f3717035d903572f074a35d1b

Pith citing papers

No inbound Pith citation observations are available.