Pith. sign in

Paper Citation Record · LEDGER

Performance Optimization of Ratings-Based Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2501.07755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07755 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:40:11.565215Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:50.092283Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:59:51.529424Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e14325e-0f9e-4519-b762-d45702d5e5fa · outbound

This paper cites Agent57: Outperforming the Atari Human Benchmark.

Performance Optimization of Ratings-Based Reinforcement Learning Agent57: Outperforming the Atari Human Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.481193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.481193Z digest=sha256:05601a914d43e5872c68a7600c462ba5bfa15dca2cee61f96cbff4b27598d86e

Observation 744c028b-0a69-4661-b68e-880abb38bbe7 · outbound

This paper cites an unresolved cited work.

Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:40:11.841478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.486158Z digest=sha256:938e9ca4ed53b42696e45a6d425be8e737c103078e570f9bccdcd1634bc92f32

Observation 52b88e9c-9b94-451d-b4e2-f47a3ca5b1c4 · outbound

This paper cites F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D.

Performance Optimization of Ratings-Based Reinforcement Learning F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.490169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.490169Z digest=sha256:24a3243f52e78672c3c37949a330d37e6a1cd61b85282d241c952a69cd1f9875

Observation 48c8d32f-7867-4ca9-9e95-242ec330554d · outbound

This paper cites Activation Functions in Deep Learning: A Comprehensive Survey and Benchmark.

Performance Optimization of Ratings-Based Reinforcement Learning Activation Functions in Deep Learning: A Comprehensive Survey and Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.494329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.494329Z digest=sha256:1edeebb0e90f955d89fb1d4452e241c2559c08a471c36fb40c45fb2a53dafaf9

Observation 30fcdbf7-88b1-4d0c-89d3-cb7d82514cfc · outbound

This paper cites an unresolved cited work.

Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:40:11.819176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.498866Z digest=sha256:15c91d40251842c104a73a9ef20f21347db363ac27b93e9c9d04a493ff326938

Observation 2d161150-d572-4476-80d5-73a31a3f35e3 · outbound

This paper cites B-Pref: Benchmarking Preference-Based Reinforcement Learning.

Performance Optimization of Ratings-Based Reinforcement Learning B-Pref: Benchmarking Preference-Based Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.503146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.503146Z digest=sha256:d1c67fa8a095c39f828ff70e003abd719f694a4d34372873f1877844e3e03272

Observation 3a26f077-8a32-4279-8be7-74a0068d1892 · outbound

This paper cites an unresolved cited work.

Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:40:11.805269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.508840Z digest=sha256:0e7a15e93fbf5cfe5f3d81dc8a038351f85435747c459c3e28e076c0ee6bffe0

Observation c61b0677-ef68-4e40-9d4e-8968334407c1 · outbound

This paper cites Decoupled Weight Decay Regularization.

Performance Optimization of Ratings-Based Reinforcement Learning Decoupled Weight Decay Regularization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.513684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.513684Z digest=sha256:8108eb4a0411b2a0708ee68c5555b7ec02d8951fbeade6bb63c73183d0b77f0f

Observation 14a7a304-0416-48c5-83c7-b67e613e19a2 · outbound

This paper cites Y.; Russell, S.

Performance Optimization of Ratings-Based Reinforcement Learning Y.; Russell, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:40:11.791462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.518411Z digest=sha256:1a9a99eec607c08900a74a22652fd91937a2ebb7872713c3e39e2255989e33fe

Observation a5c42e4e-cdb1-4573-a1f1-6f1cdbdf04b1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Performance Optimization of Ratings-Based Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.523222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.523222Z digest=sha256:1735f845b868c4b46841fc365b6cbd968cc72ac17255d06b77c55577f17f4367

Observation 0affd4da-57dc-466e-b54c-2732e2957c3c · outbound

This paper cites I.; and Kashif, F.

Performance Optimization of Ratings-Based Reinforcement Learning I.; and Kashif, F

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:40:11.777195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.528309Z digest=sha256:8626f5949c28fb866ab16c19b76f89bf2ff481d11bfd16b414824a1a031085f0

Observation 471b4842-f474-4ab4-b0ae-fe774862b4ac · outbound

This paper cites an unresolved cited work.

Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:40:11.763689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.533234Z digest=sha256:73213e757006c73cd6c427961b23f7128e3a0d76e2dc17fd9a0af99982bd6130

Observation b3e105d3-6fe1-49d4-b27a-5fee52ea17c3 · outbound

This paper cites an unresolved cited work.

Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:40:11.749548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.537826Z digest=sha256:1ec6c117e5eb78bee8cf6186fbf4ebe8f62a3ebfb557e9acbf95e0ef688c19aa

Observation 8f3f9179-5345-44d3-a094-fe598652c8c0 · outbound

This paper cites Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes.

Performance Optimization of Ratings-Based Reinforcement Learning Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.542400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.542400Z digest=sha256:d1949be9586e7f27a2820184a740865f9a3daa478cb37adae6dd1e0f13714f06

Observation c19ae490-dde3-42ff-b0ee-6cc1f0142991 · outbound

This paper cites J.; Waytowich, N.; and Cao, Y.

Performance Optimization of Ratings-Based Reinforcement Learning J.; Waytowich, N.; and Cao, Y

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:40:11.735069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.547057Z digest=sha256:09d2e325c61d69f66d196c69e2203831b6696ab53f9c62810a8ba0adb3cb4062

Observation c5d34811-a011-4313-ace6-0c3bd96c73e0 · outbound

This paper cites Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks.

Performance Optimization of Ratings-Based Reinforcement Learning Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:40:11.606335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.551559Z digest=sha256:5422846649f4a350718f868b9a36f2238d24d3c25d5bba40aec5c7240c3d700d

Observation ce75a339-4f0a-456e-b447-15176c87fa94 · outbound

This paper cites D.; Maas, A.; Bagnell, J.

Performance Optimization of Ratings-Based Reinforcement Learning D.; Maas, A.; Bagnell, J

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:40:11.719690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T20:40:11.556096Z digest=sha256:26ad03a52d1e861e32fffc95e2536cb70fd7759776084266eac611d8f103e0e3

Observation 097a20d8-e4f7-4942-87a4-8fae3024bb0b · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Performance Optimization of Ratings-Based Reinforcement Learning , " * write output.state after.block = add.period write newline

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.560318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.560318Z digest=sha256:e7ebf63fc8b22ba03d5fe340bdbe5e0538559db0772c302d805a1947a33defc5

Observation 8008686b-5ef9-42d7-b7e9-f41124a21d3f · outbound

This paper cites write newline.

Performance Optimization of Ratings-Based Reinforcement Learning write newline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:40:11.565215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:40:11.565215Z digest=sha256:cb0dd8b761ef971e59d02674f07e0238e016175b83e05d329184914d83e4db2d

Pith citing papers

Observation 198b248d-6de3-42a4-93d3-e0d195550c7f · inbound

Multi-Task Reward Learning from Human Ratings cites this paper.

Multi-Task Reward Learning from Human Ratings Performance Optimization of Ratings-Based Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:59:51.596559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T04:59:50.092283Z digest=sha256:1019c9da5293bf44c6ada7f5449388e9023e9f34dc1d281f755f6c8b1133438b