Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:40:11.565215Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2501.07755.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:40:11.565215Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:59:50.092283Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T04:59:51.529424Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2e14325e-0f9e-4519-b762-d45702d5e5fa · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Agent57: Outperforming the Atari Human Benchmark
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 744c028b-0a69-4661-b68e-880abb38bbe7 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 52b88e9c-9b94-451d-b4e2-f47a3ca5b1c4 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c8d32f-7867-4ca9-9e95-242ec330554d · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Activation Functions in Deep Learning: A Comprehensive Survey and Benchmark
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30fcdbf7-88b1-4d0c-89d3-cb7d82514cfc · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2d161150-d572-4476-80d5-73a31a3f35e3 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning B-Pref: Benchmarking Preference-Based Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a26f077-8a32-4279-8be7-74a0068d1892 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c61b0677-ef68-4e40-9d4e-8968334407c1 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Decoupled Weight Decay Regularization
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a7a304-0416-48c5-83c7-b67e613e19a2 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Y.; Russell, S
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a5c42e4e-cdb1-4573-a1f1-6f1cdbdf04b1 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0affd4da-57dc-466e-b54c-2732e2957c3c · outbound
Performance Optimization of Ratings-Based Reinforcement Learning I.; and Kashif, F
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 471b4842-f474-4ab4-b0ae-fe774862b4ac · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b3e105d3-6fe1-49d4-b27a-5fee52ea17c3 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f3f9179-5345-44d3-a094-fe598652c8c0 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Deep Reinforcement Learning for Robotics: A Survey of Real-World Successes
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19ae490-dde3-42ff-b0ee-6cc1f0142991 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning J.; Waytowich, N.; and Cao, Y
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c5d34811-a011-4313-ace6-0c3bd96c73e0 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ce75a339-4f0a-456e-b447-15176c87fa94 · outbound
Performance Optimization of Ratings-Based Reinforcement Learning D.; Maas, A.; Bagnell, J
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 097a20d8-e4f7-4942-87a4-8fae3024bb0b · outbound
Performance Optimization of Ratings-Based Reinforcement Learning , " * write output.state after.block = add.period write newline
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8008686b-5ef9-42d7-b7e9-f41124a21d3f · outbound
Performance Optimization of Ratings-Based Reinforcement Learning write newline
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 198b248d-6de3-42a4-93d3-e0d195550c7f · inbound
Multi-Task Reward Learning from Human Ratings Performance Optimization of Ratings-Based Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.