Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-07T16:24:43.688967Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 1 inbound Pith citation observation for arXiv:2604.25872.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-07T16:24:43.688967Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 100 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7a34d8fc-b13d-43e3-ac7f-067bd024a5c4 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 25482215-e86d-4465-b6c4-70102db495c1 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bd5e698c-3480-452c-9ca7-4047323ee798 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Understanding the impact of entropy on policy optimization
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d54c6f2-c896-41a5-ac7e-88e78269e706 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Concrete Problems in AI Safety
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 246d52d9-a8e3-4c28-b4bb-2de9932b416f · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Potential-based shaping in model-based reinforce- ment learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b373cf31-2bdd-4a9f-9bfa-2a15ebe38687 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient InfAlign: Inference-aware language model alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 053b26ba-67ef-4828-a9e4-f1c62adc4f01 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Dota 2 with Large Scale Deep Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 51e7cdf1-4e82-42b6-803e-d770919c1c2b · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b77832eb-85f9-4636-a428-0ba3f8298037 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The accuracy paradox in rlhf: When better reward models don’t yield better language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b0803899-d69e-443e-b11f-e67df679bff4 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Heuristic-guided reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4899a4c-28d0-444c-a2f5-3d47850b65af · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning navigation behaviors end-to-end with autorl.IEEE Robotics and Automation Letters, 4(2):2007–2014
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2df8d76b-484d-47ff-a1e4-6730ad9ade43 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient More is less: inducing sparsity via overparameteriza- tion.Information and Inference: A Journal of the IMA, 12(3):1437–1460
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bafe182-0e35-4c0b-bd22-7149147b4048 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward model ensembles help mitigate overoptimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 001ce548-e2a0-4a8a-b017-38e03f8407c4 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ultrafeedback: Boosting language models with high-quality feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b6bba76d-fb51-4eb3-8356-1311f0257689 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Maximum expected hitting cost of a markov decision process and informativeness of rewards.Advances in Neural Information Processing Systems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3195c83e-82e7-4a9b-983a-e246baccc8b9 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Exploration-guided reward shaping for reinforcement learning under sparse rewards.Advances in Neural Information Processing Systems
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61d4bf89-748f-4f6b-87dc-723058a0a7dd · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Continuous vs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 77b8edde-0e2a-42b1-b3da-fffbe35e77f9 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The perils of optimizing learned reward functions: Low training error does not guarantee low regret
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2304c86-3d52-4ac1-9fc0-d76092cecfea · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is a good foundation necessary for efficient reinforcement learning? the computational role of the base model in exploration
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d14e7f16-b218-42ac-aab0-be81f3b28c37 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient How to evaluate reward models for rlhf
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39e03859-ef9c-4662-8d30-f9a6d8735a40 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Scaling laws for reward model overoptimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bacdaebb-550e-4a6d-8693-a5602a9b3cf6 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient An alternate policy gradient estimator for softmax policies
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00c18eea-a1cd-4f85-95c7-8248d629d8c7 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Quantifying differences in reward functions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ff05225-ae09-41cd-b833-66e2010ba483 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The Llama 3 Herd of Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd9ba846-5e88-440f-81d5-a62981433e19 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward shaping in episodic reinforcement learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0cee9084-8afc-4ecf-b712-e6f53d50705b · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2b53894-57f5-4717-aae4-3f45bd19d9af · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity.Advances in Neural Information Processing Systems
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e8bd2e1e-9a1d-4d8e-bf3e-e7939a5ca7fe · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Neural Replicator Dynamics
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a9db3d6b-33f0-4820-b168-cb647a65ce88 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e768850-5d4d-4dc9-a3bf-9d122f133052 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the Emergence of Implicit Curriculum in RLVR Learning Dynamics
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4292780-bbe9-47f2-8cb6-4447e8fb1132 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79356e28-dcfb-46ca-b1ba-7c2e125cd99e · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Goodhart’s law in reinforcement learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ee1562b-2bad-4ee7-994d-05965bcfd311 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Beyond stationarity: Convergence analysis of stochastic softmax policy gradient methods
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b1c7b91-3a83-4ce4-819c-d1c6b11ad14f · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Buy 4 reinforce samples, get a baseline for free!Deep Reinforcement Learning Meets Structured Prediction ICLR Workhsop
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c313ba31-c8b0-4040-8b8d-569a560a30e6 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient A neural collapse perspective on feature evolution in graph neural networks.Advances in Neural Information Processing Systems
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f74ff2c-b947-4f68-b9c1-2b78bd4fd77a · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Correlated proxies: A new definition and improved mitigation for reward hacking
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b79c9aee-6548-486f-8dd0-87ec58fd2409 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6b3e9138-9638-4f18-9945-7bf498b74769 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rewardbench: Evaluating reward models for language modeling
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a9957a1d-b670-4804-bde9-34413deea7cb · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The influence of reward on the speed of reinforcement learning: An analysis of shaping
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5b2414b-01f3-4776-957a-58931e1a29a7 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Softmax policy gradient methods can take exponential time to converge
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3acb3b2f-9de6-41fa-b43c-54d351479c5a · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Hashimoto
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 054f8504-b2e6-4b03-a827-9ab741972c4c · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8e41e6ee-74ea-4de0-96cb-53143922b22e · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78f1a5c5-0921-403e-ae92-c241984b705b · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Elementary Analysis of Policy Gradient Methods
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcfc11c6-1ee1-4546-b85c-2d855129d2c7 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient RLTF: Reinforce- ment learning from unit test feedback.Transactions on Machine Learning Research
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c70cdc9-b1a0-4538-86f7-feb0e02f2e9c · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rm-bench: Benchmarking reward mod- els of language models with subtlety and style
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d310456e-86a5-4dd7-a59d-3dad4320a502 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient RewardBench 2: Advancing Reward Model Evaluation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da614b5e-a305-461e-a951-b4d3dc04a718 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward engineering for reinforcement learning in software tasks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c30e2b3-8a65-4ef6-aad2-cc8f10b05a47 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward functions for accelerated learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c95685b4-ae34-43b6-b82a-ae9695cbb204 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Escaping the gravitational pull of softmax.Advances in Neural Information Processing Systems
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3d0dd837-3d61-420f-bbb6-f7eb95c65c0d · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the global convergence rates of softmax policy gradient methods
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 960fb550-3502-4e22-9e08-65740ce9b899 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Leveraging non-uniformity in first-order non-convex optimization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 248c8549-954e-416d-8029-11395b13baf0 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ordering-based conditions for global convergence of policy gradient methods.Advances in Neural Information Processing Systems
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33d3b4f0-5ba3-434f-ae45-e150ee40740d · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Stochastic gradient succeeds for bandits
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 987f6b8a-5dfd-40a6-9e92-7c999a5a1dde · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Policy invariance under reward transformations: Theory and application to reward shaping
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e28f85ac-e154-4a89-a61c-b8fa25db28e3 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient 2 OLMo 2 Furious
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a170f578-7634-4aa8-bc62-0b5865648f16 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Olmo 3
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 519a7f06-fdb2-4a40-8947-523801103c68 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e0aef6b5-3da2-4eea-b445-f7fab22c60b0 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward gaming in conditional text generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30751c67-af25-4434-8dad-9db33a339f83 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Automatic differentiation in pytorch
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57dbbfe6-b229-48be-a261-b000c3d43132 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Amp: Adversarial motion priors for stylized physics-based character control.ACM Transactions on Graphics (ToG), 40(4):1–20
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dd783c71-682e-4090-993b-c31105667067 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Generalizing verifiable instruction following
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 097d2b6e-9754-4d08-8403-6e69c16330f1 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1760c012-5233-4b15-9b96-c2cf885b0f2f · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning to drive a bicycle using reinforcement learning and shaping
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad7c4af3-c797-4a68-afb8-233e215bcba2 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in deep learning may not be explainable by norms
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10457d44-eb20-4523-a1c2-b0ab66501224 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in tensor factorization
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation efc34028-c14f-4e67-a47b-ed586a5b2417 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 471d2902-ec56-4d25-a6a4-00650a010efd · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Susskind, and Etai Littwin
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b1a8b75-d679-4728-8d7b-6f986db2992e · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient What makes a reward model a good teacher? an optimization perspective
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d585a040-940a-4d8f-be86-bd62a25510d7 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Why is your language model a poor implicit reward model? InInternational Conference on Learning Representations
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cee29f65-6e2d-4f9f-aa72-c6d6a227f58c · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias.Advances in Neural Information Processing Systems
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b9d81ba-92ac-41be-82b4-16e2a588389a · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 903a3577-8159-47d5-953b-5ea29d577059 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 16b258bc-2c08-45a3-bc04-c0cba41e2226 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Proximal Policy Optimization Algorithms
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d69d36c-3d2d-4992-a032-4116c8f57c0b · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c3ce965-d6d4-4970-8910-bca64027c993 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Intrinsically motivated reinforce- ment learning: An evolutionary perspective.IEEE Transactions on Autonomous Mental Development, 2 (2):70–82
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f18a4585-de3d-4add-b1dd-e5088a989381 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Defining and characterizing reward gaming.Advances in Neural Information Processing Systems
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4685644c-8a2c-47f1-a102-03505ab7b8ff · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Starc: A general framework for quantifying differences between reward functions
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff578be5-2523-4960-adf2-91febc47327c · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The implicit bias of structured state space models can be poisoned with clean labels
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8052fc45-4fd5-4da8-9a7f-298f85c40e55 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward design via online gradient ascent.Advances in Neural Information Processing Systems, 23
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ad401cd-72af-4fef-8bb7-e731b91238ce · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning to summarize with human feedback
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ae292cf-cc5a-4eb9-afaf-454659e56a41 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rl grokking recipe: How does rl unlock and transfer new algorithms in llms? InInternational Conference on Learning Representations
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33501b71-3632-4fad-8b33-1198ac9fb464 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient All roads lead to likelihood: The value of reinforcement learning in fine-tuning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 82c9f1d3-92ad-4d9d-ba59-3a1cf0f06f06 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Understanding the performance gap between online and offline alignment algorithms
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 530bd498-9511-42bb-ac2e-50bc6c38d329 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e9020eb9-ad60-4f07-bbc6-55043e5a9723 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Perturbation analysis of neural collapse
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6f5f133-7298-4f34-8121-1faf2ea73373 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in relu networks with the square loss
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78c93807-1306-4273-b3fb-328262e3caf1 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97aadcd7-abbf-4da2-a2fc-00e035fbd907 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rethinking reward model evaluation: Are we barking up the wrong tree? InInternational Conference on Learning Representations
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a299a1bd-a796-466a-bc73-4d8fd1093a67 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8(3):229–256
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee639010-6b2f-4c5c-81d1-f0e14cb02f1c · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df978ab5-46a8-4420-8cba-df6b7d5ccc91 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Kernel and rich regimes in overparametrized models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91cb503f-62a6-4488-aa6a-22baec79bc27 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Dynamics-aware comparison of learned reward functions
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf047b6a-c400-4d90-8d3d-0e3bfbe6ea2d · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3fcb075-38e6-4eaf-b9b1-ee1850569394 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Qwen3 Technical Report
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 59af3c34-be70-43a5-9e92-5b9434de1e4d · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7c9d62b-ed80-410c-8034-fa70ad788a31 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On learning intrinsic rewards for policy gradient methods
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75db98d7-8b6b-4d13-8254-9d0ce84bccde · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rmb: Comprehensively benchmarking reward models in llm alignment
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 86b2e17f-577a-4982-8958-8f23b466519c · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Fine-Tuning Language Models from Human Preferences
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae83b582-bad8-46ba-b797-a75062ed3730 · outbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient partial rewards
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28af7302-b8f7-4f48-885d-77815d82bb5e · inbound
Multimodal Reward Hacking in Reinforcement Learning When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.