Pith. sign in

Paper Citation Record · LEDGER

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

As of 5 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 1 inbound Pith citation observation for arXiv:2604.25872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.25872 v1

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T16:24:43.688967Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact27
  • verified fuzzy69
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a34d8fc-b13d-43e3-ac7f-067bd024a5c4 · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the theory of policy gradient methods: Optimality, approximation, and distribution shift.The Journal of Machine Learning Research, 22(1):4431–4506

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.863945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:1ad10793ff4d70b90d14d43bc4b61dcf6809dcf966e2447d135ac1be9ae99f3f

Observation 25482215-e86d-4465-b6c4-70102db495c1 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:42:22.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:cee4f2f4787f093f0bb7154859666e836d6cf0f4450a2c161e5ea323e0051361

Observation bd5e698c-3480-452c-9ca7-4047323ee798 · outbound

This paper cites Understanding the impact of entropy on policy optimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Understanding the impact of entropy on policy optimization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.904768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6c8140684ac530a44fdef6d6c402fc66d59805eb6daa773c62cd159e8c0b0e5e

Observation 0d54c6f2-c896-41a5-ac7e-88e78269e706 · outbound

This paper cites Concrete Problems in AI Safety.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Concrete Problems in AI Safety

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.510706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:62e8bc8a3daa958161876a444bc04b7da0987e9e584c5ca5fe2404a84261d130

Observation 246d52d9-a8e3-4c28-b4bb-2de9932b416f · outbound

This paper cites Potential-based shaping in model-based reinforce- ment learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Potential-based shaping in model-based reinforce- ment learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.984806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d17adf9a72bff925646a59175935b2b4343b636344dbdc27cfe7998a81c07a7c

Observation b373cf31-2bdd-4a9f-9bfa-2a15ebe38687 · outbound

This paper cites InfAlign: Inference-aware language model alignment.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient InfAlign: Inference-aware language model alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.360622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:8ff4d915e6d3833a4de1834b9cd2f5ecb2b2d0d4cb852d6ba1d41700be90b9aa

Observation 053b26ba-67ef-4828-a9e4-f1c62adc4f01 · outbound

This paper cites Dota 2 with Large Scale Deep Reinforcement Learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Dota 2 with Large Scale Deep Reinforcement Learning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T22:18:16.018301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:868a707828419d8a226e18285866a12a9713a5e8acf98bf2b5f13a831ca72117

Observation 51e7cdf1-4e82-42b6-803e-d770919c1c2b · outbound

This paper cites an unresolved cited work.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-27T02:18:25.100242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d9dda44a4899edf335ebd8bca2d0e1c91d6b284d76fbaf8314c7fb8aa9120ea4

Observation b77832eb-85f9-4636-a428-0ba3f8298037 · outbound

This paper cites The accuracy paradox in rlhf: When better reward models don’t yield better language models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The accuracy paradox in rlhf: When better reward models don’t yield better language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.072863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:197cc92b91168c2fa8a120a3d035b50077adc33875371c96c731aa2383be8e6d

Observation b0803899-d69e-443e-b11f-e67df679bff4 · outbound

This paper cites Heuristic-guided reinforcement learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Heuristic-guided reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.859056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:22da285c1b811e0f0d0a119ae99bd840c9c29326938e7e3e2eeb26daa21f4434

Observation d4899a4c-28d0-444c-a2f5-3d47850b65af · outbound

This paper cites Learning navigation behaviors end-to-end with autorl.IEEE Robotics and Automation Letters, 4(2):2007–2014.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning navigation behaviors end-to-end with autorl.IEEE Robotics and Automation Letters, 4(2):2007–2014

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.013493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:5de133345255c4ff182e4c6cc42e4977cb8527566190a986801cf56b03eb96e5

Observation 2df8d76b-484d-47ff-a1e4-6730ad9ade43 · outbound

This paper cites More is less: inducing sparsity via overparameteriza- tion.Information and Inference: A Journal of the IMA, 12(3):1437–1460.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient More is less: inducing sparsity via overparameteriza- tion.Information and Inference: A Journal of the IMA, 12(3):1437–1460

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.175394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:1bb438b970ec6e479d77488c62e5262be61b27685cdcdde0c5ab3b28f8bbd42c

Observation 0bafe182-0e35-4c0b-bd22-7149147b4048 · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward model ensembles help mitigate overoptimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.068336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:3777ed2aba34aa6e8df5528d05b5b9e9db6efd9094bde915e459db7f1b4e9263

Observation 001ce548-e2a0-4a8a-b017-38e03f8407c4 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ultrafeedback: Boosting language models with high-quality feedback

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.141787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:95427b446d4421088604a78fa5ad1fb80d00f4edb24724c56c5e263a1665b385

Observation b6bba76d-fb51-4eb3-8356-1311f0257689 · outbound

This paper cites Maximum expected hitting cost of a markov decision process and informativeness of rewards.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Maximum expected hitting cost of a markov decision process and informativeness of rewards.Advances in Neural Information Processing Systems

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.885611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7a31637b10ed7f55e54743048c29c6b2414c93ba094eac89d51e53b8f83ad681

Observation 3195c83e-82e7-4a9b-983a-e246baccc8b9 · outbound

This paper cites Exploration-guided reward shaping for reinforcement learning under sparse rewards.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Exploration-guided reward shaping for reinforcement learning under sparse rewards.Advances in Neural Information Processing Systems

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.838961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:c8d57ce02cd7b56358c6824f61a0970e4eada0341044fc80103da6e714f20c28

Observation 61d4bf89-748f-4f6b-87dc-723058a0a7dd · outbound

This paper cites Continuous vs.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Continuous vs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.868010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:8b90e65898ae251b0b85f5075772af0e21d2f6c331253ddaee8343dc5f24a15b

Observation 77b8edde-0e2a-42b1-b3da-fffbe35e77f9 · outbound

This paper cites The perils of optimizing learned reward functions: Low training error does not guarantee low regret.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The perils of optimizing learned reward functions: Low training error does not guarantee low regret

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.895553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6832efa1cd96d168183d3086d849bc3a4131e61f8486faee522ee4df14a9cc78

Observation e2304c86-3d52-4ac1-9fc0-d76092cecfea · outbound

This paper cites Is a good foundation necessary for efficient reinforcement learning? the computational role of the base model in exploration.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is a good foundation necessary for efficient reinforcement learning? the computational role of the base model in exploration

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.786800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7d0e5c76ac8f8306b6efbed4e90e8de7ff79930909e76cd9103343bbe558765e

Observation d14e7f16-b218-42ac-aab0-be81f3b28c37 · outbound

This paper cites How to evaluate reward models for rlhf.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient How to evaluate reward models for rlhf

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.123615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:2faaf6c19247c9ce6f59d7e98bab0a09aa2fb075b23eaecb5b29f7f36c1d2467

Observation 39e03859-ef9c-4662-8d30-f9a6d8735a40 · outbound

This paper cites Scaling laws for reward model overoptimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Scaling laws for reward model overoptimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.805765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:15fd087f7c2a15d167ed6d5785c6af2424bd9dcaec30c96a91f46ffb4564c1ad

Observation bacdaebb-550e-4a6d-8693-a5602a9b3cf6 · outbound

This paper cites An alternate policy gradient estimator for softmax policies.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient An alternate policy gradient estimator for softmax policies

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.018960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:919446c580dbaea71748713a5a20cd0c6544f6b27a1fffe0ded5511f0f374031

Observation 00c18eea-a1cd-4f85-95c7-8248d629d8c7 · outbound

This paper cites Quantifying differences in reward functions.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Quantifying differences in reward functions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.828078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:c67aa6798e9cc1b461db27fcd60b01e71e717e5df0f77aa82a62a39bdb7fe0f6

Observation 5ff05225-ae09-41cd-b833-66e2010ba483 · outbound

This paper cites The Llama 3 Herd of Models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The Llama 3 Herd of Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.354997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:aa8a752189027387af1017069b389371c93cc734c4570f4af167edb13c0b8a72

Observation cd9ba846-5e88-440f-81d5-a62981433e19 · outbound

This paper cites Reward shaping in episodic reinforcement learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward shaping in episodic reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.109371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:978036b62a9ad5df08592e79557ade5347ee2cf69c9caba3ac684edbad63ad38

Observation 0cee9084-8afc-4ecf-b712-e6f53d50705b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.569181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b5b0dfba234a2504ac685c68aa2d2dce758344a5bf0288f83e2b356885d07542

Observation b2b53894-57f5-4717-aae4-3f45bd19d9af · outbound

This paper cites Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity.Advances in Neural Information Processing Systems

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.795897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:0009753f44039c3e62a974bb5e7c878f4cb72a28a09ec5a89495e4052133efa1

Observation e8bd2e1e-9a1d-4d8e-bf3e-e7939a5ca7fe · outbound

This paper cites Neural Replicator Dynamics.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Neural Replicator Dynamics

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.370655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:222b5bbb105427bbb16de87942125a65c4ad920267aa7b30eec6a8b6e388ef3b

Observation a9db3d6b-33f0-4820-b168-cb647a65ce88 · outbound

This paper cites Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.843538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:8f1ece416b79ad7abebe38ff26dd060ac75354d6d0a873a56994f86db305fa9f

Observation 2e768850-5d4d-4dc9-a3bf-9d122f133052 · outbound

This paper cites On the Emergence of Implicit Curriculum in RLVR Learning Dynamics.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.411893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:2a0552ce1e4f407d71fc513a48ac69e1e59855ce5695b0bb14e2f2797847637b

Observation f4292780-bbe9-47f2-8cb6-4447e8fb1132 · outbound

This paper cites Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Pitfalls of rule- and model-based verifiers–a case study on mathematical reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.516586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a5ee1c95e8992a2cd8ce482a285e165c0ae011fa27de834c942e366c26adae45

Observation 79356e28-dcfb-46ca-b1ba-7c2e125cd99e · outbound

This paper cites Goodhart’s law in reinforcement learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Goodhart’s law in reinforcement learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.119332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:bfbee86380624b947f12d51cb29841c3e1c791de8740ecb1980513bfac97a58f

Observation 9ee1562b-2bad-4ee7-994d-05965bcfd311 · outbound

This paper cites Beyond stationarity: Convergence analysis of stochastic softmax policy gradient methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Beyond stationarity: Convergence analysis of stochastic softmax policy gradient methods

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.890286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:255d4c07bd9a63ec9feeddda41ab8a578fa14344d2169aaa55b0e7d38b5d4553

Observation 7b1c7b91-3a83-4ce4-819c-d1c6b11ad14f · outbound

This paper cites Buy 4 reinforce samples, get a baseline for free!Deep Reinforcement Learning Meets Structured Prediction ICLR Workhsop.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Buy 4 reinforce samples, get a baseline for free!Deep Reinforcement Learning Meets Structured Prediction ICLR Workhsop

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.800342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d6ba5e74a15366eb3dfd85a238019b311a55d90bfecddc14de6fcb5a3e84d603

Observation c313ba31-c8b0-4040-8b8d-569a560a30e6 · outbound

This paper cites A neural collapse perspective on feature evolution in graph neural networks.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient A neural collapse perspective on feature evolution in graph neural networks.Advances in Neural Information Processing Systems

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.791469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a373daf58d554b4d9a5b1508b15d6423254cf530117009cacb780957dc15f362

Observation 9f74ff2c-b947-4f68-b9c1-2b78bd4fd77a · outbound

This paper cites Correlated proxies: A new definition and improved mitigation for reward hacking.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Correlated proxies: A new definition and improved mitigation for reward hacking

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.777821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e9b792df93b3f79186d643dbd0671ccc76d291406f97f595ed2e009896b6fb29

Observation b79c9aee-6548-486f-8dd0-87ec58fd2409 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.452354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:4a48528fb0e57d8ea6cb252cc3da9ee6ac57a9a8b3fe2bf4b616a32829fef5bb

Observation 6b3e9138-9638-4f18-9945-7bf498b74769 · outbound

This paper cites Rewardbench: Evaluating reward models for language modeling.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rewardbench: Evaluating reward models for language modeling

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.880935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:3f8fdbe140c2d94468a3617395177943cc8af216331c55f3576ba77ffe4153ba

Observation a9957a1d-b670-4804-bde9-34413deea7cb · outbound

This paper cites The influence of reward on the speed of reinforcement learning: An analysis of shaping.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The influence of reward on the speed of reinforcement learning: An analysis of shaping

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.810202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:212c7d54299b8e704a8779e1ac1f86e05aacd7dcff8c5cf7fe0027028c135805

Observation c5b2414b-01f3-4776-957a-58931e1a29a7 · outbound

This paper cites Softmax policy gradient methods can take exponential time to converge.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Softmax policy gradient methods can take exponential time to converge

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.160033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:704890dfa49b3cd2ec89f6f0c9bea06275fd5c133e565444ce8476fbcb50acf3

Observation 3acb3b2f-9de6-41fa-b43c-54d351479c5a · outbound

This paper cites Hashimoto.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Hashimoto

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.170643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:87167d3cfeb1b7e48cc0575d3908403211838cce6cacd9e38b9617589efef62f

Observation 054f8504-b2e6-4b03-a827-9ab741972c4c · outbound

This paper cites Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.689881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:d07ffdfb480f617de11e6689d5606c64cb97e9cb4928aad013614197558e4fb3

Observation 8e41e6ee-74ea-4de0-96cb-53143922b22e · outbound

This paper cites Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:17:43.404381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:0f9198b8c1b9048a82fecfd9a4e50387370de019e895934287d85e89b973077c

Observation 78f1a5c5-0921-403e-ae92-c241984b705b · outbound

This paper cites Elementary Analysis of Policy Gradient Methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Elementary Analysis of Policy Gradient Methods

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.581704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:16907347870fac2f5e230b4458ea5b9e0921f634d3dabc4008487e3d4e2e6e1b

Observation dcfc11c6-1ee1-4546-b85c-2d855129d2c7 · outbound

This paper cites RLTF: Reinforce- ment learning from unit test feedback.Transactions on Machine Learning Research.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient RLTF: Reinforce- ment learning from unit test feedback.Transactions on Machine Learning Research

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.150697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:73bd3f5b7fdd7e2f2cbbb92bd5899f686e75ca079c03a965b0d97308ae8b291c

Observation 6c70cdc9-b1a0-4538-86f7-feb0e02f2e9c · outbound

This paper cites Rm-bench: Benchmarking reward mod- els of language models with subtlety and style.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rm-bench: Benchmarking reward mod- els of language models with subtlety and style

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.050469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6663f60c5eb54242b403afcd484a82e5ebce9e77f672911408d944c0ded27f00

Observation d310456e-86a5-4dd7-a59d-3dad4320a502 · outbound

This paper cites RewardBench 2: Advancing Reward Model Evaluation.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient RewardBench 2: Advancing Reward Model Evaluation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.467386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9b700815810526184466ca317a9c795b2f5ef9e85a11307f96a40f24cae4a778

Observation da614b5e-a305-461e-a951-b4d3dc04a718 · outbound

This paper cites Reward engineering for reinforcement learning in software tasks.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward engineering for reinforcement learning in software tasks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.440658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e019833aa7003c0b1e3bda0506dfc64e14e883a135ac993219c43094c04c5e46

Observation 0c30e2b3-8a65-4ef6-aad2-cc8f10b05a47 · outbound

This paper cites Reward functions for accelerated learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward functions for accelerated learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.059742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:0d6af986d4e0298ce3445cd705d0f36dc5d1d4b32be74490adf92ec2a0f8825c

Observation c95685b4-ae34-43b6-b82a-ae9695cbb204 · outbound

This paper cites Escaping the gravitational pull of softmax.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Escaping the gravitational pull of softmax.Advances in Neural Information Processing Systems

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.999383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:4fca29dd079d749d1d7d2ccb19f8bbd4e9fcfab935b085e7f53d0d16ef4086de

Observation 3d0dd837-3d61-420f-bbb6-f7eb95c65c0d · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the global convergence rates of softmax policy gradient methods

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.146068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b6121e557bc899b0d0da2e5a5a81bd5bde5e49a33329191ac8081bab3971e19b

Observation 960fb550-3502-4e22-9e08-65740ce9b899 · outbound

This paper cites Leveraging non-uniformity in first-order non-convex optimization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Leveraging non-uniformity in first-order non-convex optimization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.782435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:788a4b12c29d2c1f612be8f92b16e009de29516201bdb0cf6171208b728280c9

Observation 248c8549-954e-416d-8029-11395b13baf0 · outbound

This paper cites Ordering-based conditions for global convergence of policy gradient methods.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ordering-based conditions for global convergence of policy gradient methods.Advances in Neural Information Processing Systems

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.872459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6e9ce490b296c1f2ef16bce5851c49d6248cfb089c092e1e29a7882d9b6e4281

Observation 33d3b4f0-5ba3-434f-ae45-e150ee40740d · outbound

This paper cites Stochastic gradient succeeds for bandits.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Stochastic gradient succeeds for bandits

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.823342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:fdc0a53e6812ed40150065d65590d3c0592ee0651ea9b7fcbcfd05f80ccfe312

Observation 987f6b8a-5dfd-40a6-9e92-7c999a5a1dde · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Policy invariance under reward transformations: Theory and application to reward shaping

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.032695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:2e5d54d1f73e792a0de2a6548f8b1dd6ff98018ad859b9c4cbd63ffd546f79b3

Observation e28f85ac-e154-4a89-a61c-b8fa25db28e3 · outbound

This paper cites 2 OLMo 2 Furious.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient 2 OLMo 2 Furious

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.431341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:168de18cc2402cad8b4580e5e9de7455170c0db73caec549292fc99f93e6a825

Observation a170f578-7634-4aa8-bc62-0b5865648f16 · outbound

This paper cites Olmo 3.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Olmo 3

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:41:19.524345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:bc2a40173c65f8a78ae6cc2819373262f1c97ffda2500955d4d687d3504a65ef

Observation 519a7f06-fdb2-4a40-8947-523801103c68 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.055485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e1b0270c8519a00d0aad72e9083dea7814ebc25f7b2a49d3a5f5a64cf28d89ff

Observation e0aef6b5-3da2-4eea-b445-f7fab22c60b0 · outbound

This paper cites Reward gaming in conditional text generation.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward gaming in conditional text generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.900348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6aaaa1bde3e87234af476e0e54e724a06b5c3db7d819b0e4231ac57e643dc5d2

Observation 30751c67-af25-4434-8dad-9db33a339f83 · outbound

This paper cites Automatic differentiation in pytorch.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Automatic differentiation in pytorch

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.082466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:78cf42e78c7ad1022c58e96713f12141e1f9ebbf40e97df12db82587999481db

Observation 57dbbfe6-b229-48be-a261-b000c3d43132 · outbound

This paper cites Amp: Adversarial motion priors for stylized physics-based character control.ACM Transactions on Graphics (ToG), 40(4):1–20.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Amp: Adversarial motion priors for stylized physics-based character control.ACM Transactions on Graphics (ToG), 40(4):1–20

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.155802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6e866a8abb39274a3548dd307a12292fb43785df94ad254725da94b5e76414bc

Observation dd783c71-682e-4090-993b-c31105667067 · outbound

This paper cites Generalizing verifiable instruction following.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Generalizing verifiable instruction following

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.037042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f3321738f7768bdf528273ccafcfc3d3c7c60a4270eb8d5d089a7031f877a5d9

Observation 097d2b6e-9754-4d08-8403-6e69c16330f1 · outbound

This paper cites Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:08.062014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e2be663b574c70de057013146831230233a716341b529bfa8e03ca66ba21a10f

Observation 1760c012-5233-4b15-9b96-c2cf885b0f2f · outbound

This paper cites Learning to drive a bicycle using reinforcement learning and shaping.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning to drive a bicycle using reinforcement learning and shaping

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.834343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:6345fec376d236dbb44f1e9026c9633e1eacc6e498a7bbec64b45bfcc9aaefbd

Observation ad7c4af3-c797-4a68-afb8-233e215bcba2 · outbound

This paper cites Implicit regularization in deep learning may not be explainable by norms.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in deep learning may not be explainable by norms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.994339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a6f69a07e8e8eccb762c12ac9ed251fea6a37716045035b64213c181f213b4b7

Observation 10457d44-eb20-4523-a1c2-b0ab66501224 · outbound

This paper cites Implicit regularization in tensor factorization.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in tensor factorization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.091425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a9b4d72b718b9de4572006d5f27540513b42a034462db69aae9b6e07ee2707f0

Observation efc34028-c14f-4e67-a47b-ed586a5b2417 · outbound

This paper cites Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.004683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:cee664f83453c60490b6623047dbd8fcf129517b8550de2da75712bd829ed1a9

Observation 471d2902-ec56-4d25-a6a4-00650a010efd · outbound

This paper cites Susskind, and Etai Littwin.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Susskind, and Etai Littwin

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.028147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:309584bac9c1f70f0102c4222b8e8c7e801681154498c823b0d56da0acdddd60

Observation 5b1a8b75-d679-4728-8d7b-6f986db2992e · outbound

This paper cites What makes a reward model a good teacher? an optimization perspective.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient What makes a reward model a good teacher? an optimization perspective

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.095922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:958f1574d636b2cecc9d64f319d29d30412c891304f3df5fd77cac2731df9391

Observation d585a040-940a-4d8f-be86-bd62a25510d7 · outbound

This paper cites Why is your language model a poor implicit reward model? InInternational Conference on Learning Representations.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Why is your language model a poor implicit reward model? InInternational Conference on Learning Representations

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.854294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:c7931e99b744ad783d78fa3c4f5866c0096ee8d4630595b4cdc1454454ed9c04

Observation cee29f65-6e2d-4f9f-aa72-c6d6a227f58c · outbound

This paper cites On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On the effective number of linear regions in shallow univariate relu networks: Convergence guarantees and implicit bias.Advances in Neural Information Processing Systems

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.773208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ae8c15d559322f1bf781da43e8a53469a67872fa650a0426f881a98fa1a2025d

Observation 4b9d81ba-92ac-41be-82b4-16e2a588389a · outbound

This paper cites Exact solutions to the nonlinear dynamics of learning in deep linear neural networks.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Exact solutions to the nonlinear dynamics of learning in deep linear neural networks

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.848707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ad1c0630e86559cf104ea75926b873beaad5398c4c15577b309166a398b824f0

Observation 903a3577-8159-47d5-953b-5ea29d577059 · outbound

This paper cites Ray Interference: a Source of Plateaus in Deep Reinforcement Learning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Ray Interference: a Source of Plateaus in Deep Reinforcement Learning

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.487655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:78e6e9c34cf38ab93cf0389ac826874de52a93484a0b1ada214e9e95e06c532b

Observation 16b258bc-2c08-45a3-bc04-c0cba41e2226 · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Proximal Policy Optimization Algorithms

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.387711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:94e16a9025a0136f7324f45e686f96aa31c498b3f4e9f73a143890d92baed776

Observation 4d69d36c-3d2d-4992-a032-4116c8f57c0b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.624857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:7dec29afcfd886a723bf5eb4ce62ba7764b95c42760d9a2547c80b4b6e7804f2

Observation 3c3ce965-d6d4-4970-8910-bca64027c993 · outbound

This paper cites Intrinsically motivated reinforce- ment learning: An evolutionary perspective.IEEE Transactions on Autonomous Mental Development, 2 (2):70–82.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Intrinsically motivated reinforce- ment learning: An evolutionary perspective.IEEE Transactions on Autonomous Mental Development, 2 (2):70–82

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.132804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f4939bade23961e642b664b8d3e429cf0b6cdc6e9057fc2459cd54f771d63474

Observation f18a4585-de3d-4add-b1dd-e5088a989381 · outbound

This paper cites Defining and characterizing reward gaming.Advances in Neural Information Processing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Defining and characterizing reward gaming.Advances in Neural Information Processing Systems

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.041345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a34944bc658f5b7fb146dde0a19898c68acf1be4c1d342fcb43018f97958f89f

Observation 4685644c-8a2c-47f1-a102-03505ab7b8ff · outbound

This paper cites Starc: A general framework for quantifying differences between reward functions.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Starc: A general framework for quantifying differences between reward functions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.045761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:e90954d851ff1538a47ad1a20e18cb9a21d6b40d38163ff658ce31ab12d15701

Observation ff578be5-2523-4960-adf2-91febc47327c · outbound

This paper cites The implicit bias of structured state space models can be poisoned with clean labels.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient The implicit bias of structured state space models can be poisoned with clean labels

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.814814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:06d2c86b4c5e05f306dc7f49d7024eb5018633d856e3121322a472f00367f1d2

Observation 8052fc45-4fd5-4da8-9a7f-298f85c40e55 · outbound

This paper cites Reward design via online gradient ascent.Advances in Neural Information Processing Systems, 23.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Reward design via online gradient ascent.Advances in Neural Information Processing Systems, 23

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.127911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9895448e0ce69aea6670c6e86d030cbacc1dfc88f5e2e7201143a74f9151adba

Observation 5ad401cd-72af-4fef-8bb7-e731b91238ce · outbound

This paper cites Learning to summarize with human feedback.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Learning to summarize with human feedback

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.876597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:4273ad73fd14ecd3f3343b75fda621c5e36299b2b5a7e866a419621d995c9d2b

Observation 2ae292cf-cc5a-4eb9-afaf-454659e56a41 · outbound

This paper cites Rl grokking recipe: How does rl unlock and transfer new algorithms in llms? InInternational Conference on Learning Representations.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rl grokking recipe: How does rl unlock and transfer new algorithms in llms? InInternational Conference on Learning Representations

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.105202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ad9e7c49a2cbd4a0dac15aa68cb159964c403eba746bc9962cecb959ea4192f2

Observation 33501b71-3632-4fad-8b33-1198ac9fb464 · outbound

This paper cites All roads lead to likelihood: The value of reinforcement learning in fine-tuning.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient All roads lead to likelihood: The value of reinforcement learning in fine-tuning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.063971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:8be2706c29b8fe4b7c59aedeff6110763cb486c5811d57c186c94a417beefa48

Observation 82c9f1d3-92ad-4d9d-ba59-3a1cf0f06f06 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Understanding the performance gap between online and offline alignment algorithms

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.645340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:f09e3992e98c0cab5abe187b3528c571857fc11c031f20c347474711e143ac0c

Observation 530bd498-9511-42bb-ac2e-50bc6c38d329 · outbound

This paper cites InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient InThe Thirty-ninth Annual Conference on Neural Information Process- ing Systems

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:41:19.591630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:abc114e6dfb69d64652878ffc20f49b3f238b4f9e5f10c5e7cd0de81667dbebd

Observation e9020eb9-ad60-4f07-bbc6-55043e5a9723 · outbound

This paper cites Perturbation analysis of neural collapse.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Perturbation analysis of neural collapse

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.137475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:774bce542745298c34b3b8db157377ee37448853105bb984d738e9b557e92f27

Observation f6f5f133-7298-4f34-8121-1faf2ea73373 · outbound

This paper cites Implicit regularization in relu networks with the square loss.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Implicit regularization in relu networks with the square loss

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:24.819104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:2eb2cca2f701aaf32cc9f178e3051f324de35fd6adbd649940d7ea320806c98c

Observation 78c93807-1306-4273-b3fb-328262e3caf1 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.698135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:9d2b7c10c16120090baf08eca6b3ba6cc8ae9025d0a5c3ddbc2e2814ab224ec5

Observation 97aadcd7-abbf-4da2-a2fc-00e035fbd907 · outbound

This paper cites Rethinking reward model evaluation: Are we barking up the wrong tree? InInternational Conference on Learning Representations.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rethinking reward model evaluation: Are we barking up the wrong tree? InInternational Conference on Learning Representations

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.114951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:bc331c5e3b89a6fd117fcc29c73c0c08cd456cad9d450d8050147c7e87e9ffbb

Observation a299a1bd-a796-466a-bc73-4d8fd1093a67 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8(3):229–256.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8(3):229–256

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.165478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a840c1704667370679cab172ac1dea91788bea3656d5ed6b5b0dd65a99aed0ce

Observation ee639010-6b2f-4c5c-81d1-f0e14cb02f1c · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.611570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:78f2b604e8e11d5b04802374a722eafd899e7a527d2f8e5a397a966da8b08b5d

Observation df978ab5-46a8-4420-8cba-df6b7d5ccc91 · outbound

This paper cites Kernel and rich regimes in overparametrized models.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Kernel and rich regimes in overparametrized models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.008874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:ac6ffc9c7fb12bb9201437ddfceecbe426dc68669729819d103f1f83d0e6021b

Observation 91cb503f-62a6-4488-aa6a-22baec79bc27 · outbound

This paper cites Dynamics-aware comparison of learned reward functions.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Dynamics-aware comparison of learned reward functions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.077423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:b4c4fed9bffcc2b2887cb72c8a2ce68403c06fb023c457c14bfa714021615271

Observation cf047b6a-c400-4d90-8d3d-0e3bfbe6ea2d · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:41:19.562482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:880f961822a18a17f75d95de8bc5293192ef8a66788f08986936934b39ff001d

Observation f3fcb075-38e6-4eaf-b9b1-ee1850569394 · outbound

This paper cites Qwen3 Technical Report.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Qwen3 Technical Report

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.605231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:2868d3612f4bdfb80912d31af43ad75723e516c5c0df3b61c221b167485ea9ca

Observation 59af3c34-be70-43a5-9e92-5b9434de1e4d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.708864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:c4a9a0dbc6119689273ee66ce0f8681e6b330193dee3f275b539f08b91fa3a44

Observation f7c9d62b-ed80-410c-8034-fa70ad788a31 · outbound

This paper cites On learning intrinsic rewards for policy gradient methods.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient On learning intrinsic rewards for policy gradient methods

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.024062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:10d776dd06d1a809055fc053913821aebd55b16b6fb8027c77fc6fdab683df92

Observation 75db98d7-8b6b-4d13-8254-9d0ce84bccde · outbound

This paper cites Rmb: Comprehensively benchmarking reward models in llm alignment.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Rmb: Comprehensively benchmarking reward models in llm alignment

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T02:18:25.087015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:108729e37b2d86bc91dfa3b7e98479c546bbdc924ec46aded16c9964c4c32131

Observation 86b2e17f-577a-4982-8958-8f23b466519c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Fine-Tuning Language Models from Human Preferences

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:41:19.598253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:a0aa65b3312da5493d2bdab44718c6b74ea4f72e071db458c29d01453c4a461d

Observation ae83b582-bad8-46ba-b797-a75062ed3730 · outbound

This paper cites partial rewards.

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient partial rewards

Reference 100

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T23:41:19.586625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T16:24:43.688967Z digest=sha256:57daebabc20b34570333a8fee9f304cd72bc3130c81126ad11551d19a4554f7d

Pith citing papers

Observation 28af7302-b8f7-4f48-885d-77815d82bb5e · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:45a619e147edd1d8ad2420d7aa0b1d568b5b4c434c86f93c479539c30370b4c8