Pith. sign in

Paper Citation Record · LEDGER

Reward Models in Deep Reinforcement Learning: A Survey

As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2506.15421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15421 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:38:51.086103Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:08:19.797150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:50.941267Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49a2ba99-8f48-418d-a2b3-2d472e72612c · outbound

This paper cites Vision-Language Models as a Source of Rewards.

Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models as a Source of Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.445678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.445678Z digest=sha256:44eaac617e8a6b6f70d0eb102c1f8a950daabc2597bb310e19b9bd8c84fb25f1

Observation 4b81b121-5b55-402d-b1e7-7a72d76bf2d1 · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

Reward Models in Deep Reinforcement Learning: A Survey Diversity is All You Need: Learning Skills without a Reward Function

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.461416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.461416Z digest=sha256:c769bf4700e4428aa38a7a639763190d10e3f9e32c1e1e7e9356246050a06f3a

Observation f002d8c3-1e2c-423d-8e3e-a11ad8568f40 · outbound

This paper cites Quantifying Differences in Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey Quantifying Differences in Reward Functions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.478051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.478051Z digest=sha256:b936747e5c9a3a6542fa28829e4ea4bdd5b90b64e887717c277ac36ac1a8a09f

Observation 0bff0728-f114-4b9d-8edc-c33ad3690e8d · outbound

This paper cites Preprocessing Reward Functions for Interpretability.

Reward Models in Deep Reinforcement Learning: A Survey Preprocessing Reward Functions for Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.544516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.544516Z digest=sha256:a8e954db73c12de0f85cb772b8e7266041a40facf4da02ed2cca613b6b87ff5b

Observation f4a188ef-28ba-45b9-b6d5-9f3237367050 · outbound

This paper cites Regularized Inverse Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Regularized Inverse Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.578945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.578945Z digest=sha256:baff6c608a72c0e71027fe38f7ad25e96e639d1c98e9d56a9348e667d67e59f0

Observation 9270cc22-8d05-44ab-bba5-0b97572f10d2 · outbound

This paper cites A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,.

Reward Models in Deep Reinforcement Learning: A Survey A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.627819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.627819Z digest=sha256:b596b72d8fc0326914f753fbb1f85259caa3b9d2b7098f0b8bb1d5c527723dd0

Observation d6570f42-3b80-4095-b1cc-01223c13b89b · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Reward Models in Deep Reinforcement Learning: A Survey Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.639257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.639257Z digest=sha256:bbe15f0302b96ec92be9aecf841c2590e7f8683099e240d3d249d9c827b692d3

Observation dc08f9b9-f160-4338-a95e-482f090b333c · outbound

This paper cites Empowerment: A universal agent-centric measure of control.

Reward Models in Deep Reinforcement Learning: A Survey Empowerment: A universal agent-centric measure of control

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:38:52.185603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.650866Z digest=sha256:c47ee7882126d3100d064f7569ef903852ba635f9908ced0410cb30af256ebc0

Observation fb6b9eb2-7d20-4713-b7cb-c33da20962a7 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Reward Models in Deep Reinforcement Learning: A Survey Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.662293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.662293Z digest=sha256:5f5187d8ca26366f85e45fa5c53543ac058799ea0b478de3710e309a4be1499a

Observation 5144f0ef-b109-4920-a17b-a2d97aab60bc · outbound

This paper cites Reward Modeling with Ordinal Feedback: Wisdom of the Crowd.

Reward Models in Deep Reinforcement Learning: A Survey Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.656840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.668075Z digest=sha256:5a18cb8434f718aec50f4b7c36a8d0830e86ac030de10e0cbc0308876e9e98f5

Observation a6529e8c-1975-4c4e-86d9-64d2cd39f9a1 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Reward Models in Deep Reinforcement Learning: A Survey ReFT: Reasoning with Reinforced Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.679115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.679115Z digest=sha256:98476f9cf6b781b1a784aa19ab117198363ee2c367fb83fc1b89c9b414300e5c

Observation 4795901f-7eab-45bf-ab6e-6c5ea4aec581 · outbound

This paper cites Choreographer: Learning and Adapting Skills in Imagination.

Reward Models in Deep Reinforcement Learning: A Survey Choreographer: Learning and Adapting Skills in Imagination

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.698388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.698388Z digest=sha256:7adb82c1831f17d2d9070b26b8510116c0eaaa855bb3926c5ae294814a411d89

Observation 61063ea3-cb44-4b2e-9e8c-769bdd690a8a · outbound

This paper cites Learning to Assist Humans without Inferring Rewards.

Reward Models in Deep Reinforcement Learning: A Survey Learning to Assist Humans without Inferring Rewards

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.792909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.792909Z digest=sha256:7090859165d29807b6cc19e417b678636e399b868ac719dca5f3c92f245ae865

Observation 598ec36a-b1d1-4312-86f5-6eeb6144bcfb · outbound

This paper cites METRA: Scalable Unsupervised RL with Metric-Aware Abstraction.

Reward Models in Deep Reinforcement Learning: A Survey METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.841430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.841430Z digest=sha256:de1539e1ae509f72c2c910ea3039fd4da8361fe00c47e61b275df2f5159cd3ef

Observation b9c43ce1-7ec8-434d-99cf-8d878f4a37cb · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.847290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.847290Z digest=sha256:c92abb785bda9fa489f086d9967c49ec296f2bf7f93173d7de9eef0421a3d123

Observation 26e7ff2b-2682-45f3-8248-ac6aaa0779a6 · outbound

This paper cites STARC: A General Framework For Quantifying Differences Between Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey STARC: A General Framework For Quantifying Differences Between Reward Functions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.852504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.852504Z digest=sha256:7efaf424e7fedeb3813f13484e38b911dd4c8f54de09f2ccb75613cc3e24e2f3

Observation fd93ee8b-7900-4168-838c-4bfeb2a8f248 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reward Models in Deep Reinforcement Learning: A Survey Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.858219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.858219Z digest=sha256:dbe0b364ba061e4d26e84cbd4117283488ee03f51d51100479dabe7025ec8c2d

Observation b80eeca1-34af-449e-9d5f-de88fe30a17a · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Reward Models in Deep Reinforcement Learning: A Survey Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.863987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.863987Z digest=sha256:11f2834c3fea0df334d50ab5fc9720ac8d2846e3279ecc26fc1bcaa7ee522b2e

Observation 199389b0-881b-4c00-9e30-6e75e105ec4a · outbound

This paper cites Hindsight PRIORs for Reward Learning from Human Preferences.

Reward Models in Deep Reinforcement Learning: A Survey Hindsight PRIORs for Reward Learning from Human Preferences

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.407764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.869505Z digest=sha256:edff94dbfc1983b72b7b3f02ed8a31fb01d496d482ccb15bac49433f40fe3471

Observation bea44e5a-ebf5-4d1a-a72c-bffd2ba218a8 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.030036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.030036Z digest=sha256:3735774ca20d8e38033bc01e254ea67aee90d7374d9b6bf4b05237b27c1d7787

Observation 81ce2191-54d4-409a-8c40-53578b8eaa83 · outbound

This paper cites DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.

Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.075704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.075704Z digest=sha256:bc4e6ecb14787b87063a399263f774abe3dcbf5862c2f1e5da9f6c5119ec2784

Observation 9dbf7c04-eef8-4c1e-83be-43a1cee1455f · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Reward Models in Deep Reinforcement Learning: A Survey Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.080835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.080835Z digest=sha256:fce5bc2003d8510360c3029b5507ac02740598030b9b3b48b2e7d258274444ad

Observation 82014311-ddbe-4e41-8db7-7b212b576b57 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Reward Models in Deep Reinforcement Learning: A Survey Maximum entropy inverse reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:38:52.167651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:51.086103Z digest=sha256:c17b3c886063548d2e73893fe5c16201cb0c9887cd1f1bc7b05561fbc7af0d53

Observation b88bbc5f-199d-4b32-961a-8cb31b19f2c7 · outbound

This paper cites Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery.

Reward Models in Deep Reinforcement Learning: A Survey Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery

Reference 1950

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.489300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.489300Z digest=sha256:244ed455953888bbebbdc91a4a1472ff2aa13e9c60e5e9abbb9d0353831755b5

Observation e4c7d4f9-2d29-479c-9d14-f6b860535c79 · outbound

This paper cites Exploration by Random Network Distillation.

Reward Models in Deep Reinforcement Learning: A Survey Exploration by Random Network Distillation

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.450710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.450710Z digest=sha256:e5652f80718593a425af4b0a92287c37e918f71e2720e3645e0ec6d398e79442

Observation ebb72092-549a-4dec-84b7-b23e74c8330b · outbound

This paper cites Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.410219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.410219Z digest=sha256:53f99b35198af59ba3f535c38bb4fcc7aafe09a36f44f0d30449cf85f2ca7576

Observation baff5cc1-6d0e-4bf8-adc7-3bcd9acfebef · outbound

This paper cites Models of human preference for learning reward functions.

Reward Models in Deep Reinforcement Learning: A Survey Models of human preference for learning reward functions

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.656340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.656340Z digest=sha256:b74ad30a4d4d6a0769715f3644560853f807f71ca0dd454f42112145f5143003

Observation d8b3e62f-02a8-49eb-b66c-a8dc62d1b0ee · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.483634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.483634Z digest=sha256:04df13f23de2bd8e48fe443126d0c95acc89e04ed5ad91c876f01dbeea239344

Observation 875992f6-9b86-4dba-a3c3-852304b4c120 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.473011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.473011Z digest=sha256:c1f0c80de541e63cd2760501612b7c3483c0f36707b7a5ae205399157398c1c8

Observation d5b5aa4f-bf1a-44fe-8b45-54f78a626e19 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.456084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.456084Z digest=sha256:3bd9cc2d4a8e6e906839027f9f8b1f175dd9cd93d1937620f778c78472fc9627

Observation 53580f90-e529-44f7-a411-577256bbb502 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Constitutional AI: Harmlessness from AI Feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.439486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.439486Z digest=sha256:1b696932ef89ce4953ab167224fa34347ee23a17ffd9c9bec75be8ca758c8bfc

Observation 64e3677c-cd60-4840-8c23-e9dacf260f04 · outbound

This paper cites Never Give Up: Learning Directed Exploration Strategies.

Reward Models in Deep Reinforcement Learning: A Survey Never Give Up: Learning Directed Exploration Strategies

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.433306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.433306Z digest=sha256:adfcc1c3b2704f52fa5bea50b04b714fc96cd16e189f8bc08684106db8fa1617

Observation 1b420d97-7414-4ab5-a1da-b8a05c7f42bc · outbound

This paper cites A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models.

Reward Models in Deep Reinforcement Learning: A Survey A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.467651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.467651Z digest=sha256:bb11e9af96a1a3c645eb687568a84056280839ced9eb3c2575ff16b8db6b93cb

Observation 1be62d4a-6672-4734-b03a-985e95e42ca8 · outbound

This paper cites Motif: Intrinsic Motivation from Artificial Intelligence Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Motif: Intrinsic Motivation from Artificial Intelligence Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.645727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.645727Z digest=sha256:3703ac7a32a4c46ed00f8a11bc0fd5298b6acbb12b57c13f5e2a2b3beb725f9f

Observation 2fe16626-45d2-4e09-a520-a493d26b41fe · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Reward Models in Deep Reinforcement Learning: A Survey LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.674235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.674235Z digest=sha256:c149f91fc328811720fd4f10135d0d646488e9f2a501008269f4948a580f49aa

Observation 5f0cf442-4de2-42dd-a4a1-b515064d66f3 · outbound

This paper cites Dynamics-Aware Comparison of Learned Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey Dynamics-Aware Comparison of Learned Reward Functions

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.286313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.959942Z digest=sha256:52c42b7e4d7fb467f21520636f7340a90bd6b7e8486b1f35f90794c5574edde9

Pith citing papers

Observation 9c30b6bc-215c-4a92-8b91-d9a93b451a39 · inbound

Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment cites this paper.

Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment Reward Models in Deep Reinforcement Learning: A Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:41.945064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T09:38:05.280629Z digest=sha256:b6f5946c1411b6e9c52cc292b28b56662187a804f51f02335122a9d4d0d5d149

Observation e53582ca-beba-4ed5-90ca-36d1efeb6910 · inbound

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Reward Models in Deep Reinforcement Learning: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.223690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:83c3154aa6aa5e9249064fd03499adaab3e886089d8fadca80ae6fcb4cf80d17

Observation 79cea2e5-be26-43e8-9720-51304d27cddc · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:47:53.425028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T19:45:26.458859Z digest=sha256:9f37598b387fc89bda9d5d74e1060fe0fcef94542a382ca5a4683b2e3bfee390

Observation ee58a84b-58e0-4b21-80ad-9659fcf303f4 · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.483419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T06:04:17.574622Z digest=sha256:4f3e1f2bae62168021b0b94a6480f1fed7fa1e38df56321995ecd46179837b61

Observation 6c8a3fc2-9855-4232-b2f2-9bf323ba31e0 · inbound

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning cites this paper.

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning Reward Models in Deep Reinforcement Learning: A Survey

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:37:27.332423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T18:08:19.797150Z digest=sha256:816d993ae333e82d61165c6902436edc6c76d23ea9448e8fffdfbb337933bf82

Observation ad523e39-1735-4f45-8802-1b24dff731b1 · inbound

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation cites this paper.

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation Reward Models in Deep Reinforcement Learning: A Survey

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:50.942774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T05:14:14.053344Z digest=sha256:f3c732fffdd36e0f943f900346f33b89d7b9bf50341c223cee3dfc894b351712