Pith. sign in

Paper Citation Record · LEDGER

Reward Models in Deep Reinforcement Learning: A Survey

As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2506.15421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15421 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:38:51.086103Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:08:19.797150Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:50.941267Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49a2ba99-8f48-418d-a2b3-2d472e72612c · outbound

This paper cites Vision-Language Models as a Source of Rewards.

Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models as a Source of Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.445678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.445678Z digest=sha256:e317ed1e5cbaa90913a30c8d3efedbf02058b4ebd3458515ba21bd64d9edaaa5

Observation 4b81b121-5b55-402d-b1e7-7a72d76bf2d1 · outbound

This paper cites Diversity is All You Need: Learning Skills without a Reward Function.

Reward Models in Deep Reinforcement Learning: A Survey Diversity is All You Need: Learning Skills without a Reward Function

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.461416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.461416Z digest=sha256:1f8d6db8d0e467762e4abc0b5d398864d7dc8160e15c5044adc3595f5f75154d

Observation f002d8c3-1e2c-423d-8e3e-a11ad8568f40 · outbound

This paper cites Quantifying Differences in Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey Quantifying Differences in Reward Functions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.478051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.478051Z digest=sha256:e8299d2cd4cdf4df62bfc565153f4fe5f8658405eb703d283814582e0d6b0c6f

Observation 0bff0728-f114-4b9d-8edc-c33ad3690e8d · outbound

This paper cites Preprocessing Reward Functions for Interpretability.

Reward Models in Deep Reinforcement Learning: A Survey Preprocessing Reward Functions for Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.544516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.544516Z digest=sha256:c7a0212ac5ddc14b7393827d9d08ad96f101ea7d5da76ef98bac191c8d4f37d0

Observation f4a188ef-28ba-45b9-b6d5-9f3237367050 · outbound

This paper cites Regularized Inverse Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Regularized Inverse Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.578945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.578945Z digest=sha256:1c4460886d011c76d8af722a2b533615e64e566154ec3ee2b73dae56a1c4c3b1

Observation 9270cc22-8d05-44ab-bba5-0b97572f10d2 · outbound

This paper cites A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,.

Reward Models in Deep Reinforcement Learning: A Survey A survey of reinforcement learning from human feedback.arXiv preprint arXiv:2312.14925, 10,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.627819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.627819Z digest=sha256:35e79c13b9bfce7ccb88f5fcd8fad9b01c3baa98a0effcddb4f9148188dc32a9

Observation d6570f42-3b80-4095-b1cc-01223c13b89b · outbound

This paper cites Preference Transformer: Modeling Human Preferences using Transformers for RL.

Reward Models in Deep Reinforcement Learning: A Survey Preference Transformer: Modeling Human Preferences using Transformers for RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.639257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.639257Z digest=sha256:d827a7ab3b9419ea0e53d440e2fd67870f08ce8662d01066a29f2eef33beb674

Observation dc08f9b9-f160-4338-a95e-482f090b333c · outbound

This paper cites Empowerment: A universal agent-centric measure of control.

Reward Models in Deep Reinforcement Learning: A Survey Empowerment: A universal agent-centric measure of control

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:38:52.185603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.650866Z digest=sha256:c2d6594d57f329873b5c00a21910e4030da63c7e212ca3de2bdd3930778734fe

Observation fb6b9eb2-7d20-4713-b7cb-c33da20962a7 · outbound

This paper cites Goal-Conditioned Reinforcement Learning: Problems and Solutions.

Reward Models in Deep Reinforcement Learning: A Survey Goal-Conditioned Reinforcement Learning: Problems and Solutions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.662293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.662293Z digest=sha256:547838ab97efe8b7f1e581d5e56f48ac8c74bac15a72a4651b41d185a034634c

Observation 5144f0ef-b109-4920-a17b-a2d97aab60bc · outbound

This paper cites Reward Modeling with Ordinal Feedback: Wisdom of the Crowd.

Reward Models in Deep Reinforcement Learning: A Survey Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.656840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.668075Z digest=sha256:dbf01c38e9e914d103a29b54bfab7a7b7b024664949e95dcd87e55f2320bb03b

Observation a6529e8c-1975-4c4e-86d9-64d2cd39f9a1 · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Reward Models in Deep Reinforcement Learning: A Survey ReFT: Reasoning with Reinforced Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.679115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.679115Z digest=sha256:172e966096e34469760cebb0a7a81ce88a2a6c4f1b3665fd8a3b336b753ee4c4

Observation 4795901f-7eab-45bf-ab6e-6c5ea4aec581 · outbound

This paper cites Choreographer: Learning and Adapting Skills in Imagination.

Reward Models in Deep Reinforcement Learning: A Survey Choreographer: Learning and Adapting Skills in Imagination

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.698388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.698388Z digest=sha256:c7b57dcb6b67215427a71c7d8a3667bbb6aa7ab73cef8f6487a91c2afd03efb8

Observation 61063ea3-cb44-4b2e-9e8c-769bdd690a8a · outbound

This paper cites Learning to Assist Humans without Inferring Rewards.

Reward Models in Deep Reinforcement Learning: A Survey Learning to Assist Humans without Inferring Rewards

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.792909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.792909Z digest=sha256:7b6c698244afa8daf02160e5a7e844ffb0617e7ec8ce3f0c243d4d4ce2f0f313

Observation 598ec36a-b1d1-4312-86f5-6eeb6144bcfb · outbound

This paper cites METRA: Scalable Unsupervised RL with Metric-Aware Abstraction.

Reward Models in Deep Reinforcement Learning: A Survey METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.841430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.841430Z digest=sha256:188cce9a594db4a6528115448f73211d076f207500b563bfbdec1382e54c6cbc

Observation b9c43ce1-7ec8-434d-99cf-8d878f4a37cb · outbound

This paper cites Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.847290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.847290Z digest=sha256:ba6546c0b09cabdaca08b553b842f67186024d914b121a2216bb4708cb8c4bd5

Observation 26e7ff2b-2682-45f3-8248-ac6aaa0779a6 · outbound

This paper cites STARC: A General Framework For Quantifying Differences Between Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey STARC: A General Framework For Quantifying Differences Between Reward Functions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.852504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.852504Z digest=sha256:9a738736013b2244b86c57ed104cd9035cac1cee5d08b8744eb6d0427e53230d

Observation fd93ee8b-7900-4168-838c-4bfeb2a8f248 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reward Models in Deep Reinforcement Learning: A Survey Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.858219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.858219Z digest=sha256:457275a87ae77c6c1742697ef01cf0f4c023dd12ea79b064608e432ed1a4ce4a

Observation b80eeca1-34af-449e-9d5f-de88fe30a17a · outbound

This paper cites Gymnasium: A Standard Interface for Reinforcement Learning Environments.

Reward Models in Deep Reinforcement Learning: A Survey Gymnasium: A Standard Interface for Reinforcement Learning Environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.863987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.863987Z digest=sha256:5fc82edf471f14af684501f5acac72d27d3fd0b026854eee3f2cb16c3493e6bb

Observation 199389b0-881b-4c00-9e30-6e75e105ec4a · outbound

This paper cites Hindsight PRIORs for Reward Learning from Human Preferences.

Reward Models in Deep Reinforcement Learning: A Survey Hindsight PRIORs for Reward Learning from Human Preferences

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.407764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.869505Z digest=sha256:f11dff684176cbd738fd5ea76d49aaec075c895bc733f62a4c6e74be9a3c12cb

Observation bea44e5a-ebf5-4d1a-a72c-bffd2ba218a8 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.030036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.030036Z digest=sha256:755295035c49c15dfb56b612a591289915274f4d27e7e80d5dcee79072b4a33d

Observation 81ce2191-54d4-409a-8c40-53578b8eaa83 · outbound

This paper cites DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search.

Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.075704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.075704Z digest=sha256:b64804a77f0aac7c23e2e5f467a303898d4e5b2d37ab50d3cbe871901d3a74d8

Observation 9dbf7c04-eef8-4c1e-83be-43a1cee1455f · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Reward Models in Deep Reinforcement Learning: A Survey Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:51.080835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:51.080835Z digest=sha256:92931f21dfb94ef19e13a9f7d14090dafc94ac65c95800b897d3970607a0ffb4

Observation 82014311-ddbe-4e41-8db7-7b212b576b57 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Reward Models in Deep Reinforcement Learning: A Survey Maximum entropy inverse reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:38:52.167651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:51.086103Z digest=sha256:f1ee503570656b90af4de2e02be8cd02a1d871dac55ac7005c504fd16ae0bbc2

Observation b88bbc5f-199d-4b32-961a-8cb31b19f2c7 · outbound

This paper cites Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery.

Reward Models in Deep Reinforcement Learning: A Survey Dynamical Distance Learning for Semi-Supervised and Unsupervised Skill Discovery

Reference 1950

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.489300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.489300Z digest=sha256:8b27a43b44726794faf689d0895adc5076dd8f4d8b270d90621da94d530dbcec

Observation e4c7d4f9-2d29-479c-9d14-f6b860535c79 · outbound

This paper cites Exploration by Random Network Distillation.

Reward Models in Deep Reinforcement Learning: A Survey Exploration by Random Network Distillation

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.450710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.450710Z digest=sha256:a41e3172dc2f2028221d68b43ece42e6d5f67ded58d35565434057c2afbb2961

Observation ebb72092-549a-4dec-84b7-b23e74c8330b · outbound

This paper cites Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Improving Multimodal Interactive Agents with Reinforcement Learning from Human Feedback

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.410219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.410219Z digest=sha256:281b1aceec00f2cd60d7171fab3ff883d9dfd4f730b928eaea71e66f6c0d517c

Observation baff5cc1-6d0e-4bf8-adc7-3bcd9acfebef · outbound

This paper cites Models of human preference for learning reward functions.

Reward Models in Deep Reinforcement Learning: A Survey Models of human preference for learning reward functions

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.656340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.656340Z digest=sha256:d9db51eebfdea2577c021baa4bde3490014575a980602a80e1729878caaeb268

Observation d8b3e62f-02a8-49eb-b66c-a8dc62d1b0ee · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.483634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.483634Z digest=sha256:aa09284e54a43d7e951ec8d9a24522616ef4705cb5ad9fa4cbb66b23b920affc

Observation 875992f6-9b86-4dba-a3c3-852304b4c120 · outbound

This paper cites Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.

Reward Models in Deep Reinforcement Learning: A Survey Learning Robust Rewards with Adversarial Inverse Reinforcement Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.473011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.473011Z digest=sha256:4c59adf3643f77105634520c7e795138651ec0ae73a4b70296fb9d077f9faac1

Observation d5b5aa4f-bf1a-44fe-8b45-54f78a626e19 · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.456084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.456084Z digest=sha256:87a5b91498e1afe82c158a1028059de73ac7a7eb11983e756119b33b5fa78efa

Observation 53580f90-e529-44f7-a411-577256bbb502 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Constitutional AI: Harmlessness from AI Feedback

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.439486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.439486Z digest=sha256:4e580d426770c60a3c3b54d239a30d2f73c139d460daf3cca2bfc6539d1bfd98

Observation 64e3677c-cd60-4840-8c23-e9dacf260f04 · outbound

This paper cites Never Give Up: Learning Directed Exploration Strategies.

Reward Models in Deep Reinforcement Learning: A Survey Never Give Up: Learning Directed Exploration Strategies

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.433306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.433306Z digest=sha256:85a855d2e9be352c8361b1cf7617439bdfa515dee98821140ec15242c37402a4

Observation 1b420d97-7414-4ab5-a1da-b8a05c7f42bc · outbound

This paper cites A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models.

Reward Models in Deep Reinforcement Learning: A Survey A Connection between Generative Adversarial Networks, Inverse Reinforcement Learning, and Energy-Based Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.467651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.467651Z digest=sha256:1feee0dcf641ee2e2be2365f76e4af64f9eb95692dd7c90bffa38a46d22f8e11

Observation 1be62d4a-6672-4734-b03a-985e95e42ca8 · outbound

This paper cites Motif: Intrinsic Motivation from Artificial Intelligence Feedback.

Reward Models in Deep Reinforcement Learning: A Survey Motif: Intrinsic Motivation from Artificial Intelligence Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.645727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.645727Z digest=sha256:340a755e01b1e7971f61b35b350a2c4260d12f99f4ee739bf394d8e68905e068

Observation 2fe16626-45d2-4e09-a520-a493d26b41fe · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Reward Models in Deep Reinforcement Learning: A Survey LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.674235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.674235Z digest=sha256:cb8902fd8c3bc91c743fdc281d3babc3deadab746325a0fbe5b660e3836fc9b4

Observation 5f0cf442-4de2-42dd-a4a1-b515064d66f3 · outbound

This paper cites Dynamics-Aware Comparison of Learned Reward Functions.

Reward Models in Deep Reinforcement Learning: A Survey Dynamics-Aware Comparison of Learned Reward Functions

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:38:51.286313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:38:50.959942Z digest=sha256:135d045e5f0cab4efe033a6265236349db2b206e1ac6845f2dfd3791706aaa2b

Pith citing papers

Observation 9c30b6bc-215c-4a92-8b91-d9a93b451a39 · inbound

Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment cites this paper.

Multi-objective Reinforcement Learning With Augmented States Requires Rewards After Deployment Reward Models in Deep Reinforcement Learning: A Survey

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:41.945064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T09:38:05.280629Z digest=sha256:15d169d91ad733cf829711b926b22efb67ef127f9ec62e18b96d42a207ca6c89

Observation e53582ca-beba-4ed5-90ca-36d1efeb6910 · inbound

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning cites this paper.

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning Reward Models in Deep Reinforcement Learning: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:24:47.223690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T00:19:48.053466Z digest=sha256:37a31b12c72a67b2e75d6917ba1d52b2ce4696d7a185f48182b64b43412140a0

Observation 79cea2e5-be26-43e8-9720-51304d27cddc · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:47:53.425028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T19:45:26.458859Z digest=sha256:b03be5d068c57ff8ba5a5122fb96df3806b7b9399441eb57730235b2c121a141

Observation ee58a84b-58e0-4b21-80ad-9659fcf303f4 · inbound

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models cites this paper.

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models Reward Models in Deep Reinforcement Learning: A Survey

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:05:06.483419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T06:04:17.574622Z digest=sha256:c29fd6aed9f1a0588100ed2e9d98976bedc6a12ef7efe539f553bae7ae6c22f9

Observation 6c8a3fc2-9855-4232-b2f2-9bf323ba31e0 · inbound

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning cites this paper.

AudioProcessBench: Benchmark for Identifying Process Errors in Audio-Grounded Reasoning Reward Models in Deep Reinforcement Learning: A Survey

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:37:27.332423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T18:08:19.797150Z digest=sha256:a4858058f8396df107cd9e1f271838d2152e6aa839d9e56a97f6c0416a08a895

Observation ad523e39-1735-4f45-8802-1b24dff731b1 · inbound

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation cites this paper.

PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation Reward Models in Deep Reinforcement Learning: A Survey

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:50.942774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T05:14:14.053344Z digest=sha256:8f35cb7a924d8f44b85cd8af76e368aa212bdc4394cc7a3c4ca0b00edc80b443