Pith. sign in

Paper Citation Record · LEDGER

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

As of 21 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2506.02553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02553 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:32:31.776781Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:07:30.092931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.514103Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1610d62-8496-4528-b307-6534336b9c74 · outbound

This paper cites GPT-4o System Card.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective GPT-4o System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.119855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.119855Z digest=sha256:378232681f266d68c8d1c363e23ad15874b65a56e84027f0c5462a14e6ae8697

Observation 073987cd-41cf-4941-a263-ca7b0bcbdc3c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.178328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.178328Z digest=sha256:bd96628573a35824d3a2bab64abb80b4c920b3dc070dfcb805fca5e05b63ad46

Observation d847781b-9233-4555-8232-c1e126abbad6 · outbound

This paper cites Qwen2.5 Technical Report.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Qwen2.5 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.272952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.272952Z digest=sha256:0a76fe503ebb75034a66b746b83c8b8e621ff98ca50b7cb053e7c756f3ffcb69

Observation fac76498-5a66-438f-8048-394162622066 · outbound

This paper cites The amazon nova family of models: Technical report and model card.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective The amazon nova family of models: Technical report and model card

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:33.438977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:28.391741Z digest=sha256:7f17ced61b7306f0cff0379bbc9d1a4294a4e768bb8407d7757f417afbbf199e

Observation 8cfe9e74-aa5e-4339-a953-161932231c2a · outbound

This paper cites OpenAI o1 System Card.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective OpenAI o1 System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.487274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.487274Z digest=sha256:24a805b1af4caf8a83164bb9f358b0265f7c195be4d3c489348c198f7ad428a1

Observation 838fd653-fbef-4b82-b5c1-7cc2b920e801 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.616012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.616012Z digest=sha256:219a2fc444994e224c347cb7dfa995fe92c3d78a39e39573cf57c33071ec6262

Observation 206e65fa-4c05-4e23-83ff-e6b70f85acf8 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Proximal Policy Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.675759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.675759Z digest=sha256:3d09bfa035c55a6608a1abae6f2ba5677c3eae04cb1930a770535d42e6c0f8ef

Observation 7a22d8ee-ed50-44dc-9392-7f7f9866d0bb · outbound

This paper cites Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021, 2020.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Learning to summarize with human feedback.Advances in neural information processing systems, 33:3008–3021, 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.759565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.759565Z digest=sha256:f3c6268ec6be43aed24d1c3144238fd411fe6d1360ce57e8ab62ecb2310da085

Observation db41d348-05c2-4c17-b586-f1c328306197 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.849369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.849369Z digest=sha256:096fc6862ac07a8e7b41728ff9f9bfcf394f5d2091880724f7d7638e926f9728

Observation fa568fdd-33ba-47a7-8f3a-c6a07f2bdd31 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:28.936948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:28.936948Z digest=sha256:08007804679ee26448900d6f1e6b76717562f7678d1b0ce3a870a690dbcecde1

Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.011230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.011230Z digest=sha256:894aea899f54496e83cf1bc7574c0d0137a8e80bf5a81ca7b5ba701c3a56da1b

Observation 92d7bb58-110a-4dc1-b1a8-d6fdf26249d2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.157791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.157791Z digest=sha256:f6dc8dc5ed8a0eab70a2358f45933a4a7f7593f21ad0b975506496b817f59e2c

Observation ac4235ba-1686-4383-89d7-670ef940fb6a · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.285863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.285863Z digest=sha256:c8be7fee467e26352be1c2dd7b1af375c9beb76c4e0116bf37bbfdd42acb56ab

Observation c31fd329-18ae-4c4a-af35-5ef3063ac485 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.379167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.379167Z digest=sha256:fb500401b39019db783c7c47c96c84ad4b547a9dd403d795e452f5156017f04d

Observation 3d4739b4-8fbb-4f16-9d85-ca8afd503afd · outbound

This paper cites Dense Reward for Free in Reinforcement Learning from Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Dense Reward for Free in Reinforcement Learning from Human Feedback

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.474618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.474618Z digest=sha256:1df4785cc82084c127c58a14e6379aa9d36362896baea35993d7298d55d17e4b

Observation 9dc7d6f1-aa41-45d1-a570-87e7f284de1d · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Process Reinforcement through Implicit Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.517440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.517440Z digest=sha256:d2d1fd4f8910b6b14326b5f8bcb1780235ce6a2cca82006ab27a22fc8276681b

Observation d1e21099-45e5-456f-9a22-a89f33de763d · outbound

This paper cites RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T11:32:32.289654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:29.571593Z digest=sha256:15fe3e00c3ab1964daf552f089c7466a4cc87c67bea41593d61ab3f89ff4d1ba

Observation 226e8b7d-71ea-4786-aa9b-ea8cb8881666 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.652497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.652497Z digest=sha256:402a3792a67bd16a3fc364824cb09ab07c59a90659f9ea77ea0734d1af316dc7

Observation 0f06674a-aab8-4c2f-ace5-3d413506452c · outbound

This paper cites Reinforcement learning: An introduction.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Reinforcement learning: An introduction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:33.106505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:29.736929Z digest=sha256:77622e544127facb3ca84f2b077577a4364b12a1a7ea481bdd2b8dba2d0decde

Observation e3eae744-e475-416a-b86c-3b11a18d485e · outbound

This paper cites Analysis of on-policy policy gradient methods under the distribution mismatch.arXiv preprint arXiv:2503.22244, 2025.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Analysis of on-policy policy gradient methods under the distribution mismatch.arXiv preprint arXiv:2503.22244, 2025

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-07T11:32:32.175493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:29.838106Z digest=sha256:6bd2eac6407c520add7ef34b92d03dea59ab18d053a38f21e4ceb91d546c64d6

Observation ebc8704f-7ad7-42cb-b105-420a41cba4df · outbound

This paper cites High-dimensional continu- ous control using generalized advantage estimation, 2018.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective High-dimensional continu- ous control using generalized advantage estimation, 2018

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.986611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:29.935825Z digest=sha256:1fd43fcc6c2ae41b305ec85f8adf06468889e3a0fda24b2e278adb89f930e3c3

Observation 229e0672-ac94-41c0-aad3-7d737231e735 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.026404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.026404Z digest=sha256:2541f4ee49652bf6cd413d74a7f82ba705d143f16c248e56c75d9584859d00d4

Observation faea92d1-0529-408c-8cf3-62950293087a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Gonzalez, Hao Zhang, and Ion Stoica

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.149593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.149593Z digest=sha256:8e6bb27b0db208fe3b4dc634f0d3219fedea9ed6b1071ebe012e08fc49aec39c

Observation 8d855c2e-f274-4eab-97d1-d23f6041ab13 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.869318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:30.236850Z digest=sha256:e5d4ff43a378c834c542dbb228b570edc1762827f52aef29d4c9e75a8a3e2546

Observation 3d37a5d8-7487-4031-94a2-c63c0b08b245 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.314881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.314881Z digest=sha256:8c4b9efeed93a5ca2a36b42aa60acc82f80331f53b5cf32e6b465db701996329

Observation 06c03cc1-5efd-4109-ad8b-30f0011fe3dc · outbound

This paper cites Multi-turn Reinforcement Learning from Preference Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.406644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.406644Z digest=sha256:4ad14da66cd39df95539217a3118e785edc92073719e048e978260dc97f36d95

Observation 4ef9d1d0-0259-4636-8775-80a7ff65f737 · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.490860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.490860Z digest=sha256:fdd45de215633186806ea22b2638a641289af6153ca365c67760302582009868

Observation f04a9702-4b69-44c0-ac3b-86915a75c762 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.549942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.549942Z digest=sha256:77269dcaf308a6f315ce85d773d3161180019b2191112570cba0a948944609e4

Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.615635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.615635Z digest=sha256:34cbf7ab3e7132885b56b5a399194616c5787ad61b62dbee8894e4230472fce7

Observation 46c7a99b-9411-4c48-9a88-2d440dd9bf5d · outbound

This paper cites PlanGenLLMs: A Modern Survey of LLM Planning Capabilities.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective PlanGenLLMs: A Modern Survey of LLM Planning Capabilities

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.682383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.682383Z digest=sha256:bc0e263b655a69b40da23521576e332dd3b7d6522fea2feefac18efe63ad5c45

Observation d316c10c-5751-4cb2-a82f-c5351ec05453 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.761569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.761569Z digest=sha256:944997a737d4a2a33afb4e404018b1321a27ec4c8451326f044587c4affc7883

Observation 4c96c6b6-cb23-4153-9448-1ad05c0b0ae9 · outbound

This paper cites Interactive evaluation for medical LLMs via task- oriented dialogue system.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Interactive evaluation for medical LLMs via task- oriented dialogue system

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.721910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:30.841930Z digest=sha256:52b1920e0720a04bf8676b2b4f4b2938f811b67023b99fbd43c4928d25952061

Observation 7ed50a14-fea9-4537-9fe4-15c531531cc2 · outbound

This paper cites Let’s verify step by step, 2023.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Let’s verify step by step, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.927433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.927433Z digest=sha256:18b7e163ee9ac41492b6899213ac9f53c23bad02da374168a74f170d027bfb45

Observation 7a0a2145-8133-4678-ac20-be71fc5c50ae · outbound

This paper cites Unraveling rlhf and its variants: Progress and practical engineering insights.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Unraveling rlhf and its variants: Progress and practical engineering insights

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:32:32.554729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:32:31.056825Z digest=sha256:c260f291f88e8313b2c9ed40c83721ee5fa800751cb3ee64de24291f561e24f9

Observation c4d243be-85b3-4deb-8dd7-850e0b14cfb7 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Solving math word problems with process- and outcome-based feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.151918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.151918Z digest=sha256:123e0996ee64354477b3242f1683f64e33d68336f93d06fc9c3134ef810a59b0

Observation ce13e47a-3bf7-42d3-bbf8-efa2c0227fb7 · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Fine-grained human feedback gives better rewards for language model training.Advances in Neural Information Processing Systems, 36:59008–59033, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.223415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.223415Z digest=sha256:cdaffc6fbce2852d76278db81bebfea113e1cdc281b3b1bc96dc4967e150966e

Observation 86e6e972-8ff4-4c20-a920-2dfcddb4ecf9 · outbound

This paper cites Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.310922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.310922Z digest=sha256:fa3cc49d6d67dd65a07d881ef64cc5c773a5fb14d5e50e9f9e5f6d13c4a5a15b

Observation 70da07b1-0cf1-44f1-ba5c-0c3613c1b6ab · outbound

This paper cites TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.365581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.365581Z digest=sha256:7ae8a7c19eed279c3f3bce7d1a0c9368e90fe1a9a7e9db4f619bdef1b57ed87d

Observation 005c73bf-7ead-4a34-a89f-163671031cb4 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.431529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.431529Z digest=sha256:f6fabfa7e3771982eb2bd8c79c247338ac3ad797c1e276fa09d4081a47de3337

Observation 867e3e6f-9b30-47c1-b77b-cd076a11bcef · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.527587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.527587Z digest=sha256:a62f5b6354aa0792e5dbb9202b6f90f0a06ef86950c98f2048a779a6397baf2f

Observation eb286a80-c534-4ce2-98f5-35563562af88 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Chatbot arena: An open platform for evaluating llms by human preference

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.601560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.601560Z digest=sha256:11446b928649b1dfc815a8e1026fc59cd4752c888609986a6113cd757916928f

Observation b472d78c-92ea-4f16-8e06-1c1a56c278ee · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.693251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.693251Z digest=sha256:4247f4446f44a253fbb35c151defaaa71b503acdb46e8d417a4f8f9f8e64f16e

Observation 71bc59d4-7290-4a87-b902-305f8be53e54 · outbound

This paper cites MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.776781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.776781Z digest=sha256:7285ba422f9a5cdba4793ba42ba2e1194722174094147f26305bf2cc5a058b46

Pith citing papers

Observation 458b6a9e-4f12-408c-853b-786e6f2ba5ed · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 252

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.517098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:68beb9aa2ab53ae2340aeba31fa3d0a6b95dc15994026eb4c597362fff6f94fb

Observation cff38d40-0384-4dfd-b50f-4b7a1be48da5 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.092931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.092931Z digest=sha256:f3de731555e6f6c808b32208557f78ea4aafc5faf8dfc17e8dc936a86d6393d6

Observation 389afa80-1000-430a-9628-73a046081c5c · inbound

Relative Score Policy Optimization for Diffusion Language Models cites this paper.

Relative Score Policy Optimization for Diffusion Language Models Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:56:29.690474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T03:47:42.196931Z digest=sha256:01a8d7a40821f696c0d510d8c4dd4acfe1567d84de8b4e10225e2452edb9a70c