Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:04.281553Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2505.23363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:04.281553Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T04:07:26.224919Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:28:44.589552Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b7e148b4-963c-40be-a943-e8a2fed01237 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c568d27-ee89-4730-a83b-59ff45722b99 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d90448-4963-40a1-833c-85fc9c28b0a5 · outbound
Discriminative Policy Optimization for Token-Level Reward Models InternLM2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e671e9a0-ca6e-4148-a433-b62f4c7b745d · outbound
Discriminative Policy Optimization for Token-Level Reward Models J., Sun, H., Holt, S., and Van Der Schaar, M
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6ba23d50-d678-4642-91c7-7ec6c4b59821 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da094103-4f8a-4beb-8fa2-20f6db544879 · outbound
Discriminative Policy Optimization for Token-Level Reward Models ULTRAFEEDBACK : Boosting language models with scaled AI feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07a6f9bb-ba1a-4091-8c3e-58fca04d9e37 · outbound
Discriminative Policy Optimization for Token-Level Reward Models The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734fa4d1-f489-4b57-8976-a661bec6c2f0 · outbound
Discriminative Policy Optimization for Token-Level Reward Models X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 111759b4-88e0-4ade-a945-12fa880d79ae · outbound
Discriminative Policy Optimization for Token-Level Reward Models Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02d4b2e-4849-4d69-bd77-6a3da47ce89c · outbound
Discriminative Policy Optimization for Token-Level Reward Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2685d004-6370-4d52-9047-1c961b29abff · outbound
Discriminative Policy Optimization for Token-Level Reward Models The curious case of neural text degeneration
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0436e1a4-38b4-4bdf-8626-b935f7155199 · outbound
Discriminative Policy Optimization for Token-Level Reward Models ORPO : Monolithic preference optimization without reference model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee5fa32-c2f2-44f1-a026-c14aa9a32e3a · outbound
Discriminative Policy Optimization for Token-Level Reward Models J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4730ea76-6a7a-4f55-a2d8-0e6386a5c80f · outbound
Discriminative Policy Optimization for Token-Level Reward Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3209566-044e-4de5-8b0f-5693da8fce9c · outbound
Discriminative Policy Optimization for Token-Level Reward Models Adam: A Method for Stochastic Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b39bb8-df3e-4cbd-90c0-8bee13c667ef · outbound
Discriminative Policy Optimization for Token-Level Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18e4999-6eab-447b-abe9-4e904ac4bb1e · outbound
Discriminative Policy Optimization for Token-Level Reward Models Let's verify step by step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89291a9a-b254-46cb-887d-b3fe6adc8750 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Focal loss for dense object detection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef3dacc7-bcbf-4059-a2c0-c7a75b53c247 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Focal loss for dense object detection
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0b6d3c-de52-481b-ac72-eb4fa9306ae6 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Simpo: Simple preference optimization with a reference-free reward
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 214dbf5c-601e-4e4b-91e4-491fe735d9e0 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c524e7a1-7f5c-4e45-a482-a2f202171da8 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Introducing OpenAI o1 , 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f5a6936-3892-45ca-925a-64d68330879a · outbound
Discriminative Policy Optimization for Token-Level Reward Models O1 Replication Journey: A Strategic Progress Report -- Part 1
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d82617-0a9f-47b7-87b3-f1362e2cc132 · outbound
Discriminative Policy Optimization for Token-Level Reward Models From \ r\ to \ q *\ : Your language model is secretly a q-function
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ee7fcb0-bfd1-48ce-8c61-cc88102eecec · outbound
Discriminative Policy Optimization for Token-Level Reward Models D., Ermon, S., and Finn, C
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5128d80d-a047-4e3d-ae92-b960352d8f3c · outbound
Discriminative Policy Optimization for Token-Level Reward Models D., and Arora, S
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36377019-a0e4-4574-8a3e-bbf66a2990f6 · outbound
Discriminative Policy Optimization for Token-Level Reward Models High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2781492-3e3f-4302-a97d-0b64e2decd7a · outbound
Discriminative Policy Optimization for Token-Level Reward Models Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39d473c-6b70-4b08-8953-e33ab90a1394 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b5b61b-beb6-46a2-8fd5-9a793fcbb820 · outbound
Discriminative Policy Optimization for Token-Level Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f976e7a7-1230-42cc-8514-de3a68daed93 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0c314e-2943-48c8-8d05-f2b3eafab842 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17a6c26-1c7b-4fd9-a6f8-f5312e440d53 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · outbound
Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9f67d1a-d148-4a31-b69e-96b9941485e9 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ac999b2-4ec7-4251-9977-52ee69dabee3 · outbound
Discriminative Policy Optimization for Token-Level Reward Models A., Ostendorf, M., and Hajishirzi, H
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b694f6c-7c0f-42b2-8c8e-1dd3336370c5 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Qwen2.5 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966c4a33-ee29-426f-9d66-0eab766d6dd0 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Preference-grounded token-level guidance for language model fine-tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 902bba15-32bb-41e4-a1ca-dc6fe9861718 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Free Process Rewards without Process Labels
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76cfcab-45fa-46b7-8a3e-87ae51ec8b98 · outbound
Discriminative Policy Optimization for Token-Level Reward Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ea10ff-3702-4fc1-9508-d064fd29c2be · outbound
Discriminative Policy Optimization for Token-Level Reward Models DPO meets PPO : Reinforced token optimization for RLHF
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 815cec98-71d5-46dc-9839-4fb328ba13ca · outbound
Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ebdcbd6-9030-4eb5-8cc6-d3d52ebde259 · outbound
Discriminative Policy Optimization for Token-Level Reward Models write newline
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64672a6b-4d2b-4641-bb7a-29a6eee19258 · inbound
BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation Discriminative Policy Optimization for Token-Level Reward Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.