Pith. sign in

Paper Citation Record · LEDGER

Discriminative Policy Optimization for Token-Level Reward Models

As of 20 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2505.23363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23363 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:04.281553Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:06.600251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:28:44.589552Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7e148b4-963c-40be-a943-e8a2fed01237 · outbound

This paper cites Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s.

Discriminative Policy Optimization for Token-Level Reward Models Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.170010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.170010Z digest=sha256:b749f5fb6647c784b86bd9cdb60e5adc5b452bd3b53e077ef2c189c3739851d9

Observation 0c568d27-ee89-4730-a83b-59ff45722b99 · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.223597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.223597Z digest=sha256:e891bfe42f45f91b8fef5aa03e4cb70e9217219935d276439625cdbb7d0c1044

Observation 47d90448-4963-40a1-833c-85fc9c28b0a5 · outbound

This paper cites InternLM2 Technical Report.

Discriminative Policy Optimization for Token-Level Reward Models InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.308256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.308256Z digest=sha256:7598cacd0b7972bc1e5f392e095639aeb903566d339c9cc1836bb77e82b85dfe

Observation e671e9a0-ca6e-4148-a433-b62f4c7b745d · outbound

This paper cites J., Sun, H., Holt, S., and Van Der Schaar, M.

Discriminative Policy Optimization for Token-Level Reward Models J., Sun, H., Holt, S., and Van Der Schaar, M

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.731373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:00.382055Z digest=sha256:0d96c2aa5750016472272892b2909feef688fb927eb385fca9c057364362c97a

Observation 6ba23d50-d678-4642-91c7-7ec6c4b59821 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Discriminative Policy Optimization for Token-Level Reward Models Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.499583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.499583Z digest=sha256:283afe970bdd48c855547c31b0309206995683dd15b162c8e0b4f22b7ddf07e7

Observation da094103-4f8a-4beb-8fa2-20f6db544879 · outbound

This paper cites ULTRAFEEDBACK : Boosting language models with scaled AI feedback.

Discriminative Policy Optimization for Token-Level Reward Models ULTRAFEEDBACK : Boosting language models with scaled AI feedback

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.462280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:00.571204Z digest=sha256:6e26e9f3aaff9aa3f0d869f4eb457dca0b0c6487e3152d2c019ed44133ca3c5e

Observation 07a6f9bb-ba1a-4091-8c3e-58fca04d9e37 · outbound

This paper cites The Llama 3 Herd of Models.

Discriminative Policy Optimization for Token-Level Reward Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.686329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.686329Z digest=sha256:082bfb21f15c3cee982d7fa8a734aae5eac5c7cdebd0f3621228c009705cb29b

Observation 734fa4d1-f489-4b57-8976-a661bec6c2f0 · outbound

This paper cites X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P.

Discriminative Policy Optimization for Token-Level Reward Models X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.197135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:00.786266Z digest=sha256:dd58c0ee9373b7da24ec70776ecc9d926b2655661041d2f5081fa830cc157a44

Observation 111759b4-88e0-4ade-a945-12fa880d79ae · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Discriminative Policy Optimization for Token-Level Reward Models Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.869018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.869018Z digest=sha256:4106f68caeaad2af66dce6eb8ce1d6d37c75b9af32c9b3eb8196856f37d10279

Observation f02d4b2e-4849-4d69-bd77-6a3da47ce89c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Discriminative Policy Optimization for Token-Level Reward Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.940715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.940715Z digest=sha256:532976758de4c4b10108e36ed6e97cc800ef589fcc76191da15b5073e33680d3

Observation 2685d004-6370-4d52-9047-1c961b29abff · outbound

This paper cites The curious case of neural text degeneration.

Discriminative Policy Optimization for Token-Level Reward Models The curious case of neural text degeneration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.029667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.029667Z digest=sha256:a1d91b9282fd4b6dec33b49b852183ef890b3d89be901d3861420b8e2f64f001

Observation 0436e1a4-38b4-4bdf-8626-b935f7155199 · outbound

This paper cites ORPO : Monolithic preference optimization without reference model.

Discriminative Policy Optimization for Token-Level Reward Models ORPO : Monolithic preference optimization without reference model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.108360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.108360Z digest=sha256:f940bf2593df5544298b5295bcf20e1e6e8d0e95f1c10c1f13fab6a2673adec4

Observation aee5fa32-c2f2-44f1-a026-c14aa9a32e3a · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Discriminative Policy Optimization for Token-Level Reward Models J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.244539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.244539Z digest=sha256:1a7ba66440a98b478eda89f2f6a16c9f8b0653d4ea82c4ba4fce8f340ecd999a

Observation 4730ea76-6a7a-4f55-a2d8-0e6386a5c80f · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Discriminative Policy Optimization for Token-Level Reward Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.360490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.360490Z digest=sha256:3a54946a039607a7d079d928460bb6cc31fec1c8c88284472c083c09ac45c9fd

Observation b3209566-044e-4de5-8b0f-5693da8fce9c · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Discriminative Policy Optimization for Token-Level Reward Models Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.479481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.479481Z digest=sha256:f00dc8c030841a5ddf779bd51e935fbb7b9db47f92a4afd10215e2818e7589bf

Observation e0b39bb8-df3e-4cbd-90c0-8bee13c667ef · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Discriminative Policy Optimization for Token-Level Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.634524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.634524Z digest=sha256:02df53bea89db1e335b5f46f5686a52515b9d1722c4c4e0834b776d63210ed9c

Observation a18e4999-6eab-447b-abe9-4e904ac4bb1e · outbound

This paper cites Let's verify step by step.

Discriminative Policy Optimization for Token-Level Reward Models Let's verify step by step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.756993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.756993Z digest=sha256:13fb86aeb7886d500f09cc334d44de470435b3c372f0fff7526f1e1f3f10dbe1

Observation 89291a9a-b254-46cb-887d-b3fe6adc8750 · outbound

This paper cites Focal loss for dense object detection.

Discriminative Policy Optimization for Token-Level Reward Models Focal loss for dense object detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.017255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:01.873518Z digest=sha256:d767b5b0b2a0bf2909b15d415bd7687d0f3881f810d034a6bc9107a8fefc7cfc

Observation ef3dacc7-bcbf-4059-a2c0-c7a75b53c247 · outbound

This paper cites Focal loss for dense object detection.

Discriminative Policy Optimization for Token-Level Reward Models Focal loss for dense object detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.958225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.958225Z digest=sha256:9898601b604e1ae17b88299331f078ad62bf6b4894dab2be90bac99b08b65540

Observation 4f0b6d3c-de52-481b-ac72-eb4fa9306ae6 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Discriminative Policy Optimization for Token-Level Reward Models Simpo: Simple preference optimization with a reference-free reward

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.853693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:01.963540Z digest=sha256:fe03aaaa03945e3a55a0bb4d2a3cb6297f15f04b9b70f0488403a84b967574dc

Observation 214dbf5c-601e-4e4b-91e4-491fe735d9e0 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Discriminative Policy Optimization for Token-Level Reward Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.050249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.050249Z digest=sha256:2c28b2d839e8a13259be9a4ebb4ac6f5675d3a7914a71a611586100fad591c0b

Observation c524e7a1-7f5c-4e45-a482-a2f202171da8 · outbound

This paper cites Introducing OpenAI o1 , 2024.

Discriminative Policy Optimization for Token-Level Reward Models Introducing OpenAI o1 , 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.616125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:02.123565Z digest=sha256:8a6e948b902b7a6bfba9f420a411938aaa888c300a11f5dde84d5635c033b33a

Observation 2f5a6936-3892-45ca-925a-64d68330879a · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Discriminative Policy Optimization for Token-Level Reward Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.180756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.180756Z digest=sha256:474b4d7ab72f3e9fbae5b330ab6fb84c9ce07f75f9b352500a421704dd2c134b

Observation b8d82617-0a9f-47b7-87b3-f1362e2cc132 · outbound

This paper cites From \ r\ to \ q *\ : Your language model is secretly a q-function.

Discriminative Policy Optimization for Token-Level Reward Models From \ r\ to \ q *\ : Your language model is secretly a q-function

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.445722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:02.308993Z digest=sha256:2604ab7a0e930e685ffa2295c9891b3a8ac8a2d893db0e58399d9d33c41c9d8e

Observation 8ee7fcb0-bfd1-48ce-8c61-cc88102eecec · outbound

This paper cites D., Ermon, S., and Finn, C.

Discriminative Policy Optimization for Token-Level Reward Models D., Ermon, S., and Finn, C

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.382625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.382625Z digest=sha256:d99b6a1e8962827e2fb35d2521f234bd5a2c18ac1dfd9ba418462ffaf13e8d11

Observation 5128d80d-a047-4e3d-ae92-b960352d8f3c · outbound

This paper cites D., and Arora, S.

Discriminative Policy Optimization for Token-Level Reward Models D., and Arora, S

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.451697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.451697Z digest=sha256:113e2c2f89b265e31356b40887dbc83452eb8963749f09135c5dc541b77fd521

Observation 36377019-a0e4-4574-8a3e-bbf66a2990f6 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Discriminative Policy Optimization for Token-Level Reward Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.526618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.526618Z digest=sha256:7d825a670413d0950d3a78a9ac4ff9d5ed54cf03cb99f267338aafcf7dfb1371

Observation f2781492-3e3f-4302-a97d-0b64e2decd7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Discriminative Policy Optimization for Token-Level Reward Models Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.679044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.679044Z digest=sha256:e7e0763281255dddd9882e09101119c1aafeed9eb0209bf1e9a0af5aa8d3bdec

Observation b39d473c-6b70-4b08-8953-e33ab90a1394 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Discriminative Policy Optimization for Token-Level Reward Models Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.787047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.787047Z digest=sha256:c4614f088b74f0d075835dca4831f4d3f36e0a96f8d98861ec1af5c43a2c7893

Observation a5b5b61b-beb6-46a2-8fd5-9a793fcbb820 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Discriminative Policy Optimization for Token-Level Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.872772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.872772Z digest=sha256:9bb104ceb9dffa79bd89e2308ed611f152619f66772961d6190f5c5bb174ced2

Observation f976e7a7-1230-42cc-8514-de3a68daed93 · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.002314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.002314Z digest=sha256:d28c653df573aa5a8c952ff1a03fd37f5784e7dd5c7ecf20dba177e030416512

Observation 1f0c314e-2943-48c8-8d05-f2b3eafab842 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Discriminative Policy Optimization for Token-Level Reward Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.101409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.101409Z digest=sha256:a3666809cfcef82be3023dc5f48dfb57a0f04a4bbbee89746db6e2b325dc1eeb

Observation a17a6c26-1c7b-4fd9-a6f8-f5312e440d53 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Discriminative Policy Optimization for Token-Level Reward Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.258558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:03.213515Z digest=sha256:ecf9fd0146195d20edfa71eb46eafb11008be07baa00e4ebdd1177685cb4c306

Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.339487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.339487Z digest=sha256:bd7a5da33254f80505234a5e10e0b7d01200f7b72cbaddad288d1d051330266f

Observation f9f67d1a-d148-4a31-b69e-96b9941485e9 · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.449654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.449654Z digest=sha256:93cf45cba0330a526236d63712cde81c6ed2c2aa259732e37f119b25867e8034

Observation 8ac999b2-4ec7-4251-9977-52ee69dabee3 · outbound

This paper cites A., Ostendorf, M., and Hajishirzi, H.

Discriminative Policy Optimization for Token-Level Reward Models A., Ostendorf, M., and Hajishirzi, H

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.557005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.557005Z digest=sha256:7b3535e1cc898d0d0c9d3134756dadb8890be2a46d15faa35f344e4dea390e6a

Observation 2b694f6c-7c0f-42b2-8c8e-1dd3336370c5 · outbound

This paper cites Qwen2.5 Technical Report.

Discriminative Policy Optimization for Token-Level Reward Models Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.688731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.688731Z digest=sha256:5ce1995ab9ede227798c09ec8012bd09d2ec6845824b5097b0624619302ead5e

Observation 966c4a33-ee29-426f-9d66-0eab766d6dd0 · outbound

This paper cites Preference-grounded token-level guidance for language model fine-tuning.

Discriminative Policy Optimization for Token-Level Reward Models Preference-grounded token-level guidance for language model fine-tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.104059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:03.798895Z digest=sha256:fef4a9e48d25617ed4de7d5ef65c84cefc595cb4243583497c975df3b5f41781

Observation 902bba15-32bb-41e4-a1ca-dc6fe9861718 · outbound

This paper cites Free Process Rewards without Process Labels.

Discriminative Policy Optimization for Token-Level Reward Models Free Process Rewards without Process Labels

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.880457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.880457Z digest=sha256:d391bcf31bbb2e15d97589e448854a945206676bfe00c91ce4a86c444e78b029

Observation a76cfcab-45fa-46b7-8a3e-87ae51ec8b98 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Discriminative Policy Optimization for Token-Level Reward Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.969381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.969381Z digest=sha256:80c065f28b9412738fa9a68d4d887ac584d4b0c27e31e0eae2c52d351706b852

Observation 44ea10ff-3702-4fc1-9508-d064fd29c2be · outbound

This paper cites DPO meets PPO : Reinforced token optimization for RLHF.

Discriminative Policy Optimization for Token-Level Reward Models DPO meets PPO : Reinforced token optimization for RLHF

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:04.908276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:53:04.065233Z digest=sha256:e14eb046c5642985302b77a08970cb4fbd26f24ea4b0d6396b7f328c54f32829

Observation 815cec98-71d5-46dc-9839-4fb328ba13ca · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:04.169622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:04.169622Z digest=sha256:2703bd492c3a917713555e49a8f79b26e585d5936adfaca4a50afd1253361af1

Observation 6ebdcbd6-9030-4eb5-8cc6-d3d52ebde259 · outbound

This paper cites write newline.

Discriminative Policy Optimization for Token-Level Reward Models write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:04.281553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:04.281553Z digest=sha256:e8d53cdc4c7e37bcd1fa74b8f4be06e40f6d7ca79656196d4a23665190c3bbdf

Pith citing papers

Observation 6a48b8fa-7c4f-46de-8db1-6b88a540deda · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design Discriminative Policy Optimization for Token-Level Reward Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.600251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.600251Z digest=sha256:799780d704ef9890f2c14755fec43d8d8119c20c5203b97677f2724832ba5992

Observation 64672a6b-4d2b-4641-bb7a-29a6eee19258 · inbound

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation cites this paper.

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation Discriminative Policy Optimization for Token-Level Reward Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.591020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T04:07:26.224919Z digest=sha256:4dd966c38bb09f4de9105aa9864d240bad689771adcfe96e855cea528fcf76a1

Observation 959f098a-23bb-411c-9ad1-170ba632a6f2 · inbound

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization cites this paper.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Discriminative Policy Optimization for Token-Level Reward Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.070055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.070055Z digest=sha256:92d44a818ada470a433def28d6e9144416ec049e534ede3a5d2cb44a9948b377