Pith. sign in

Paper Citation Record · LEDGER

Discriminative Policy Optimization for Token-Level Reward Models

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2505.23363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23363 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:04.281553Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:07:26.224919Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:28:44.589552Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7e148b4-963c-40be-a943-e8a2fed01237 · outbound

This paper cites Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s.

Discriminative Policy Optimization for Token-Level Reward Models Back to basics: Revisiting REINFORCE -style optimization for learning from human feedback in LLM s

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.170010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.170010Z digest=sha256:208e2fe4663e621b0b48b3b2f751268d241c573fbcab753caff6bcbbdb537ff5

Observation 0c568d27-ee89-4730-a83b-59ff45722b99 · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.223597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.223597Z digest=sha256:f81394bb44f63ac84fa72efd9b31dd92169c1adb21f3f10d677fc7ab584ac118

Observation 47d90448-4963-40a1-833c-85fc9c28b0a5 · outbound

This paper cites InternLM2 Technical Report.

Discriminative Policy Optimization for Token-Level Reward Models InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.308256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.308256Z digest=sha256:4713ad0162fbf618697f35f5a24d99c2c01b6b300d0be51ba14f3ee45acd507e

Observation e671e9a0-ca6e-4148-a433-b62f4c7b745d · outbound

This paper cites J., Sun, H., Holt, S., and Van Der Schaar, M.

Discriminative Policy Optimization for Token-Level Reward Models J., Sun, H., Holt, S., and Van Der Schaar, M

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.731373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:00.382055Z digest=sha256:4da96d30905f8caa32cb17ffc9da9a69f313ceba709d40c48142f9275a6e75d2

Observation 6ba23d50-d678-4642-91c7-7ec6c4b59821 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Discriminative Policy Optimization for Token-Level Reward Models Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.499583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.499583Z digest=sha256:4bea6af04b26e21eaa8b5d9535156842e4ac46794bdbced421f9ab4c2c29cfae

Observation da094103-4f8a-4beb-8fa2-20f6db544879 · outbound

This paper cites ULTRAFEEDBACK : Boosting language models with scaled AI feedback.

Discriminative Policy Optimization for Token-Level Reward Models ULTRAFEEDBACK : Boosting language models with scaled AI feedback

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.462280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:00.571204Z digest=sha256:322c08d6de1122265b6fbeb107f884c64e79bb0e3e853d44dd74dbd939850d91

Observation 07a6f9bb-ba1a-4091-8c3e-58fca04d9e37 · outbound

This paper cites The Llama 3 Herd of Models.

Discriminative Policy Optimization for Token-Level Reward Models The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.686329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.686329Z digest=sha256:d31f6305030770166c59b80f5f60001115cbb5f5362924245869a5e918629245

Observation 734fa4d1-f489-4b57-8976-a661bec6c2f0 · outbound

This paper cites X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P.

Discriminative Policy Optimization for Token-Level Reward Models X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.197135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:00.786266Z digest=sha256:274d53328efb36bea3889ae66444545f5ec2dc4502a4abac3f0aa2f9e84e53d5

Observation 111759b4-88e0-4ade-a945-12fa880d79ae · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Discriminative Policy Optimization for Token-Level Reward Models Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.869018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.869018Z digest=sha256:8634de2b6625164ab6f198a022bb5324e56b2b2d1030a36b66955f4b2ed1b3c8

Observation f02d4b2e-4849-4d69-bd77-6a3da47ce89c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Discriminative Policy Optimization for Token-Level Reward Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:00.940715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:00.940715Z digest=sha256:7335e84cf8c220bf2f99d3ad4e2f07c16deba8a1a81778dfbe6fc55713b34c92

Observation 2685d004-6370-4d52-9047-1c961b29abff · outbound

This paper cites The curious case of neural text degeneration.

Discriminative Policy Optimization for Token-Level Reward Models The curious case of neural text degeneration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.029667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.029667Z digest=sha256:4bbe05146f0499c4cb1f79eb89e97498221a59f924c4d126d8e1dd57bd044878

Observation 0436e1a4-38b4-4bdf-8626-b935f7155199 · outbound

This paper cites ORPO : Monolithic preference optimization without reference model.

Discriminative Policy Optimization for Token-Level Reward Models ORPO : Monolithic preference optimization without reference model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.108360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.108360Z digest=sha256:534bc9f33c1a9d424c3d5271ae687440fe82b8211d0c4f6aeb860d9d4abc7308

Observation aee5fa32-c2f2-44f1-a026-c14aa9a32e3a · outbound

This paper cites J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W.

Discriminative Policy Optimization for Token-Level Reward Models J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.244539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.244539Z digest=sha256:c5c34ef76ea7f1fd3ea4b95176ad3698c2c42d8884b17bc9217586502b03b624

Observation 4730ea76-6a7a-4f55-a2d8-0e6386a5c80f · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Discriminative Policy Optimization for Token-Level Reward Models O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.360490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.360490Z digest=sha256:872fa05106c5d7589b674c74c988a80872954165a2c798a5e75833a2e492f99d

Observation b3209566-044e-4de5-8b0f-5693da8fce9c · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Discriminative Policy Optimization for Token-Level Reward Models Adam: A Method for Stochastic Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.479481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.479481Z digest=sha256:5fa2996a4ef476d79bc9dd6232c43e060a0041a01ce39d9cf2016778fcf756d1

Observation e0b39bb8-df3e-4cbd-90c0-8bee13c667ef · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Discriminative Policy Optimization for Token-Level Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.634524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.634524Z digest=sha256:eeed617eccc779038a59de5107cd5c0083b611404a7ac64712dfe6957a4a7d40

Observation a18e4999-6eab-447b-abe9-4e904ac4bb1e · outbound

This paper cites Let's verify step by step.

Discriminative Policy Optimization for Token-Level Reward Models Let's verify step by step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.756993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.756993Z digest=sha256:94d3f4aea154f9e67a0b36d1b4c1577bf639190f2ebc11abf4f88b60101d7599

Observation 89291a9a-b254-46cb-887d-b3fe6adc8750 · outbound

This paper cites Focal loss for dense object detection.

Discriminative Policy Optimization for Token-Level Reward Models Focal loss for dense object detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:06.017255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:01.873518Z digest=sha256:262c0e383f9a14114f493f0d0aa7de762cc358ccd7e6ae9262c051985eba11c9

Observation ef3dacc7-bcbf-4059-a2c0-c7a75b53c247 · outbound

This paper cites Focal loss for dense object detection.

Discriminative Policy Optimization for Token-Level Reward Models Focal loss for dense object detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:01.958225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:01.958225Z digest=sha256:f261d5868bd31d85b1a166454ff4454ac29ee9718acb21b311cd0813e087d7ac

Observation 4f0b6d3c-de52-481b-ac72-eb4fa9306ae6 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Discriminative Policy Optimization for Token-Level Reward Models Simpo: Simple preference optimization with a reference-free reward

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.853693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:01.963540Z digest=sha256:0fd9c9d47beb7987bb806ab1dec80c8ed04874b635f5012788e7f6b4c204a793

Observation 214dbf5c-601e-4e4b-91e4-491fe735d9e0 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Discriminative Policy Optimization for Token-Level Reward Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.050249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.050249Z digest=sha256:feb46fd5fa0ab8729177eec7f64de27d63767c510b5919281ed85227a1df8b2c

Observation c524e7a1-7f5c-4e45-a482-a2f202171da8 · outbound

This paper cites Introducing OpenAI o1 , 2024.

Discriminative Policy Optimization for Token-Level Reward Models Introducing OpenAI o1 , 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.616125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:02.123565Z digest=sha256:3f0d75b0f4e8acbac767bb635bb910d3ac28a5f221353643dcb5fb8491da5a28

Observation 2f5a6936-3892-45ca-925a-64d68330879a · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

Discriminative Policy Optimization for Token-Level Reward Models O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.180756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.180756Z digest=sha256:0654a2577e6a4de78d3587b03612e40b7e0db01277c752da404efbd6a7cf8b64

Observation b8d82617-0a9f-47b7-87b3-f1362e2cc132 · outbound

This paper cites From \ r\ to \ q *\ : Your language model is secretly a q-function.

Discriminative Policy Optimization for Token-Level Reward Models From \ r\ to \ q *\ : Your language model is secretly a q-function

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.445722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:02.308993Z digest=sha256:8455fdff04117747bbce7fcd18be721c80b0cc087fb86ceb79772acef60fb59a

Observation 8ee7fcb0-bfd1-48ce-8c61-cc88102eecec · outbound

This paper cites D., Ermon, S., and Finn, C.

Discriminative Policy Optimization for Token-Level Reward Models D., Ermon, S., and Finn, C

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.382625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.382625Z digest=sha256:8ec4a9ef1bc2326d079683c54991b5bbf69701151aa3ff73cbfe5b4d402b469c

Observation 5128d80d-a047-4e3d-ae92-b960352d8f3c · outbound

This paper cites D., and Arora, S.

Discriminative Policy Optimization for Token-Level Reward Models D., and Arora, S

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.451697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.451697Z digest=sha256:ab217b029bc1ccd8565baf06b754cffe4898bc4fc156c4245fc2c0e3c18dc214

Observation 36377019-a0e4-4574-8a3e-bbf66a2990f6 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Discriminative Policy Optimization for Token-Level Reward Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.526618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.526618Z digest=sha256:3e0c1898812a1e007168b3305dc936a75fbcb7ce278704306817dcc99e2f499e

Observation f2781492-3e3f-4302-a97d-0b64e2decd7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Discriminative Policy Optimization for Token-Level Reward Models Proximal Policy Optimization Algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.679044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.679044Z digest=sha256:41f5cacf8d6864e1b39ee401fbd2188aad76b395bcd6f765af4e43745fe828be

Observation b39d473c-6b70-4b08-8953-e33ab90a1394 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Discriminative Policy Optimization for Token-Level Reward Models Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.787047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.787047Z digest=sha256:5287415aea20c77bb9d16a46ef806394e585f4e1e49932fd22ce850d42e148b0

Observation a5b5b61b-beb6-46a2-8fd5-9a793fcbb820 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Discriminative Policy Optimization for Token-Level Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:02.872772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:02.872772Z digest=sha256:9f3f51b6735fe67155e1ba74a89575a18c649d168b0483edef31057f5948e937

Observation f976e7a7-1230-42cc-8514-de3a68daed93 · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.002314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.002314Z digest=sha256:74b0dca2d8d763398c9261bf2530e0bcc1dad24cde2bfb91b62f40294267825b

Observation 1f0c314e-2943-48c8-8d05-f2b3eafab842 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Discriminative Policy Optimization for Token-Level Reward Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.101409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.101409Z digest=sha256:bbb339586772392955510e0550ca6884fb5d49493f7275d96566217ebb22250e

Observation a17a6c26-1c7b-4fd9-a6f8-f5312e440d53 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Discriminative Policy Optimization for Token-Level Reward Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.258558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:03.213515Z digest=sha256:daf2b2f01921d3ee5bdccb44863c5e4ca34d5b081321baca48a538595057d8e3

Observation 93f23228-ef50-42f8-802c-ad68b2f37369 · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Discriminative Policy Optimization for Token-Level Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.339487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.339487Z digest=sha256:079c3867fb048c1c47c227fcc4d8927f8e285c5f38be97c72f7b9a6387055d48

Observation f9f67d1a-d148-4a31-b69e-96b9941485e9 · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.449654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.449654Z digest=sha256:0dbe6fbd3e2a28342bff9d10986b1d1a92d13184d4ee2533077ac30f28383b66

Observation 8ac999b2-4ec7-4251-9977-52ee69dabee3 · outbound

This paper cites A., Ostendorf, M., and Hajishirzi, H.

Discriminative Policy Optimization for Token-Level Reward Models A., Ostendorf, M., and Hajishirzi, H

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.557005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.557005Z digest=sha256:85834df7126b7928198aee797e4b53e6a5c8bcbc4288e81c3f98bf85dc5c0912

Observation 2b694f6c-7c0f-42b2-8c8e-1dd3336370c5 · outbound

This paper cites Qwen2.5 Technical Report.

Discriminative Policy Optimization for Token-Level Reward Models Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.688731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.688731Z digest=sha256:b0cd49c88d60fcd8bfbbdc686684e3bd6dc66ab6c41415c34dffce5cf742f7a8

Observation 966c4a33-ee29-426f-9d66-0eab766d6dd0 · outbound

This paper cites Preference-grounded token-level guidance for language model fine-tuning.

Discriminative Policy Optimization for Token-Level Reward Models Preference-grounded token-level guidance for language model fine-tuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:05.104059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:03.798895Z digest=sha256:b67fce1af33f0b84eee719e9309090691761686ada29325e7ffe6b25f685491e

Observation 902bba15-32bb-41e4-a1ca-dc6fe9861718 · outbound

This paper cites Free Process Rewards without Process Labels.

Discriminative Policy Optimization for Token-Level Reward Models Free Process Rewards without Process Labels

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.880457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.880457Z digest=sha256:85cce2061f89092961290356e575d316c5980f100f2d0d6060843b8b551f93ae

Observation a76cfcab-45fa-46b7-8a3e-87ae51ec8b98 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Discriminative Policy Optimization for Token-Level Reward Models Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:03.969381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:03.969381Z digest=sha256:72722f95e049891edaab67e2b98722ece50583d2f3d6ea363a925e55019b3054

Observation 44ea10ff-3702-4fc1-9508-d064fd29c2be · outbound

This paper cites DPO meets PPO : Reinforced token optimization for RLHF.

Discriminative Policy Optimization for Token-Level Reward Models DPO meets PPO : Reinforced token optimization for RLHF

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:04.908276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:53:04.065233Z digest=sha256:a7d689fc9a5f0d6960ff9a5da547424f9a9c26660d1d13251e8f053a21ffecd5

Observation 815cec98-71d5-46dc-9839-4fb328ba13ca · outbound

This paper cites an unresolved cited work.

Discriminative Policy Optimization for Token-Level Reward Models Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:04.169622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:04.169622Z digest=sha256:eba1b6cac900552ea423f35b2eb24f173cbfe010ea7b16a105674d859bdcb3d1

Observation 6ebdcbd6-9030-4eb5-8cc6-d3d52ebde259 · outbound

This paper cites write newline.

Discriminative Policy Optimization for Token-Level Reward Models write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:04.281553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:53:04.281553Z digest=sha256:bf8565ac0fd4125edeffbf2e2f063458ab54772b4597e8dcb25fecd649e20498

Pith citing papers

Observation 64672a6b-4d2b-4641-bb7a-29a6eee19258 · inbound

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation cites this paper.

BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation Discriminative Policy Optimization for Token-Level Reward Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.591020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T04:07:26.224919Z digest=sha256:167b62364e174dd5c0e8b2da716c841266f9fe425b15b98b83194daafcbe2870