Pith. sign in

Paper Citation Record · LEDGER

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 3 inbound Pith citation observations for arXiv:2507.17515.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17515 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:51:25.450746Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:14:19.399777Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.582742Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 952b4499-3103-4a5f-9bad-4a9ad8609a59 · outbound

This paper cites GPT-4 Technical Report.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.370171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.370171Z digest=sha256:427724ceab5e811f09185b630561ac500287964fdfddcdeacf7c0344b5c5a965

Observation ad92ed78-b216-439c-a5d4-87b4e42ae997 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.381797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.381797Z digest=sha256:15f5f2c880c3a6ebe668b53504ac77f2627aa781d9b4cedc17dd0063629f0931

Observation 8e04d0d6-e7a9-4627-84d0-9a1e08554e4c · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Skywork Open Reasoner 1 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.393581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.393581Z digest=sha256:bd7b512bbad645bad35a9c5bc1098eb5fd2b7e4fd45164733e36383a8618fe91

Observation 04f15869-2750-4dc7-b7ce-191bdb2543a1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.396830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.396830Z digest=sha256:0ec886fe75dc679b3bb2f3533368202b28a998201b141cfc2d5ded48bed6319d

Observation 2b7340c7-0fdd-4941-8604-8eeb0c96109e · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.403341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.403341Z digest=sha256:00c678282d6d6b0545bce24d39e51d3d96e5820fc4ae20eef2cc71dd48d985e0

Observation b7c407c9-2d4f-4be5-ac1f-0a661227b2ad · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.407063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.407063Z digest=sha256:754f1f4370af5c7eaf5484c2c21fc209cfad57b3e3a9b851f99ea6b9c7a9868e

Observation e667bb08-d015-43f3-8595-c926925194cd · outbound

This paper cites Inference-time scaling for generalist reward modeling.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Inference-time scaling for generalist reward modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.410064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.410064Z digest=sha256:ff1204e0e0c18d55c1f07879d763755e6fecd8c9aa91fdc6f5dc57c63a001c8d

Observation e9f9cf97-8a36-4916-b2fa-b57603f5e83a · outbound

This paper cites Decoupled Weight Decay Regularization.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Decoupled Weight Decay Regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.413138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.413138Z digest=sha256:63c22b0b89eb8262e6c2f2e93d530e108bf0301e51e1f884a4d4aa24aa4459e1

Observation 3bf7c51d-776f-471c-bb18-ed29a60eb10f · outbound

This paper cites BPR: Bayesian Personalized Ranking from Implicit Feedback.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models BPR: Bayesian Personalized Ranking from Implicit Feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.416391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.416391Z digest=sha256:10485235b634d9a578c36e07c698a98bed45cf208826e3622f85d454af407bea

Observation 2c42e392-b627-4cc1-bfde-6e24ddda2fc1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.419490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.419490Z digest=sha256:6103e1c9cb506be76ddcdb77bd95c9ef9bd8dd1b730aae79ae29a8815332b61d

Observation d8941b72-95c4-4586-a154-9027845a5560 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.423521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.423521Z digest=sha256:90e20d0c1994d1d5aea841687d7cd72b29d3e234d93f9c02ff64c667e50f30a7

Observation 5b2a5dc6-36d6-4581-814e-d8c97300f54d · outbound

This paper cites Qwen2.5 Technical Report.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Qwen2.5 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.426843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.426843Z digest=sha256:e87602dbc23024f3f3c1e19c42515568853a7957cd707b63d13118cc3a5425e7

Observation 989ca361-9d54-4bfb-9821-11b471c1ccc4 · outbound

This paper cites Qwen3 Technical Report.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.430191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.430191Z digest=sha256:a39b204abded8d45edcbcfd25722d1484ceb7601f32ac3f4b1cc7d355dc5f426

Observation 07c88954-0191-4d73-9788-9c3ebca59d6d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.433932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.433932Z digest=sha256:20e14265c6bf45512207f1d45308c62655ee2920941d777d859d947ba2b8f452

Observation 2e281e93-06b0-4e24-8a8f-50013c59e1b2 · outbound

This paper cites Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.437067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.437067Z digest=sha256:11f527dc4f5e8904f5ae8749647ad447463bcf1e9705f8d89b66c208342f5ea6

Observation 3f7246bd-1778-43cc-aa36-525c58dfdc26 · outbound

This paper cites Learning to Reason without External Rewards.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Learning to Reason without External Rewards

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.440378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.440378Z digest=sha256:a80ab39d8db71a842e59ba71c953ec05735618e23349185a9ce75cf4d5cd826f

Observation 0d4aa580-82cb-4bd6-ab26-895ceb6edbdb · outbound

This paper cites RMB: Comprehensively Benchmarking Reward Models in LLM Alignment.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.443699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.443699Z digest=sha256:d72e263c34acd4b87bb0b8873399879148cc0ea094bc59bb3eb61ea06ee00173

Observation d9cb06d2-5012-4061-985a-434b37d473ca · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Reinforcing General Reasoning without Verifiers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.447032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.447032Z digest=sha256:3da23f8b2e5e3ce7ce8783cc0f1569157a66c2093bdaa9c4a9a646768709ca2d

Observation 45e30223-ad4e-4469-8a6d-66eca0c5373e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Fine-Tuning Language Models from Human Preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.450746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.450746Z digest=sha256:e9469b295bbe800ec6bec5aa0f5854639d50104e3e2fcfbe88ba2175e51413a2

Observation a7440627-50e2-427f-a08f-f9acecb18f6d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Training Verifiers to Solve Math Word Problems

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.378319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.378319Z digest=sha256:63e5be240aed312e5b6018f1c13dbcbc0c6a609333dc43aa05df7e5f31721835

Observation 6a8794c5-e013-4c25-839e-2ec8f30c545c · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.400168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.400168Z digest=sha256:4988a65a5d77457785c304fe82d3d400443270a8d86ddbae563c462ab2b07c2f

Observation cb4fb983-6682-461f-8570-5f87f50e3973 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.374231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.374231Z digest=sha256:2887ec31b4059f4c545018d3b4552f9cf35a67aefceed76e96bb79e6b5ccd1b8

Observation 6559628d-565c-4fa1-b7d0-9c0e26d7c9de · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.389900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.389900Z digest=sha256:f244883ddc5530c7807d9b3ea11eeef503348e7d07099cf87e4ce79da0ba9461

Observation be06d7d8-bd0f-4fba-8ed6-6c9a0d21f315 · outbound

This paper cites The Llama 3 Herd of Models.

URPO: A Unified Reward & Policy Optimization Framework for Large Language Models The Llama 3 Herd of Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:25.386031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:25.386031Z digest=sha256:ba3f30ecdc09b1a35f8b8b811d1df9175c1d038bb1c913476bb90f68504707da

Pith citing papers

Observation 27b9735b-8a94-4cb8-bfd3-6a74ed075534 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:19.399777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:19.399777Z digest=sha256:33b819b0d9463ff94e90b204e5a3ea09fad5277ac609300374f8363e7e03efb5

Observation 26c944a4-d31b-42fc-ac1f-727d2d44be7e · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.696133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:612a4c3b3d031ee42cb7f9ea6f78be842b1c216c0d5b605c7ea82feba0e76063

Observation 46e8ad6d-eb00-4005-812f-c4a412b6898d · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation URPO: A Unified Reward & Policy Optimization Framework for Large Language Models

Reference 199

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.584112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a11558d42c6e58c830d9314c0c7ae3136e086fc40993d399ab71d2856a1d2720