Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:51:25.450746Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 3 inbound Pith citation observations for arXiv:2507.17515.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:51:25.450746Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:14:19.399777Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:56:13.582742Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 952b4499-3103-4a5f-9bad-4a9ad8609a59 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad92ed78-b216-439c-a5d4-87b4e42ae997 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e04d0d6-e7a9-4627-84d0-9a1e08554e4c · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Skywork Open Reasoner 1 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f15869-2750-4dc7-b7ce-191bdb2543a1 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7340c7-0fdd-4941-8604-8eeb0c96109e · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c407c9-2d4f-4be5-ac1f-0a661227b2ad · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e667bb08-d015-43f3-8595-c926925194cd · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Inference-time scaling for generalist reward modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f9cf97-8a36-4916-b2fa-b57603f5e83a · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Decoupled Weight Decay Regularization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf7c51d-776f-471c-bb18-ed29a60eb10f · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models BPR: Bayesian Personalized Ranking from Implicit Feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c42e392-b627-4cc1-bfde-6e24ddda2fc1 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Proximal Policy Optimization Algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8941b72-95c4-4586-a154-9027845a5560 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2a5dc6-36d6-4581-814e-d8c97300f54d · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Qwen2.5 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989ca361-9d54-4bfb-9821-11b471c1ccc4 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c88954-0191-4d73-9788-9c3ebca59d6d · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e281e93-06b0-4e24-8a8f-50013c59e1b2 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7246bd-1778-43cc-aa36-525c58dfdc26 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Learning to Reason without External Rewards
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d4aa580-82cb-4bd6-ab26-895ceb6edbdb · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RMB: Comprehensively Benchmarking Reward Models in LLM Alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9cb06d2-5012-4061-985a-434b37d473ca · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Reinforcing General Reasoning without Verifiers
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e30223-ad4e-4469-8a6d-66eca0c5373e · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Fine-Tuning Language Models from Human Preferences
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7440627-50e2-427f-a08f-f9acecb18f6d · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Training Verifiers to Solve Math Word Problems
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a8794c5-e013-4c25-839e-2ec8f30c545c · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb4fb983-6682-461f-8570-5f87f50e3973 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6559628d-565c-4fa1-b7d0-9c0e26d7c9de · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be06d7d8-bd0f-4fba-8ed6-6c9a0d21f315 · outbound
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models The Llama 3 Herd of Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b9735b-8a94-4cb8-bfd3-6a74ed075534 · inbound
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c944a4-d31b-42fc-ac1f-727d2d44be7e · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 46e8ad6d-eb00-4005-812f-c4a412b6898d · inbound
Trust Region On-Policy Distillation URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.