Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:29:01.735774Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 15 inbound Pith citation observations for arXiv:2505.12929.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:29:01.735774Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:52:02.124869Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T14:27:07.373132Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5526fc05-a027-4cc4-87c6-ddaddea30bdd · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5406fdbb-4131-4219-b078-e436f69483f6 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a56aad7-552c-4c90-a4cc-73f0a93fef00 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849d660e-a703-481c-a394-af141c2fa03e · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Monte carlo tree search boosts reasoning via iterative preference learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8918163f-31d7-4a78-a55c-f9f82be8e9ec · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Alphamath almost zero: Process su- pervision without process
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9da1503d-f1d6-4bd1-982f-54a9245f8d6a · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Let’s verify step by step
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1598d75-938b-4f66-afa4-48079d76a431 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6bca8f1d-3e62-4948-9d66-ab6019275a4f · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05325b6c-6c2f-4ed9-b818-fd78ae52ee50 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b6a274-cf18-4d2e-a71f-f91342c50a10 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Understanding R1-Zero-Like Training: A Critical Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead49eaa-8eb3-46c2-b3bd-ebb996ea348c · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86589244-9023-4876-b9c9-048fc525436b · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Deep reinforcement learning from human preferences
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eb4ed842-0e90-48ff-81a3-24067ce60bff · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Fine-Tuning Language Models from Human Preferences
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff14750-b579-4fef-8a0f-808dd0f1ef5c · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Learning to summarize with human feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 19befe5b-abac-4da2-b8f0-3019168b228a · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Training language models to follow instructions with human feedback
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6f80dba7-3e47-4faf-9b55-d81fbd7307fb · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0eb6887-8369-4124-bfd6-94ab3d1cf39f · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Language models are few-shot learners
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d9dff6a3-44c1-4cc3-99c9-622bed19f749 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs GPT-4 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8721c4e6-4de0-400d-88e9-f525a519c028 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5494e24-675c-4573-9b1e-238c749465f7 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a37f115-ddb2-40c6-8dbd-f4c0800be69e · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs The Llama 3 Herd of Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5ae833-4b33-4c01-a70a-33b0528dfe75 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Qwen Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 592b6e06-1238-4e85-a6d8-5d22647b2ba2 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb40335f-8ec3-4346-905a-1d47c932203b · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Qwen2 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c936a1ed-f186-4dc2-beef-6ff0e6f89a4d · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Gemini: A Family of Highly Capable Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b64ba0f-edad-4a25-93c6-609c5fca5bf0 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10b5731-abd3-4506-93b6-4454f4999754 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs The Claude 3 model family: Opus, sonnet, haiku
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 70e37f1d-b0ac-4b72-90ed-508ec8d1f208 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Direct preference optimization: Your language model is secretly a reward model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f45a84af-8051-43c6-9fe8-9a24b9753eb3 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs From r to Q*: Your language model is secretly a Q-function
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 886ec526-8660-4ab8-8745-dd8ab8c30c4d · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Is DPO superior to PPO for LLM alignment? a comprehensive study
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation af302d94-5822-44ae-9f6e-1e4bc9e9000c · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs DPO meets PPO: Reinforced token optimization for RLHF
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ca71bb59-21f8-4a8c-a71f-2b463adc45f9 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Sutherland
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 358cdf93-5346-44f3-85dc-ec67d831372a · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs ORPO: Monolithic preference optimization without reference model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b03a277f-c07d-4d39-9843-a34a48e52781 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2598b004-24fb-4140-9093-fafc0e0d9354 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs SimPO: Simple preference optimization with a reference-free reward
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c609134d-aa82-41d7-b185-fd9961827394 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Chain-of-thought prompting elicits reasoning in large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 382cad17-3053-43f5-b5a6-843a7bd4733f · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c454f95f-1eff-4265-8ec9-d772c025551e · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56f0902-ac21-4a42-b8e3-70542325af25 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9b3875-e80d-41f1-99ae-44df00a5b9fa · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a2ab91-fd47-4fbf-be6b-6453c3378dc6 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678ca9be-48c7-41f6-9b1e-2c2002903985 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d493bbd5-ac1e-424f-ad86-3f04f77d726a · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Attention is all you need
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7c8b96d-491a-4a48-9c27-be32b1f6ca01 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Hybridflow: A flexible and efficient RLHF framework
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 778f7272-2792-4308-8ac7-820d26be5555 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs On Memorization of Large Language Models in Logical Reasoning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd308106-049f-4688-bc4d-c98e4417cb56 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf63c99-bd86-4c73-b8d2-e1b74458d0fb · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs What is the name of this book? Touchstone Books Guildford, UK, 1986
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 166b8592-9003-4ddd-b4ad-2f844a56be24 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Meta-logical problems: Knights, knaves, and rips
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dc81e272-2755-4b6f-8cba-3e83b7eeea38 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1f3f9c9a-d5b5-41ca-b292-f76012178bd4 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Solving quantitative reasoning problems with language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2c0a9a57-5372-4268-9459-0029ae0f16da · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bcb8349-f2d1-4a3a-a853-68f3fa746092 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Approximating KL divergence
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 92e56b54-cf11-41e8-9be7-e9594d427974 · outbound
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs Lily is a knave or Lily is a knight
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e4be3531-b861-4d71-b485-7ec46ae5d5d7 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c743253d-9e0d-4fdb-aa86-b63a49ff7c4e · inbound
Sharpness-Guided Group Relative Policy Optimization via Probability Shaping Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 02ac3733-2f36-4d5f-8fbe-047402cdfd01 · inbound
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cad0b166-4da6-4db6-a43a-72de0c98a855 · inbound
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648c269a-14bd-455c-a677-c86f9c05996b · inbound
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4c3acee1-4909-4b4f-968c-40287fac6581 · inbound
Holder Policy Optimisation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ea8c1185-787b-401f-98a2-5f755ff490cd · inbound
Holder Policy Optimisation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6a685f5c-6fd3-43b6-8ad2-5f37239a67b5 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b10868ed-2f85-4319-9de5-65ce2c902b99 · inbound
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 131054e5-b69f-4763-ab6b-3bc8dd993145 · inbound
Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd100329-ae05-43c5-ba78-30f6b5c19ba2 · inbound
Not only where, But when: Temporal Scheduling for RLVR Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 76058b5a-fb62-4cd7-ac8a-b0925351b827 · inbound
Trust Region On-Policy Distillation Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 239
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d296df5d-1a88-43d9-bdd4-bdf8708b5267 · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation af7bf6e4-7168-4bd3-8cb3-3bc0f4d082bd · inbound
What are Key Factors for Updates in RL for LLM Reasoning? Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation de480b0e-a72d-4343-9e30-f44cbd760865 · inbound
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.