Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:01:06.335633Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 25 inbound Pith citation observations for arXiv:2506.18254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:01:06.335633Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:04.646543Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:29:38.225670Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ce1e0782-368f-4239-bf9f-e56f7cfbb7bd · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bf481e-c9db-43c0-9005-3ff1c6c63510 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd05f7ee-8c26-440d-9aff-3e0904018d19 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 915432fb-4374-430c-ab42-05abd501686e · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97457065-455a-4e5c-b1cc-e45db9ccbe59 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b093e1a4-50d4-4928-b73b-a30a6c323577 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers OpenAI o1 System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c83b5507-431e-4b0f-8ebe-693c7648aaa6 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Generative Reward Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a89878a5-ce4a-4ccd-a40a-f1246badded1 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 55a7c6c6-2f04-4b3b-9f3b-ea12f01412be · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58473a10-9b6e-429b-8fd1-3e51499dcf72 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers HybridFlow: A Flexible and Efficient RLHF Framework
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb696aa5-43c6-40a4-a78f-a3e0591b508b · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Gemma 2: Improving Open Language Models at a Practical Size
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d7910cf-5702-4846-914c-ac909ebe0710 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Qwen3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba65bfeb-b3e1-4126-9b51-18e832c67aef · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c556a74e-9b47-48f6-9d73-113596c00028 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf019ec0-544f-4ace-a61c-286476ddd018 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Reinforcing General Reasoning without Verifiers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed8ba76-2a25-42b4-b29a-c53e311cf65c · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers TTRL: Test-Time Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 542646f7-09b1-40cf-9517-e33f32fbd12e · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers doi: 10.1016/S0031-3203(96)00142-2
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a32af9-867a-4e69-a782-255be97d715d · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Training Verifiers to Solve Math Word Problems
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3f062b-a510-4fb4-a47e-b4f2e32cd0a3 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Solving Quantitative Reasoning Problems with Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1605fa3-3497-438e-af25-46e2dd63d94d · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers TheoremQA: A Theorem-driven Question Answering dataset
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70abe72-b100-4552-9a3a-9f1fec8ff71d · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c3a115a9-4e4a-4d45-8caa-b5e6e63d6d05 · outbound
RLPR: Extrapolating RLVR to General Domains without Verifiers FullStack Bench: Evaluating LLMs as Full Stack Coders
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79755e74-e6e5-458f-b5dd-2e89aacd48ca · inbound
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8969b9a4-7393-42e5-95c1-4d046f3c9338 · inbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1750ea-d176-45cf-a88a-bc2dd7c10d68 · inbound
An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e66c7c-7c58-4988-8903-f54b84e49500 · inbound
Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86079210-c35f-4451-aea5-043af72958f2 · inbound
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 054c9827-e6dc-4db5-8108-e7bc3b7f6fd8 · inbound
TRINITY: An Evolved LLM Coordinator RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b790b15-16be-4b3e-9cfc-f1531b9f4188 · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c47942a3-ba62-454a-aa9f-655e9224bdec · inbound
Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3a2be3a4-1eb2-4e55-bd07-8f2556bdeb5d · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation edb6943d-e58d-4e84-8861-e882d01b06da · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 72cb19f3-c96d-4734-b47f-68a9b6dacd23 · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 901cfdcf-ce73-4be3-b531-742fd2ea79ef · inbound
Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation de309758-8d6d-4c05-9cdc-a0fcca03b0d7 · inbound
Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d5447ff0-8b0c-4b8d-ab76-5f6809095f65 · inbound
LamPO: A Lambda Style Policy Optimization for Reasoning Language Models RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2571351c-abdd-42c0-88d9-539d5988e6cb · inbound
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 01487990-7a6a-4622-869a-f7e4c6e99a89 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 67d565f6-b0ee-485c-81a4-c78fa1535f7f · inbound
Trust Region On-Policy Distillation RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7b35d5da-b98f-4acc-a2b8-543a0c429259 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e2ad807-18ad-4438-a044-3e8a0c8d781a · inbound
Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b3e06ca5-3c98-46a5-a471-ab315d6f3d88 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91a78fa8-cb7c-40a4-be49-3ff58b70e6b9 · inbound
Sakana Fugu Technical Report RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3ecd9f5c-fb34-47fb-93e2-e73d4881b216 · inbound
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24dbb89-1a3f-4931-9a84-47e8a977714b · inbound
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4403b786-7262-4849-adda-33fb400c4398 · inbound
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.