Pith. sign in

Paper Citation Record · LEDGER

RLPR: Extrapolating RLVR to General Domains without Verifiers

As of 19 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 25 inbound Pith citation observations for arXiv:2506.18254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18254 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:01:06.335633Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:04.646543Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:29:38.225670Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce1e0782-368f-4239-bf9f-e56f7cfbb7bd · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

RLPR: Extrapolating RLVR to General Domains without Verifiers The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.238527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.238527Z digest=sha256:3ad996b4393f73549d993592c50fded4ed65079dddd2fe9d36ee95a958d25a3f

Observation 83bf481e-c9db-43c0-9005-3ff1c6c63510 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RLPR: Extrapolating RLVR to General Domains without Verifiers DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.262860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.262860Z digest=sha256:ff32cdd301f3faf4841c664355c921c40226e4b632b0d8a0abbb2a2e853d8405

Observation dd05f7ee-8c26-440d-9aff-3e0904018d19 · outbound

This paper cites The Llama 3 Herd of Models.

RLPR: Extrapolating RLVR to General Domains without Verifiers The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.274204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.274204Z digest=sha256:c334da263800fc2e25467dc4a9de96d37e1b839bc03b78700acc4e87ef035731

Observation 915432fb-4374-430c-ab42-05abd501686e · outbound

This paper cites Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge.

RLPR: Extrapolating RLVR to General Domains without Verifiers Andreas Hochlehnert, Hardik Bhatnagar, Vishaal Udandarao, Samuel Albanie, Ameya Prabhu, and Matthias Bethge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.278630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.278630Z digest=sha256:6c2c4ee759627636f8b99fd5399b9e3ea5d0953d5213ba93b0295392dcaa37ac

Observation 97457065-455a-4e5c-b1cc-e45db9ccbe59 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

RLPR: Extrapolating RLVR to General Domains without Verifiers REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.282949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.282949Z digest=sha256:0dfb81123678292ff16e5cd3077e0a220fbb02eebf8e47b0783248b1333ece33

Observation b093e1a4-50d4-4928-b73b-a30a6c323577 · outbound

This paper cites OpenAI o1 System Card.

RLPR: Extrapolating RLVR to General Domains without Verifiers OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.287522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.287522Z digest=sha256:a80bf7644f8af4ac1d3877fc6f20d77b16f45b6ccd4cb5669dca0b6ca9e78f9f

Observation c83b5507-431e-4b0f-8ebe-693c7648aaa6 · outbound

This paper cites Generative Reward Models.

RLPR: Extrapolating RLVR to General Domains without Verifiers Generative Reward Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.296452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.296452Z digest=sha256:6564b8847ba1b09dbdccbf368c9ab1a95583131c85c797160af0fe8d71790cff

Observation a89878a5-ce4a-4ccd-a40a-f1246badded1 · outbound

This paper cites David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R.

RLPR: Extrapolating RLVR to General Domains without Verifiers David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:06.766493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:01:06.300979Z digest=sha256:e494d2f8fd17985be46e48ed11918ad383964805ab0058c590441e3a36fbd3d2

Observation 55a7c6c6-2f04-4b3b-9f3b-ea12f01412be · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

RLPR: Extrapolating RLVR to General Domains without Verifiers GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.305112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.305112Z digest=sha256:e6b7ccb972b104b2a36cc9e897efda23c7ce48f0e9e9b7374ce6b70a6083816d

Observation 58473a10-9b6e-429b-8fd1-3e51499dcf72 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RLPR: Extrapolating RLVR to General Domains without Verifiers HybridFlow: A Flexible and Efficient RLHF Framework

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.309221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.309221Z digest=sha256:483f3b73464b10229b9cb9ff0def93f8a9418b1c506cdd2919b9d7b5ed742d78

Observation eb696aa5-43c6-40a4-a78f-a3e0591b508b · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

RLPR: Extrapolating RLVR to General Domains without Verifiers Gemma 2: Improving Open Language Models at a Practical Size

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.313376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.313376Z digest=sha256:0fd7170d3c98c59979066331ad043a55330726a333bc69172c0dc1f2c0d7d04c

Observation 7d7910cf-5702-4846-914c-ac909ebe0710 · outbound

This paper cites Qwen3 Technical Report.

RLPR: Extrapolating RLVR to General Domains without Verifiers Qwen3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.317339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.317339Z digest=sha256:db5e05e73eb2dac2f49a1573c63a424ae308981e3b148e5cb5b3ce399f3be1b1

Observation ba65bfeb-b3e1-4126-9b51-18e832c67aef · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

RLPR: Extrapolating RLVR to General Domains without Verifiers MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.321596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.321596Z digest=sha256:c8e46801b82250b406c0da635a4b875b5351661e26aa405ec29a98e71b8b7c00

Observation c556a74e-9b47-48f6-9d73-113596c00028 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

RLPR: Extrapolating RLVR to General Domains without Verifiers SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.326330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.326330Z digest=sha256:b23fcb809bb16ff8440d979b188e41e11046f47b08126a0ffa778cc35e07d991

Observation cf019ec0-544f-4ace-a61c-286476ddd018 · outbound

This paper cites Reinforcing General Reasoning without Verifiers.

RLPR: Extrapolating RLVR to General Domains without Verifiers Reinforcing General Reasoning without Verifiers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.330755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.330755Z digest=sha256:b01062d9f368408c4f569049a2964e2cbb5fd02562e28d7d6f030f3cb2974f44

Observation 8ed8ba76-2a25-42b4-b29a-c53e311cf65c · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

RLPR: Extrapolating RLVR to General Domains without Verifiers TTRL: Test-Time Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.335633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.335633Z digest=sha256:02378b89e8a7c51f4b65a0a37cf12b68362dd6eceef2669288e392b9822a2c16

Observation 542646f7-09b1-40cf-9517-e33f32fbd12e · outbound

This paper cites doi: 10.1016/S0031-3203(96)00142-2.

RLPR: Extrapolating RLVR to General Domains without Verifiers doi: 10.1016/S0031-3203(96)00142-2

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.243921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.243921Z digest=sha256:e25b5544dfaf0a92b213a0bbea34b1050b2c624c85ce8b1127e6d128ccbc49d6

Observation d9a32af9-867a-4e69-a782-255be97d715d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RLPR: Extrapolating RLVR to General Domains without Verifiers Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.258076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.258076Z digest=sha256:0a306ad7029180d0d5d79a0c511964cfdd7aefe01c0aace7ade4ab639950af5d

Observation ff3f062b-a510-4fb4-a47e-b4f2e32cd0a3 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

RLPR: Extrapolating RLVR to General Domains without Verifiers Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.292222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.292222Z digest=sha256:0be7bff4d973503a6c2bbee914640e1a4a5ddb9fa57ebadb09251063ac8fc9e4

Observation c1605fa3-3497-438e-af25-46e2dd63d94d · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

RLPR: Extrapolating RLVR to General Domains without Verifiers TheoremQA: A Theorem-driven Question Answering dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.253262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.253262Z digest=sha256:592a15573ddf0a484189255911d8a5cf1b9fd576008ae64111379179a3eb58fa

Observation b70abe72-b100-4552-9a3a-9f1fec8ff71d · outbound

This paper cites an unresolved cited work.

RLPR: Extrapolating RLVR to General Domains without Verifiers Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:06.780875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:01:06.267871Z digest=sha256:b6af2a434f05743c925176e9cd5c4ec690b7dc834a42eb01effdb2eb56501b8f

Observation c3a115a9-4e4a-4d45-8caa-b5e6e63d6d05 · outbound

This paper cites FullStack Bench: Evaluating LLMs as Full Stack Coders.

RLPR: Extrapolating RLVR to General Domains without Verifiers FullStack Bench: Evaluating LLMs as Full Stack Coders

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:06.248686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:01:06.248686Z digest=sha256:fb1139c1d1360870037e5ae3bbc13c0508d0233814f99c3da6becc19871aaead

Pith citing papers

Observation 79755e74-e6e5-458f-b5dd-2e89aacd48ca · inbound

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning cites this paper.

Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:53:04.646543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:53:04.646543Z digest=sha256:6ed53cefe94b5b30ee213103c71f8a433e7abd658393cddc1aba94bdd7001837

Observation 8969b9a4-7393-42e5-95c1-4d046f3c9338 · inbound

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models cites this paper.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:17.022564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:17.022564Z digest=sha256:b45c648fcbaa4bc3ae4bf18abd171991db334c975de28329966e81a167dbb5e4

Observation 8c1750ea-d176-45cf-a88a-bc2dd7c10d68 · inbound

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal cites this paper.

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T23:25:50.287093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:25:50.287093Z digest=sha256:62b4aac3f208ffd5bf53a3591b75856d4e51990aa8242851d78dbc52847d45c8

Observation 06e66c7c-7c58-4988-8903-f54b84e49500 · inbound

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction cites this paper.

Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:07:12.004152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:07:12.004152Z digest=sha256:a7d40bf97ff28d3914447dfe91b68a95685bbbc36f12c674a6eef06673ca67e3

Observation 86079210-c35f-4451-aea5-043af72958f2 · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:46:34.115346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:7c619c1dc47290795749f5ec9606b4c0d33a1a742a424d89dda2cdcd7965c86e

Observation 054c9827-e6dc-4db5-8108-e7bc3b7f6fd8 · inbound

TRINITY: An Evolved LLM Coordinator cites this paper.

TRINITY: An Evolved LLM Coordinator RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:21:24.781976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T01:19:10.420266Z digest=sha256:45409b2d41a9bdc7c6693cf5e6608e37419b2baf3c01295c029e180aaba347b3

Observation 2b790b15-16be-4b3e-9cfc-f1531b9f4188 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.963670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:9f23e7d66bc4370132dca88f4d4a8cdd7eac1f90ffd90e6dc31b86825cab7eb6

Observation c47942a3-ba62-454a-aa9f-655e9224bdec · inbound

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards cites this paper.

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.178591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:49:00.343580Z digest=sha256:361f26873951e1d98ecc0a0143e293796b4c311e64af0a47d8ecbf513d839076

Observation 3a2be3a4-1eb2-4e55-bd07-8f2556bdeb5d · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:55.287022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:f81e06ae5314d401d3fc99045e4d89d49a47ac1b43558cc5a8d891b0c51b6060

Observation edb6943d-e58d-4e84-8861-e882d01b06da · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.503384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:b7815e1a219b81eb49f968c359b1304672770b17e50ff947dee537cfa9a3b6f6

Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · inbound

G-Zero: Self-Play for Open-Ended Generation from Zero Data cites this paper.

G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:24.558508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T03:39:40.780801Z digest=sha256:6aac74e6fe033dea2bdf5decacc274fe5162c0f8498fbc79ff6cd5641fc189ab

Observation 72cb19f3-c96d-4734-b47f-68a9b6dacd23 · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.998128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:040a6b179072434a8256890c3fbc02300daa8eb110948c5b402862d5a04a0d98

Observation 901cfdcf-ce73-4be3-b531-742fd2ea79ef · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:05.185587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:f64563567a678cd730a07bf990d1e7af8ad219dfee8c67340ad9883846dcdbc9

Observation de309758-8d6d-4c05-9cdc-a0fcca03b0d7 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:59.347610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:3fef7a4a1db614c996b56f4a6035598437ca1a5b432bd96000e2f8e751d33ddd

Observation d5447ff0-8b0c-4b8d-ab76-5f6809095f65 · inbound

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models cites this paper.

LamPO: A Lambda Style Policy Optimization for Reasoning Language Models RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:13:58.419624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T05:11:35.785227Z digest=sha256:bf4a0ef88290b8361ab0b6593f38d00ff39b0f710f8e9f63e6efa5f1916b6e86

Observation 2571351c-abdd-42c0-88d9-539d5988e6cb · inbound

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination cites this paper.

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:47.141445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T23:08:05.597810Z digest=sha256:fbd9306104f7d059051d25e0ba71e51adbea08f53fcd68975c73b892497af952

Observation 01487990-7a6a-4622-869a-f7e4c6e99a89 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.047073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:6b43d8f3e413987e91f34fd142f7f7aa300e9374189624e04b3f688e8ef51f6e

Observation 67d565f6-b0ee-485c-81a4-c78fa1535f7f · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.380190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:a0441708478f78f8250aacf8a2495557f2f69f36723e373fb569be5b228d737c

Observation 7b35d5da-b98f-4acc-a2b8-543a0c429259 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:40.808650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:03144c06adaee77ea7549c97b5174d9f0ba552e37e8776c574326d240714e6d9

Observation 1e2ad807-18ad-4438-a044-3e8a0c8d781a · inbound

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short cites this paper.

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.542547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T17:18:03.459289Z digest=sha256:fef4a62139ae79184fb5cd7e3186f753ff1279524755e39f7d451d88e2ed52ce

Observation b3e06ca5-3c98-46a5-a471-ab315d6f3d88 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.006584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:cfbe1af0d2c20436532449ae9c5fde13b698dd6fde365abd60028e906410a46f

Observation 91a78fa8-cb7c-40a4-be49-3ff58b70e6b9 · inbound

Sakana Fugu Technical Report cites this paper.

Sakana Fugu Technical Report RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 243

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:29:38.227309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T14:22:37.596720Z digest=sha256:2c61ab90efca84cd536d85abf8fa794ec26ae90c57157c59a57bb149b1686777

Observation 3ecd9f5c-fb34-47fb-93e2-e73d4881b216 · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:04.176578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:04.176578Z digest=sha256:712cacd9f7f700c86734d9cc6b13d3b08a3d0e986ed449276234232f853a40d7

Observation b24dbb89-1a3f-4931-9a84-47e8a977714b · inbound

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning cites this paper.

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:53.821375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:53.821375Z digest=sha256:b83a16fa40422765cb8a5f1cd2073de32faa828a4fcf9ff4c72d92faa5a7f8d8

Observation 4403b786-7262-4849-adda-33fb400c4398 · inbound

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design cites this paper.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.704541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.704541Z digest=sha256:345e453e707486d6b8838199b5134a82a9935d405aa8961038e2359b8b3e6297