Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2506.18254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:53:04.646543Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:29:38.225670Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 79755e74-e6e5-458f-b5dd-2e89aacd48ca · inbound
Can One Domain Help Others? A Data-Centric Study on Multi-Domain Reasoning via Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8969b9a4-7393-42e5-95c1-4d046f3c9338 · inbound
StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c1750ea-d176-45cf-a88a-bc2dd7c10d68 · inbound
An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e66c7c-7c58-4988-8903-f54b84e49500 · inbound
Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86079210-c35f-4451-aea5-043af72958f2 · inbound
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 054c9827-e6dc-4db5-8108-e7bc3b7f6fd8 · inbound
TRINITY: An Evolved LLM Coordinator RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2b790b15-16be-4b3e-9cfc-f1531b9f4188 · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c47942a3-ba62-454a-aa9f-655e9224bdec · inbound
Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a2be3a4-1eb2-4e55-bd07-8f2556bdeb5d · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation edb6943d-e58d-4e84-8861-e882d01b06da · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 85edc4c9-1f3d-4cd6-af42-7efe2ab62f6f · inbound
G-Zero: Self-Play for Open-Ended Generation from Zero Data RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72cb19f3-c96d-4734-b47f-68a9b6dacd23 · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 901cfdcf-ce73-4be3-b531-742fd2ea79ef · inbound
Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de309758-8d6d-4c05-9cdc-a0fcca03b0d7 · inbound
Selective Off-Policy Reference Tuning with Plan Guidance RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5447ff0-8b0c-4b8d-ab76-5f6809095f65 · inbound
LamPO: A Lambda Style Policy Optimization for Reasoning Language Models RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2571351c-abdd-42c0-88d9-539d5988e6cb · inbound
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01487990-7a6a-4622-869a-f7e4c6e99a89 · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67d565f6-b0ee-485c-81a4-c78fa1535f7f · inbound
Trust Region On-Policy Distillation RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b35d5da-b98f-4acc-a2b8-543a0c429259 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e2ad807-18ad-4438-a044-3e8a0c8d781a · inbound
Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b3e06ca5-3c98-46a5-a471-ab315d6f3d88 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 91a78fa8-cb7c-40a4-be49-3ff58b70e6b9 · inbound
Sakana Fugu Technical Report RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 243
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ecd9f5c-fb34-47fb-93e2-e73d4881b216 · inbound
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24dbb89-1a3f-4931-9a84-47e8a977714b · inbound
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4403b786-7262-4849-adda-33fb400c4398 · inbound
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design RLPR: Extrapolating RLVR to General Domains without Verifiers
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.