Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2503.23829.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:43:44.467130Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 4bcdf1b4-c21a-4413-bc33-70e03eb18358 · inbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e14f3f10-e97e-4664-b51d-4925cd83e528 · inbound
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e84209f-bbff-45ee-a5d2-77e813fae459 · inbound
General-Reasoner: Advancing LLM Reasoning Across All Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c61492-c207-493b-b743-8d840017c1c7 · inbound
Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation facf22d4-454e-4613-b858-c10da0a74038 · inbound
Reinforcing General Reasoning without Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf464214-97c7-4498-990d-9e23b52bfaea · inbound
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645593a1-35a9-47e3-9965-95a705d71bbf · inbound
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96aebd1-5915-490a-b807-7b216794f814 · inbound
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4c3211-c002-4827-8936-5f75068aeef6 · inbound
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f20c883-72af-4be5-b670-ffd717f77947 · inbound
Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f5da95-4b79-4e1e-ba65-f83acba760b6 · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45d137f3-ca87-4e2f-b01a-e084297c6e34 · inbound
From Emergence to Control: Probing and Modulating Self-Reflection in Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7745fa08-d56a-4cc0-aa9c-f06996f5d17d · inbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65d37c96-ba03-4096-9c57-857ff75d8237 · inbound
Energy-Based Transformers are Scalable Learners and Thinkers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ff0fa0-99bd-4b98-b2c2-b7454de116c4 · inbound
One Token to Fool LLM-as-a-Judge Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6983d6-f37e-4f3d-acf4-257c3df171fc · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0ae26018-5913-4345-bf2f-319ccf727136 · inbound
The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b93d9e6-be98-401e-aee1-b81ff7a36c70 · inbound
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459fb54b-206d-4965-8a55-b0e54227814b · inbound
Reverse Browser: Vector-Image-to-Code Generator Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 500c22dd-a511-42f3-b0b0-879be3b0da6b · inbound
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21974947-640f-41f2-956a-4a4e42720073 · inbound
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5678ad83-14b1-4b82-b552-9a4d16097830 · inbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6d6e0de-6812-4916-84ef-98aeff15b891 · inbound
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 49959a33-24e4-4b0d-861a-15ade233c683 · inbound
DVPO: Distributional Value Modeling-based Policy Optimization for LLM Post-Training Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 553adc0d-e76a-4293-a47a-83b6d66855ed · inbound
CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 53947a8e-5330-48d8-88c6-f47d8edd3061 · inbound
Specificity-aware reinforcement learning for fine-grained open-world classification Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1570ade8-c9af-4dc6-bf9f-63b133f018d8 · inbound
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3b7189d9-5807-4d9e-aa2f-658142edfee5 · inbound
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62ac229-1b51-47de-b77c-2b9ccaf8fd5b · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb70e96-f794-4e69-a379-9ec69c09355e · inbound
Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 369ad831-a096-4da0-8abc-13cba5e99796 · inbound
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5803ac88-fbb8-4ee0-bba0-c31b7b85402d · inbound
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b3786b5d-2427-4e79-8e10-970e17cfd6c1 · inbound
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1dd83977-cef3-48d3-8789-8ceb9697ab7a · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6e20c37c-f7c7-41eb-ab24-00109f9543f2 · inbound
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2ca204c-6241-40e0-afc0-1f11f80191cb · inbound
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1f31b556-3816-4be0-bc71-8924d6b7bc72 · inbound
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bfe80b35-bcae-4429-a91f-e2caf0a75a60 · inbound
CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f305aa37-c9eb-46a2-9aa9-e80b02291d7d · inbound
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0569465f-718a-4fb8-91cd-7afc15e533d5 · inbound
GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69cbd603-7104-40b4-853a-246a243078a3 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 55788ff5-3d72-4a1a-888c-f2014b89a2f6 · inbound
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bb171b7-5dd2-4e15-b5ba-78bbaeb0644f · inbound
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dcf008c4-a9b1-4f38-9e35-7bf8114e9a0c · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61f6e4fe-1eae-489a-950e-6014c1940867 · inbound
Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c369fed-3427-4cc4-84b1-b2f2437d9bc2 · inbound
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66b68307-6bdc-4878-ab54-31e38b2cac55 · inbound
MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e9b7ae-8a61-4a64-b77f-673325afb5da · inbound
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2707c228-3d18-4c3a-96b0-dd1d011bc224 · inbound
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2611e7-fe7e-4729-8abe-f6948bf32ead · inbound
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c91d0b00-6ba4-407d-9136-bf043f5db81e · inbound
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.