Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.02835.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.359664Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T23:57:29.062496Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e582f9d6-11e4-4597-8b15-4d8479bce771 · inbound
Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8f6167-1acb-41c8-b638-8ff1fd1d1978 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aecb0807-75d4-4cd1-9858-fbf31d88585b · inbound
Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · inbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · inbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb148a72-3212-457f-8a67-f271635f405f · inbound
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e55930e-bd20-4101-a527-91a1f656848d · inbound
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · inbound
Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e4116c8-75f2-4c32-9b37-a1a6eb7e9797 · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dc8ab51-8e38-4a81-9a20-817baa9ba334 · inbound
Reward-Aware Trajectory Shaping for Few-step Visual Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38709681-ae0f-4c6b-be99-0233d59b65eb · inbound
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 876803e7-4eab-4ff8-8f2a-be4e5e29c7ac · inbound
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61e56765-442a-4be9-8d1b-b0dac92a37be · inbound
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe7449d0-f8bb-4331-ae5c-70692f8d852c · inbound
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fce63959-381a-47b7-97f9-ea0711fd5167 · inbound
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2321188-055e-40e6-8c19-fc000d9dd49c · inbound
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc2c86d0-8265-4c27-8395-a029444a507c · inbound
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4eadc967-8a10-4d7a-aed1-4aec89d466db · inbound
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb54247d-0abd-426b-9a4a-ad51fb3942ac · inbound
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 078cae58-a97a-4fe8-8ac4-a95bd8ec0d75 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 277
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f68406-5f9f-496c-b3a2-8f6e11c5da4b · inbound
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.