Pith. sign in

Paper Citation Record · LEDGER

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.02835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02835 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.359664Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:29.062496Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e582f9d6-11e4-4597-8b15-4d8479bce771 · inbound

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games cites this paper.

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:04.359664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:04.359664Z digest=sha256:90e0d0918635dc7565b16e66221a89c1524d34069d7ceafcbfc86f89d874021e

Observation 6c8f6167-1acb-41c8-b638-8ff1fd1d1978 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:22.005729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:22.005729Z digest=sha256:57b02fba2853f3356dcb3b080fa7094b6cabb52e25e93170fdd286230413e145

Observation aecb0807-75d4-4cd1-9858-fbf31d88585b · inbound

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning cites this paper.

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.180265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T13:00:27.232079Z digest=sha256:c031119df702fc5108661df51d7be978b8f685cfcf713bd92c0553dfb1478a1b

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:f2ea15529c6370807da9f806a54d80f7c8d315ce8fc9c2a1e873c6b9f0ad4283

Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.958836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.958836Z digest=sha256:329631c843313b71363e86d1b53f91bf6cc1bfa568f85874ddd9debfd50d7cc1

Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.091160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.091160Z digest=sha256:ab55fee43d32c901860bc83278c7302d4a40f9e8693d917f5106b2f4d2ab5a5c

Observation fb148a72-3212-457f-8a67-f271635f405f · inbound

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models cites this paper.

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.232604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.232604Z digest=sha256:b89ac40566bda66742a41751402410d3e587c0ee146b5dc8cc712b60347562e9

Observation 0e55930e-bd20-4101-a527-91a1f656848d · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.297932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.297932Z digest=sha256:81c47038f16f1a72d08fe848a23693b956fa0d57e00dad7612a06be4f57badc9

Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.856440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.856440Z digest=sha256:fa80b103ec910ed088b3f3c40e636305a99e35b9b99a4519402a523fbe006fc7

Observation 8e4116c8-75f2-4c32-9b37-a1a6eb7e9797 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.984367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:3c251e59ce3cc779a0a958bb0235c11f65ef7b3d39c7d1981155f7ab31a9a402

Observation 3dc8ab51-8e38-4a81-9a20-817baa9ba334 · inbound

Reward-Aware Trajectory Shaping for Few-step Visual Generation cites this paper.

Reward-Aware Trajectory Shaping for Few-step Visual Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.293870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T11:43:55.499474Z digest=sha256:4d495b9573ec785ca42639f39e15eaa43f58a846d1c5fd2bb977a505611ba937

Observation 38709681-ae0f-4c6b-be99-0233d59b65eb · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.799793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:9f44ea2a769e532eb3a3f517cd5e5809e08e9031ec236ef79da9856c96b568cf

Observation 876803e7-4eab-4ff8-8f2a-be4e5e29c7ac · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.471032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:320ca9cef1fe70a37e3d589965df0bd0d740e47e8ba000c37bc6040c3590c611

Observation 61e56765-442a-4be9-8d1b-b0dac92a37be · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.874973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:7c380ba820b78c5fc45ebd33f01cf3ac8957cb73802d80c9ee65c91efa5c9ecb

Observation fe7449d0-f8bb-4331-ae5c-70692f8d852c · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.508835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:cb7f0e2fbc18334e991cfa22668d03f0277fdfdb131fe90b30f8380c2155f77d

Observation fce63959-381a-47b7-97f9-ea0711fd5167 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.385559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:6a3286e47f4923916e8084d01aa4e5f60aebfdf336250b4da917ec0e5ab5d887

Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.224742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:1dd5efc056ebe87e2848804daadc46c8d48c0523fc7db7bfbcf159605dc4d981

Observation c2321188-055e-40e6-8c19-fc000d9dd49c · inbound

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs cites this paper.

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:52.586888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:57:57.791351Z digest=sha256:aca7fcfe504eec0e04997cc6c3ddb65948d51098baf5663170a74f1348e5601f

Observation bc2c86d0-8265-4c27-8395-a029444a507c · inbound

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization cites this paper.

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:12.728691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T07:19:57.063802Z digest=sha256:ecd0d1bbe37b082d7be023723b38f10f75a87eae010b0d5a3bcfbe97fa88696b

Observation 4eadc967-8a10-4d7a-aed1-4aec89d466db · inbound

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning cites this paper.

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.573718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T22:53:55.407461Z digest=sha256:92bbfd55129af37356976d6685b1c67d344596eff6ee0e6e3348166968a48900

Observation fb54247d-0abd-426b-9a4a-ad51fb3942ac · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:29.063754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:d7b1fa78eefd977107c499c34cc8d94a387bc6f918da39b633c2a5ba6ee415ad

Observation 078cae58-a97a-4fe8-8ac4-a95bd8ec0d75 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 277

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:ccf27f3ad449f1bf11c180ad16df1eb807e507c05aa5eb97dfffe669358f8a27

Observation 59f68406-5f9f-496c-b3a2-8f6e11c5da4b · inbound

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification cites this paper.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T14:08:11.243120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:08:11.243120Z digest=sha256:576837ec45c16d94d4faf187e2d1d0c00c06e76fafe1de2837ad6d9136078c30