Pith. sign in

Paper Citation Record · LEDGER

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2505.02835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02835 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.359664Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:29.062496Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e582f9d6-11e4-4597-8b15-4d8479bce771 · inbound

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games cites this paper.

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:04.359664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:04.359664Z digest=sha256:803ef8786f7a4d5c754521d65f39db6904c53b967d223eb7d387f207ec44b35e

Observation 6c8f6167-1acb-41c8-b638-8ff1fd1d1978 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:22.005729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:22.005729Z digest=sha256:ceba7132f931717f3d4d8940abe25c149909fab94d1fa1c62c5b6ba80b765956

Observation aecb0807-75d4-4cd1-9858-fbf31d88585b · inbound

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning cites this paper.

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.180265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T13:00:27.232079Z digest=sha256:e28744c1b7b6cca72316fea6742433f635dc063d1322ea6bad65eb2bc2c02d30

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:ac434a0ec8612d77ee879a21e11731ba7ece3bb6a3bae67ca77da315fc154605

Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.958836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.958836Z digest=sha256:6e9367c816c501660e0ba53bc5043f94fa246467a4d08a12e1fe32160b5427b2

Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.091160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.091160Z digest=sha256:10a4fd66885b653b8416e6a22c8b00281b48133eb45d75a239756ea7d8eecae4

Observation fb148a72-3212-457f-8a67-f271635f405f · inbound

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models cites this paper.

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.232604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.232604Z digest=sha256:d0cc52487d928ee8966b41ce0be374c5441e804d121ca52c98e5d33e8de44da2

Observation 0e55930e-bd20-4101-a527-91a1f656848d · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.297932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.297932Z digest=sha256:a142ced7177cf1a3653592d613481515efbe941f5bd130f83e374a2b0dedf770

Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.856440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.856440Z digest=sha256:e4fd6be033077a71da317e3b7072f63bf1c34744a8ebfff59ba52f3652ac99d9

Observation 8e4116c8-75f2-4c32-9b37-a1a6eb7e9797 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.984367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:21cc14de9e57df991289f984ce44d4c533adbba4dbfbb6c1a7537dca5b7b81a7

Observation 3dc8ab51-8e38-4a81-9a20-817baa9ba334 · inbound

Reward-Aware Trajectory Shaping for Few-step Visual Generation cites this paper.

Reward-Aware Trajectory Shaping for Few-step Visual Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.293870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:43:55.499474Z digest=sha256:89726c2f8fd20e5d912c041f69f4a4f58b587fbcd6d138f7b1848d97372fa191

Observation 38709681-ae0f-4c6b-be99-0233d59b65eb · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.799793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:2ff66c55a011714cec12f003fbef58b1cc50cf0d0b4c96f00694b55fd2978f3d

Observation 876803e7-4eab-4ff8-8f2a-be4e5e29c7ac · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.471032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:4ddc26b743cd4320aebef18445c05191971132949c817895493860ababa0c68c

Observation 61e56765-442a-4be9-8d1b-b0dac92a37be · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.874973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:01ed70568b7d34b7c6d1c8916c9beb43b6764a25f5bb6818fe1814797d92628d

Observation fe7449d0-f8bb-4331-ae5c-70692f8d852c · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.508835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:87c8e51112e1b91bfac6ad1fb89aec15e5364c3c57661c8f72519033fb469457

Observation fce63959-381a-47b7-97f9-ea0711fd5167 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.385559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:6986c28c32218b8c166bbb86ae7fb7a4a10e65643b852f9dc104394377086f87

Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.224742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:917d48e580e44e2181cef77f2b3ff21351d1aae978bfe95183b7e5facc16ce3d

Observation c2321188-055e-40e6-8c19-fc000d9dd49c · inbound

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs cites this paper.

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:52.586888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:57:57.791351Z digest=sha256:39768baa9d5f210a8cee9b23865f9d36ead5d98a87e2d09804c024e63cc06118

Observation bc2c86d0-8265-4c27-8395-a029444a507c · inbound

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization cites this paper.

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:12.728691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:19:57.063802Z digest=sha256:1f4dd72c6577ee3dd76e6215129b4d921cbf569eaffe1a3237f6429cbf4ad820

Observation 4eadc967-8a10-4d7a-aed1-4aec89d466db · inbound

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning cites this paper.

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.573718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:53:55.407461Z digest=sha256:dfc9806fc328e00b1b00e0c80a570622c3a97ac5a74080d08f8b03e13a679f29

Observation fb54247d-0abd-426b-9a4a-ad51fb3942ac · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:29.063754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:950e5d5b80c02df25a35f4cdfb068357f643edb4af7749be23a6fdadc5e795cc

Observation 078cae58-a97a-4fe8-8ac4-a95bd8ec0d75 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 277

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:4703aa78b65d0a1aa324ab034b1d6b9cd7044f7f2a2252d8ee32e57e486f24eb

Observation 59f68406-5f9f-496c-b3a2-8f6e11c5da4b · inbound

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification cites this paper.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T14:08:11.243120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:08:11.243120Z digest=sha256:c7f9517d5efc87d59a6e27f4f97aced8d0bde6f1f802a58009ff0bcce7fe58cc