Pith. sign in

Paper Citation Record · LEDGER

Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2502.14191.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.14191 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:19.719944Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bbd64e70-7770-4536-8d60-02c2d38e2c38 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.167091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:54750413d8c5d0ef50e8428f7f0a88b9263bdc30c22e1d83ffa54c6958db0f62

Observation 1319239b-82a9-44d7-822c-8d284f2c19fd · inbound

RewardBench 2: Advancing Reward Model Evaluation cites this paper.

RewardBench 2: Advancing Reward Model Evaluation Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:22:16.608330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:18:03.965711Z digest=sha256:91ced24caa07ab827f16cf80b28699e0a60ac533d5113eecc36dc0c109d61350

Observation 4f95da42-0fbe-49f0-b7a7-b9d955e24ffa · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.719944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.719944Z digest=sha256:1e85b18e15615323df727528f1af8b9a53858039aaae0a223dc3a51d907ed7d7

Observation e2908560-f838-4916-a340-625aab08f338 · inbound

Activation Reward Models for Few-Shot Model Alignment cites this paper.

Activation Reward Models for Few-Shot Model Alignment Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.505501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.505501Z digest=sha256:edbcfba112d15f54b615c7b43f55431dd877339d25b51ce81a20d20cda269736

Observation 5cf78738-ebe9-47ea-8cec-8709444bdde9 · inbound

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations cites this paper.

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:27.091213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:27.091213Z digest=sha256:64f5e85656d44c7af4e32f7f5fa6963f1e3b3586b19454963db1abdec096070a

Observation 34d46e00-8265-46f7-9b5d-b82fb86938ac · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.967292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.967292Z digest=sha256:ac10fdd7747d4182bf70571c870d561348fdba6eb5e575b92ee6f43f8dfb1bad

Observation 9f2c75f4-4926-474b-a789-d457d18ac3e6 · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:32.626573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:32.626573Z digest=sha256:ae86a4a49fa0c34412615d8f58e345d84d185fcc595b6742cff7ffab7b3d2b46

Observation 109f52c6-f545-4e49-bd84-c8daf21b16f3 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.849378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:b5a4c5633070a27a0d32dbeb7ef2ddb41aac90e3eb76ad1c9f191135c80282aa

Observation ad5abe45-5d39-4f6f-90c2-04f93480da09 · inbound

Lost in Translation: Do LVLM Judges Generalize Across Languages? cites this paper.

Lost in Translation: Do LVLM Judges Generalize Across Languages? Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.856802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T02:18:17.579110Z digest=sha256:80d79f439d4db5336af9fd1e2318d0d893a9889802790896a4c55a162dd6d000

Observation a4daae0a-28c5-47f7-bf5e-79d00ad236eb · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.893380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:76f53a25b08b0e39e53960ed5ca8440ae211cbf2819971dc8a8f3c94089cd359

Observation 0301289b-53ba-42bb-a78a-9253bb63f7ef · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.512494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:19a6602da94f86effa62e6370ddafb9f1bd87b8753c41c16f1cd06e5e4c428f8

Observation def1d5e6-910a-4f5e-8dea-a6d3c493f332 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.316951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:aca22845f9896e050916c0a01ce91fa7fc1b1423153c25cc97d309c1d9f3bb2f

Observation 5a2bc0ff-d920-4e03-8077-660d03fc6295 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.318015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:e2bff807762686460bf4fdc5789ca0ed9e925713af7028dd379c775b7931663a

Observation 5c822b84-a066-493a-a62b-baf5fcfee596 · inbound

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents cites this paper.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:05:29.027260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T05:59:13.631078Z digest=sha256:18082ada06414635707e14e8b3023dfabeeccc980bf0a891815b81b46fd64bea