Pith. sign in

Paper Citation Record · LEDGER

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 2 inbound Pith citation observations for arXiv:2509.00484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00484 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:37:00.301396Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:12:18.190583Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:45:58.423192Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy54
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4f6cb42-b89a-43b4-b278-7082640825dd · outbound

This paper cites Phi-3 technical report: A highly capable lan- guage model locally on your phone, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Phi-3 technical report: A highly capable lan- guage model locally on your phone, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:12.259469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:55.685164Z digest=sha256:732e3353785399d5bee90e989b7b70e6f5df208574101d43fca6f26d63559f41

Observation e94cf7cb-74d9-45c1-87d5-acfb09d28294 · outbound

This paper cites Claude-3.7-sonnet.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Claude-3.7-sonnet

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:12.073331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:55.776050Z digest=sha256:a92dc5a392c809cff047c60e21d2f2d39f4c0b29714435f30671936d163d7267

Observation c7ca1495-48b3-4dfe-b77c-81d475bd9ab4 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Qwen2.5-vl technical report, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:11.928079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:55.881039Z digest=sha256:0004e7638fdc5b6e2ce831d6e3bdd33d6abfb46c4fa0de32bca60d2a3250c3ee

Observation 5915fc07-4818-4255-af09-a49ab2e26d7f · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Mllm-as-a-judge: Assessing multimodal llm-as-a-judge with vision-language benchmark

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:11.766552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:55.992017Z digest=sha256:afa42a2b3c97f26361cb9fef8546e7b3ea5a1ca0c223c0221b6dd3289410d99c

Observation ff2e0499-c770-4847-af77-bff4b4272419 · outbound

This paper cites From captions to rewards (carevl): Leveraging large language model experts for en- hanced reward modeling in large vision-language models,.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding From captions to rewards (carevl): Leveraging large language model experts for en- hanced reward modeling in large vision-language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:11.523436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:56.109886Z digest=sha256:3d54c1423f356fc0c02b2e7e2055d813f8004803f825957e1faeb601584d32de

Observation 458599d7-441e-49c8-9027-5eaadcfd82f2 · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video under- standing.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Mmbench-video: A long-form multi-shot benchmark for holistic video under- standing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:11.375281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:56.240599Z digest=sha256:7f2a4b29ca20d0bab873351c2be6452ae2177fe829e5814ced194e21a6bb3a04

Observation 35785f6a-bbaa-4cc6-a1bf-9492b07eec96 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in 9 video analysis.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in 9 video analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:11.235012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:56.367900Z digest=sha256:c77f4d7a98336c401b8da920192b0153439d50b2d9a4be0789d4d184e201c7e1

Observation 815a0060-ae0c-4654-93e8-1f18529321d6 · outbound

This paper cites Gemini 2.5 flash, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Gemini 2.5 flash, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:11.009206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:56.552682Z digest=sha256:91892e198f1e62404016eef1258425f16a355e524a5cb707fd05bef4297b80d4

Observation 9a0723b1-0a7a-4fcf-90d8-62cd592b9499 · outbound

This paper cites Gemini 2.5 pro, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Gemini 2.5 pro, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:10.806655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:56.711262Z digest=sha256:cbc19fbf1f06124457f4bb6b3aebd7c8e00d372afc012deda6471f6acaa54fc7

Observation f147d787-2081-4768-bfc0-2a5f8439d302 · outbound

This paper cites The llama 3 herd of models, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding The llama 3 herd of models, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:10.635204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:56.892327Z digest=sha256:871bf8789e4969797bb2a97136c026ab5ced7e0a1ab976185b692844ab31e20f

Observation 46713bf0-5af4-40a9-a194-5bd26af5e271 · outbound

This paper cites Mmworld: Towards multi- discipline multi-faceted world model evaluation in videos,.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Mmworld: Towards multi- discipline multi-faceted world model evaluation in videos,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:10.417542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.078804Z digest=sha256:58c1272be2358344e896f67d42adb31e82801222becd3b793d41adbd6eca7a45

Observation 6b80a8a4-4a30-46fc-b486-6b10d4d7e303 · outbound

This paper cites Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Video-mmmu: Evaluating knowledge acquisition from multi-discipline pro- fessional videos, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:10.146323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.173200Z digest=sha256:1623c79a5b6bfde7bfeccbe054869b8c7e41d42a19adcca9e2b01131cba90b89

Observation 0a1592f8-5a7b-4c9e-9f7f-3d6f256a387f · outbound

This paper cites Flex-judge: Think once, judge anywhere, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Flex-judge: Think once, judge anywhere, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:09.988035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.228365Z digest=sha256:cc76cd8eb6b6f1895a51a9e861d3f740b1c1a376262e6be9866cdedfbba9d85c

Observation 464475ca-6ce4-4df4-ba44-0c3398e8f56f · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Smith, and Hannaneh Hajishirzi

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:09.666185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.349655Z digest=sha256:dc31c81edd6a26cd7af78ddb2d05a38ac0ebe754b66c7fa265d86bf980f5eafc

Observation 7c133eb4-541b-4797-93b6-8be641df03cd · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Vhelm: A holistic evaluation of vision language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:09.379423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.467162Z digest=sha256:bfe91ff591ecba1d1f9e51bac29541d0ddd1da2d863de0d724f151480f67a8bd

Observation 98ca7ad7-047c-4738-b074-a748151bdfd1 · outbound

This paper cites Llava-onevision: Easy visual task transfer, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Llava-onevision: Easy visual task transfer, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:09.103447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.548201Z digest=sha256:f7a6efe3a122346283b419830f9828b8cc6c3ff52d848d662ecb983a32372f87

Observation 5e0806f5-8935-4239-b0a0-74bb98530c8e · outbound

This paper cites Aria: An open multimodal native mixture-of-experts model, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Aria: An open multimodal native mixture-of-experts model, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:08.862974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.633679Z digest=sha256:4bb501a043a76928eedf54dc37e313a5c131c8ac2cb925bc0d1da2a0a2bde027

Observation 27ba31a1-08a6-4816-86bb-1ece49a1a1ed · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:08.614241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.723527Z digest=sha256:8ac56c5b8ad43b4fa3e9b3c58af9d7eed57e5d7c091c3a9f194b543947b1c0f3

Observation e2e7aee3-320c-4b6e-acb2-dca4c9ec5f35 · outbound

This paper cites Vl-rewardbench: A challenging benchmark for vision-language generative reward models.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Vl-rewardbench: A challenging benchmark for vision-language generative reward models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:08.362228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.752949Z digest=sha256:9120dd4abaf00eea3bb1c96e77663ed8fae54abc4379d22d621f58c7b3fe8017

Observation 0f641bec-ef5d-4e35-b963-eb12b8510092 · outbound

This paper cites Holistic evaluation of language models, 2023.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Holistic evaluation of language models, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:08.084903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.836035Z digest=sha256:ba9f89716b42eb6d5506abea65320a83e9ac7aa084a48fd714e76161d1ddc4ea

Observation b3d25fb7-0624-470d-aa62-27254b37cfd4 · outbound

This paper cites Video- safetybench: A benchmark for safety evaluation of video lvlms, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Video- safetybench: A benchmark for safety evaluation of video lvlms, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:07.780040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.896073Z digest=sha256:e50236bb95db4c14783940a764aa8f40639c8a83dd311f08cf61c25879a870f6

Observation 98d78dd1-154c-4bec-a7e5-98f16ac30b4b · outbound

This paper cites Rm-bench: Benchmarking reward models of lan- guage models with subtlety and style, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Rm-bench: Benchmarking reward models of lan- guage models with subtlety and style, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:07.501117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:57.968346Z digest=sha256:e48acf6f3a22404e9eda73647fc92270a22be1e02ba8159808af0c9866e3f51c

Observation 84c079c1-9c8e-45f8-88ab-dc640f1875a1 · outbound

This paper cites Videogpt+: Integrating image and video encoders for enhanced video understanding, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Videogpt+: Integrating image and video encoders for enhanced video understanding, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:07.237524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.052845Z digest=sha256:3eadd7a7a465aa4687e62c1029b2ad7e4f6edb5669638eb95c7a934069cabd66

Observation 0e34abe0-122c-4574-b29c-40ec34abdf74 · outbound

This paper cites Smith, Hannaneh Hajishirzi, and Nathan Lambert.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Smith, Hannaneh Hajishirzi, and Nathan Lambert

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:06.920168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.086443Z digest=sha256:519e4d741aadd24c886301d0bee143ddabed4cd5d668511eb0aa0b7c75b53fbc

Observation c6042620-dd95-4a61-8c82-05f9793f1ca9 · outbound

This paper cites Hello gpt-4o.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Hello gpt-4o

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:06.677472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.146528Z digest=sha256:5e6742ab96ffd3622fb75eef79a32b29ed6a07700579b8a83343289fc7199f10

Observation 6bf86790-0c14-48f3-b2db-25ecedcd3c29 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intel- ligence.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Gpt-4o mini: advancing cost-efficient intel- ligence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:06.448524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.200367Z digest=sha256:a3d0808570b9086ef025a6137da88acd4e28b7a4c6eea00194fd4870147d4882

Observation 6f9cb145-28d3-4e6d-85a5-1bcc0d88753b · outbound

This paper cites Training language models to follow instructions with human feedback.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Training language models to follow instructions with human feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:58.250334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:36:58.250334Z digest=sha256:aa2c19e0c8d490b58d348ca26934fa8814e919907b0320e9561b1cd4b60f021f

Observation 890b82f9-9b86-4d9c-88a5-17fbb0b0c2ff · outbound

This paper cites Vibe-eval: A hard eval- uation suite for measuring progress of multimodal language models, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Vibe-eval: A hard eval- uation suite for measuring progress of multimodal language models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:06.158425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.299362Z digest=sha256:e13b282a025a2d629d195e2f4db472d2b52dc0204efbcee83be36609c514acf1

Observation e5b36f29-b815-469a-890e-d4050996fa38 · outbound

This paper cites an unresolved cited work.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T13:37:05.867135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.353204Z digest=sha256:94c3ccfd1232ad9c0254b79b79d8eecfcfae30320594cbc340381d8878094cce

Observation d3f52317-6b7d-44eb-b8b5-3ed8ba05952a · outbound

This paper cites an unresolved cited work.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-05T13:37:05.688820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.400638Z digest=sha256:90dd128d74c343fbd86262331af887cc5162cc063a3b2d719d7d014672a07452

Observation 83617b98-f60b-49d7-b30a-d613979447b1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Direct preference optimization: Your language model is secretly a reward model

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:05.473882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.443575Z digest=sha256:d576ddfa29b01280df32f39f9a7f3bcd81296736d59d132d070db891c7169afa

Observation 36f21a8a-9fe1-45cb-9dec-68ae68bd6cbc · outbound

This paper cites Scaling llm test-time compute optimally can be more effec- tive than scaling model parameters, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Scaling llm test-time compute optimally can be more effec- tive than scaling model parameters, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:05.255319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.489898Z digest=sha256:e6918041e2f4183fad372b9f6110a0261b68d817c7dd10e5f1959f73b6622419

Observation 7126e7c9-d619-4a1a-9a8f-c9fb22f570b3 · outbound

This paper cites Aligning large mul- timodal models with factually augmented rlhf.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Aligning large mul- timodal models with factually augmented rlhf

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:05.057724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.571167Z digest=sha256:dba2bf8ab0a7b03ce0796d7b358a481db18c4d5bbe64c2006a7558b43598685a

Observation 67911071-2d2e-487b-9976-26b7bc9df062 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:04.828772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.657428Z digest=sha256:4a19ad3695a2cd319e843bf6a0ddae95c57672a2680fd60540bad5c1c86355f3

Observation 8977e9be-1c7b-417f-9469-318b75a1d9ea · outbound

This paper cites Visualprm: An effective process reward model for multimodal reasoning, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Visualprm: An effective process reward model for multimodal reasoning, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:04.645184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.728999Z digest=sha256:1f8a175855625ff6b874b85585283a199fca755166124ddece7aa96aaf728986

Observation d15c199f-cb9d-4b60-b885-5527adb7502c · outbound

This paper cites Skywork-vl re- ward: An effective reward model for multimodal understand- ing and reasoning, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Skywork-vl re- ward: An effective reward model for multimodal understand- ing and reasoning, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:04.470201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.764352Z digest=sha256:d9e090ae948b3a0b16e71951a6b23446f8fd701004fca0689e1a3a6622b6c5d8

Observation fadbf3fa-85e4-4c10-8c66-5827de2b86d9 · outbound

This paper cites Videohallucer: Evaluating intrinsic and extrinsic hallucinations in large video-language models,.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Videohallucer: Evaluating intrinsic and extrinsic hallucinations in large video-language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:04.289915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.825448Z digest=sha256:c74a5ec852cc8d808a046b1a2392396c397a90c3f245480c6d6453afe56e2376

Observation 764544fd-5763-42c9-9c73-f8c0d29e2473 · outbound

This paper cites Internvideo2.5: Empowering video mllms with long and rich context modeling, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Internvideo2.5: Empowering video mllms with long and rich context modeling, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:04.111328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.883194Z digest=sha256:0dc6faca45d7bfdfa84eff2051fc22562bf171cba05ec4736595b1a2773a0e9c

Observation 261d4f87-715b-453d-ba4e-37e0ce250b29 · outbound

This paper cites Unified multimodal chain-of-thought reward model through reinforcement fine- tuning, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Unified multimodal chain-of-thought reward model through reinforcement fine- tuning, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.977219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.926700Z digest=sha256:35551773a9b758858ea13fac05a6d1ce033a060b964d0678f0538c14b83988c4

Observation be6f6f8b-4e21-4054-91a3-b5eb23fa6389 · outbound

This paper cites Unified reward model for multimodal understanding and generation, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Unified reward model for multimodal understanding and generation, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.829886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:58.989270Z digest=sha256:6ff532547c528cd0092d6ee6f2144a7d650995594fdc9e9d739cb84f1d9e8ab3

Observation b1340748-d163-4569-8cf2-7f7e62c61eba · outbound

This paper cites reword- bench: Benchmarking and improving the robustness of re- ward models with transformed inputs, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding reword- bench: Benchmarking and improving the robustness of re- ward models with transformed inputs, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.641007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.128058Z digest=sha256:c32d420684388de429817f79fd11a15aa462e917a8da6c60c784ac3258935643

Observation a3b188c3-f4bb-49fc-8167-943cb1724c22 · outbound

This paper cites Llava- critic: Learning to evaluate multimodal models.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Llava- critic: Learning to evaluate multimodal models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.469920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.239966Z digest=sha256:872ac1971a39a88b91eb5e70f0a786bd84e4468d00c9af99bfc9bfad3372dc54

Observation 4d72ee90-a197-4dfe-9da4-becf901f93c4 · outbound

This paper cites Thinking in space: How mul- timodal large language models see, remember, and recall spaces.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Thinking in space: How mul- timodal large language models see, remember, and recall spaces

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.316493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.353315Z digest=sha256:48d7ba444e4622c8158f0b980f08efde0184cccfe64482fa110fddbf51629ff7

Observation 5475c51c-cf2d-499c-9648-ceeaf6c12bd3 · outbound

This paper cites Minicpm-v: A gpt-4v level mllm on your phone, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Minicpm-v: A gpt-4v level mllm on your phone, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.192057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.433486Z digest=sha256:3e672f7ed46f0c872b1da36435fe922926cde379e0ea0bb63b3dc4b32b24a9b3

Observation 0acb75a3-7223-4a1d-bdf4-067cc6752f96 · outbound

This paper cites Multimodal rewardbench: Holistic evalua- tion of reward models for vision language models, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Multimodal rewardbench: Holistic evalua- tion of reward models for vision language models, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.057857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.513332Z digest=sha256:877aa7813628eb8e6bd9162cf3ec4d858d2f519a281896c94ccebfa852932215

Observation 601e0b30-c8e5-4ebb-bdbf-6d04c41722a4 · outbound

This paper cites mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding mplug- owl3: Towards long image-sequence understanding in multi- modal large language models, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:03.000377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.571232Z digest=sha256:496083047a6d1f4cfe067bdff24f16e73ce36b6eeb029a683e350f34b9c8e5ba

Observation fff5b201-f05a-416a-a0cd-f7f81c4c8b98 · outbound

This paper cites Internlm-xcomposer2.5-reward: A simple yet effec- tive multi-modal reward model, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Internlm-xcomposer2.5-reward: A simple yet effec- tive multi-modal reward model, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:02.910320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.661912Z digest=sha256:c21146c07791f5c981a786171acb4402a3f30f2c137e89bbd1303832207d35f7

Observation 6e28fccb-51f0-498a-b7ec-7ccd3f04d21f · outbound

This paper cites Video instruction tuning with synthetic data, 2024.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Video instruction tuning with synthetic data, 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:02.800246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.755444Z digest=sha256:0e4be263463a2fcedd722eb7cb87d3d1e530b445ff319eb20b255ca080399e00

Observation d100f179-064a-4808-bcab-4546d9d6392c · outbound

This paper cites R1-reward: Train- ing multimodal reward model through stable reinforcement learning, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding R1-reward: Train- ing multimodal reward model through stable reinforcement learning, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:02.665649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.815536Z digest=sha256:06813133b62c69d51c61c3a8fb1e070a4f8b3bf3353ee922b1b1e9f457f64dd5

Observation 4d94a7f3-44f2-43b8-ab82-8e8bb70059ee · outbound

This paper cites Mm-rlhf: The next step forward in mul- timodal llm alignment, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Mm-rlhf: The next step forward in mul- timodal llm alignment, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:02.524756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.848574Z digest=sha256:ee7b806c546a2aa78fcdee01698ce21d892ada92df4beceef6a8bc2bb9b29eaa

Observation de729c4a-47e1-433b-b0ad-a1ff6cee003c · outbound

This paper cites Mmvu: Measuring expert-level multi- discipline video understanding.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Mmvu: Measuring expert-level multi- discipline video understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:02.304836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.879659Z digest=sha256:443d6be634bbb2e0e32b38db3e4ccb3b0ace1287310eb890790ec97bbe9cfacc

Observation 9d797eb8-6c3d-49e7-9467-33eb1fe8a745 · outbound

This paper cites Generative rlhf-v: Learning principles from multi- modal human preference, 2025.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Generative rlhf-v: Learning principles from multi- modal human preference, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:02.086954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:36:59.941132Z digest=sha256:43877efc000144c6f1af292bae11b61d584017544d39a69d0d8ab915871f5d58

Observation d19521cb-91bd-47fc-89d7-b44065f3fdc1 · outbound

This paper cites Input Frames.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Input Frames

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:01.739502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:37:00.007216Z digest=sha256:9b291c140b7dcb4d920d06f7acac62626f901822b04c5e174745ed54e608268e

Observation e337c64c-0f18-4685-95e0-eaae06bf12b8 · outbound

This paper cites When placed in water, there is a violent reaction.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding When placed in water, there is a violent reaction

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:01.387620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:37:00.097372Z digest=sha256:aee9f25871db12b622d70957e4f3269c6cc93c537fd9b5f87324724bac02a59f

Observation d92b6859-9d44-4444-94ba-8e7f4ae697f7 · outbound

This paper cites - Silver (\\(Ag\\)):\n - Silver is a very unreactive metal and does not react with water under normal conditions.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding - Silver (\\(Ag\\)):\n - Silver is a very unreactive metal and does not react with water under normal conditions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:00.984353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:37:00.166054Z digest=sha256:1a63fc62a41b1557825756a6653103c6545f0b4b4b99edadeebd9c8250e1047f

Observation d9ae3455-a6ed-4695-9bc7-0d6a12cc0ed0 · outbound

This paper cites - Iron (Fe) reacts with steam (not cold water easily in a simple setup like this video) and silver (Ag) is a noble - metal that does not react with water under normal conditions.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding - Iron (Fe) reacts with steam (not cold water easily in a simple setup like this video) and silver (Ag) is a noble - metal that does not react with water under normal conditions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:00.820297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:37:00.236795Z digest=sha256:74f929d4cf5bd559576c0b0dc57a1462eb669a9568d665357f4b6bf03ce82495

Observation ec53e020-5ad6-4f07-a864-0be5f63ac85e · outbound

This paper cites Also, when phenolphthalein is added (the pink - colour change indicates a basic solution), which is consistent with the reaction of alkali metals with water.

VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding Also, when phenolphthalein is added (the pink - colour change indicates a basic solution), which is consistent with the reaction of alkali metals with water

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:37:00.514904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T13:37:00.301396Z digest=sha256:550ada27c41cdd721739293022d700a7e148e4b47223d29f317cf5d415a821e7

Pith citing papers

Observation b67f1f99-ad3a-43e0-975b-10d945926009 · inbound

Social Caption: Evaluating Social Understanding in Multimodal Models cites this paper.

Social Caption: Evaluating Social Understanding in Multimodal Models VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T09:12:18.190583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:12:18.190583Z digest=sha256:55b28c1577e9956b040d7cd9a89488825c7aef21cde331d0338c58b6cda4b0e3

Observation 85a1b947-7ccc-4cf7-9702-476b1f28f9cd · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.426040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:8f26c4b620ce24d75ac3591f5feebe0fb61d52a79503f230109ef448a98638f4