Pith. sign in

Paper Citation Record · LEDGER

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

As of 17 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 21 inbound Pith citation observations for arXiv:2501.12368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12368 v2

Coverage vector

measured 100 of 121 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:41.276230Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:24:50.665855Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:27:25.514916Z

Reference resolution

100 of 121 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 58d466b8-4661-4ac4-94a4-70d1320a4cf4 · outbound

This paper cites write newline.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.733873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.733873Z digest=sha256:44c90befbb81cbc8615990f8c2ee0bd616035fd49ed2bab1d25e4ea85d5243e0

Observation abb2e5e7-7488-4e7a-bace-6d5eed9066cd · outbound

This paper cites Pixtral 12B.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.740331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.740331Z digest=sha256:3f00570f065c676b024a7ddc4c45449f742708dd73c447a808061762bb187792

Observation 0a0a2297-cc46-4993-baea-55277e0d3952 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.744554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.744554Z digest=sha256:22a11d5440622bba0ea2fa97e3672c34c5994ffa2b8e340756505571b9fa6eeb

Observation 308b8b09-15f2-4a16-9a00-79617e86b914 · outbound

This paper cites Hello gpt-4o, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Hello gpt-4o, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.784548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.784548Z digest=sha256:edfeba3df12b0fb298a01c07c55f47e62fee87c88b606dbf0975ccccd6affcd6

Observation f416b15a-134f-45bc-8468-f36dbd519e90 · outbound

This paper cites Claude 3.5 sonnet model card addendum.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Claude 3.5 sonnet model card addendum

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.873733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.873733Z digest=sha256:83d929b698724a024550dda7f78188566f798a3b8050f5fa55b41f3f4313c976

Observation c3362012-a413-41e9-b6c8-ebd3cccdfa13 · outbound

This paper cites A general language assistant as a laboratory for alignment, 2021.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model A general language assistant as a laboratory for alignment, 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.932884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.932884Z digest=sha256:83b6e1c7027956c0f1e20872561dac30ca36b3570e8fa0fa4d4ed592daa917d6

Observation cff2761c-efe5-445b-a9a7-5421abd75456 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022 a.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022 a

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.989250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.989250Z digest=sha256:ef6f05d8cc4d9f31a2f628659d26aa6c239eae1069079e770cbc53b800d344f1

Observation f1936c73-93a5-4010-b1fe-c97c4465e134 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.992578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.992578Z digest=sha256:8f7e3165fc078dfa6757628dad58fefdc06b3202ae02eb877efd7a0ed973217d

Observation 0c81af2e-f9ce-4efc-98ba-6970f4a382be · outbound

This paper cites InternLM2 Technical Report.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model InternLM2 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:39.996148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:39.996148Z digest=sha256:d6a755bcebc3d28ad6e8f9ac20f53a570f08c651171f0853627386fa4a83c411

Observation a76663c8-4a8f-4508-9668-446e09af7a64 · outbound

This paper cites Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:18:42.352989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:39.999709Z digest=sha256:3f15288ed49ab6aab662ab626582cc42fcca1f82ca2f5f0d6aa3c01f6dff2522

Observation d20653c7-597e-4d09-8ceb-611f7f2522b6 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.003741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.003741Z digest=sha256:f141f8bdcd91031dfb17e19495f7f02613b6a2783e0633a0def7603a6807cfaa

Observation 4371e3e0-06ad-4ccc-aa02-a9927339681a · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.008109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.008109Z digest=sha256:8e81991c53467afe3e860212e42bea9c8e3384ea424603e8cade5b31bab7b5f7

Observation ba719efe-0197-4a54-91de-5dc6c461edcb · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.012622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.012622Z digest=sha256:e6ca167ffb04034c5ea55e489be381a4532d8fc521b0b03234e5095b8f48fb17

Observation f76dc3b8-f695-4eaf-9c72-ed9ca8af8034 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.016166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.016166Z digest=sha256:e811fcb24b856c52e1a69ec2eb543aa37ea7fdbed47b90700c7d6dbd4a2c4a78

Observation fdbc36e7-7bd3-4698-8f73-99efc9412160 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.020319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.020319Z digest=sha256:0db7fe653c33f633119d1d2c36ae5e1773327a94b5a865e6e3200095e741e7dc

Observation a9f7a83e-3c1a-491f-b7cb-c54ce9cbffa4 · outbound

This paper cites Gonzalez, and Ion Stoica.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Gonzalez, and Ion Stoica

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.023606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.023606Z digest=sha256:98fa9e23639ac79adf08de0933e3fe2a13cf27da7f09579ab3d8eab2ef01ae46

Observation 37398d8d-7290-4226-9e4f-f32733487964 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Training Verifiers to Solve Math Word Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.026561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.026561Z digest=sha256:8b6eb5e3fcb275f28d48edcb23103d1eb61bcef9bf393b35da2ac58a54bf77b5

Observation 101de2ce-124e-4cdb-bbc9-9483cb8b18f7 · outbound

This paper cites UltraFeedback : Boosting language models with scaled ai feedback, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model UltraFeedback : Boosting language models with scaled ai feedback, 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.030331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.030331Z digest=sha256:6bf0b9169ee786cb5f81af7f67c83552a2275a13280c1651146736b9b245bf41

Observation 08df7141-6ad8-49f9-a21b-05ed2eff8acc · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Safe RLHF : Safe reinforcement learning from human feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.033523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.033523Z digest=sha256:6ac84891a9538a6647140a3ce2695a1ccb275bb30e80c2b606f9b84168640317

Observation 07750255-4877-4540-a978-a9ddea935e20 · outbound

This paper cites Nvlm: Open frontier-class multimodal llms, 2024 b.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Nvlm: Open frontier-class multimodal llms, 2024 b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.036520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.036520Z digest=sha256:bbb42c5661b1a0c496f983509f7b7530c21c8743543d2df6c11e14e9b983f415

Observation 442d64e1-b8ce-4923-8602-cc3b41276550 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.040215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.040215Z digest=sha256:9d18eb3934a0a6751672a68319860dde08cf2f505b86e6ff62c0360deeba1acb

Observation 612bf3fb-f15e-4b05-85a7-cc4a788cbf1d · outbound

This paper cites Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.043652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.043652Z digest=sha256:b8fa842a26b73a635b45b551ab191ec019c3c994c2f2bef9bda0e05bc75a99eb

Observation c589c456-a654-476e-8903-177fb877283d · outbound

This paper cites Enhancing Large Vision Language Models with Self-Training on Image Comprehension.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Enhancing Large Vision Language Models with Self-Training on Image Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.048630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.048630Z digest=sha256:fc8a279cc921e19ba8ffd39c9d0d449bd73a6afe92f8d8e23d1e865aff114d92

Observation e6e9e2d1-9db3-4af9-8c46-fbdf91f4c118 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.051995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.051995Z digest=sha256:79ba11c395c2d3e75b3b16a57f61b46cfb3ca30cdfef78be0aba3573d0114ebe

Observation aa988a50-a441-4095-ad9d-cae9b58ebecb · outbound

This paper cites Understanding dataset difficulty with V -usable information, 2022.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Understanding dataset difficulty with V -usable information, 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.063012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.063012Z digest=sha256:1cceaf5abac58aa12ada883db2c944dbfb7457755f965681f920d9f2f06fc7e8

Observation 09bbe186-d714-4269-9934-d6eec117ef55 · outbound

This paper cites Multi-modal hallucination control by visual information grounding.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Multi-modal hallucination control by visual information grounding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.088351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.088351Z digest=sha256:890f931278ee0e9a55719fa84e00032671ca9db431dd984fe77351d7e7c537dd

Observation f5dca67c-1523-4ecc-b463-6978652be895 · outbound

This paper cites Gemini-2.0-Flash https://deepmind.google/technologies/gemini/flash/, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Gemini-2.0-Flash https://deepmind.google/technologies/gemini/flash/, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.133384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.133384Z digest=sha256:c09e249cc68ca6da6223cec11be6b399c35f2e737f1bae58c708f2fc914f6014

Observation 96fee390-a7a4-460f-ad96-88ef3d2a3164 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Reinforced Self-Training (ReST) for Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.187180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.187180Z digest=sha256:009964c50344de22495fc663a2c22effb2e5ccd449ae88b98aca64986aa85d2f

Observation 0af39565-f021-4e93-8634-6d84e6b00ba6 · outbound

This paper cites M-RewardBench: Evaluating Reward Models in Multilingual Settings.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model M-RewardBench: Evaluating Reward Models in Multilingual Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.289337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.289337Z digest=sha256:ceedfb86e5c945015101b61715eee05c78277cd26fed9e456a02840ab23ff2ff

Observation a048fb1b-5180-445b-9358-f086841cb616 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.302326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.302326Z digest=sha256:52b3dfbabe09c5ef458aa8873027211abda8fd6a15914d02854420532c75d45e

Observation 5ab04001-84f7-41e3-84a3-84d0aec55dd9 · outbound

This paper cites ChatGLM-RLHF : Practices of aligning large language models with human feedback, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ChatGLM-RLHF : Practices of aligning large language models with human feedback, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.305716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.305716Z digest=sha256:e088f6d4e081bb321847dec4d76ebcf56bb370aafa463d2626adbfc4570b5d34

Observation 5c3e020b-5eb3-4fad-b5f0-ef71c55ce7e2 · outbound

This paper cites GPT-4o System Card.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.308909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.308909Z digest=sha256:d2f3a8a40e28d5b421f60264ce028733b48c370f01a6ad4b59dcb69df006e57b

Observation cfc4f9b0-36ae-42ec-829d-49e72306cffa · outbound

This paper cites Smith, Iz Beltagy, and Hannaneh Hajishirzi.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Smith, Iz Beltagy, and Hannaneh Hajishirzi

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.312036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.312036Z digest=sha256:be09695e6ee8b9107ea6f4328036f56560ac88ccc09e346393f7b2d136d549d6

Observation 515c7b57-8268-48ba-b500-97056909de4c · outbound

This paper cites Modality-Fair Preference Optimization for Trustworthy MLLM Alignment.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.315614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.315614Z digest=sha256:06f1a0c879a0c7d3e547ff6a3600b87cbecb7e70dd2d98f867f4c07ad3a0c9ec

Observation 4099f462-0776-4890-a203-4333cab0c352 · outbound

This paper cites RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.319693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.319693Z digest=sha256:efbef31522053cb5aa603411ca428ecc7cb006d4d8dc11133d4955772fe20e4e

Observation 4518ad0f-f119-4320-b745-b8c985a49f7a · outbound

This paper cites MiraData : A large-scale video dataset with long durations and structured captions, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MiraData : A large-scale video dataset with long durations and structured captions, 2024

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.323358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.323358Z digest=sha256:8bf041a7b8001b60f08920f2275b8543738a1dce9cc78a2c1baa722b39e7de91

Observation 49243fb7-a149-4e87-85ce-394bb0ddaab3 · outbound

This paper cites DVQA : Understanding data visualizations via question answering.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model DVQA : Understanding data visualizations via question answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.326686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.326686Z digest=sha256:204a01a4bd5f63ae7d808c002b38f2d847e0390ce60b6601051f042cff54d197

Observation 2e59b7f5-0657-40b7-8318-4f87ee28cb06 · outbound

This paper cites A diagram is worth a dozen images.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model A diagram is worth a dozen images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.330107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.330107Z digest=sha256:9bb265d59f387b589f6afe138db84c3c35325012d2101decb6d36b4d24ad11e4

Observation 5fb7bcf8-7a8d-41fe-892b-b1944122cbad · outbound

This paper cites Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.334228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.334228Z digest=sha256:4b8468c1d81304941982419b1a1b64bdc6d0988955c49d15c0072382993d6fa1

Observation 79cb0ae7-ba2d-4a4c-9e3b-3e7fff583507 · outbound

This paper cites Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling, 2023.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling, 2023

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-08-10T17:18:40.337511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.337511Z digest=sha256:eab6737d8a5a35fdaf9dbbf7dcba1873ed25f1b29033f5f9df4d2cccb285e467

Observation aa318f91-e11a-414e-bf44-60a684de6eb6 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.341393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.341393Z digest=sha256:4b58544bb07f9c45cb7f547c65b01fa6529a0933008847e4538e20c59b913b1b

Observation 6f5c59c3-738b-47c2-b346-5277a2cd6459 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RewardBench: Evaluating Reward Models for Language Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.346017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.346017Z digest=sha256:9e9200c1052c074bb34b0b619d14bbdf03bffe37aa1255019e1ab3301fe5982b

Observation cb5a167b-871e-413b-b7ca-4b33653c03ec · outbound

This paper cites LLaVA-OneVision : Easy visual task transfer, 2024 a.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model LLaVA-OneVision : Easy visual task transfer, 2024 a

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.349859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.349859Z digest=sha256:a8cab12baf8908e654a432e52cda16f83064c5a3f8b78ceab8f924dd8ba878ad

Observation 1c825286-f9a1-427e-be53-46c0f39f8f43 · outbound

This paper cites VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.353566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.353566Z digest=sha256:a569c4d6f426ede8e8d5d52e2e28b5b7303f0a9f07ca6ead896d0027c27fd2b6

Observation 1188d9ef-0642-4a17-9831-73a078874489 · outbound

This paper cites VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.357027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.357027Z digest=sha256:0a0c6df1c2f9169d10be1ceb84c2c1399ed91c40d44281cd4b95879f293a82c4

Observation 285d6ba9-b35a-448c-b8c0-b9ae15f66b35 · outbound

This paper cites Super-CLEVR : A virtual benchmark to diagnose domain robustness in visual reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Super-CLEVR : A virtual benchmark to diagnose domain robustness in visual reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.360992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.360992Z digest=sha256:755a200e1af8f75110c08b0a94425fb584964df20a01627d59c8331bb89eedb0

Observation 7408baa0-c25c-497a-9826-7969589fceff · outbound

This paper cites Let's Verify Step by Step.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Let's Verify Step by Step

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.365616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.365616Z digest=sha256:e273d19a68abf80291aaeee7201f2511210409a12b5b85f96b383ad65e7853ac

Observation a7d25e91-ad11-431e-b94a-e505c2517387 · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.380507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.380507Z digest=sha256:45ef14c64622137e1fc60c885dbe3a16c41361d829fd83fe8c4c75b00f45724c

Observation e77e1440-cd08-48f9-ada1-a3249b3b1051 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.418284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.418284Z digest=sha256:308dcba930ea6aa2b625cc3a51fe0215afdd3681924d25b5be9c7208d7bee257

Observation 4881fd41-fa58-4911-8ffc-e987a128768e · outbound

This paper cites OCRBench : on the hidden mystery of ocr in large multimodal models.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model OCRBench : on the hidden mystery of ocr in large multimodal models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.479180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.479180Z digest=sha256:50db9fa079226919e7ebe5e47b876646b7807f711ad562879f79cbcd9cf6d6c6

Observation 8995f9b6-2354-4758-aae2-58978eb21bde · outbound

This paper cites POINTS1.5: Building a Vision-Language Model towards Real World Applications.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model POINTS1.5: Building a Vision-Language Model towards Real World Applications

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.522131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.522131Z digest=sha256:3d5e414166a08b929eaefe039df833ec6e6df984756d5b1e80c74494ddae5aa4

Observation c25ada88-2319-46cf-873e-afa009daf17e · outbound

This paper cites RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.563918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.563918Z digest=sha256:e902dd999f66dda425983edfc346c2c8ac8262112fd548f557006b37f580de34

Observation f488368e-1c0a-421e-b1ed-9a0c9a8a92ab · outbound

This paper cites MMBench : Is your multi-modal model an all-around player? In ECCV, 2025.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MMBench : Is your multi-modal model an all-around player? In ECCV, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.568167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.568167Z digest=sha256:82c1c5e4742e2221f0c0c2c6031c83fef09bd44cdfaf32c5e12a876768022ac0

Observation 7a8cf238-711c-4185-a867-8e79beac2d7c · outbound

This paper cites MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.575703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.575703Z digest=sha256:6c74a41fc5b71c9b166cda6c225a9d1a5b78f7787344011343af38336a1e33a7

Observation 5b9f5248-de18-4fda-8e3f-797c009e6805 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.579427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.579427Z digest=sha256:eef7d48fc2793a61c3f25e53ecb813f2b3a6d2fa8587980aa9ab485d04b0c58c

Observation 78288ff8-3bd9-472c-990b-3f5217db9ca7 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.583349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.583349Z digest=sha256:6e418d6970ab1c9f0f06f67cc76f9863d3572f625a00e9c3ba62024c35533084

Observation 3ad19d6e-c128-4786-bf34-099f23e33c51 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.587351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.587351Z digest=sha256:e0e1ecb9b99d5cb18991fb31b5a1900f442b1e3925d80630865f849f2fa28bcd

Observation 219b968c-698f-47ae-aa01-efe93dbd34e2 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.590732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.590732Z digest=sha256:c65b39a0a45f848645993dc963a3ae042b3591427da59e1254545f56b9e4cfcf

Observation 7957d2ce-2237-420d-9ad5-bb17187ffd6f · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.594307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.594307Z digest=sha256:690bb856919f849d33905b09ee4dbec821bab65b5137d3da0dfa9410567360cf

Observation cdc4050a-55aa-4029-b674-5a00e3cd283f · outbound

This paper cites Ovis: Structural embedding alignment for multimodal large language model.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Ovis: Structural embedding alignment for multimodal large language model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.598612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.598612Z digest=sha256:9fb873e0cdb6844a65250fd1c5e86e01bef7c67be8cd8ad1b22a4a3fee6b0438

Observation 1efae7e5-581b-43bc-89f4-76ecca9cbd5e · outbound

This paper cites BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.601712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.601712Z digest=sha256:8b651bb2c0b2213606c6251f5900fccd421cb23b7a0a98aaee4993e35e25e0af

Observation e949c44f-8a1d-4595-a56b-36ad88e827eb · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.604882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.604882Z digest=sha256:35781df6e8f75f13351b8cb0ca280bc6d3cdfbf71b532a70065d8c9808641e48

Observation b5b622ca-28c3-4b42-9a7c-22c64cb84bff · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.607907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.607907Z digest=sha256:d4541cd5a9cccc900280910d2aaea18ff6d10f1dfbcdd8258e8748119d9e22db

Observation 478c6912-1417-4b61-8c53-f9c31d908cac · outbound

This paper cites an unresolved cited work.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.610991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.610991Z digest=sha256:74751a917804443a8caf786d3203c57fad6dce9be049a303504fdd075995ca5e

Observation 20922ed5-4389-4107-afb0-163b5a88e166 · outbound

This paper cites CLIP-DPO : Vision-language models as a source of preference for fixing hallucinations in lvlms.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model CLIP-DPO : Vision-language models as a source of preference for fixing hallucinations in lvlms

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.613542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.613542Z digest=sha256:149e11d9e8db38ef829711075bbc2ee40f702a8e65f014071f3d1b4900cc8969

Observation 875b6e08-ce30-4bf4-b665-dd00535debef · outbound

This paper cites Training language models to follow instructions with human feedback.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Training language models to follow instructions with human feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.628592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.628592Z digest=sha256:6960bfb82d79d2c4d0a067d5c1661d805a07f20dcf1a2f1bf8986d89de00de7d

Observation de65f01f-f3d3-4c6b-844f-54778dde1ff8 · outbound

This paper cites Strengthening multimodal large language model with bootstrapped preference optimization.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Strengthening multimodal large language model with bootstrapped preference optimization

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.679361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.679361Z digest=sha256:4f1a3bd3303d4f1d2a453ab6f505a8e1869938051d070adb8f49dd9fce897ff0

Observation 90818076-b40e-4a5c-b924-ee2dc2f1b00b · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.682766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.682766Z digest=sha256:1b748cc41d8a1d18888fdb84096307ffe00979526c5099d0e90c9253df00550e

Observation 30b34e5c-83b3-406b-9389-5919afe1a63a · outbound

This paper cites Learning transferable visual models from natural language supervision.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Learning transferable visual models from natural language supervision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.686089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.686089Z digest=sha256:2d043c62d3dc04ead9c4e10d8495a0f4cd1b3031cf29c787ec07aec8a9dd2dc7

Observation 436b17c0-ef21-4f14-ac51-d3522ffac504 · outbound

This paper cites Direct P reference O ptimization: Your language model is secretly a reward model.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Direct P reference O ptimization: Your language model is secretly a reward model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.690249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.690249Z digest=sha256:66ff2def61dc9599ff56338e7a84abc32e361b7b7367309620886780da4e6a4d

Observation 57fcecd5-43c9-4c9f-b8b7-760c3e770cd1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Proximal Policy Optimization Algorithms

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.694381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.694381Z digest=sha256:c757b385c51b79f274191fd39cd6d5397fd726742dd726b55ffe2613cadde161

Observation 2aa2f087-19f3-4dc7-8096-2c346e092f35 · outbound

This paper cites High-dimensional continuous control using generalized advantage estimation, 2018.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model High-dimensional continuous control using generalized advantage estimation, 2018

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.697547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.697547Z digest=sha256:c6ab92cc47b2e923060aed46d6644ef8c25351d87b3a69ba2552a0c6dc723b59

Observation a4b4be7d-6b44-4480-99be-4062f0d309e7 · outbound

This paper cites A-OKVQA : A benchmark for visual question answering using world knowledge.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model A-OKVQA : A benchmark for visual question answering using world knowledge

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.700334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.700334Z digest=sha256:1752708e8de8247df8974910c6d36fe2d69b4f1b97dc9af38c4a1deeedd5b0d1

Observation 11fe6856-b491-4455-a76b-99b7f905281f · outbound

This paper cites SenseNova.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model SenseNova

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.703155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.703155Z digest=sha256:9ef85f9c43ca9320af4040239dbc9f8beb4a6475c79444374de5c34a4e1838a8

Observation 20aac059-6c36-4088-b322-d116401722ea · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.705967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.705967Z digest=sha256:94d20fc9403c207bc501945e23b657a7f6e9acb28bf156f20c4b210962837165

Observation ad5ce93d-80b4-4605-a06b-220c229a09d1 · outbound

This paper cites KVQA : Knowledge-aware visual question answering.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model KVQA : Knowledge-aware visual question answering

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.709141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.709141Z digest=sha256:30144ca602c835c286c4ef47a3770729aa4c1d4c41476c488d3dee1e14f5b0aa

Observation 60fdbdef-aa39-408c-bea1-f7f04cf74be7 · outbound

This paper cites an unresolved cited work.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.712595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.712595Z digest=sha256:be722d4f588ef4af52f7475bf88d7955727837a278410becd246589231f1f474

Observation fbdc80d1-d44c-4094-b355-3effe39c1038 · outbound

This paper cites Skywork critic model series.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Skywork critic model series

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.752880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.752880Z digest=sha256:a84ece5f464640db46755c8c36c42bc68bfca67a0833f76d497d1e7da2757f11

Observation f5edfb6b-7ae8-42c9-a183-44a3b654b76d · outbound

This paper cites Towards vqa models that can read.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Towards vqa models that can read

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.774050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.774050Z digest=sha256:2e5280d6ca3218b6afefec85338138f237a62cabb37383c0b6bfe7302cc6ac80

Observation 267e7d76-4d72-4135-a056-8a1a22a9c20e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.819168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.819168Z digest=sha256:06851d2c8c7740ea211586d509f3640c43a52ffb69153f96e4fe2290be521405

Observation 0cbd0ccb-bf5e-4f99-873c-1821f29e00b6 · outbound

This paper cites MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.878653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.878653Z digest=sha256:5c67742d7510e826d707ceace54ade734ff55c85b145e5b92465c9262ef5274f

Observation e867eb92-46ed-4f6b-9234-3fac4742bcee · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.924422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.924422Z digest=sha256:a916912dcde733809c4be496df7d73c4698d98433c11356a0a495ff005f0e46b

Observation 98ce674b-7312-45ac-a662-cf2e0c98db37 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024 a.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024 a

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.042071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.042071Z digest=sha256:3d42d3f27ec2937e201b3dcc1fd48e94599bb95f4c3075f8855e4a763996b3e2

Observation 039d1596-53f5-4aa0-b7bd-4e749315c2df · outbound

This paper cites The llama 3 herd of models, 2024 b.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model The llama 3 herd of models, 2024 b

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:43.293231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.045129Z digest=sha256:a4639310a03e99eaf65dcb8549f71da331f974be02e3560b2b20acbcf5d731a1

Observation 1271437d-7acb-49df-90eb-120a72575184 · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024 c.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024 c

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:43.282973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.048480Z digest=sha256:e752d2bfc25421b65a0537f75577657edc96f904e2fd7902694e3778792540e6

Observation 6d858170-afe5-4fe8-ab8c-005917015ef1 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Solving math word problems with process- and outcome-based feedback

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.052281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.052281Z digest=sha256:8f8747814d430c5681550ff8922331115dd26f1966f3cd717c1b04848b86ef74

Observation 5d005002-d81d-490c-bf4b-5ddfabecd607 · outbound

This paper cites RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-10T17:18:41.947365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.056178Z digest=sha256:c49326fead80e0831d121d14a5243ba2412e4a9364aa36f80b61b0e9a9b9fb9a

Observation e7092d19-35bc-4643-861b-ceaa4b504d6c · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.059884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.059884Z digest=sha256:96b2e444b2313b3911ab1b2b6d749c09c636982701802e2eca8e13bf0f9af746

Observation 26683e97-4e56-44a1-b3be-b18a534c2d7b · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.063611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.063611Z digest=sha256:a5c09b001185a5dca3b17351b4b9c0ab37414f842a2392950324052b38848575

Observation cb318e60-307e-4a39-9f51-7a5e4e143791 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.067215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.067215Z digest=sha256:abe8541eb8a7b6da301e4108217399d617e217720353e9faf81985df31901212

Observation 41efb4af-7311-4954-9367-87bf666197c7 · outbound

This paper cites Self-Taught Evaluators.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Self-Taught Evaluators

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.070735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.070735Z digest=sha256:9c2784d34e9f1f6d379cef97823ef84b146b2bbd6cc7436e1955afb45b1948e5

Observation 29e9a852-64f1-484d-8d9a-abe2624e34f2 · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.073776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.073776Z digest=sha256:99311720833ddc5613678cc81bbe3f6975e12a6431d3448c76cc99413df08b0b

Observation f90cb740-cdbf-4b82-b858-58f83c056816 · outbound

This paper cites FunQA : Towards surprising video comprehension, 2024.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model FunQA : Towards surprising video comprehension, 2024

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:43.191462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.077019Z digest=sha256:14dbadb143bbfb2a1d03a824a43e8c3a1301bca46cade614db2eeb8402f14e22

Observation e6dbfc2c-f213-427f-ab31-493ba0e3dfa8 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.080692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.080692Z digest=sha256:be03c89e7ad87d67e046223c2a2212582ac10108daf843f7f35068de06707070

Observation e1cacfed-d5a6-4c9a-857a-9f80a2ba00dc · outbound

This paper cites Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:41.089878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:41.089878Z digest=sha256:5ac673376bda5e7375eecd477c80e737e767cc9382e3fa6b0cce03329e242a1a

Observation 7a82b675-5cce-414e-ab4f-4594492fb324 · outbound

This paper cites SUTD-TrafficQA : A question answering benchmark and an efficient network for video reasoning over traffic events.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model SUTD-TrafficQA : A question answering benchmark and an efficient network for video reasoning over traffic events

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:43.093560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.145353Z digest=sha256:b6cddd22200df2c730d3309d2f2cb5ebc973add6d6bdef635b0b3a8fd6ff9b3f

Observation 731f932f-5711-447a-83bc-95fbdd1e9fd8 · outbound

This paper cites Inf outcome reward model.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Inf outcome reward model

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:42.898205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.199723Z digest=sha256:72b2140f3bea2848c8c8044c031f54de1dbaf656b5d09dd23fc4839068af6d58

Observation 0a0836de-eb21-4637-92f4-08c00797e542 · outbound

This paper cites Regularizing hidden states enables learning generalizable reward model for llms.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Regularizing hidden states enables learning generalizable reward model for llms

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:42.804105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.269132Z digest=sha256:00837c4eb1069c314aad3db37b330c9df552f238ccbae4be842b3551d13cf000

Observation adf2a856-1306-4899-b6f5-f6335a026060 · outbound

This paper cites MiniCPM-V : A gpt-4v level mllm on your phone.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MiniCPM-V : A gpt-4v level mllm on your phone

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:42.793363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.272937Z digest=sha256:5d151ba2a171a08a118dfa46a53c3c1266900be04c312a2906a62a2f532184e7

Observation a2b79ee4-74e8-4746-80e5-7c071fa66362 · outbound

This paper cites RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T17:18:42.782420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-10T17:18:41.276230Z digest=sha256:b3ddf5aea793733f5ff92d449d27c3598ce31b6083eb436a1a5fb2b161949502

Pith citing papers

Observation 7f27ffdb-7d2c-42c1-b6e2-3c18214a288d · inbound

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step cites this paper.

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:35:25.944906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T11:35:25.894465Z digest=sha256:90b6a86e3159c5809a972f9567e33c57876d6963c0c1fd99a2ddc4532cb3fb35

Observation d4118f63-dd22-44cc-943a-d3c28a5d7eef · inbound

VideoRoPE: What Makes for Good Video Rotary Position Embedding? cites this paper.

VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:35.213153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:06:35.213153Z digest=sha256:ef13575e13e4ee25ea9a1fdd895e192b7264e84b9b0612acaaef00274a64aac7

Observation 3224bb3e-e772-43d0-8163-b0098b5be323 · inbound

Visual-RFT: Visual Reinforcement Fine-Tuning cites this paper.

Visual-RFT: Visual Reinforcement Fine-Tuning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:16:16.607976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T22:16:16.528682Z digest=sha256:8f5dfacc9e901eb3d5272cee12fad905ffe0e92d660f3ae5c57ac8cb4841e327

Observation 7a26ca15-e149-4653-b3f7-d1fbdafb6563 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:44:30.624665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:770bfc904764320ef1f0aacec06caf0c5e10b51218e9b99f7c58de95d6e8493d

Observation 4a697026-1be9-46ed-b0f5-473b729784fc · inbound

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model cites this paper.

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.508448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T01:13:57.368874Z digest=sha256:4a9b7cb0c3f43a7c3d5d31bee70056f264283951b3109abd404ff389a2a059a1

Observation 8b0068df-7793-450c-97e6-2e5c84a92d2c · inbound

Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning cites this paper.

Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:24:50.665855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:24:50.665855Z digest=sha256:e0e818e3f115d3e24d5a0623630938a0c537a7504f5d87d29b01f91fe84a9e5b

Observation ff096c44-57f5-4d46-8d45-e53e287d1424 · inbound

MR. Judge: Multimodal Reasoner as a Judge cites this paper.

MR. Judge: Multimodal Reasoner as a Judge InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:52.977818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:18:52.977818Z digest=sha256:3ec5acca5f722d8670e2b2471fc37b06b0346f390b74f3b4d9fa86fce2f0278f

Observation 4023cf70-a35a-48ef-8c9b-7675d116fb4a · inbound

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision cites this paper.

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:17:23.265015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:17:23.265015Z digest=sha256:ca81a437209696254a75e162a4aa97dc619d8936d47d7d499314eaab28d4b41c

Observation af863026-8903-46e6-adf7-f264d3c4b139 · inbound

Visual Agentic Reinforcement Fine-Tuning cites this paper.

Visual Agentic Reinforcement Fine-Tuning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.512442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.512442Z digest=sha256:116745c1b4d506b923ada3621ace222c77a44f67d70252b9e57f77e66998f64a

Observation 4b405f5b-deab-4f49-88d5-95f077449607 · inbound

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models cites this paper.

ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:24.804904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:24.804904Z digest=sha256:a112ecba91c330a4d2fa0c63b28e9617e14008d1c1abc49287588eb851ae7ec3

Observation 62de62e2-b0ec-4e6f-8e41-c84fe1f9368e · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.112648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.112648Z digest=sha256:6f6dbe2e7b58f49a76cd7ec6cc442dec352af57f63ee1c68eaf7d8e400d1dd32

Observation 38988312-0153-4080-8c93-3efae3d66d08 · inbound

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents cites this paper.

UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:34:16.325870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:34:16.325870Z digest=sha256:5966b18f3a171c327fc82a74c1744c17ba0193231f420dd2a86332abc681841f

Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.748441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.748441Z digest=sha256:ccdd50176e41c94c027eda1c3494b1f158a59980b5f129015547e2eaff60182a

Observation a8581371-e20e-49bd-b5d3-ab2602db9964 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:55.357084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:55.357084Z digest=sha256:4d82cdd4dde801c5e44e297591c15929433268e8e8fbe7b25105c442623e21d7

Observation 2d5e40fb-ea29-4f77-b6e6-136cf38cf54c · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.936440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.936440Z digest=sha256:2bdf3b833308adb5d08e39f5ceeaff96059102687e1f7b00d9ce98961fa379fc

Observation abd59930-c41a-4f52-a69c-cdfaf9472107 · inbound

AdsQA: Towards Advertisement Video Understanding cites this paper.

AdsQA: Towards Advertisement Video Understanding InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T20:20:36.958513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:20:36.958513Z digest=sha256:893c12b5cb32fc862a34631316e1988d998147ce40740af02eb37fb6ef6bec1c

Observation 8ac583ed-ae6f-4ae6-84c8-ee1303df5c55 · inbound

Visual-ERM: Reward Modeling for Visual Equivalence cites this paper.

Visual-ERM: Reward Modeling for Visual Equivalence InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:19:57.941133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T11:19:42.002790Z digest=sha256:92ac2e78eb905548e31391fc3c8780a31d91303843b85de319303ca97f474a99

Observation 4c2c0386-ae29-4357-96f2-a12ded5569e9 · inbound

DRM: Diffusion-based Reward Model With Step-wise Guidance cites this paper.

DRM: Diffusion-based Reward Model With Step-wise Guidance InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:13:59.364370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T22:12:39.225557Z digest=sha256:ba612b8912a76ca9068299823732fcabd7eeff1f579e3970d2b837c8a127e2cf

Observation 3ef200f6-b981-4852-870a-3bae331782cf · inbound

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving cites this paper.

DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:27:25.516872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T18:58:26.701929Z digest=sha256:c0064110a35084962fa9c52223581c51fc32d66e4d1c6a1224fb0bdee5a103f5

Observation bd273c75-6f75-4bf8-ac52-952a1847c676 · inbound

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation cites this paper.

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T14:00:00.388339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:00:00.388339Z digest=sha256:ada0625bc33b8019e4d525d3ba91c0249f62148d6f9cf21348be063ab3a63a79

Observation 225973e3-99ce-490d-bfa7-bbe563a68028 · inbound

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences cites this paper.

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T14:37:30.113159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:37:30.113159Z digest=sha256:162a545521f829a0a68226d06aa0f5b74083894e135fa2c2d76e443e62eb126b