Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:41.276230Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 100 of 121 outbound references and 21 inbound Pith citation observations for arXiv:2501.12368.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:18:41.276230Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:24:50.665855Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T22:27:25.514916Z
100 of 121 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 58d466b8-4661-4ac4-94a4-70d1320a4cf4 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb2e5e7-7488-4e7a-bace-6d5eed9066cd · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Pixtral 12B
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a0a2297-cc46-4993-baea-55277e0d3952 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308b8b09-15f2-4a16-9a00-79617e86b914 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Hello gpt-4o, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f416b15a-134f-45bc-8468-f36dbd519e90 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Claude 3.5 sonnet model card addendum
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3362012-a413-41e9-b6c8-ebd3cccdfa13 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model A general language assistant as a laboratory for alignment, 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff2761c-efe5-445b-a9a7-5421abd75456 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022 a
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1936c73-93a5-4010-b1fe-c97c4465e134 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Constitutional AI: Harmlessness from AI Feedback
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c81af2e-f9ce-4efc-98ba-6970f4a382be · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model InternLM2 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a76663c8-4a8f-4508-9668-446e09af7a64 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Decompose and Leverage Preferences from Expert Models for Improving Trustworthiness of MLLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d20653c7-597e-4d09-8ceb-611f7f2522b6 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4371e3e0-06ad-4ccc-aa02-a9927339681a · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba719efe-0197-4a54-91de-5dc6c461edcb · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76dc3b8-f695-4eaf-9c72-ed9ca8af8034 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdbc36e7-7bd3-4698-8f73-99efc9412160 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f7a83e-3c1a-491f-b7cb-c54ce9cbffa4 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Gonzalez, and Ion Stoica
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37398d8d-7290-4226-9e4f-f32733487964 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Training Verifiers to Solve Math Word Problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 101de2ce-124e-4cdb-bbc9-9483cb8b18f7 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model UltraFeedback : Boosting language models with scaled ai feedback, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08df7141-6ad8-49f9-a21b-05ed2eff8acc · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Safe RLHF : Safe reinforcement learning from human feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07750255-4877-4540-a978-a9ddea935e20 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Nvlm: Open frontier-class multimodal llms, 2024 b
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 442d64e1-b8ce-4923-8602-cc3b41276550 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 612bf3fb-f15e-4b05-85a7-cc4a788cbf1d · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c589c456-a654-476e-8903-177fb877283d · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Enhancing Large Vision Language Models with Self-Training on Image Comprehension
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e9e2d1-9db3-4af9-8c46-fbdf91f4c118 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa988a50-a441-4095-ad9d-cae9b58ebecb · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Understanding dataset difficulty with V -usable information, 2022
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09bbe186-d714-4269-9934-d6eec117ef55 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Multi-modal hallucination control by visual information grounding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5dca67c-1523-4ecc-b463-6978652be895 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Gemini-2.0-Flash https://deepmind.google/technologies/gemini/flash/, 2024
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fee390-a7a4-460f-ad96-88ef3d2a3164 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Reinforced Self-Training (ReST) for Language Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af39565-f021-4e93-8634-6d84e6b00ba6 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model M-RewardBench: Evaluating Reward Models in Multilingual Settings
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a048fb1b-5180-445b-9358-f086841cb616 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model V-STaR: Training Verifiers for Self-Taught Reasoners
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab04001-84f7-41e3-84a3-84d0aec55dd9 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ChatGLM-RLHF : Practices of aligning large language models with human feedback, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3e020b-5eb3-4fad-b5f0-ef71c55ce7e2 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model GPT-4o System Card
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc4f9b0-36ae-42ec-829d-49e72306cffa · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Smith, Iz Beltagy, and Hannaneh Hajishirzi
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515c7b57-8268-48ba-b500-97056909de4c · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4099f462-0776-4890-a203-4333cab0c352 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4518ad0f-f119-4320-b745-b8c985a49f7a · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MiraData : A large-scale video dataset with long durations and structured captions, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49243fb7-a149-4e87-85ce-394bb0ddaab3 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model DVQA : Understanding data visualizations via question answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e59b7f5-0657-40b7-8318-4f87ee28cb06 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model A diagram is worth a dozen images
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb7bcf8-7a8d-41fe-892b-b1944122cbad · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79cb0ae7-ba2d-4a4c-9e3b-3e7fff583507 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling, 2023
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa318f91-e11a-414e-bf44-60a684de6eb6 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5c59c3-738b-47c2-b346-5277a2cd6459 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RewardBench: Evaluating Reward Models for Language Modeling
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5a167b-871e-413b-b7ca-4b33653c03ec · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model LLaVA-OneVision : Easy visual task transfer, 2024 a
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c825286-f9a1-427e-be53-46c0f39f8f43 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1188d9ef-0642-4a17-9831-73a078874489 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 285d6ba9-b35a-448c-b8c0-b9ae15f66b35 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Super-CLEVR : A virtual benchmark to diagnose domain robustness in visual reasoning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7408baa0-c25c-497a-9826-7969589fceff · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Let's Verify Step by Step
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d25e91-ad11-431e-b94a-e505c2517387 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77e1440-cd08-48f9-ada1-a3249b3b1051 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4881fd41-fa58-4911-8ffc-e987a128768e · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model OCRBench : on the hidden mystery of ocr in large multimodal models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8995f9b6-2354-4758-aae2-58978eb21bde · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model POINTS1.5: Building a Vision-Language Model towards Real World Applications
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c25ada88-2319-46cf-873e-afa009daf17e · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f488368e-1c0a-421e-b1ed-9a0c9a8a92ab · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MMBench : Is your multi-modal model an all-around player? In ECCV, 2025
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8cf238-711c-4185-a867-8e79beac2d7c · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9f5248-de18-4fda-8e3f-797c009e6805 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78288ff8-3bd9-472c-990b-3f5217db9ca7 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad19d6e-c128-4786-bf34-099f23e33c51 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 219b968c-698f-47ae-aa01-efe93dbd34e2 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7957d2ce-2237-420d-9ad5-bb17187ffd6f · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc4050a-55aa-4029-b674-5a00e3cd283f · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Ovis: Structural embedding alignment for multimodal large language model
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1efae7e5-581b-43bc-89f4-76ecca9cbd5e · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e949c44f-8a1d-4595-a56b-36ad88e827eb · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b622ca-28c3-4b42-9a7c-22c64cb84bff · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 478c6912-1417-4b61-8c53-f9c31d908cac · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Unresolved cited work
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20922ed5-4389-4107-afb0-163b5a88e166 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model CLIP-DPO : Vision-language models as a source of preference for fixing hallucinations in lvlms
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875b6e08-ce30-4bf4-b665-dd00535debef · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Training language models to follow instructions with human feedback
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de65f01f-f3d3-4c6b-844f-54778dde1ff8 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Strengthening multimodal large language model with bootstrapped preference optimization
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90818076-b40e-4a5c-b924-ee2dc2f1b00b · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b34e5c-83b3-406b-9389-5919afe1a63a · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Learning transferable visual models from natural language supervision
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 436b17c0-ef21-4f14-ac51-d3522ffac504 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Direct P reference O ptimization: Your language model is secretly a reward model
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57fcecd5-43c9-4c9f-b8b7-760c3e770cd1 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Proximal Policy Optimization Algorithms
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa2f087-19f3-4dc7-8096-2c346e092f35 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model High-dimensional continuous control using generalized advantage estimation, 2018
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b4be7d-6b44-4480-99be-4062f0d309e7 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model A-OKVQA : A benchmark for visual question answering using world knowledge
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11fe6856-b491-4455-a76b-99b7f905281f · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model SenseNova
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20aac059-6c36-4088-b322-d116401722ea · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5ce93d-80b4-4605-a06b-220c229a09d1 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model KVQA : Knowledge-aware visual question answering
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fdbdef-aa39-408c-bea1-f7f04cf74be7 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbdc80d1-d44c-4094-b355-3effe39c1038 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Skywork critic model series
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5edfb6b-7ae8-42c9-a183-44a3b654b76d · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Towards vqa models that can read
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267e7d76-4d72-4135-a056-8a1a22a9c20e · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cbd0ccb-bf5e-4f99-873c-1821f29e00b6 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e867eb92-46ed-4f6b-9234-3fac4742bcee · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98ce674b-7312-45ac-a662-cf2e0c98db37 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024 a
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039d1596-53f5-4aa0-b7bd-4e749315c2df · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model The llama 3 herd of models, 2024 b
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1271437d-7acb-49df-90eb-120a72575184 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy, 2024 c
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6d858170-afe5-4fe8-ab8c-005917015ef1 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Solving math word problems with process- and outcome-based feedback
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d005002-d81d-490c-bf4b-5ddfabecd607 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e7092d19-35bc-4643-861b-ceaa4b504d6c · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26683e97-4e56-44a1-b3be-b18a534c2d7b · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb318e60-307e-4a39-9f51-7a5e4e143791 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41efb4af-7311-4954-9367-87bf666197c7 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Self-Taught Evaluators
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e9a852-64f1-484d-8d9a-abe2624e34f2 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model HelpSteer2-Preference: Complementing Ratings with Preferences
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90cb740-cdbf-4b82-b858-58f83c056816 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model FunQA : Towards surprising video comprehension, 2024
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e6dbfc2c-f213-427f-ab31-493ba0e3dfa8 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1cacfed-d5a6-4c9a-857a-9f80a2ba00dc · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a82b675-5cce-414e-ab4f-4594492fb324 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model SUTD-TrafficQA : A question answering benchmark and an efficient network for video reasoning over traffic events
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 731f932f-5711-447a-83bc-95fbdd1e9fd8 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Inf outcome reward model
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0a0836de-eb21-4637-92f4-08c00797e542 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Regularizing hidden states enables learning generalizable reward model for llms
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation adf2a856-1306-4899-b6f5-f6335a026060 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model MiniCPM-V : A gpt-4v level mllm on your phone
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2b79ee4-74e8-4746-80e5-7c071fa66362 · outbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model RlHF-V : Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7f27ffdb-7d2c-42c1-b6e2-3c18214a288d · inbound
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d4118f63-dd22-44cc-943a-d3c28a5d7eef · inbound
VideoRoPE: What Makes for Good Video Rotary Position Embedding? InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3224bb3e-e772-43d0-8163-b0098b5be323 · inbound
Visual-RFT: Visual Reinforcement Fine-Tuning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7a26ca15-e149-4653-b3f7-d1fbdafb6563 · inbound
Unified Reward Model for Multimodal Understanding and Generation InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a697026-1be9-46ed-b0f5-473b729784fc · inbound
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b0068df-7793-450c-97e6-2e5c84a92d2c · inbound
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff096c44-57f5-4d46-8d45-e53e287d1424 · inbound
MR. Judge: Multimodal Reasoner as a Judge InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4023cf70-a35a-48ef-8c9b-7675d116fb4a · inbound
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af863026-8903-46e6-adf7-f264d3c4b139 · inbound
Visual Agentic Reinforcement Fine-Tuning InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b405f5b-deab-4f49-88d5-95f077449607 · inbound
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62de62e2-b0ec-4e6f-8e41-c84fe1f9368e · inbound
Generative RLHF-V: Learning Principles from Multi-modal Human Preference InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38988312-0153-4080-8c93-3efae3d66d08 · inbound
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · inbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8581371-e20e-49bd-b5d3-ab2602db9964 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 138
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d5e40fb-ea29-4f77-b6e6-136cf38cf54c · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd59930-c41a-4f52-a69c-cdfaf9472107 · inbound
AdsQA: Towards Advertisement Video Understanding InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ac583ed-ae6f-4ae6-84c8-ee1303df5c55 · inbound
Visual-ERM: Reward Modeling for Visual Equivalence InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4c2c0386-ae29-4357-96f2-a12ded5569e9 · inbound
DRM: Diffusion-based Reward Model With Step-wise Guidance InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ef200f6-b981-4852-870a-3bae331782cf · inbound
DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bd273c75-6f75-4bf8-ac52-952a1847c676 · inbound
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 225973e3-99ce-490d-bfa7-bbe563a68028 · inbound
Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.