Pith. sign in

Paper Citation Record · LEDGER

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

As of 18 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 3 inbound Pith citation observations for arXiv:2505.24164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24164 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:37:50.854503Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:24:08.154904Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T18:13:44.857313Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a205b59b-d940-44ec-b20b-b1ce05113763 · outbound

This paper cites GPT-4 Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:42.783123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:42.783123Z digest=sha256:73b63e1a5386b14b2a2a2fa565a6d1f0f3ccfffe1d5feb2ec8d7135da56e616b

Observation 5b68f802-f117-4d20-9c57-9b2f6fc45502 · outbound

This paper cites Qwen Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:42.873209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:42.873209Z digest=sha256:9afb4f1eecf676950b28ebafe193c77bc8eb2ddf1db7b83b8453d57f567dd926

Observation bc5380aa-495b-403f-b824-c951af462fd2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:42.957030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:42.957030Z digest=sha256:99b7d572819a739c0411c6e9fd01e449226423bfd11cbb1a99c203beb18211d1

Observation a354a0fe-2f23-4469-9fb5-6d57164bfc21 · outbound

This paper cites MapQA: A Dataset for Question Answering on Choropleth Maps.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MapQA: A Dataset for Question Answering on Choropleth Maps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.047242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.047242Z digest=sha256:fd953b847af1577e7bc8c9f1c416d084839a1e8d7a60dc4741fdc5f79bbe4570

Observation 0ba50aa3-3999-460d-a51b-f97816fe33cf · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.118516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.118516Z digest=sha256:652a928de0f73acbc7190893e7521eeb7b2bb6d83c35bbc1359707751e45a1b5

Observation c960d2c9-ada4-4dc6-b172-2fc19031cb0f · outbound

This paper cites UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.218009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.218009Z digest=sha256:1274ed6b906522c2d0d3b46b66b02a51e44c9378f8e95bddb62f8ffaac6b7ec9

Observation e4615098-7cfd-4dfc-abb7-03ccc8f23df0 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.332355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.332355Z digest=sha256:0d37f526153de6a4c14a9c41fa2ab4b9dbab6f49ce046cca1045f74c5f416773

Observation 910e8922-4bdd-4d7c-95a5-cd450e5e09d3 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.https://github.com/Deep-Agent/R1-V, 2025.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models R1-v: Reinforcing super generalization ability in vision-language models with less than $3.https://github.com/Deep-Agent/R1-V, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:56.305112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:43.437869Z digest=sha256:5a82159824d471c8a74beb31c0ec1357b399aac01971717a7e5c71f3fdac03b3

Observation ad424351-a347-45cf-9118-0edb67921a5c · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.536170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.536170Z digest=sha256:91e75669467814054c036528adb5084c62ae73ad7f64435c912f1bb8edaca9b8

Observation 82dd815f-cc4c-4aac-a85e-a3dd84af6992 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.644290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.644290Z digest=sha256:9f697ba0d670e9c79eaa8bd63206ad760b91d7c816e4866580f12a8c1cb977f7

Observation d1e9d65b-6059-4ad6-be0b-c45f35d4a376 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:56.125085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:43.735687Z digest=sha256:8d1996e153ae88a0c782bcf07088a3723f64573eb8ae7b2e309cb1d697e015f1

Observation 3531540a-cfb7-48ce-8f1b-fe894f167470 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.949527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:43.800014Z digest=sha256:15b848918929499dbc25ed6f2c7ff842fad9748fcfa6f567d0059b19e19a0a7c

Observation 271734de-2d98-44a3-b87c-25be0653fdd4 · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.748517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:43.845422Z digest=sha256:bd5fac9d4f4c8ce842159c55bd6fc99682384e04268180e57bea1fe366305fe1

Observation d29e0a8c-db58-4aaa-aa98-07b0f9a8dac0 · outbound

This paper cites Bert: Pre-training of deep bidirec- tional transformers for language understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Bert: Pre-training of deep bidirec- tional transformers for language understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.599809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:43.920685Z digest=sha256:6f8d199aae066518e2faf6da99f358cefcef307fb525ae1788404630fce04f13

Observation 1b27fa32-f58d-4a23-b944-4dd3a2c638d1 · outbound

This paper cites On Path to Multimodal Generalist: General-Level and General-Bench.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models On Path to Multimodal Generalist: General-Level and General-Bench

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:43.985549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:43.985549Z digest=sha256:23c083c8dbd0600782a11583c3c5904b20a960029a3beb32f0fd3337c2582c37

Observation fa40b409-6c78-442b-84d4-fc1701f09a5a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.072119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.072119Z digest=sha256:2f287126c2543881f19611a99c9b305c0d1175d03b1e4dc58c79a6f8d6520fae

Observation 2d9e2b31-4beb-472a-9052-bcc9a9d5f27c · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.110624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.110624Z digest=sha256:5b94fac58e15b74be08d710682ec920fe24d3f367be4c945c80ac2ccf75677b2

Observation 701d8e6c-81b1-4d7a-85e4-20d51bce3fd2 · outbound

This paper cites Cantor: Inspiring multimodal chain-of-thought of mllm.ACM MM, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Cantor: Inspiring multimodal chain-of-thought of mllm.ACM MM, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.458081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:44.208714Z digest=sha256:0cc7af5b935d47fc958c109da7dd74649e7f3fa4cdef3adf6735ffce68d0d5a0

Observation ba33b5e7-45dd-4ffc-b33e-d25f34b1d1ed · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.289867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:44.295082Z digest=sha256:636c26e83f48c0debfdad9bc525800625c4a0229818a2052cac396f5eb39714e

Observation 0812ec8e-1fea-492a-9148-2110ca2e4d0d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.393355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.393355Z digest=sha256:476b7a044ef850bac0f4fc9638523e4c75e1008137e8750772eb7919ff5dd56b

Observation 9848067d-40d2-44bd-a3a3-aac8c40b6c7f · outbound

This paper cites Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Free Video-LLM: Prompt-guided Visual Perception for Efficient Training-free Video LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.473998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.473998Z digest=sha256:0f2031064e5ab9f2d0891736b79bec0f0563d8f2be6346a061ec97e882f528a5

Observation 3fa536d7-b889-42b8-afff-03f25b2f2e80 · outbound

This paper cites GPT-4o System Card.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.516429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.516429Z digest=sha256:cd7590c5375007ca46c9a2cbebaa3ebd4599a5a826110f4ce6a39e145102c509

Observation d1dc0482-7b39-4693-b682-578cc52782db · outbound

This paper cites Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.565913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.565913Z digest=sha256:b366fca341a7be9a4dc5deef832c1871404c9cad64296ac1ac0706f25c18e096

Observation 045da288-d805-43d5-872b-449c83ab6d74 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:55.115246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:44.642885Z digest=sha256:0af43264989431b43b4ae1bbeaa10950d149a8acd1bc36177688be2302bf2a07

Observation bacc674a-5048-48a6-8529-75e556c96760 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.703110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.703110Z digest=sha256:d3be94fefbf70957c9f92301b40f93289f08cc9529c3b692c8e2303dd6146bb4

Observation 0140bf2a-a5ea-4d8a-a649-d9bf3e65de77 · outbound

This paper cites A diagram is worth a dozen images.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models A diagram is worth a dozen images

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.951630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:44.746200Z digest=sha256:e0ee03f2c9e0f0a35b8e1403b4628062fa253b4ace2927b1622f68e839545031

Observation 080d880f-5ecb-4032-a0cc-3c6fb3a37e2f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.784518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.784518Z digest=sha256:d6c0331236e6106a837944abdced36690f2fa1ee16faaccc0591c73f3f35cc0c

Observation fb4f7a2e-b358-47af-b11e-7b046b6e82c5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.857093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.857093Z digest=sha256:2189b128e639cfe89aefed75581d78e72d19129a741b39bb4a922930f6c3f843

Observation 2ab70df7-2694-4d02-9eac-22b516c053c1 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.794908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:44.902929Z digest=sha256:40c6ce1d0e22e0835fdc920d470ffc3bf1ca16e67ea8ee478333d753a40c5058

Observation 059e7857-0160-4348-a977-ab5a86d5b89e · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:44.976247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:44.976247Z digest=sha256:5188cb67e78d09baad9a72394a0a3a35a7edea609f3fa40e2719f542f9f02ea9

Observation 311990ad-275b-4c1d-943d-b49c5c529f35 · outbound

This paper cites DeepSeek-V3 Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:45.063410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:45.063410Z digest=sha256:8cd98759e8f2136021c3aba0e28f0b676129e56851e88e12326ca1e065aa2800

Observation e3d474fd-64fa-4b36-b4b2-e8e5f73f7bed · outbound

This paper cites Visual spatial reasoning.TACL, 2023.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Visual spatial reasoning.TACL, 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.625675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:45.149681Z digest=sha256:1ada542c1b4a12b7e4d1b12da790bf8215aef63e90791058b851e5102ee1a286

Observation f8a42cc8-70a8-4716-91f5-9b8cd2cf7515 · outbound

This paper cites Improved baselines with visual instruction tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Improved baselines with visual instruction tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.441783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:45.152810Z digest=sha256:1aa89ed60b2271735e18f7211b57d9655f456009392bb9bf6bb3aa3dce4e2571

Observation 848b80bc-c73f-4377-93be-7c36b161b31b · outbound

This paper cites Visual instruction tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:45.234152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:45.234152Z digest=sha256:91490b4ed23f70069ea9552e670a29ec24a8abb1b5de52d1ebf291bb888c10ee

Observation f937c7b9-8cfc-4f61-9ad7-50ae0c9f615b · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? In ECCV, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.262844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:45.396245Z digest=sha256:233b7ed083a857ce45aab9b416232ab5f18848f2782b05dd29a4e5914f4ba675

Observation aec24713-eebd-41a3-a6c3-f838f33c3b54 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:54.130873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:45.618007Z digest=sha256:0fdc8660ddf0448980f8aa4b67cea91dfe4b154d2115d8d8a4c225bfd24af43c

Observation 49741e44-f24e-421a-9897-31100a89f9e4 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:45.787946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:45.787946Z digest=sha256:b711abc75faaf86bb62ac014cc42a3ace4713ad7d77b10d9859bb9d7ab6e4f09

Observation 14856869-d85c-4176-9161-29c064ad3d38 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.983355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:45.960301Z digest=sha256:eae10a8ac4c17a55cbd5fe58950ba8d9bd5254613a986195db76e58741883500

Observation 6a495537-6ea1-4de9-9765-6701e7bbcf4e · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.808532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:46.128342Z digest=sha256:45ffb22d821ed789ebd8beca30a8607dbbd985811ec561d5c7a5a56a5dcc7f4c

Observation 4a02b294-cac4-4826-b864-84262d5b913c · outbound

This paper cites Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Cheap and quick: Efficient vision-language instruction tuning for large language models.NeurIPS, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.634049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:46.363496Z digest=sha256:ed240ba96b00bf854ec12350037e93b38ce90332260941e2dbbba8fda7e182dd

Observation 7d0232be-36ae-4f25-a436-a166fe6a183e · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.509499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.509499Z digest=sha256:cdefdcc98f9b7ad9f33feeefc9d2bba057d8a15e7a3dac082fe402bfdc6a52a6

Observation 90b0c37c-04c1-4dc9-93d3-0e3233b263b0 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Video-rag: Visually-aligned retrieval-augmented long video comprehension.arXiv preprint arXiv:2411.13093, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.617047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.617047Z digest=sha256:ea43fccabb1aed6af07f9a11f6a3df7d9175abdbc022c2a0eea0c6c15c914b42

Observation a972f16f-ab27-413d-8770-72947fbea608 · outbound

This paper cites MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:46.790491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:46.790491Z digest=sha256:998f044e71086f0f7e265faa8821ed5ec401ecdbd0033ac6f937b89a684c0f15

Observation db05e5a2-683a-422b-9098-eb214b87e391 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Generation and comprehension of unambiguous object descriptions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.474140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:46.908031Z digest=sha256:74d6e9d9d3f82d2e2eb09cd121fbce87a43a42facab3d0f9ea347eb56f4f2e98

Observation 6aa570e5-7757-4a4d-9507-a2c62a29d32a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:47.054610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:47.054610Z digest=sha256:51106d9d1b5a4040de25496d7b395d19228ba6296414c06d3a43a64e8ad330ba

Observation feb65134-920d-4bf9-8fbb-cd50e46492e0 · outbound

This paper cites Infographicvqa.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Infographicvqa

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.318799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:47.219388Z digest=sha256:0a3e30477b453a2308acdabd9943e61dbe4e67c4f18d62008e7ef916a60468c2

Observation eb606e72-c086-4dac-bfbc-73847b3c9101 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Docvqa: A dataset for vqa on document images

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:53.165749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:47.344366Z digest=sha256:960765ad18f8e31690bc2fe09a2a79b8e9ae53d1c0db82b1645a34c0985fe439

Observation a7e68bb9-1055-4b61-88d0-f5951e67213b · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:47.537408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:47.537408Z digest=sha256:7df44978afbbce59f0428c83d1a019f122be0e51085fa39e3ed6d70e02aa4165

Observation 9d097eab-a669-4ed7-9c41-8431050b7f92 · outbound

This paper cites Training language models to follow instructions with human feedback.NeurIPS, 2022.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Training language models to follow instructions with human feedback.NeurIPS, 2022

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.920910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:47.706893Z digest=sha256:33ea66127a33c598e8340fbdf5cd18c29a8f95f07935bb6721bf419d70eb4844

Observation e11f3118-abc1-47d5-8ca4-0dd272df0855 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Learning transferable visual models from natural language supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:47.865340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:47.865340Z digest=sha256:33089ba26de8b321e7627336154816ca67b4aedda52e0d54044a3bfd99212f1b

Observation c0ff8c5d-9f1b-41e6-911a-855075ae6886 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.NeurIPS, 2023.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Direct preference optimization: Your language model is secretly a reward model.NeurIPS, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.695218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:48.030928Z digest=sha256:0e6cf2f48cfe9e8ef8ab815ad983dc6d0844d5ffdf3c3c65d8bc8b67bf901ef3

Observation 879f13ae-d0b9-4b5c-a844-38f843aaf72a · outbound

This paper cites Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.199482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.199482Z digest=sha256:e7a301a2c9854b562d2cd08b9668574048db7ff3d729e156fce7329edf0c793e

Observation 791d5a95-3c79-490a-b88c-e20176d1b599 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.318387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.318387Z digest=sha256:3255c5fbce433c554e020359aced2f4485ac0422b3c2824962ff4c8ffd641d08

Observation a8abd1ea-e136-48d5-a244-bb3c22e22062 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Proximal Policy Optimization Algorithms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.437328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.437328Z digest=sha256:639f336fcfd6b2decc913a6e5735d970cb8b0bfd5e01f5bc5834210add25c923

Observation d0afd5ec-f21b-4d35-8aca-16b8b8b998d0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.531778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.531778Z digest=sha256:21eae3d56c4f314485a0c2c6d687f0b20c05d2c81b3f13c58c5d8fd680e77cdf

Observation 75ff2723-9024-4e4a-883d-f23746d7593b · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.622594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.622594Z digest=sha256:216cb830945418bf7eb0cb5fc5cd3845d5adb0d7b7648520e89846f34c9480ef

Observation 2323e1fd-fc9b-4b47-9264-b649af1f9014 · outbound

This paper cites Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuray.arXiv preprint arXiv:2502.05177, 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.691734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.691734Z digest=sha256:4a733fa7406ff8517def5b82aeec8a01b8a72066d6c5d28448da2df011f66be8

Observation 020874f4-2c31-4234-a8ee-10e3083fa6cc · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.782938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.782938Z digest=sha256:26809d5c4f8c8b136ee9425536768e2415e5f50407cf9d677ea93a5fcffd85e0

Observation 13ea1e57-dedf-46a7-90b6-3388db6d6b98 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.877517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.877517Z digest=sha256:b4a993d28319f6c90c0176a6e6e56da4f457ac025655f4109c6922a7bcfb2543

Observation 61cfd93d-8772-489c-9a56-e7c2efee99cf · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:48.974699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:48.974699Z digest=sha256:fbe6270a34b8d1457ec38edc452e1302d304a2952dee7058baae0930827eb6cd

Observation e22b9b99-6fa5-4feb-9213-419fb69ecb85 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.073537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.073537Z digest=sha256:d21a382c9cb361ff57f94333158e3eb4b06458820b739165ab03c17d5c51a6a8

Observation bc7baed4-4b75-4a71-ad43-1f2c7c87d3b2 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.199326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.199326Z digest=sha256:769976e793b1ac9b8079ac230075990cb011bd4cf46976739c2921724facad82

Observation 9b330a79-d238-4eb8-889c-aa84aefb3fce · outbound

This paper cites Qwen2.5-1M Technical Report.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Qwen2.5-1M Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.263799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.263799Z digest=sha256:d88130f530e391f73b48991d48eda51c32ff779ae9e222abe4679f30b7e38154

Observation 9297dae2-6215-44be-8af7-ec6c6cb9706f · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.357574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.357574Z digest=sha256:9d96d8690000a4ecb6cfe28e15045aac857a2a31977680869d7cb89aa5aa8ce8

Observation 575783b0-dd21-4941-b5eb-464ca8c32aea · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.453135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.453135Z digest=sha256:e699579a48aaf2f6cda4155bfdca3c8aceb936a0be5ccb51f032ded60f0cd944

Observation 6ee38d23-7e1c-47d9-9d2f-37e72a156375 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.581401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.581401Z digest=sha256:f6afee5499ac2c6e543087ec3f08d218bcb3599db16d6cd65a0a669e7ec1f5c9

Observation 20fefdf2-e833-4a4e-aee3-f6208b8e0f46 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:49.696896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:49.696896Z digest=sha256:fcc798626e378a83cf116f7745be2690dd818253c457ed16e25ac5f21a3089ac

Observation 7099775a-08ff-495f-82ff-6a60aa94f3bf · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.485583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:49.827448Z digest=sha256:2b727d08e751f46d35c5fa415b8a958db82ed177857f1d42ac68030a6ff58ba8

Observation 7976af05-89f0-42e2-a8c6-20aa80e3b659 · outbound

This paper cites Sigmoid loss for language image pre-training.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Sigmoid loss for language image pre-training

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.301424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:49.936532Z digest=sha256:301290a23fccea99a56574335167d8537989e6840be5a095d3b08107b1a29a12

Observation 9b77ae93-91de-49b5-8e52-1196ea0c1596 · outbound

This paper cites Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.037158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.037158Z digest=sha256:dad338695348f039cef38f2631d6f6d72b64d619458482cdf056bcdc71597d87

Observation cb348a27-c885-4a0a-9667-9fb0ed96adda · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.123928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.123928Z digest=sha256:562eb04db86a30ff3c9006772a6a64cfcf29ee3e5f7dfc97c2ade3812a7a15e2

Observation 71428098-f6cc-4eec-83ea-f68eedcd226a · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:52.132720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:50.210810Z digest=sha256:4990aa8834d191435813bd2e7b2854178f2c2729c64a6f5ed547238cf296191c

Observation 5984fb9c-150c-4bbd-9f4c-cb0289374dcf · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.273762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.273762Z digest=sha256:c862a8ca721398c3ad6ee5237f2f4783234846e7e6c650cb89ef71904710d6cc

Observation 5be7b6d3-807f-4516-908c-2ae4f0eb2d20 · outbound

This paper cites Enhancing multimodal large language models complex reason via similarity computation.AAAI, 2024.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Enhancing multimodal large language models complex reason via similarity computation.AAAI, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:51.942715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:50.386378Z digest=sha256:c07285d4a11623ec7d8864d786b0ffe0846ad58272e175dad816a243407d726b

Observation 3b8c0bd6-739c-4e66-bd74-6d7741264b85 · outbound

This paper cites MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MultiHiertt: Numerical reasoning over multi hierarchical tabular and textual data

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:51.749703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:50.496263Z digest=sha256:0b906b757e0f0730bc50f2e31f495977ec45ea49f9b8f190f62cc421f5e1ff81

Observation 627beefb-2d62-4fc4-9721-e296d26b16a3 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.564213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.564213Z digest=sha256:c3ecc24e3d20a593565e408445770690b92b8878ca159da3a92943b5a826283e

Observation b1c18b3a-35d8-48a1-9128-c02701f41d19 · outbound

This paper cites Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.658662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.658662Z digest=sha256:dd5ddeda92cfbb6e21ec6928a97c2b876cd685eb65afbc19f217805e8f839e0a

Observation ea093f35-2b5b-40ce-9134-c4dd340489b8 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:50.749824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:50.749824Z digest=sha256:2be20783d7cdf99903c04728ba3db05d23744f4d81effcdfe2ac339be55a5253

Observation 9105b61a-f351-4ab7-be17-3628b1c0f762 · outbound

This paper cites Genimage: A million-scale benchmark for detecting ai-generated image.NeurIPS,.

Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models Genimage: A million-scale benchmark for detecting ai-generated image.NeurIPS,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:37:51.554229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:37:50.854503Z digest=sha256:b74285b49c83e922f565053729b5929a7dfccef801b863df29d23fa9704f6666

Pith citing papers

Observation d2cec0f4-2c7c-45ac-8c25-2662a84eb841 · inbound

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation cites this paper.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.154904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.154904Z digest=sha256:b0bd3afe6e06a410c49f4442d5ef7f7032746a690e2927c6539848f009b3e67d

Observation 47793b81-619c-4a58-94a6-a23602ca3b68 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.940627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.940627Z digest=sha256:b01ec46ee46c61b0d3d7ff2c036333de47f6b5b667bb904e3a075712eb143c2e

Observation 41f95256-3d3a-4c1f-a884-73c828d10848 · inbound

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward cites this paper.

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:13:44.864443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T18:13:44.320002Z digest=sha256:56d8397d151dd91c3675a6be996074f8c15f360d2f8d068ef3d31268af91849e