Pith. sign in

Paper Citation Record · LEDGER

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

As of 10 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2506.04280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04280 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:20.044416Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:54:43.474266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:12:50.323154Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5391bfd9-4e4c-4283-b50a-84b1f44de3e1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.093166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.093166Z digest=sha256:a21851e3705f6e362095392b7dc78dbe6a9b1014460f335b3f4f48e88e7ae793

Observation c72cb348-886c-46ff-9438-363f60a1b434 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.125992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.125992Z digest=sha256:f0c567a7b8023f2d462c11d37e9f7b4f6d0dc538be92f16f1b56dcda551ca6a7

Observation c92bdff2-402b-46bf-92af-c5ffe95e425b · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.149385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.149385Z digest=sha256:e60e3568da8e798be9405e1941d4760620040d192d53df1cf56814c4bedb76fd

Observation c657ea92-7c09-48a2-9414-40d49c3df2a0 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.168229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.168229Z digest=sha256:c5fc7f93d3434699e2780a960e12208bf7d152f84824d851f8df29501d7c19d4

Observation 449cf1fb-3f45-442d-9e5a-69ee1fe9b17a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.206037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.206037Z digest=sha256:6fac503449c6844151363b127df2f489574d62f303e3094ff61fb7dcaf78ccca

Observation b2332d1c-679a-49e9-9264-e55557eb3bbf · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.314570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.221144Z digest=sha256:86a46f4b16062868792890dec1eb7f3fabababcad6869180a13c4f70b432acf1

Observation aee44f84-b50e-4f52-8e75-f1d730020eeb · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.232426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.232426Z digest=sha256:b334989f6eef6fffcc12b4cf9f51c14b134a8686ddca8bd08375e6a37ef863c0

Observation 4f0b9bfc-e67f-4693-b269-3faff3a84983 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.246975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.246975Z digest=sha256:5fe5732816c611957e0870241f103ec01fbfe2384a8649cd08cf8f5bb59b06b1

Observation 3769c6c6-8602-491c-8a98-64c15c0dde53 · outbound

This paper cites Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:05:21.666096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.279701Z digest=sha256:9d312c7876f1c8b0861e284f5fd28da14ffe9e12d2413ce851596bf19fd9afc5

Observation b83ffd3f-dd39-4fec-9365-ec175c24b1c8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.286471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.286471Z digest=sha256:f684affd9aaea1b9ce411872a06ec267d348446c149ad3e1bf4e54718e350720

Observation 87f573ed-96f7-49dc-b444-396ddc047ac6 · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.291846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.291846Z digest=sha256:862e9666915f06eba336c15727f0af6595eab4114d0440f1370bf93c0f3371a7

Observation 2a463a9b-c158-4504-b449-7030e94417d5 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.067313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.302473Z digest=sha256:a47f2afcaea88c089c0cdaac78e3a3d226b9b039aa3e277bf3c7cbf1361b453d

Observation b7a365eb-80de-45f2-8806-193219db6b68 · outbound

This paper cites OpenAI o1 System Card.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.318506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.318506Z digest=sha256:1049ef4a01f5abd17b9673f75d7e11aa135bc7ab8e811b89ac46f41632761aa9

Observation 3783139e-c5f8-4251-832e-b1894c57c50b · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.352234Z digest=sha256:185487e9a4c21eab7661d74092d15ff1bf03ed573192237ab38d8a36177cfffa

Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.357259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.357259Z digest=sha256:c282bc481d5144ae34ab2915b9fd89d70d4735f9e098f031c681ef162e35c3a0

Observation 7a7f698c-e082-4f45-9278-6ee28fb1844e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.373052Z digest=sha256:64cb24628d534f8a16a6926508247f9dcde834a94d9bdcfe6e7f87ddf980c902

Observation b7622b24-9377-4fb8-bc24-f16a898b13f7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.396582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.396582Z digest=sha256:a4f97a28542387fa97b96d1a3fb59ba0a1cff139ca3c50b0d1f32d52a1c6c2b0

Observation e7821c05-6750-4e3b-91b5-d3fde00a313e · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.405502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.405502Z digest=sha256:42dd43e6c3372190e8622d31243a631e9cc53f091f4771531b3801605b6b16ff

Observation e2f30b36-44ac-44b6-a637-bb39ff2dd03a · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.423016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.423016Z digest=sha256:88429520daac26f78a9fb05802020042e8dc25c532fc189f7ba9efbf6c2c8322

Observation d1523ac9-2457-485e-81b5-2463cdc7d8fe · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.986394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.386373Z digest=sha256:34fcb9726f46205d71c7f0e6a18cbd4ed3b0c5b15aaf8fc9e262917bdf5c1414

Observation 813a2978-95cb-40c9-b6a3-adacfaeb7e29 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.889599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.445145Z digest=sha256:60d03a5f6f2ae4e16b2d8dbaecc755b0004ca15c85c51fd745e9630529bb29fc

Observation 55e9a714-8997-4266-9b20-f54251a95a86 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.462282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.462282Z digest=sha256:8e36f6cc89512ebe516d6d3f8de930c0bc4085b5b752181e720db74d6c159d57

Observation 57dc166e-b0ca-4809-8d63-1189cd44c09b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.472918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.472918Z digest=sha256:b7e9c0cfcbc987238ba1417362a6252352a81d727050f2964ba5f1d22a614cfd

Observation 934bdc89-a9a3-4304-b26a-b210f382595f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.927157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.432019Z digest=sha256:3a4bb49dc732818f96f67913f08c035c8ddc7da6ce9a6b757efcb6e0811d0890

Observation 383301d0-48b3-47a5-a9fa-aff5a789d17f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.810498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.491198Z digest=sha256:60af62398b5305a75f6bbe003f0fb17b62c6bddae6f774313966ef4a7bfd03c1

Observation 28d0fb4b-8fb2-40ad-906b-aae4f0d84da1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.499218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.499218Z digest=sha256:4535c92d2b664c534fae04ad002863e089a77d8143921b61d28b1d7ba37f903c

Observation 0d9ba4f4-81b9-478f-84f5-54dd0f3419fd · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MileBench: Benchmarking MLLMs in Long Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.523326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.523326Z digest=sha256:5918ccfdadb8a4522271cf2cc7b6dd75f1092a5f08ba8ab7b0674de76b82fb01

Observation 7ac23c4c-fbae-43fb-bea4-f92333afc950 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.483720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.483720Z digest=sha256:db0d4f2c3af97d2715e062ea328171170dc39dd694785e9826b07564d27a4a13

Observation bed07209-e2ff-4850-9771-6ba94f4008fc · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.744082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.546625Z digest=sha256:74a5d122a178119544cc209380e6bb856a7e1ce57489d4058010cd9f9811ac09

Observation 2a7da2ce-8183-4bf8-b05e-358ed8c1ce6e · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.555184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.555184Z digest=sha256:9c8cb6fd8fd96e99b16ee71f67e33abcbe163f357ae29359559c87c83a2ffd9b

Observation 2c0f76dd-9160-4dd7-8fd6-571bc8943ae7 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.565467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.565467Z digest=sha256:decab7f9c9d1a6a753a8e733fc67ecde31617ab69fc30bab29f9c2d273f51389

Observation 991b90ed-dccd-416d-b568-5e449d05da9f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.580744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.580744Z digest=sha256:be718925d7f48b9fcc69e3cd7a20778603c08ce17727f5bb8a28af8c6118675d

Observation 072f0cd2-45ba-4135-88d6-0ead62ec9e8c · outbound

This paper cites NLVR2 Visual Bias Analysis.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark NLVR2 Visual Bias Analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.537088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.537088Z digest=sha256:9d6e6a3745ce5c8acab1138077cf0c65cc296b423cec347f3df5d53d46ad5afa

Observation 6290e526-21b7-4b54-ad8e-4f63cb734c17 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.613480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.613480Z digest=sha256:edfe16bbf2c0ef5b1304995dc92f2c42e4fe090cd72d92cdaff9e840138a9cb6

Observation 0c55dbbe-cfdc-41e9-8a47-d5c02a572196 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.676752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.622195Z digest=sha256:2ee6cd208be5865159f30b3c8bc295a11c38f25b86c6423510a3bf589f42a468

Observation 4a5f715c-91b1-4750-bea6-34f1d1fa0f3d · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.643578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.643578Z digest=sha256:be48c335e269b6fec4f39edc1c1edbf23094d1c463abea0981459e794821e848

Observation c7c041f0-6fd7-491e-b307-15c86595814e · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.654435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.654435Z digest=sha256:b0bf499d8c6cda39c75b7ab11b69a66870ab11f1b14fc75c922ce6218a95b046

Observation daac543d-7b2b-4ab2-a972-1effe3b51c83 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.601376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.601376Z digest=sha256:7e180679db0e230cdc955f0047860d78d5bd502cee408aa532e13c3e045a678c

Observation 0921162d-cacd-4676-966b-647ed987a204 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.683399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.683399Z digest=sha256:8e3438137f520bfcee07b258274faadc58e36f55f6ac4eec0176a0a6316de60d

Observation 7278bfc9-b506-4303-a6a8-b0c457cedc3d · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.704763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.704763Z digest=sha256:da397d7d1f39bdd5a5436a87a10162e337d86db3fd3b912c9cb8afd0b6de7dcf

Observation 7e2399a7-c682-4cf5-9248-a7ad6c9ad0bf · outbound

This paper cites Qwen3 Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.713959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.713959Z digest=sha256:85281d11ef2a20d1972db2f56ceadfa364890c2a8ac440afb6b71202ed3e30a1

Observation 4f95da42-0fbe-49f0-b7a7-b9d955e24ffa · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.719944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.719944Z digest=sha256:bd31cc6a71a753bb38fccc433626e208de002b04f7a8644f89e79dbbd0dbdf62

Observation 406e5081-45d4-448a-93c0-300a95aa87e7 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.667340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.667340Z digest=sha256:7db180232557d7f41d262b58ba54be6e2a44448dcd2da9548a268e56f15bb013

Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.748441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.748441Z digest=sha256:9fd320d91f51bfdaaa0996ea8ea7c63d6a54995bdf5da980e60ab7cced30a263

Observation b207ca52-e644-403c-a340-c47554341d74 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.759069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.759069Z digest=sha256:a1b81bd557d2818f2a7398f7db59186fb968eede4331c0ecb512cbb968cdba91

Observation 04bcbfa7-9314-4b0c-83ac-9331cbf40313 · outbound

This paper cites WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.768115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.768115Z digest=sha256:c2a97e12ce3de46de8b0126ca90dde2c076047fa5d53489a63c758ced6498ddb

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:f2ea15529c6370807da9f806a54d80f7c8d315ce8fc9c2a1e873c6b9f0ad4283

Observation 674a23ef-a859-4fa0-abfb-8c3dd55cf2d9 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.737053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.737053Z digest=sha256:1846a85e0b53f7f1781582c2d659d2fe64af79161eabe6a9c76e8a7adabbcae7

Observation 374273db-e653-4a63-945d-d745994c824d · outbound

This paper cites MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.808950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.808950Z digest=sha256:bdaa2f05be318cbb57c2f871bcf680b0291aa3dbf54ca9524b267edfd5e62093

Observation 242fbcc3-9161-4688-add2-97f2000fa578 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.821194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.821194Z digest=sha256:e1a0ca16319f157c1bbfe25881414a372189fef4db2272f52fa58c5b8f10dd97

Observation 48a225fb-5b1d-4cf3-8168-4a210622fcce · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.790349Z digest=sha256:884e55aa91dbca26c89e8220d22e0341f8b303a731c478a2b7e7faf23cf2a347

Observation 1e7286e3-27d0-49ca-b3d0-4040e9b2d71e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.541576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.834814Z digest=sha256:1140e67893d238f793b2f84f38dad4c1c4f28e7c0c300ebcf23ac09856ca83ca

Observation 9afa9931-c062-408b-a212-7cdd30230c4c · outbound

This paper cites reasoning step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark reasoning step

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.504739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.856925Z digest=sha256:4d23b65e3681d1b62329f67c5d920f092148df11a4d8bd8a68631d397bb054c5

Observation c25669a8-8996-44f7-be32-3c9ab4488f32 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.311098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.900673Z digest=sha256:34d1ca39756f317bbe7b85d34917c82e35c2f49525d3bb5d8de34d307cf43397

Observation 09ab0e9d-67e1-47e7-9269-464b8458d6a1 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.259112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.946132Z digest=sha256:ab76002130884cbb00c95880c664e7410dd8bd9b83219b903f7155ba345fac1a

Observation 2cb14538-63db-429e-98ba-48aa6e23b6e2 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.465449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.961296Z digest=sha256:e76e7f98e7577121b5770a5ac860a0c3d848253762832cf0e51c4b59b42cae69

Observation c130aa28-c9d8-4f8c-93ac-78e6f1e19275 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.395288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.971376Z digest=sha256:bacf607dcf2e48b44cc5b289b9521303cd9cba0ab4c6ceac59da30af7c258a37

Observation b7038d16-669d-460c-a5e0-4ef8f49abba8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.363949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.980113Z digest=sha256:9861dd7b1647ee9462a4a960ac877064acfc49f4e8be57d8f080e9d3b5a71a6a

Observation 56aba37e-d952-4cb6-ad57-00930f296933 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.199841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.988591Z digest=sha256:585d6ae1be18389fdcb6a93eec9f8abb53b3cfb091297eff699eca95fedd40e2

Observation d4e85cc4-f8a5-487e-b582-eee8aea80311 · outbound

This paper cites Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.116840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:20.013797Z digest=sha256:02b6bb438687b43005d2fc5ef83d7768313b22ccd5d8c451d07a969cdde21153

Observation 78a0e081-f8d1-4af4-bf4f-192fa85a8888 · outbound

This paper cites Rank higher the responses that least misrepresent these relationships.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Rank higher the responses that least misrepresent these relationships

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.066928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:20.026221Z digest=sha256:ab340ff2f9976d3c9d191ab86bddab3a433a645bd2d6cba4fd6ffb534e213ba0

Observation 88f71158-6ef9-430c-88b4-7f57a2d86032 · outbound

This paper cites Responses should avoid inaccuracies in describing the characteristics of the objects present.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should avoid inaccuracies in describing the characteristics of the objects present

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.020746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:20.034083Z digest=sha256:44f3c1666200b36f978713bba1b8713d0580fa7e1545cb1400c390ed5617e556

Observation edc4972b-8872-4d0a-b8bc-92ad31259023 · outbound

This paper cites Equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Equally good

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:21.962642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:20.044416Z digest=sha256:fba1325ddd97e4e0e4099b7a51e39bb947375646c35b5e510abcdec47afcdf8b

Observation 7994fa18-e622-4d8b-b9ee-64d433200258 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.509807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.509807Z digest=sha256:a76caade1d12b80ec372e35a43b2520ccd5bcce5eb2bffb6b98a3d52b2ae3b3e

Observation 3823083d-0afc-468d-bd06-d4c92fc52bff · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.114146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.114146Z digest=sha256:d11f0f7b0545698582948a6d97c87ff6ec91726bca5b28023458359c4903b8b9

Observation c013ae08-676f-4166-8fbe-e7b9fac4652c · outbound

This paper cites InProceedings of the IEEE/CVF international conference on computer vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InProceedings of the IEEE/CVF international conference on computer vision

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:23.155612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T11:05:19.262204Z digest=sha256:b3fc9e4eba181b4be29f3762c0d9b3a919e09d1d12de33a3bba71ff16569ab38

Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.193447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.193447Z digest=sha256:ee8a074fccec1fcd54e8dbb22686e14d0a8895d32ba222f8bfc284b43add0a92

Pith citing papers

Observation 5c6864df-b047-4bee-bf80-fedc21632611 · inbound

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation cites this paper.

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:43.474266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:43.474266Z digest=sha256:7bf9be35f941c37b471fd5fa309018d79e60fbd32b5562574155aef7cf87645e

Observation d334f9d5-9cff-48f6-91d2-3fba26525cc2 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.779717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:de5ce4b6cd7f2f3d24701820679b2a1ac9f84c43002b841440af88cbfcc21d62

Observation 92fa0042-066d-4efb-aa58-7d57f2915315 · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.830119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:70c4c11e370574e17e9d0913a6055a86d49cf4965e739929c5f803db322d3845

Observation 8425e803-0d62-4c85-bd60-4fbd56364dac · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:02.750336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:260bef26308c27b7998ca2c032ed3d687af130f62e4ad1c4bab4aaf83cb60685

Observation a1c2534f-54a4-4485-8371-831cfd49b2d7 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.324575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:14d1b432f768a85102f4098371240de73bc6d45937a339238ac3a2b74a0452b3