Pith. sign in

Paper Citation Record · LEDGER

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

As of 8 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2506.04280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04280 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:20.044416Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:54:43.474266Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:12:50.323154Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5391bfd9-4e4c-4283-b50a-84b1f44de3e1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.093166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.093166Z digest=sha256:5f2caf4597a7dcafde279e8872c814f2f2be8a6e6980706da06dc1af4b1e2193

Observation c72cb348-886c-46ff-9438-363f60a1b434 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.125992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.125992Z digest=sha256:80676d64c3a57ed126351692cc2930344b572734632ee54cc5464adeedc3a7ae

Observation c92bdff2-402b-46bf-92af-c5ffe95e425b · outbound

This paper cites Qwen2.5-VL Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.149385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.149385Z digest=sha256:26fde2e33d41350a5bba7fd12d24149dfaf954d0058fb05bc4510be44025e101

Observation c657ea92-7c09-48a2-9414-40d49c3df2a0 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.168229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.168229Z digest=sha256:d60b134e3af6f64d2c0a619c3348ec9002ef4edf0678e8e41b6eaed09e2deb0b

Observation 449cf1fb-3f45-442d-9e5a-69ee1fe9b17a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.206037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.206037Z digest=sha256:2f455a6a5f5ed0997d01919f98a4791fd1af9c243ee3d8d6340d3364373abb61

Observation b2332d1c-679a-49e9-9264-e55557eb3bbf · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.314570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.221144Z digest=sha256:5eafc21fd4bfac39a43a97059050c9845e9b1221a3784727552301f8f0fff430

Observation aee44f84-b50e-4f52-8e75-f1d730020eeb · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.232426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.232426Z digest=sha256:f27568249133d6d9fe2be16c433161bc9393d527a5bffbd7bd6398b8f4390251

Observation 4f0b9bfc-e67f-4693-b269-3faff3a84983 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.246975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.246975Z digest=sha256:77fafa4f5f6ee8af8a06ee18fca887c4746037a6202577acb4f5e82dac0b54ad

Observation 3769c6c6-8602-491c-8a98-64c15c0dde53 · outbound

This paper cites Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:05:21.666096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.279701Z digest=sha256:2d78cb20cb54e24170d8a1738718af8f47d6637a056c76a3a5a4b79ea05f69e5

Observation b83ffd3f-dd39-4fec-9365-ec175c24b1c8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.286471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.286471Z digest=sha256:cefea4983756b46c9208d9adc9139601da1354198564a08a676491ed29546273

Observation 87f573ed-96f7-49dc-b444-396ddc047ac6 · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.291846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.291846Z digest=sha256:40944f8d8052e7602009a4458a173d445986fdab259531e5a420a67c762eb027

Observation 2a463a9b-c158-4504-b449-7030e94417d5 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.067313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.302473Z digest=sha256:8bd50b8e857370235b3bdcaeacb606dbe4ad8f6c97fe8fa4d93e8bcb08ba20a1

Observation b7a365eb-80de-45f2-8806-193219db6b68 · outbound

This paper cites OpenAI o1 System Card.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark OpenAI o1 System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.318506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.318506Z digest=sha256:d41fbed7f844ad0a834334ef9cd16fc33784f264fe08f4419dc94e4ae417e4fb

Observation 3783139e-c5f8-4251-832e-b1894c57c50b · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.352234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.352234Z digest=sha256:89ba878e56857ee3d81eb5dc8c8ce1bb83f1b2a95287e835cec9e5553066c2c8

Observation 110bcaec-d601-4d7f-807d-775cf6756cf5 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.357259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.357259Z digest=sha256:8592147a7ec89b77b67316ae1580b0a0b499b20d8d1805970833b6e91bb7928a

Observation 7a7f698c-e082-4f45-9278-6ee28fb1844e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:23.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.373052Z digest=sha256:3701f9415179fbe2a7fbd3f6ca8eeaa4dc023a348ca379ad8bb7dba497aed3d5

Observation b7622b24-9377-4fb8-bc24-f16a898b13f7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.396582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.396582Z digest=sha256:4fd43aa48d0dc56b9c6de6fac3091f87c34aac045314e539d4c4b7cf0fdce40b

Observation e7821c05-6750-4e3b-91b5-d3fde00a313e · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.405502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.405502Z digest=sha256:f8c6cf711133476cbc6aba71f04ac334a5d063b75c25a0c7c8b38bbbe3ba54f2

Observation e2f30b36-44ac-44b6-a637-bb39ff2dd03a · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.423016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.423016Z digest=sha256:bc87dd0b66cbf960cfdefc40f1582a663c9829a79f75ff7ebdaa939b6c226fd8

Observation d1523ac9-2457-485e-81b5-2463cdc7d8fe · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.986394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.386373Z digest=sha256:ad298656aa22f10fa584a3b0e9ddc4fd3a440765a60583a8a38383412234d7af

Observation 813a2978-95cb-40c9-b6a3-adacfaeb7e29 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.889599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.445145Z digest=sha256:25a29ea395227ba1e8e128a001ab60060e582f589ecf63ad9b314e98436e6663

Observation 55e9a714-8997-4266-9b20-f54251a95a86 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.462282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.462282Z digest=sha256:90e94c9e1162817fe484a23f4960065e98ea6fa87c6e2e795ef8eefd78342c6f

Observation 57dc166e-b0ca-4809-8d63-1189cd44c09b · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.472918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.472918Z digest=sha256:a9f374ab897d62465824199e7b515856c496b2f989a3e8dbdec7329e776fbb46

Observation 934bdc89-a9a3-4304-b26a-b210f382595f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.927157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.432019Z digest=sha256:bfdc4a54eef53ac6a24ec6f8130b1ae4343f8b29d464ac3a5368ba3b87af84af

Observation 383301d0-48b3-47a5-a9fa-aff5a789d17f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.810498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.491198Z digest=sha256:da2c33623f97ac146d1a0ab09c867ab9540a9a949c353ad154860ccee9196785

Observation 28d0fb4b-8fb2-40ad-906b-aae4f0d84da1 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.499218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.499218Z digest=sha256:40ff57a46e5d130ffcbe42b4b099d323304adcf40ebca4e37ee522395c4041ab

Observation 0d9ba4f4-81b9-478f-84f5-54dd0f3419fd · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MileBench: Benchmarking MLLMs in Long Context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.523326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.523326Z digest=sha256:188590c3a0704d3aff1bde7af0410318c137380366073b34140b0bf6f78c3f9f

Observation 7ac23c4c-fbae-43fb-bea4-f92333afc950 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.483720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.483720Z digest=sha256:86635fe51ec0769c7b604900881c70f7e84adea5fe889dd1350bedf46ad5dbd9

Observation bed07209-e2ff-4850-9771-6ba94f4008fc · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.744082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.546625Z digest=sha256:360c33fb96c14ca960d0669defa82c816e310248c74a3af7c0e1bdee2842ef95

Observation 2a7da2ce-8183-4bf8-b05e-358ed8c1ce6e · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.555184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.555184Z digest=sha256:3df2ed2787effa36f325c18228911a4a1aecbcd30e3a20b761ffe68a35c73f47

Observation 2c0f76dd-9160-4dd7-8fd6-571bc8943ae7 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.565467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.565467Z digest=sha256:bb44f8b8cdebae7a0aa52acac1646b608d1697976c6336fd509d0fb7edb558c2

Observation 991b90ed-dccd-416d-b568-5e449d05da9f · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.580744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.580744Z digest=sha256:e2a1f8ad6e3c01e6506d34572b09f122b02ac274d9f63a3f7f755f93d06e73bb

Observation 072f0cd2-45ba-4135-88d6-0ead62ec9e8c · outbound

This paper cites NLVR2 Visual Bias Analysis.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark NLVR2 Visual Bias Analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.537088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.537088Z digest=sha256:323106a7dbe63552a4008dda6da8448d620e227f0d688d3bf6c5e6dfcbaf1dff

Observation 6290e526-21b7-4b54-ad8e-4f63cb734c17 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.613480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.613480Z digest=sha256:03e149ffbd55ee1b480d0b20a91feac383d5b1372b53d7febcabe318c2672f11

Observation 0c55dbbe-cfdc-41e9-8a47-d5c02a572196 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.676752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.622195Z digest=sha256:170a05bf01e613dcb5a8158c38881a25492635d29835c006ab5c1fad3e417349

Observation 4a5f715c-91b1-4750-bea6-34f1d1fa0f3d · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.643578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.643578Z digest=sha256:1b69522aa6e7f2374f810d6127292e090804c80a147bac40df0c75f30750635b

Observation c7c041f0-6fd7-491e-b307-15c86595814e · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.654435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.654435Z digest=sha256:365a04cafdfd6325c47644ba3dd71848c90057a1ca37a0ff02f80837a46e46a4

Observation daac543d-7b2b-4ab2-a972-1effe3b51c83 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.601376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.601376Z digest=sha256:8882e170da765cc8f1c588f65d3d614af737059705d159a1862fec1e3abf1ea6

Observation 0921162d-cacd-4676-966b-647ed987a204 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.683399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.683399Z digest=sha256:618e7619a9a0c7bd460eb57ede0d5c93a3bffb5cec7628bc0daa995f58456bae

Observation 7278bfc9-b506-4303-a6a8-b0c457cedc3d · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.704763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.704763Z digest=sha256:49a76cbbc21193aadf4ba1d72ae452a05a40ac9d9e8fd53510e996779f101981

Observation 7e2399a7-c682-4cf5-9248-a7ad6c9ad0bf · outbound

This paper cites Qwen3 Technical Report.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Qwen3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.713959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.713959Z digest=sha256:090676aa4a78b111067d91a94c6273939b3d9ce333b941ff56c4863b9dcfe43e

Observation 4f95da42-0fbe-49f0-b7a7-b9d955e24ffa · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.719944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.719944Z digest=sha256:1e85b18e15615323df727528f1af8b9a53858039aaae0a223dc3a51d907ed7d7

Observation 406e5081-45d4-448a-93c0-300a95aa87e7 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.667340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.667340Z digest=sha256:1db5dc778436e38c3f878fd01e00e94eec617e60d85aedde4ac16b9e61d6cc5f

Observation 869ae997-430e-4c73-89a3-0a3e0ab35c59 · outbound

This paper cites InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.748441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.748441Z digest=sha256:9bb31676aca53f2925d486b9456bc10926059ffc81ce95dabebe29d66213e4d5

Observation b207ca52-e644-403c-a340-c47554341d74 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.759069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.759069Z digest=sha256:c69f505d019def4cce8173b9060cf65a7f3625048c65d94f022d203d4b7ce77c

Observation 04bcbfa7-9314-4b0c-83ac-9331cbf40313 · outbound

This paper cites WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.768115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.768115Z digest=sha256:ea3299f71c0d5db8c75177aa7eefd2d7c60464731b4bf4687abae1c2ad7aa2ba

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:ac434a0ec8612d77ee879a21e11731ba7ece3bb6a3bae67ca77da315fc154605

Observation 674a23ef-a859-4fa0-abfb-8c3dd55cf2d9 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.737053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.737053Z digest=sha256:4de89a17f50daf72a59a6b470d5c5c70e204d40c8a2f1310c11a7ce205c9a4fe

Observation 374273db-e653-4a63-945d-d745994c824d · outbound

This paper cites MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark MiCEval: Unveiling Multimodal Chain of Thought's Quality via Image Description and Reasoning Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.808950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.808950Z digest=sha256:5a3ece1eee4f369e19709a63be2b71006f07f15915709c9c50e8777121c86a20

Observation 242fbcc3-9161-4688-add2-97f2000fa578 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.821194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.821194Z digest=sha256:5e0ad159c8cabf8ef9001139c45a7ba2db40f7eef0e1d25cb56181eb7382ccc7

Observation 48a225fb-5b1d-4cf3-8168-4a210622fcce · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Multimodal Chain-of-Thought Reasoning in Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.790349Z digest=sha256:1b71d961b2b7895092294212aff69035b3731c4e0c25baa7a3c5df6735407a71

Observation 1e7286e3-27d0-49ca-b3d0-4040e9b2d71e · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.541576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.834814Z digest=sha256:ac0b34f882d4cc507bc80ab702dd2fb7e5fd1599a1e9eec4717310395a6636e0

Observation 9afa9931-c062-408b-a212-7cdd30230c4c · outbound

This paper cites reasoning step.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark reasoning step

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.504739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.856925Z digest=sha256:feb31495a95d02cdcdeae4cf9976bcdcaf353e7e97be75f0aa356ef8cc4d71d8

Observation c25669a8-8996-44f7-be32-3c9ab4488f32 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.311098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.900673Z digest=sha256:6584ffb4ab8091d2c4b1650ab43e202567d30019c585224c0312e68a5e3a037c

Observation 09ab0e9d-67e1-47e7-9269-464b8458d6a1 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.259112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.946132Z digest=sha256:f9d70fac86f79b2104d2fceb81760a04220111cbb1c597fee7c2fb66936828c3

Observation 2cb14538-63db-429e-98ba-48aa6e23b6e2 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.465449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.961296Z digest=sha256:cd9d0aa45ed89664f6e64df04cdcc76087cbd647dcf54745aae1289730c3a680

Observation c130aa28-c9d8-4f8c-93ac-78e6f1e19275 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.395288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.971376Z digest=sha256:c8d1922710379b32e64166f5a77cc23abe4a8e375b02dffc06b932ef74966526

Observation b7038d16-669d-460c-a5e0-4ef8f49abba8 · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:05:22.363949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.980113Z digest=sha256:bf3d6644687ec13b37c0428fbd8f0a2ddc0187cfe94df8dfdd205d3f78076aef

Observation 56aba37e-d952-4cb6-ad57-00930f296933 · outbound

This paper cites equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark equally good

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.199841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.988591Z digest=sha256:372ed5b68a7b5497f87bde81645686436554713124f365b2a87143538531f2df

Observation d4e85cc4-f8a5-487e-b582-eee8aea80311 · outbound

This paper cites Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should minimize the mention of objects not present in the ground truth answer, and inaccuracies in the description of existing objects

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.116840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:20.013797Z digest=sha256:be67962f9e5b392b0943ea5d90ea411ad3a46fd5d3fb0026dd4f08d38813224f

Observation 78a0e081-f8d1-4af4-bf4f-192fa85a8888 · outbound

This paper cites Rank higher the responses that least misrepresent these relationships.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Rank higher the responses that least misrepresent these relationships

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.066928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:20.026221Z digest=sha256:e62ade3a9da7f3597567df7f0cca11b0ea8020add0c97fe9fb140200a9cfe62d

Observation 88f71158-6ef9-430c-88b4-7f57a2d86032 · outbound

This paper cites Responses should avoid inaccuracies in describing the characteristics of the objects present.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Responses should avoid inaccuracies in describing the characteristics of the objects present

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.020746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:20.034083Z digest=sha256:09d9dc1ddaafd93d50e54a742d08b2c6fd0363dbf96d36076a20681a0956c55f

Observation edc4972b-8872-4d0a-b8bc-92ad31259023 · outbound

This paper cites Equally good.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Equally good

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:21.962642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:20.044416Z digest=sha256:5d9d158d34ce9396bbe2ec49a0edff4b9beee32062b59f5864d4839908df233b

Observation 7994fa18-e622-4d8b-b9ee-64d433200258 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.509807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.509807Z digest=sha256:868650b1a5ec0a053902baac01e4086096c97c522557a6287931481e3526a23a

Observation 3823083d-0afc-468d-bd06-d4c92fc52bff · outbound

This paper cites an unresolved cited work.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.114146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.114146Z digest=sha256:473aca545f6ce803b91a2051ab198bd0c7faac8ba849a992a154c61415b01978

Observation c013ae08-676f-4166-8fbe-e7b9fac4652c · outbound

This paper cites InProceedings of the IEEE/CVF international conference on computer vision.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark InProceedings of the IEEE/CVF international conference on computer vision

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:23.155612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:05:19.262204Z digest=sha256:24a16c0ae9f179e2c2b462ed87fba067466c467ce1e0f03f32a5ef47cf4a44ea

Observation 2760d9a4-0362-4c1d-a434-7a40e9ef31e0 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.193447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.193447Z digest=sha256:204d3c8e237dff4b4d517cf6d312509eeb71f0259b2f935c5639a570fb72b65d

Pith citing papers

Observation 5c6864df-b047-4bee-bf80-fedc21632611 · inbound

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation cites this paper.

Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:43.474266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:43.474266Z digest=sha256:6616d4dedb87935aadd489d2ff0580a2f6d6cc8b3610c185e539836d57dad52c

Observation d334f9d5-9cff-48f6-91d2-3fba26525cc2 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.779717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:e6df194a838b691393ebcc2358565f712dd0f41015bb63ef179527647baefa0a

Observation 92fa0042-066d-4efb-aa58-7d57f2915315 · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.830119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:c2ccb160c2014b02659e2176b83f29924a41f0d9d32bbd246663fea57bc1bf8d

Observation 8425e803-0d62-4c85-bd60-4fbd56364dac · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:02.750336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:610f01e902134276f68362ef00c4bd8307744211d3c0a7a62d5c715b98744c6f

Observation a1c2534f-54a4-4485-8371-831cfd49b2d7 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.324575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:09d518546b7c8863437e52cd7bffca64c10222bc7c77098b6b733d31dae7d74d