Pith. sign in

Paper Citation Record · LEDGER

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2407.15272.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.15272 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:54:45.667418Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T06:20:36.463031Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bddcacd3-b19a-4486-bce2-0921f67279f1 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 233

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.464952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:0b8a626b91fc523f904958d90fd26b993e1745be12f457d65740a4005ddd3622

Observation fefc0e29-5e57-4c20-95df-6b8133599db9 · inbound

LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation cites this paper.

LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T20:54:45.667418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:54:45.667418Z digest=sha256:f539a84e00eaa51a50f964e0334accbffb20dcb395b65b7d85cb634b2e1109ca

Observation ef4b137d-6cb2-4f60-8e99-e537413b3477 · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.591924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.591924Z digest=sha256:5514d55eaf027571f7a0d17d8fa17f59d50a1a3be185ec7dcca2e4f15e2323d9

Observation ef8efdad-a3f1-460d-bb10-a93245101439 · inbound

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! cites this paper.

Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No! MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T19:06:46.219688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:06:46.219688Z digest=sha256:5942da061956bca024844b95c09e146e76557d7ec1295e5f3cbbde3d823c1704

Observation d1f390e2-0e1d-4697-a7f6-530f2d2faf57 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:09.946184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:09.946184Z digest=sha256:78803c94c014e5af78c911d97f75cb0cdf6d09f734d92dc7e3343981b0ea1b64

Observation 3148e369-a2de-4a8f-aab3-1ace73c11c2d · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.466045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.466045Z digest=sha256:fba2704553f36282a330a60a71a0e27c01dc0d966a41d3322276934a0bdb76e8

Observation 90dc752f-6c24-4de5-932f-02fdb89d9861 · inbound

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding cites this paper.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:29.679885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:29.679885Z digest=sha256:61fb8dd1068616f75d1d621470db39f46c5b75370bb35cfa117a48f2686373a8

Observation ca0dc0d1-1b8c-4076-b211-6d6fc480ed29 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.176871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:a98848c5558e31aaf11fb5aa3507b83e392e4df49320057f7c333bf1138d22d1

Observation 4543bf77-6c6c-41a9-b0f3-ac94b6292f88 · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.526702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:7d66ef9f5954cbb4870d3c1cd6bb9208821cb362cce097c670a03dcefbcff1ba

Observation 8afc926f-646d-4ecc-867d-b7266e2a1970 · inbound

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration cites this paper.

Logit-Attention Divergence: Mitigating Position Bias in Multi-Image Retrieval via Attention-Guided Calibration MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:01.703230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:26:45.182597Z digest=sha256:81d067dcec35923adcd3c62a142ef49b9133ce01ea8879697c5597aa8ba39717