Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2406.12742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.12742 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:14:09.693716Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T00:02:50.478421Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6a0b118-e0be-4c3d-9bfa-3a2839f44dd0 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 154

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.202763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:2eb5fd483414438afbee74d69f74a20f18a0ab883a8a32a4562586d7b732b47a

Observation c998bba7-4b75-4325-9045-4da573f3b1b5 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:09.693716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:09.693716Z digest=sha256:f21e0594811f8288c46a1ce7f565945ec39aacdd0d6d92cb20b529bdb04118f3

Observation 00d1e608-90c7-4f33-8808-8973a85ceb53 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:23.728013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:23.728013Z digest=sha256:51cef38e27031c2ae61b2f8d08688ef6b1801318a5a548fc39233bac54229e22

Observation 8ba2266b-47ea-4432-87f9-5ebfd913950d · inbound

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning cites this paper.

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:23.068331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:23.068331Z digest=sha256:a9321ed62557b4e5a5cc2a8b77fe7ac8e6e12781fde54c5f6d003d9c04041536

Observation 0980ce6d-691a-4e29-ba97-c736a7cd9863 · inbound

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models cites this paper.

FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language Models Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:52:14.967780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:52:14.967780Z digest=sha256:0cc6a50e7209e8abcd5dfcf35644dc8a35ff069233d9158e1f04c6d7c24ee4d8

Observation 92f13c2a-8aa9-4fef-8940-59b03860dc44 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 182

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.978903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:a3c7dc4f92e8db2dc74087df115ebcdd59c530a891d61f20c5b2f0273d5fa64e

Observation aec1ebc3-760e-41fc-b7d4-a9a163d4f9d7 · inbound

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding cites this paper.

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T11:13:16.978503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:13:16.978503Z digest=sha256:947fb32bf53ae3eb8de165d88a992b374180bc5d65e8963ac3870c937d438c21

Observation b2425791-05f2-48b0-b1f6-befefdef3029 · inbound

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model cites this paper.

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:03.146864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:40:47.562861Z digest=sha256:27fea290555a0ec074b4a32bc82ef2f0c137c4ad2cf4c2c3f2901046dc7d64a9

Observation bd96174b-9006-4678-a4c6-49ff1f505736 · inbound

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning cites this paper.

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.480515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T23:15:56.598968Z digest=sha256:7a73c6c2a97745d8c6213a83e9df383bc79249f1a874fed848c48c080985b991