Pith. sign in

Paper Citation Record · LEDGER

MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2406.11833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11833 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:11.997396Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.009666Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation abd895d6-bc80-4a86-baba-d8f1cd9a72bc · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.787020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:234a4a5140ad3d720b5b2dff83571bf7f4f3806f0fc07039231acba62267fa78

Observation abff02ee-3b32-4204-bb6a-1b944bf634f0 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.472394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:795d5980b0d38f6f81ef537e83f19b978c5f36cf82d7c22985969542ae957d69

Observation 2b48b8d1-9ec8-44bc-98dc-14d5758f7f76 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.741943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:5ae738f23f5cd4dab438b49b83b8c3ffda846e80ee254ee04fe784d9fdda3b2b

Observation 1ab510f6-ba8f-4d3a-9464-a203b0137c7a · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:53:26.349112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:e9df26fef59c1a76811c04c3b492dc1534680bc919cf9f1227d9b318567789b4

Observation 09a076c8-e187-4eee-a58b-a509c02c1004 · inbound

Medical Large Vision Language Models with Multi-Image Visual Ability cites this paper.

Medical Large Vision Language Models with Multi-Image Visual Ability MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:11.997396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:11.997396Z digest=sha256:5b3119d89a92ceed74145af34fb45946f3fe9f9464b91c5f58e481650aa57bf6

Observation 39bb3852-249a-4e9c-83a5-81c5995c78b5 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.308787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:28e6de1a74f4582d7fcb4cc88928a79148bd5e5d601694ae81f662bf28ac6045

Observation 09821de1-a06d-4d42-8d64-0d39800c32f7 · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:06.523108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:06.523108Z digest=sha256:ec02d98e7f4a737d08c95d1010461f2659836dd6e9670d75802626b039c2f627

Observation 5695cc77-06fc-4aca-a66e-e9a988521821 · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.346839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.346839Z digest=sha256:018cd1b5ee28b299305774b74489e4d746192fbd9d20085b706baa0e192b2fc7

Observation 8a8da736-404d-457d-8ec5-ad7898f24989 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.873449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.873449Z digest=sha256:97490c9b0fd3f36bd93a3406bc6712a618b72d7e337fdccaec86b98d083d8108

Observation 699acb2b-98e2-4393-8f77-c1e5e8648c94 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:23.165410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:23.165410Z digest=sha256:cd047931221b6793d9830534f9323dec48ddcac2b5f18af8567fe13f7f964f3d

Observation 5e10f521-ca79-48b0-a676-b458bc36ebd3 · inbound

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents cites this paper.

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.373649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:31:15.525331Z digest=sha256:271ba757cbc23ad8902169a251f7fb277b20dc7e9610f65f5b231524c15b51ff

Observation be3073c6-1515-4b5c-8202-190bdd5e2c93 · inbound

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models cites this paper.

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.269872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:38:01.208136Z digest=sha256:583d731ec1f129c46b877c15fa2973e7609158b9ce7eaf5e1a66341e7a02a875

Observation 16216f9a-7d11-4345-8490-85ecae3c69a4 · inbound

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory cites this paper.

MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.012795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:02:58.640918Z digest=sha256:e1aa2edf52ea48ce7720d97e7130dd9620a6c26f842da07efec2fb104e125449