Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2404.16790.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:09:15.070161Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-05T04:30:40.718071Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation eb1b67c9-cb8f-4cf3-9bed-c91a0b6ad9eb · inbound
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7911f6b-4b0f-4c65-abcf-bbfa683b9ea9 · inbound
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7e5b5d7-16fe-4807-a4fc-faf20d7c2657 · inbound
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 230
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e991f0-6e74-440d-ac84-0659174b7709 · inbound
Detailed Object Description with Controllable Dimensions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99cd16a7-4900-44ee-8ea3-b69c39b6ea7c · inbound
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47953182-0762-4380-aaaf-0871ed770f91 · inbound
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48a5e20-08fe-4e3b-8a18-2b2305f3b4ce · inbound
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71bc3799-1b16-4aa9-9020-0ae1e1c36bc5 · inbound
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf601f4f-cc32-4b4f-884f-34104f92a460 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7fed4a12-5eaf-40e0-bb7f-58a14df4529b · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 611c7e4c-0ec4-4ffb-99de-ecdc32dfe686 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ffa1d7f-aaee-47f1-9efa-8d2f643abf52 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 334ef63f-17c2-494d-8997-515465fa4dae · inbound
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · inbound
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e98ba9-e124-4b41-b4d7-cd1e9217cb1a · inbound
SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb312b7c-9db3-4237-b007-03eaee6fbdc8 · inbound
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6110f5-c63c-467d-9361-79217ce2c16e · inbound
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues? SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5bc9a1d-66e4-4615-8dc2-80065ca6c88b · inbound
OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ebc8c25-c5e4-4cd8-9ae5-9f2d8400e533 · inbound
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a1a634-b303-4b6e-b6c7-e8b03cf126b3 · inbound
FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87c0ae9-e60d-4481-af81-960fba18264f · inbound
Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab25a3d4-ae8e-46a7-9ad8-26a6507a2431 · inbound
Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 203
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24a4638-6649-4dad-9272-65e048d80726 · inbound
MiMo-VL Technical Report SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0599d4f-1320-4427-957d-93ceb82ad4d3 · inbound
Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e46d437-d4fa-4afc-800a-916d564b420e · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0d8dc1-58fc-44ab-9453-12b949975ae9 · inbound
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e446c60f-acc8-4f68-9a7d-af9aa250ad29 · inbound
MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a51221-4347-4ac5-88fe-8b51366e9d8b · inbound
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92bbbf6b-91b7-40b3-8fd0-8423d9c3780a · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 93a3a951-c912-49f2-8512-585b1fc62165 · inbound
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b56a09-3792-4a48-af5c-26eeedbf0746 · inbound
DeepEyesV2: Toward Agentic Multimodal Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b799e30-2925-4d39-a890-2e38cc6736f3 · inbound
FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2f8dcbc-db99-45ae-9962-7e9141e36ba9 · inbound
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9ed6f8f7-af64-4e22-8ed1-c90afa19f5ab · inbound
Learning More from Less: Unlocking Internal Representations for Benchmark Compression SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b8e29e-dfbe-4f72-81af-26fe4762df7f · inbound
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40648cf-11bb-41eb-af3e-203b45d85a7d · inbound
Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3e929891-5d7f-4ccf-9fc2-81d28e7affe8 · inbound
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4adad9f9-fad9-43a7-8d8e-15396f9d3750 · inbound
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b66fcbb-5d88-4ff7-abdd-cff4dad2ba57 · inbound
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f451cf29-e6b8-4dd3-a0e5-1b08896683eb · inbound
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75727995-6101-4240-89b7-3e9c522d0f1c · inbound
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b0fe877f-5bf5-449c-9e76-3a59ace3047f · inbound
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation abb05a70-1fbc-41c8-a357-413b306d3a7e · inbound
DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 16914475-6daa-40bb-ad3e-c59290015fe3 · inbound
DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 579364dd-983f-4c76-bfd3-ebb0c2944b74 · inbound
Deep Pre-Alignment for VLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation efcc01c1-daba-4837-90d9-1bdce7d19ce4 · inbound
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 46758ddf-3d42-469b-872d-32c4c5cb52da · inbound
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0ecaf041-da59-4b17-bf3e-18d4eb31c305 · inbound
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 929228b3-17f2-4f93-a933-9619e559e750 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 171
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c4026ce-6994-4e23-8df0-612519f719b0 · inbound
ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 238a0bfb-d25a-41b8-90ea-c3601466b7c3 · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8e6216d3-e66c-4f9a-9e92-8eb022bc2b9a · inbound
DataComp-VLM: Improved Open Datasets for Vision-Language Models SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 149
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73ee3721-3cd8-460a-a73f-6e71170c2f86 · inbound
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c79f30b-a7c0-4f88-904a-6c8d6a29cdf5 · inbound
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8adfbc4-0697-4892-8d1e-9bce6b39468f · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 143
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6899fd5-23ac-494e-848d-f0574a043eac · inbound
ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.