Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2403.20271.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:55.390063Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.285599Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 5c953bd2-e379-4792-a718-1b9d1fd8a12d · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ebdb2ac-b14b-4722-991f-169e278f991f · inbound
LPOI: Listwise Preference Optimization for Vision Language Models Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8038ddba-2872-40b6-9f42-843263afc2fe · inbound
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aa6eeb3-2245-4d99-8f1c-4cef689c819a · inbound
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734693f4-cdbe-494d-a851-2c3d8c130a24 · inbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88536864-af88-40b0-8df6-7f58519084cf · inbound
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13065974-bb15-4c14-9943-db6377dce38b · inbound
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f185046-efe2-4236-9ccb-86527044b2be · inbound
The high-speed X-ray camera on AXIS: design and performance updates Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b819030b-abb2-474b-a43c-1fdb5303f7b7 · inbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eb49cf9-3455-43e8-8fe9-ea4402bdc796 · inbound
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8333f38-25bb-48dd-b407-e8f5f10b7d6e · inbound
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 84cc4755-f6aa-404c-a487-402e84ec102e · inbound
VABench: A Comprehensive Benchmark for Audio-Video Generation Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 099aca5f-e86d-4736-b73d-a26ef6ccb138 · inbound
Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc407873-50c7-4d54-a77b-68db0bdaf033 · inbound
E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08b66f69-779f-4a4c-84da-280aaae95099 · inbound
Less Detail, Better Answers: Degradation-Driven Prompting for VQA Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 831e340b-8378-43ec-b7c0-8f747c970060 · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b2302e69-4835-4f13-a134-b540e1896a1c · inbound
WOW-Seg: A Word-free Open World Segmentation Model Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ea1e2af6-8aae-4be1-b89c-5dd4ddf80e68 · inbound
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 412f60c4-fc1f-48c4-9595-47af92b3b169 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9ceaa00c-f8bb-40b8-93d5-11c989a28ae9 · inbound
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.