Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2307.04087.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T11:32:42.920231Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T15:07:03.935337Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d649401e-77da-4d9e-89be-f84674c165fd · inbound
Otter: A Multi-Modal Model with In-Context Instruction Tuning SVIT: Scaling up Visual Instruction Tuning
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 536d51be-36d6-4cb8-99dd-54a5b20de514 · inbound
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets SVIT: Scaling up Visual Instruction Tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 10ba115d-d398-4bb6-99ff-8555551e5cff · inbound
Improved Baselines with Visual Instruction Tuning SVIT: Scaling up Visual Instruction Tuning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a1aeacf-975c-4c5d-9467-ad30e2b5fedc · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration SVIT: Scaling up Visual Instruction Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d5bd206d-b5b7-4509-850d-b019a7a13b6b · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI SVIT: Scaling up Visual Instruction Tuning
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 819c52b1-606b-467a-9204-110350fc9fd3 · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks SVIT: Scaling up Visual Instruction Tuning
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 60743011-f516-4957-8578-263f3ccdb2dc · inbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model SVIT: Scaling up Visual Instruction Tuning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb44d84a-7d91-45f8-a19d-299aaddecf62 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training SVIT: Scaling up Visual Instruction Tuning
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 23f91ff9-793f-4653-aeb7-d5669ae8fd77 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites SVIT: Scaling up Visual Instruction Tuning
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a92866f-6f73-46a2-8710-fe58ee55fee2 · inbound
MiniCPM-V: A GPT-4V Level MLLM on Your Phone SVIT: Scaling up Visual Instruction Tuning
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d1a283ce-0177-4dc8-a8da-bf383082fd9a · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 18ffe440-31d2-4533-a0b3-3d5af3f2de93 · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark SVIT: Scaling up Visual Instruction Tuning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9176bb75-69a3-44f5-b177-b97c2950872f · inbound
NVILA: Efficient Frontier Visual Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2c31d624-3712-4f64-bc7c-72ca9ff80519 · inbound
DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities SVIT: Scaling up Visual Instruction Tuning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13722c96-2eb2-41b1-a5c3-905c9dec1085 · inbound
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM SVIT: Scaling up Visual Instruction Tuning
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05c6122e-8d8d-44a8-bd5b-e427e1374c22 · inbound
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion SVIT: Scaling up Visual Instruction Tuning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f62fce6-9aa9-49d0-be01-6a51329e294c · inbound
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3be2cd5-5769-47ea-b6ce-7c20742fe64a · inbound
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities SVIT: Scaling up Visual Instruction Tuning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86c7702-3db9-44f1-ab53-65377a19e076 · inbound
Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c04e7633-c62d-43a3-ba52-725bb69e0a6d · inbound
GLAD: Generalizable Tuning for Vision-Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3013a8-c5cd-43e0-bdfe-d4519226568b · inbound
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark SVIT: Scaling up Visual Instruction Tuning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdb9d7e0-39e1-48e3-879a-fda3f07ec52d · inbound
Improving Large Vision and Language Models by Learning from a Panel of Peers SVIT: Scaling up Visual Instruction Tuning
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc141185-86d2-427e-8b9e-36bf929837b3 · inbound
Representation learning from OCT images SVIT: Scaling up Visual Instruction Tuning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a0308b5f-a2a2-417c-b789-1feb6cd8f6f3 · inbound
Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models SVIT: Scaling up Visual Instruction Tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f63dd66d-a399-45fe-96fa-9ec78afd9d6b · inbound
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens SVIT: Scaling up Visual Instruction Tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d25ec62-a039-4f2b-8215-c6e822ac182c · inbound
Balancing Image Compression and Generation with Bootstrapped Tokenization SVIT: Scaling up Visual Instruction Tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db848888-a092-4d1f-9fa6-ed77bb9ad570 · inbound
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning SVIT: Scaling up Visual Instruction Tuning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe6631cc-985d-41ee-a96b-da517e43bb84 · inbound
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design SVIT: Scaling up Visual Instruction Tuning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.