Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2312.07533.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:20:08.455261Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T12:09:48.871684Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 5aa228ac-1987-4675-bf6f-e0fe34447313 · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI VILA: On Pre-training for Visual Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75b31be1-6db3-41b0-bb1b-15da1e94e251 · inbound
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception VILA: On Pre-training for Visual Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e62c1d3d-15ac-4a15-ab2e-037fa7173fb2 · inbound
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training VILA: On Pre-training for Visual Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 84ad2ef8-c196-4b86-b7ff-8a212b9abe75 · inbound
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning VILA: On Pre-training for Visual Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 61e5d36f-238b-4510-8d74-148bc6e0df71 · inbound
OpenVLA: An Open-Source Vision-Language-Action Model VILA: On Pre-training for Visual Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a8e2401-59df-4aae-86f0-e0c8011f7464 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models VILA: On Pre-training for Visual Language Models
Reference 227
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 588643b7-9054-49ca-8bae-8de9a918f330 · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos VILA: On Pre-training for Visual Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 229832af-0b80-48e4-9f71-cbd0c6446fa8 · inbound
Vision-Language Models for Edge Networks: A Comprehensive Survey VILA: On Pre-training for Visual Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b238b0-36c6-48f8-a9b6-8fcd4efb3e7f · inbound
Affordance Benchmark for MLLMs VILA: On Pre-training for Visual Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7893db75-165d-4948-a513-86e9cacc1672 · inbound
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics VILA: On Pre-training for Visual Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bbeed42a-e77b-437f-b766-638f6d431e27 · inbound
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models VILA: On Pre-training for Visual Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c15058f-a19e-4f63-9807-89b51860cc96 · inbound
A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents VILA: On Pre-training for Visual Language Models
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d85d33-4de6-46c1-9662-ac35ff987eda · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VILA: On Pre-training for Visual Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08768bc4-df80-43f8-a3c7-0cdc7de55359 · inbound
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model VILA: On Pre-training for Visual Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5abc18a0-57c1-41e4-9b75-a1c7165fae67 · inbound
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning VILA: On Pre-training for Visual Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1354d313-04a2-4e86-a48a-d63ce58b7a6e · inbound
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models VILA: On Pre-training for Visual Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3d657b-2e0a-4e38-abd9-018bed994755 · inbound
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion VILA: On Pre-training for Visual Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 564b90a6-2952-4def-9ddf-883dc07b490e · inbound
Estimating the Empowerment of Language Model Agents VILA: On Pre-training for Visual Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5e48f3-1ec3-40f7-91cf-92a9edad9b97 · inbound
Video Reasoning without Training VILA: On Pre-training for Visual Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa2b176-0507-42e7-9483-4c2c5a486be4 · inbound
SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL VILA: On Pre-training for Visual Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03fd888e-7c74-439c-9582-e6d7ba631392 · inbound
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding VILA: On Pre-training for Visual Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 86d6886c-7282-433f-990e-073e791dc387 · inbound
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment VILA: On Pre-training for Visual Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bf203738-4503-43dc-b8d6-5d69041428fb · inbound
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models VILA: On Pre-training for Visual Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d3678fd2-7160-450c-8434-235fd90be38d · inbound
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction VILA: On Pre-training for Visual Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e138bcd-5b29-4938-bf6d-82a014ab1786 · inbound
What Limits Vision-and-Language Navigation ? VILA: On Pre-training for Visual Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 49b95122-0aff-4e75-a057-6ad370c89577 · inbound
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VILA: On Pre-training for Visual Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b8304d8-1782-40b8-a388-98484736660b · inbound
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VILA: On Pre-training for Visual Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118b7e57-d14b-4d7a-8c91-97878b70e66a · inbound
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models VILA: On Pre-training for Visual Language Models
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f1e41ae-a87e-4c4d-89e3-d6cd7dc11e84 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VILA: On Pre-training for Visual Language Models
Reference 290
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 13fe428f-8c6a-4bb8-a1c8-1a4f4c257cc0 · inbound
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views VILA: On Pre-training for Visual Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c355570-673b-4952-9a46-c4e119daec73 · inbound
Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse VILA: On Pre-training for Visual Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d1a90f2f-7dbe-4c44-974c-2ec2accbd19b · inbound
ESC: Emotional Self-Correction for Reliable Vision-Language Models VILA: On Pre-training for Visual Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a61cdb0-9ed5-417f-95a6-a18341fcdc43 · inbound
Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution VILA: On Pre-training for Visual Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.