Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2108.10904.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:14:03.546300Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:10:05.334285Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c79ce03d-e652-4e1a-9612-2a20441e7ebf · inbound
Florence: A New Foundation Model for Computer Vision SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3e8c198-2f4b-45a5-9024-463d8f5a0c97 · inbound
Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16095bf5-2c75-47be-9efb-9967ab6dc02c · inbound
Flamingo: a Visual Language Model for Few-Shot Learning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e62c38a-c5ba-4473-8fc7-b04d20a8b76c · inbound
CoCa: Contrastive Captioners are Image-Text Foundation Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1611710a-17ef-4a83-a512-ce11cd81fa89 · inbound
A Generalist Agent SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eed7dac8-d443-494c-8abe-df58544933c5 · inbound
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 150303d5-51e5-4690-8ec8-b4083d854709 · inbound
Inner Monologue: Embodied Reasoning through Planning with Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c47735b6-4b4a-4bc9-a0c3-540ba7b778eb · inbound
PaLI: A Jointly-Scaled Multilingual Language-Image Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46221846-5e44-4ae2-81cf-afb669f6962e · inbound
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 140417e2-36d2-4f70-aac2-476191209903 · inbound
A Survey on Multimodal Large Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03da7090-bb13-4508-8d66-f508abdd6740 · inbound
Agent AI: Surveying the Horizons of Multimodal Interaction SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 290
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3724561-48ee-48d5-9eca-06119d69a7ea · inbound
PaliGemma: A versatile 3B VLM for transfer SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8354bd25-586f-4f95-91f4-16e04aef353f · inbound
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16d8b0d0-1aa6-4358-8cfd-db03e911ef78 · inbound
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77bd1637-6379-4ef3-8d25-928d3b36ff97 · inbound
Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e924db-1927-410b-88be-ec97b3d2c838 · inbound
FREE: Fast and Robust Vision Language Models with Early Exits SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49ce8a1a-37bd-4601-aaeb-0f09e6ec81a7 · inbound
CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f16b24-5e42-4983-ad77-b0bbfb6583fc · inbound
SensorLM: Learning the Language of Wearable Sensors SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92bfa76-840a-4600-bc97-171cba9da0f1 · inbound
Vision Generalist Model: A Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 940e16ad-a44d-49b6-aea4-fc124200a517 · inbound
Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 354e4939-3157-42f6-b4cd-52aaf15c6c4d · inbound
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c065b3a-771b-4689-aae5-c083f0391f89 · inbound
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7feecfc5-30d9-4655-bd45-b597bfe9ae9d · inbound
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b31139e-1020-4c3a-ba56-e8306c567ce0 · inbound
Foundation Model Driven Robotics: A Comprehensive Review SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a718101-06aa-46d8-a604-39b2520da10e · inbound
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 214
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57360d68-7f3d-4b8d-bc83-b0b18d956ae8 · inbound
Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5700a91-3ab5-44d5-b93b-1c8fa37b2bb9 · inbound
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a81d2586-752a-418f-a8ad-eab72abd079d · inbound
Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e42c1d0-96fb-429a-bbdd-c8fbb123c2e5 · inbound
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08bb4adf-603c-420a-b56e-f78030cdd1a4 · inbound
WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7830cbe1-7e0c-47c9-b9f5-715b7dc3aaad · inbound
MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d9a4388-4867-484a-a86d-89a600935572 · inbound
RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02ec7dbd-701c-485e-bbcb-d116d23e87f8 · inbound
Let ViT Speak: Generative Language-Image Pre-training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f7df252e-0aa0-48c8-8a74-640ef2e21b82 · inbound
Let ViT Speak: Generative Language-Image Pre-training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8422f70-2647-45b4-9bfb-232a5327bb6f · inbound
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3f559bf9-22c8-41cd-a996-5ea9fef6de58 · inbound
ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c548dcb9-23c9-4aeb-95e7-9b654439f2dc · inbound
WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2d6d632-ba4e-4dca-a174-d876a3603909 · inbound
KidRisk: Benchmark Dataset for Children Dangerous Action Recognition SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fef56a8-4df4-4064-bd39-324d808df81d · inbound
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87bd443d-8580-4a0d-8b66-6712e854a642 · inbound
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.