Pith. sign in

Paper Citation Record · LEDGER

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 54 inbound Pith citation observations for arXiv:2108.10904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2108.10904 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 54 of 54 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:31:00.250761Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:10:05.334285Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c79ce03d-e652-4e1a-9612-2a20441e7ebf · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:38:09.573674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:d173ac5267ec3d8028ecc0e6ed0e3638e0fc213bd8a6916872c7a8a11b416f60

Observation d3e8c198-2f4b-45a5-9024-463d8f5a0c97 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.756707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:dbeb44f85aa0f1a9fc24d8c313bf753e8dbf903ee6260d6d900828385a66996d

Observation 16095bf5-2c75-47be-9efb-9967ab6dc02c · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.338088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:f08404fd325e1a775e6f7b0a3e17dd5ce47309aeb05a5a8757b2bd344d00db9a

Observation 1e62c38a-c5ba-4473-8fc7-b04d20a8b76c · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T10:53:08.385289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:f58f94269ee45e3fead030ae54203c9a11433159dbc1fec3d4bf9d3425f726f3

Observation 1611710a-17ef-4a83-a512-ce11cd81fa89 · inbound

A Generalist Agent cites this paper.

A Generalist Agent SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:24:49.969674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T06:24:49.833638Z digest=sha256:8d3f63dde5a28be0882aa36576c7767674d7f3f63ab54b90439ccf639af58af1

Observation eed7dac8-d443-494c-8abe-df58544933c5 · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.252855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:ea5ba971419159569f1acd5be35594c13d92f62734751dd51996173dba43fb24

Observation 150303d5-51e5-4690-8ec8-b4083d854709 · inbound

Inner Monologue: Embodied Reasoning through Planning with Language Models cites this paper.

Inner Monologue: Embodied Reasoning through Planning with Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:10:45.889197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T20:10:43.912935Z digest=sha256:1f612f753ad0566000eb625633e4781f2a0150b502b59bd10ff264188a7dc80d

Observation c47735b6-4b4a-4bc9-a0c3-540ba7b778eb · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.091876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:0caa51cd2447cc3006708ec67a33a6a93da22ace86f2e8c46d301411c6fc7976

Observation 46221846-5e44-4ae2-81cf-afb669f6962e · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:07:42.982914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:5f8ef70b772c40d458c0972a906ded58850fae121853ca2733cda3e4e9feb2cf

Observation 140417e2-36d2-4f70-aac2-476191209903 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:56:42.300910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:f2eb604c6b4a59777deedffd469f8f413844f737942933fee2cbb0a41c2011c2

Observation 03da7090-bb13-4508-8d66-f508abdd6740 · inbound

Agent AI: Surveying the Horizons of Multimodal Interaction cites this paper.

Agent AI: Surveying the Horizons of Multimodal Interaction SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:25:59.818454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T14:25:58.876978Z digest=sha256:880a4fcbfecea6781f4f47cd26aa65a9c90deb8e06df4f0456ca5a589676daa7

Observation c3724561-48ee-48d5-9eca-06119d69a7ea · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.208526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:1cec6eb2ca9ceac621a4f9e2e9cfadcc7784af7a0b86b95c5a2227ca461f0e8c

Observation d040aaa5-320b-4575-bda2-f6a7102b49eb · inbound

Multimodal Autoregressive Pre-training of Large Vision Encoders cites this paper.

Multimodal Autoregressive Pre-training of Large Vision Encoders SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.832620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.832620Z digest=sha256:db58e335a834fbb8e6186f650d2c661eb9c5cfcdeff32e12a29ecbc7ca6bccf1

Observation 7c79ed25-29e0-4b4f-a0ea-a4ffcf36494a · inbound

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning cites this paper.

Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:06:05.090407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:06:05.090407Z digest=sha256:6dd92064ed02c5952dbcf24ed1a63006063e0205b52b491b8284eca6bdefc3c5

Observation fa6ca9c6-ae4c-4d06-b0e6-3e7c3c229a8d · inbound

Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation cites this paper.

Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:38.766732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:38.766732Z digest=sha256:b0a0fb73c22cfca9c568ee3f9abc66b236c3adffedef48c34fbae3d17a0d98d9

Observation d03c0505-7d79-4c8e-a6d6-0fb2112b0a24 · inbound

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning cites this paper.

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:52.513342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:47:52.513342Z digest=sha256:f77854550991686af3dbea527b0a9171ddb83f092893fb33c863bc21910aa643

Observation 4b09d050-6986-48eb-8369-c5ac505dc0ce · inbound

Probing Visual Language Priors in VLMs cites this paper.

Probing Visual Language Priors in VLMs SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T22:55:54.246726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:55:54.246726Z digest=sha256:fa7cd206c476baf732b3db0cd058963602e13a57f896ac9691a8ec63d03cad2b

Observation 0d8b25b6-e8e4-4b23-92d2-57ddf246af6e · inbound

Dr. Tongue: Sign-Oriented Multi-label Detection for Remote Tongue Diagnosis cites this paper.

Dr. Tongue: Sign-Oriented Multi-label Detection for Remote Tongue Diagnosis SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:02:44.047931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:02:44.047931Z digest=sha256:0e754d74ca40d4964948f9296e3ea2d809033ca1c88edec320bbe7b69b2454f7

Observation 4e91617f-0f05-4724-9d82-84e584d16b05 · inbound

Visual question answering: from early developments to recent advances -- a survey cites this paper.

Visual question answering: from early developments to recent advances -- a survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:46:28.019525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:46:28.019525Z digest=sha256:36e522e3e4f6de187a661e79e90d48676ff3a36e808bf5d1d2eb83a964edda80

Observation 8631da16-cc6f-4f78-b72b-077e1824f250 · inbound

Multi-aspect Knowledge Distillation with Large Language Model cites this paper.

Multi-aspect Knowledge Distillation with Large Language Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T16:19:29.033903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:19:29.033903Z digest=sha256:1cd237d62046db4c8d3ab38b954dcec707169b2fa2496c32d979b7340f841fa3

Observation 7573674a-1a43-400a-adce-1d9c637b4657 · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.238053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.238053Z digest=sha256:8eb2e2cdce7f23fa1eaefc10574ad6b23b4e0a63701eb0a20bf37ff8c51852ab

Observation 8354bd25-586f-4f95-91f4-16e04aef353f · inbound

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding cites this paper.

Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T20:14:03.546300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:14:03.546300Z digest=sha256:3c869105f634af20494a233e7a6a8367113befc9afa3b70666488c0d3e11f410

Observation defa57c1-c932-44b8-b1d4-1991733ba014 · inbound

A Large Vision-Language Model based Environment Perception System for Visually Impaired People cites this paper.

A Large Vision-Language Model based Environment Perception System for Visually Impaired People SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:00.250761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:00.250761Z digest=sha256:e7f58a23f9657c643a7627b3345d735bb77323dde3440fa0d60eac6a27b4486e

Observation 7a70b7a4-9f90-4c71-a775-567669f16683 · inbound

ULFine: Unbiased Lightweight Fine-tuning for Foundation-Model-Assisted Long-Tailed Semi-Supervised Learning cites this paper.

ULFine: Unbiased Lightweight Fine-tuning for Foundation-Model-Assisted Long-Tailed Semi-Supervised Learning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:53.650330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:53.650330Z digest=sha256:f0c45f1119df0911d4d77e758fab09e398fd6242cda7211eef75a431845c2a74

Observation d9749ed0-4144-48d1-ad61-70dc9e887f97 · inbound

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training cites this paper.

Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:49:28.103319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:49:28.103319Z digest=sha256:bd6adf9db7820e2d22567ff821396c90f39c26ddbd75096764e3dea265195a15

Observation 38ad1e65-ca51-43e6-965b-2405c02f1d40 · inbound

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning cites this paper.

TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:52.172544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:52.172544Z digest=sha256:a7bb86232cfab1a2f759880da28615e2e06976815f127152b784bf55c4b68ded

Observation 16d8b0d0-1aa6-4358-8cfd-db03e911ef78 · inbound

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval cites this paper.

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:40.666585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:40.666585Z digest=sha256:a1358cef72cc248ae16c0c7b0fbaba09d6f70db23c801cafcaf295664021cfe2

Observation 77bd1637-6379-4ef3-8d25-928d3b36ff97 · inbound

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model cites this paper.

Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:41.477507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:41.477507Z digest=sha256:f284dede8e585d4d49b994db720efeca6d49d97ab9cb73549925f664988ec1ba

Observation b0e924db-1927-410b-88be-ec97b3d2c838 · inbound

FREE: Fast and Robust Vision Language Models with Early Exits cites this paper.

FREE: Fast and Robust Vision Language Models with Early Exits SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:45.667153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:53:45.667153Z digest=sha256:a1e5154c2898357d87190f0fec98170d7ae634a51b66b9342dcad55936738d7f

Observation 49ce8a1a-37bd-4601-aaeb-0f09e6ec81a7 · inbound

CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization cites this paper.

CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:28.993874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:28.993874Z digest=sha256:83c5d3d9ec4e2476df86996e59d37fa7be2869051bf5c8123c3c368a624b2962

Observation 32f16b24-5e42-4983-ad77-b0bbfb6583fc · inbound

SensorLM: Learning the Language of Wearable Sensors cites this paper.

SensorLM: Learning the Language of Wearable Sensors SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:06.395146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:06.395146Z digest=sha256:ddff300a39fc2684df962fb02324467a993af6bd240d5c73debddec043da2836

Observation b92bfa76-840a-4600-bc97-171cba9da0f1 · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:02.444346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:44:02.444346Z digest=sha256:e3be5acf528f757a66280539171387aa32b9230ec4af836a3e9c1b98d15a42c3

Observation 940e16ad-a44d-49b6-aea4-fc124200a517 · inbound

Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data cites this paper.

Bootstrapping your behavior: a new pretraining strategy for user behavior sequence data SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:51.039188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:51.039188Z digest=sha256:91c4f64a711e991f38bf59fb989202ad71fa30e2a187d2ee86cb0de15d2cd647

Observation 63d1dae5-5b86-4d74-9c08-fee914173724 · inbound

EMUSE: Evolutionary Map of the Universe Search Engine cites this paper.

EMUSE: Evolutionary Map of the Universe Search Engine SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T19:48:42.393029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:48:42.393029Z digest=sha256:296fdb23bd4d2e499dcb0f3d2481aaed3a56b3c6b8d5318058d944e92ab38d24

Observation 354e4939-3157-42f6-b4cd-52aaf15c6c4d · inbound

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges cites this paper.

Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:03.339629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:03.339629Z digest=sha256:ab6190a4146450df92fb0a20d7eade48d120c32289714f55703296058e685020

Observation 8c065b3a-771b-4689-aae5-c083f0391f89 · inbound

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach cites this paper.

From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:45.349914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:45.349914Z digest=sha256:3d1659b9ea1b39532f0bb6bf66a57d459135772a092e1e6f3529c910e2ddc31e

Observation 7feecfc5-30d9-4655-bd45-b597bfe9ae9d · inbound

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models cites this paper.

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:54:10.057658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:54:10.057658Z digest=sha256:ba0b55e8f116f73cfdb2b192e6a102a698366122328d11c73e430a2bf73d31a8

Observation 9b31139e-1020-4c3a-ba56-e8306c567ce0 · inbound

Foundation Model Driven Robotics: A Comprehensive Review cites this paper.

Foundation Model Driven Robotics: A Comprehensive Review SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:43:53.077148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:43:53.077148Z digest=sha256:3029aebdd7ccf5de2591e2eafec0b3f4c774ddee93d4f578130deb1a5236ca43

Observation 2a718101-06aa-46d8-a604-39b2520da10e · inbound

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey cites this paper.

Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 214

Resolution
unresolved
no resolver link, observed 2026-08-06T11:20:07.002898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:20:07.002898Z digest=sha256:cba62909f5324932668eba0f35b6c451cf6d130b7a82488b3f92bb80957223a9

Observation 57360d68-7f3d-4b8d-bc83-b0b18d956ae8 · inbound

Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models cites this paper.

Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T23:46:58.076344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:46:58.076344Z digest=sha256:f1a5cf89d0ed7b54575ffaa004e3c97bca5b7172e39160751c39c01921c3e803

Observation c5700a91-3ab5-44d5-b93b-1c8fa37b2bb9 · inbound

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline cites this paper.

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:38:37.414051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:38:37.414051Z digest=sha256:6e00e4a4c5f9edfe1d899f783483e1ae9a09752350c5ac1996976f3ffb872195

Observation a81d2586-752a-418f-a8ad-eab72abd079d · inbound

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework cites this paper.

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:45:36.961944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:45:29.400398Z digest=sha256:d0f1977f24eef393c001fad41bd97d895bd799de0ebbf97837ca638dc98a4428

Observation 5e42c1d0-96fb-429a-bbdd-c8fbb123c2e5 · inbound

Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method cites this paper.

Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T17:14:57.572731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:14:57.572731Z digest=sha256:a79170b304cf8eca95f679206eb47ef9e000616513a3c67eeb53725fd35eda50

Observation 08bb4adf-603c-420a-b56e-f78030cdd1a4 · inbound

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval cites this paper.

WRF4CIR: Weight-Regularized Fine-Tuning Network for Composed Image Retrieval SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:50.008483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:31:53.371412Z digest=sha256:21ef4ac62a44d0105eef1f07300c269995fc2622a5bed4f313f7281aaf468d1e

Observation 7830cbe1-7e0c-47c9-b9f5-715b7dc3aaad · inbound

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images cites this paper.

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.372269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:33:02.946839Z digest=sha256:5d07bddf3136816ce9999f1bfde0ad4fd018c5e54ed7ae4b6b272bbaa5d0b2b3

Observation 6d9a4388-4867-484a-a86d-89a600935572 · inbound

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation cites this paper.

RIHA: Report-Image Hierarchical Alignment for Radiology Report Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:41:25.994420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T10:00:25.846913Z digest=sha256:67953473556ca87dbd3f3d712d2b6baf4527ef249ea7f49962a32ec8b746e7e5

Observation 02ec7dbd-701c-485e-bbcb-d116d23e87f8 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:01:11.367780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T18:56:51.627714Z digest=sha256:3b955c1c0b40b9acea5887f44e48e0783c70e0763cf39e63a9d1f6036a6d7a62

Observation f7df252e-0aa0-48c8-8a74-640ef2e21b82 · inbound

Let ViT Speak: Generative Language-Image Pre-training cites this paper.

Let ViT Speak: Generative Language-Image Pre-training SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:28.761944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T07:35:07.825460Z digest=sha256:c7112e8c63391f8a6afebd7557b63665d295336d56377d73df50c6bd33a4832e

Observation b8422f70-2647-45b4-9bfb-232a5327bb6f · inbound

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments cites this paper.

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:44:57.697014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T17:40:33.082748Z digest=sha256:383e3330800924fd8cae771694423b2b8d3dd03cf94e2416fd291c26517958d5

Observation 3f559bf9-22c8-41cd-a996-5ea9fef6de58 · inbound

ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation cites this paper.

ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.089296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T09:45:35.383450Z digest=sha256:08159a41b7698f9c30c458e8978af1ef39f8a6b901ef209b604e0d1cbf311c3f

Observation c548dcb9-23c9-4aeb-95e7-9b654439f2dc · inbound

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning cites this paper.

WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:58.489179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T00:53:11.223341Z digest=sha256:2c0a6d3a1675f20b7bf87cbf218dada4bf97f607b15be3e5cd17239c92dd06f4

Observation a2d6d632-ba4e-4dca-a174-d876a3603909 · inbound

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition cites this paper.

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:10:05.335820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-25T21:38:32.529040Z digest=sha256:bc6ee7e0065181b34b6efcf181febdc2bc8d81a20427b775851bed38c3dfef46

Observation 8fef56a8-4df4-4064-bd39-324d808df81d · inbound

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models cites this paper.

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.507790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T07:07:00.265141Z digest=sha256:c3d05ffc4d7996045bfcaf962a344865287f4c38c067ba4a924e4f8cf7b37a68

Observation 87bd443d-8580-4a0d-8b66-6712e854a642 · inbound

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models cites this paper.

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models SimVLM: Simple Visual Language Model Pretraining with Weak Supervision

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.354415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-02T20:46:16.298898Z digest=sha256:8e26f78b7a34ae18f913148f8595a2bda78bcdd6827b9359e367023e14eee36f