Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:06:24.676479Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2501.09446.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T20:06:24.676479Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T02:57:46.941141Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T02:16:26.845312Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e1776f9-a3f9-45c4-8a19-f3093e62184f · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Vqa: Visual question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1953bcb1-fd59-4b1f-8313-22cd0a0cdf91 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Obfus- cated gradients give a false sense of security: Circumventing defenses to adversarial examples
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 707a7573-b362-4d61-8b22-3a887b380ca3 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Image hijacking: Adversarial images can control generative models at runtime
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ce72e92f-bf34-4f07-8a9f-df8aa1fdf273 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05fa27be-7460-4582-8546-3e277b3515bf · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Jax: composable transformations of python+ numpy programs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a96207da-6122-4dbe-896b-83229788ba33 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards evaluating the robustness of neural networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2cbaa20a-ad8f-4af7-b0b4-87e6a423bef5 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 48b260d2-c1f2-4246-86e5-5068391c03e4 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Reliable evalua- tion of adversarial robustness with an ensemble of diverse parameter-free attacks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e93064c1-2041-4a2d-aeb2-4354a36c88ad · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0850d889-3fce-4a80-800b-8db8ec7accf4 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Robust physical-world attacks on deep learning visual classification
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 757c5b33-d177-4def-8f28-88b475a703ad · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da759870-ac22-4f77-8e9f-c620861829d1 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Dat- acomp: In search of the next generation of multimodal datasets
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68faee73-b6c2-4a4d-be86-451e438c806a · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788f875e-b592-407a-a22f-0735ed5c26e1 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Explaining and harnessing adversarial examples
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ee209832-6359-4738-afb7-b08e83a87b18 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f706e098-4935-48f5-aaf6-3466c55ca202 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d052b8ce-af60-4f6e-a058-3dda08431dad · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Vizwiz grand challenge: Answering visual questions from blind people
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f472f45-7593-4d83-ae68-a92d9f8a03cc · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness LoRA: Low-rank adaptation of large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4971c88-7bc8-4620-a63f-35387a0a2163 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 55d586a0-9a53-4bff-bf59-45f2111d4027 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Collecting a large-scale dataset of fine-grained cars.(2013)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b7ccd5bd-2c60-4193-b74e-29a7c531c02b · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness What If We Recaption Billions of Web Images with LLaMA-3?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73bd220d-a2ae-4084-a384-1a02dfbb3eb9 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness An inverse scal- ing law for clip training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eea723d3-132a-4530-b8ae-21d10ba7601e · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Evaluating object hallucination in large vision-language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0e166030-19d9-4a5c-b291-3aa1e18bcf1d · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Scaling language-image pre-training via masking
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7dce995a-6755-49a4-b97c-b111622e3521 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Microsoft coco: Common objects in context
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248c01ea-a96a-49b3-995d-e45a0ae1282e · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Improved baselines with visual instruction tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ef4277ab-62ff-4b1a-b74b-6eef7ce7000a · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Visual instruction tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5d57b010-6409-4b93-9338-ac143c8aa848 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness AutoDAN: Generating stealthy jailbreak prompts on aligned large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 213a7dd0-0c0c-403a-981f-87062b9f6291 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94369438-5983-40ff-826c-3163b35467b8 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards deep learning models resistant to adversarial attacks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5944b3f7-8e89-4ca7-b0c4-327b9ca5cce5 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Understanding zero-shot adversarial robust- ness for large-scale models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d556e3a8-f0d7-439f-a070-0d34b539c37a · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Deepfool: a simple and accurate method to fool deep neural networks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdf56dc0-8e55-4392-8ed8-56becba3d80a · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Bag of tricks for adversarial training
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f12e9d07-da67-49af-9d46-e5fc31267973 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc0083e-1950-48ec-869b-95fdb093f31f · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 29cce8ce-5de9-4ca8-8843-04b3bd982534 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Visual Adversarial Examples Jailbreak Aligned Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc9fe5d-90ee-4e2d-95bc-428aa1ce8f2c · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Visual adversarial examples jailbreak aligned large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6b7bd383-f61b-4a93-879d-b3fda75be958 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 26aeb469-7347-48ec-9aee-c7d26101fb77 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Learn- ing transferable visual models from natural language super- vision
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 82a9f85a-2e5b-4e8a-a2f7-5997a119c15c · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c0b25ae-eb55-48ca-905e-2a9518b4b3e2 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness on the adversar- ial robustness of multi-modal foundation models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07f5391d-2958-48de-be57-a17e25675266 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece5ae28-d52b-4aab-b852-f61b17148b5e · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Adversarial training for free! NeurIPS, 32, 2019
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 72eee1f7-2fd3-4259-b926-a77f6e53ff14 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards vqa models that can read
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b70138a-de70-458b-9c42-0742c26f3845 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0c19ab-51f6-4a05-b6e2-23655501e6b0 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness In- triguing properties of neural networks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 51cfb025-1c29-4539-8f49-511e486369b2 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Robustness May Be at Odds with Accuracy
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56ae7a6-cc41-4e8d-ae8c-a165d86614a5 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Cider: Consensus-based image description evalua- tion
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53a731f3-6e49-4c44-9fc6-2f7dcd64e6c4 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Revisiting adversarial training at scale
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6db1a99e-93fb-4d9f-a784-2ea8cbbe4510 · outbound
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 57a563a4-ff85-4f1e-bad9-216cae047c51 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness On the safety concerns of deploying 14 llms/vlms in robotics: Highlighting the risks and vulnerabil- ities
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c6283cf5-d438-48d9-aee3-b1a49af7f99f · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Feature denoising for improving ad- versarial robustness
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 16f13be1-37b3-4f63-a5f0-faf8bdbfc405 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Coca: Contrastive captioners are image-text foundation models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 732067d3-7c24-4928-b2c3-dbf6bdb75365 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards robust resnet: A small step but a giant leap
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4e81124-00cf-49c2-bff3-08dbdddc8ecc · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Attacks which do not kill training make adversarial learning stronger
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec52100a-d665-4119-83df-c3dd69b5c902 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Geometry-aware instance-reweighted adversarial training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 22610ae9-0f3f-4e3e-bb3d-a7a28ae35c7f · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a0fbdf1-167c-4dfc-8fe4-5e62dee3e157 · outbound
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b813acf5-d0c6-4524-b03a-fbfe460883b9 · inbound
Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad91002-c2e8-4cf7-879f-29230efe9357 · inbound
Investigating Adversarial Robustness of Multi-modal Large Language Models Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0372334e-33ec-463c-9677-8bba24ef033d · inbound
Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.