Pith. sign in

Paper Citation Record · LEDGER

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 3 inbound Pith citation observations for arXiv:2501.09446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09446 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:06:24.676479Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T02:57:46.941141Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.845312Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1776f9-a3f9-45c4-8a19-f3093e62184f · outbound

This paper cites Vqa: Visual question answering.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.091010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.504262Z digest=sha256:2ed02bcb8215533b916188ed90efeb6621e56e582c0e0b1eb7894e5d89bdd34c

Observation 1953bcb1-fd59-4b1f-8313-22cd0a0cdf91 · outbound

This paper cites Obfus- cated gradients give a false sense of security: Circumventing defenses to adversarial examples.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Obfus- cated gradients give a false sense of security: Circumventing defenses to adversarial examples

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.084288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.507367Z digest=sha256:e716e0052548aee602d81337519b58f05219699a88314cdbc5a43be5f0aaae30

Observation 707a7573-b362-4d61-8b22-3a887b380ca3 · outbound

This paper cites Image hijacking: Adversarial images can control generative models at runtime.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Image hijacking: Adversarial images can control generative models at runtime

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.077462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.510904Z digest=sha256:b6505c7604184e59d90dfa5042f21fcd109458dd546feaae73c3771b438d3184

Observation ce72e92f-bf34-4f07-8a9f-df8aa1fdf273 · outbound

This paper cites Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.514368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.514368Z digest=sha256:7f386b3caea3f48ef9f4533d564108ca14f1333ec36c68845a33cec3f4095883

Observation 05fa27be-7460-4582-8546-3e277b3515bf · outbound

This paper cites Jax: composable transformations of python+ numpy programs.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Jax: composable transformations of python+ numpy programs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.069948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.517311Z digest=sha256:3a67fda623bbd601123f924d682816d8de71e69250260315f73ab44b62f6b299

Observation a96207da-6122-4dbe-896b-83229788ba33 · outbound

This paper cites Towards evaluating the robustness of neural networks.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards evaluating the robustness of neural networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.062423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.520684Z digest=sha256:502648bfbcf89c2fdf7c448fcd014e9cb8f448eebd6641ffebc3f05df82e0396

Observation 2cbaa20a-ad8f-4af7-b0b4-87e6a423bef5 · outbound

This paper cites Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.055620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.523755Z digest=sha256:ed478034416ff08d23ef0f680d91dbc17b89baa61df88e1036f46cd0e7ca4874

Observation 48b260d2-c1f2-4246-86e5-5068391c03e4 · outbound

This paper cites Reliable evalua- tion of adversarial robustness with an ensemble of diverse parameter-free attacks.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Reliable evalua- tion of adversarial robustness with an ensemble of diverse parameter-free attacks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.047481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.526408Z digest=sha256:748f66a0d2c473b7e5c53343fe91ed191c4988b53377d37c4b3e15d0f09701a1

Observation e93064c1-2041-4a2d-aeb2-4354a36c88ad · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.040227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.529305Z digest=sha256:ab0a849c0915d40f3c269b399ea50270f7ac0b6a0cce8300c99718891deaf546

Observation 0850d889-3fce-4a80-800b-8db8ec7accf4 · outbound

This paper cites Robust physical-world attacks on deep learning visual classification.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Robust physical-world attacks on deep learning visual classification

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.033084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.532237Z digest=sha256:98f6ddd6a1bd3bdc62498236bd831c2c09d59bd3f670273f81e6c5f114d6c359

Observation 757c5b33-d177-4def-8f28-88b475a703ad · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.534593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.534593Z digest=sha256:4e909d9150b1bfe38b07c27986389acae88d8e7201a435ce8cb40e8450f06cdd

Observation da759870-ac22-4f77-8e9f-c620861829d1 · outbound

This paper cites Dat- acomp: In search of the next generation of multimodal datasets.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Dat- acomp: In search of the next generation of multimodal datasets

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.025313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.537200Z digest=sha256:27f3f6a08982b0467af41b4df18b0eecdd6e9f1bd98ed3f4717e19f898eb8db8

Observation 68faee73-b6c2-4a4d-be86-451e438c806a · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.539812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.539812Z digest=sha256:91fae3917a1cf810f7a7ac71248ce9a6d720d5192e4ae9197d5ac961ceab3415

Observation 788f875e-b592-407a-a22f-0735ed5c26e1 · outbound

This paper cites Explaining and harnessing adversarial examples.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Explaining and harnessing adversarial examples

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.016699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.543078Z digest=sha256:9b4a2f4abb18970fa2756d49f00f12e860ac3821bc2d27c76e6df239e235ad9e

Observation ee209832-6359-4738-afb7-b08e83a87b18 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:25.009177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.545967Z digest=sha256:6355818ec68543ee0a2512eaea4330cb2533de0b7122967cec6d7df17006f41d

Observation f706e098-4935-48f5-aaf6-3466c55ca202 · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.548987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.548987Z digest=sha256:8cfa6af1596a9df11630338e96df9f95c80cdd111106e04a074e1a8a3300d826

Observation d052b8ce-af60-4f6e-a058-3dda08431dad · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Vizwiz grand challenge: Answering visual questions from blind people

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.552414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.552414Z digest=sha256:ca8b3e37149df38e25e66bf60a1a5566804cbb0bace1895bd8d418589d0dad1d

Observation 0f472f45-7593-4d83-ae68-a92d9f8a03cc · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness LoRA: Low-rank adaptation of large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.998152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.556159Z digest=sha256:1fe40dbbb0e8a4eaa49a88a7128e56f2bc8258deb1596c50a8855589ca68a960

Observation c4971c88-7bc8-4620-a63f-35387a0a2163 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.989646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.558546Z digest=sha256:d4dd15ac000d220f4ce82ff1a69923eda87dcceedfb0fc6b7f0e54ba50ce9725

Observation 55d586a0-9a53-4bff-bf59-45f2111d4027 · outbound

This paper cites Collecting a large-scale dataset of fine-grained cars.(2013).

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Collecting a large-scale dataset of fine-grained cars.(2013)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.982411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.561631Z digest=sha256:623662905008f80551fb1009bada5c1b3b61b156c38d3f1712eed659b257cb61

Observation b7ccd5bd-2c60-4193-b74e-29a7c531c02b · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness What If We Recaption Billions of Web Images with LLaMA-3?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.564806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.564806Z digest=sha256:b6e244311974aa9467bfaa2ccf633f1a5c363f92f07b5fa289e31428f35a842d

Observation 73bd220d-a2ae-4084-a384-1a02dfbb3eb9 · outbound

This paper cites An inverse scal- ing law for clip training.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness An inverse scal- ing law for clip training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.974131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.567817Z digest=sha256:0e342dc10a25d2a7c8e67c09f88089c60bbe94f477ee2f8cea71c267f407289b

Observation eea723d3-132a-4530-b8ae-21d10ba7601e · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Evaluating object hallucination in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.966911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.571074Z digest=sha256:28fe1a8ce48db1e7bb547b7e7a55aa90e42e954f78d8b60ee70d3eba24bdbf21

Observation 0e166030-19d9-4a5c-b291-3aa1e18bcf1d · outbound

This paper cites Scaling language-image pre-training via masking.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Scaling language-image pre-training via masking

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.959996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.573933Z digest=sha256:d29afaec532b01abd6d7936f9fb97fc80b84fd7c92f1be3f16ef7e36adaf07ee

Observation 7dce995a-6755-49a4-b97c-b111622e3521 · outbound

This paper cites Microsoft coco: Common objects in context.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Microsoft coco: Common objects in context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.576699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.576699Z digest=sha256:53da0aa95e572e03a924a341edeabbb9a18c1b1f3aff558fa85b618ecb6c90c1

Observation 248c01ea-a96a-49b3-995d-e45a0ae1282e · outbound

This paper cites Improved baselines with visual instruction tuning.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Improved baselines with visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.949057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.579201Z digest=sha256:70696a55f5802ed204f5aaa7cda7ac3027b6b736bf3e7736fdae2f50bd58bf3d

Observation ef4277ab-62ff-4b1a-b74b-6eef7ce7000a · outbound

This paper cites Visual instruction tuning.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Visual instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.941936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.582151Z digest=sha256:95b99daa15baa79f76508631c3c5da0ba5e6df9ecf2c85f6de3819d36d91e6e9

Observation 5d57b010-6409-4b93-9338-ac143c8aa848 · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.933651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.584530Z digest=sha256:6276ef1be063737776cccb921750dd7b1c8612dc436ee15365785cea04ec998d

Observation 213a7dd0-0c0c-403a-981f-87062b9f6291 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.586907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.586907Z digest=sha256:29828b9242fc721fee991fae5eb72b53574a261040537fefc783a095e274e11f

Observation 94369438-5983-40ff-826c-3163b35467b8 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards deep learning models resistant to adversarial attacks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.922818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.589981Z digest=sha256:57c1f36853dca0f6165f6ec03d565e5e2109243107223585126e60aad2cbc40c

Observation 5944b3f7-8e89-4ca7-b0c4-327b9ca5cce5 · outbound

This paper cites Understanding zero-shot adversarial robust- ness for large-scale models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Understanding zero-shot adversarial robust- ness for large-scale models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.915549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.593330Z digest=sha256:93248fe68ceebb3ca8782d48e4e3d5ff244142f29cad7ea8e7a45119baa938ba

Observation d556e3a8-f0d7-439f-a070-0d34b539c37a · outbound

This paper cites Deepfool: a simple and accurate method to fool deep neural networks.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Deepfool: a simple and accurate method to fool deep neural networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.595517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.595517Z digest=sha256:b679df665cc42934cd2b3182ed2d8c2a6fbbc91ac9188997ffbeb052e06c556e

Observation cdf56dc0-8e55-4392-8ed8-56becba3d80a · outbound

This paper cites Bag of tricks for adversarial training.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Bag of tricks for adversarial training

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.904859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.597881Z digest=sha256:ab6937e4461f054592963b19adda76305986c27f594675d0fbb1aebe37e5667d

Observation f12e9d07-da67-49af-9d46-e5fc31267973 · outbound

This paper cites Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Learning To See But Forgetting To Follow: Visual Instruction Tuning Makes LLMs More Prone To Jailbreak Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.601160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.601160Z digest=sha256:301f8a1b2d2f652c050b1c534c176aacf638767661981bb729ca7464dda7bde3

Observation dbc0083e-1950-48ec-869b-95fdb093f31f · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.897410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.604997Z digest=sha256:e29ededfb8bd3a46d5833b6d07d953416081ecdb1cb4cf2edd6c98570de361c1

Observation 29cce8ce-5de9-4ca8-8843-04b3bd982534 · outbound

This paper cites Visual Adversarial Examples Jailbreak Aligned Large Language Models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Visual Adversarial Examples Jailbreak Aligned Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.607796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.607796Z digest=sha256:92821b96c0d4f48bb1fb1abef37fb83174fe6b2bbe58055230b8dc9d6475b779

Observation fcc9fe5d-90ee-4e2d-95bc-428aa1ce8f2c · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Visual adversarial examples jailbreak aligned large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.890475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.611286Z digest=sha256:1b290a8fba784927d52ea03a84ca8eff4072d14fb2a1a3252822697d8b14e8fc

Observation 6b7bd383-f61b-4a93-879d-b3fda75be958 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.883310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.614535Z digest=sha256:891cb33cef7351d671ff117abb5a5c3938b17320adfaab3844aab4b5def62906

Observation 26aeb469-7347-48ec-9aee-c7d26101fb77 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Learn- ing transferable visual models from natural language super- vision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.875761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.617015Z digest=sha256:124f503e8b2f5d72e91ced0e2e8db4308734581805f8852fde9dd6ee0456b62d

Observation 82a9f85a-2e5b-4e8a-a2f7-5997a119c15c · outbound

This paper cites Failures to Find Transferable Image Jailbreaks Between Vision-Language Models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Failures to Find Transferable Image Jailbreaks Between Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.619802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.619802Z digest=sha256:cdd5e1a86e4541bb040a6c075f37790b4c141652d100f1ac56883fc58732a075

Observation 0c0b25ae-eb55-48ca-905e-2a9518b4b3e2 · outbound

This paper cites on the adversar- ial robustness of multi-modal foundation models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness on the adversar- ial robustness of multi-modal foundation models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.867179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.623177Z digest=sha256:d1f7a6b5023f39265b0b5596be5a970a797fb08d528cd34b30ba51fded61f929

Observation 07f5391d-2958-48de-be57-a17e25675266 · outbound

This paper cites Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.626038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.626038Z digest=sha256:5e2c88e34b0fe8ed1fb002c0c7b5ba601070ed153e0d9a2cdb2bebe9d62020b0

Observation ece5ae28-d52b-4aab-b852-f61b17148b5e · outbound

This paper cites Adversarial training for free! NeurIPS, 32, 2019.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Adversarial training for free! NeurIPS, 32, 2019

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.858541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.629344Z digest=sha256:92327693feb499e8af748a50ef4786397609fe2f62aac74f51c0b128589d0bf0

Observation 72eee1f7-2fd3-4259-b926-a77f6e53ff14 · outbound

This paper cites Towards vqa models that can read.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards vqa models that can read

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.850576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.632406Z digest=sha256:9397b58cbbf4552e1d9de071c7af6591de7fbcf75db3f539c3b7397a9f87ac3f

Observation 1b70138a-de70-458b-9c42-0742c26f3845 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.636036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.636036Z digest=sha256:d77add6ebd1d3e60c64504437dca2b5eb9bba90180a44ca7642c06a88eb71342

Observation 8d0c19ab-51f6-4a05-b6e2-23655501e6b0 · outbound

This paper cites In- triguing properties of neural networks.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness In- triguing properties of neural networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.843355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.638728Z digest=sha256:cea058632b8af3e3748c676b1d7e1f6cf25b97f16f689ed0f352dab0c98d14ad

Observation 51cfb025-1c29-4539-8f49-511e486369b2 · outbound

This paper cites Robustness May Be at Odds with Accuracy.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Robustness May Be at Odds with Accuracy

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.641502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.641502Z digest=sha256:17f024a42acf7cc897ec7a2c58518427669971747d0933dbb5749125e2dd92f4

Observation e56ae7a6-cc41-4e8d-ae8c-a165d86614a5 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Cider: Consensus-based image description evalua- tion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.644618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.644618Z digest=sha256:07249fff6c29428870f4366e56c294d2752bf75c565f3e162d43745fa92b7f17

Observation 53a731f3-6e49-4c44-9fc6-2f7dcd64e6c4 · outbound

This paper cites Revisiting adversarial training at scale.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Revisiting adversarial training at scale

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.831539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.647293Z digest=sha256:f5b4785abadf5214e063a5a5194f14bc78024de83da3fdeef45cc0aa4a9b5b85

Observation 6db1a99e-93fb-4d9f-a784-2ea8cbbe4510 · outbound

This paper cites Zico Kolter.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Zico Kolter

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.823804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.650544Z digest=sha256:6ae4bd60585ae95cc200ebaa0a17d42993475cc546f41c3e6cf2e6e325a7ba6c

Observation 57a563a4-ff85-4f1e-bad9-216cae047c51 · outbound

This paper cites On the safety concerns of deploying 14 llms/vlms in robotics: Highlighting the risks and vulnerabil- ities.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness On the safety concerns of deploying 14 llms/vlms in robotics: Highlighting the risks and vulnerabil- ities

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.815350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.654010Z digest=sha256:4f6a2cb84be0dcf6cbaa8833c1bef8f6a59f3b670a8cbefa1a4e6d36b29db7f0

Observation c6283cf5-d438-48d9-aee3-b1a49af7f99f · outbound

This paper cites Feature denoising for improving ad- versarial robustness.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Feature denoising for improving ad- versarial robustness

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.806475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.656846Z digest=sha256:f6222a4bcca5754e58acb583392dc0b68b776a58fba1c50bf3836214a1905ed5

Observation 16f13be1-37b3-4f63-a5f0-faf8bdbfc405 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Coca: Contrastive captioners are image-text foundation models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.796927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.659428Z digest=sha256:6213840b098080f98bab9b882083f86429f5c6ef278983fcdc40c915c48a3a32

Observation 732067d3-7c24-4928-b2c3-dbf6bdb75365 · outbound

This paper cites Towards robust resnet: A small step but a giant leap.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Towards robust resnet: A small step but a giant leap

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.789140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.662936Z digest=sha256:8fcdadd5cb51b911db37f1f12a96ee561b9300aac4c7782dbebcff38034dd866

Observation d4e81124-00cf-49c2-bff3-08dbdddc8ecc · outbound

This paper cites Attacks which do not kill training make adversarial learning stronger.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Attacks which do not kill training make adversarial learning stronger

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.781588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.666337Z digest=sha256:f7369cfe83a10504ec593e55b305944005315e6bc57eff619ca0540da15e050d

Observation ec52100a-d665-4119-83df-c3dd69b5c902 · outbound

This paper cites Geometry-aware instance-reweighted adversarial training.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Geometry-aware instance-reweighted adversarial training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:06:24.773519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-10T20:06:24.669605Z digest=sha256:f862ebfe3d99199b0380efd277d38eede2fd0551ed14920fd603d8b18e735f11

Observation 22610ae9-0f3f-4e3e-bb3d-a7a28ae35c7f · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.673382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.673382Z digest=sha256:396cbb791f7caaf21d2639ab8781d5308d99a8f68fbff42acd067fa5e0b9b33f

Observation 5a0fbdf1-167c-4dfc-8fe4-5e62dee3e157 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T20:06:24.676479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:06:24.676479Z digest=sha256:c47913b6abaa9d402c56cc6bf16ae6e7b0d5b0e28d785aa93e38f2098121c2d0

Pith citing papers

Observation b813acf5-d0c6-4524-b03a-fbfe460883b9 · inbound

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models cites this paper.

Grounding-Driven Attack: Improving Encoder-based Adversarial Transferability against Large Vision-Language Models Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:46.941141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:57:46.941141Z digest=sha256:29f555139b41eac69dbc1ecaafb4dd5edead22e01983dd2a64169889371dbb80

Observation 0ad91002-c2e8-4cf7-879f-29230efe9357 · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.573666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:02ab4f08a260471f51440302a3a49a33fe0c3a37e6d66e5fd946fc2cab1d842e

Observation 0372334e-33ec-463c-9677-8bba24ef033d · inbound

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models cites this paper.

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.847659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:04:30.654255Z digest=sha256:71c111dc56c22a4d3c75e78196baca07593aeef3ea4444a66f192875b661c7ed