Pith. sign in

Paper Citation Record · LEDGER

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

As of 10 August 2026, this Paper Citation Record lists 100 of 130 outbound references and 0 inbound Pith citation observations for arXiv:2502.06390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06390 v2

Coverage vector

measured 100 of 130 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:38:17.453227Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 130 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 099f1e85-e121-4dd7-82fa-771525c9fef7 · outbound

This paper cites Efficient multimodal large language models: A survey,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Efficient multimodal large language models: A survey,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.028360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.028360Z digest=sha256:d77187e023b5e671d383f9be7310ea8859274c88308efa53a3e6a5ec6d14cbc2

Observation b736f5bf-c1a3-4c91-8983-9ffb7500cb1d · outbound

This paper cites Vision-language models for vision tasks: A survey,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision-language models for vision tasks: A survey,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.033499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.033499Z digest=sha256:146fb50f5c895abe99f7b7889dfde9a546139b8a9d4ca193072ae90e2c27cec9

Observation 15d7f862-3768-4384-9218-bfcad2b6fdf1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Learning transferable visual models from natural language supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.038157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.038157Z digest=sha256:ab6c3ee5c4d5213095176d264d4dd61d7512fd1c5535a45d8367bc30caf9e630

Observation 1f70bd67-0e81-4ffe-9075-4d11bb96dfdc · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision language models in autonomous driving: A survey and outlook,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.042884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.042884Z digest=sha256:0f220c48cc561f8e7b25ff68397d19f8c9865e9fc08c41e28537d5bf051d9974

Observation d46bf3ae-44a0-48e9-9bb1-4a01bf7215d1 · outbound

This paper cites Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.047450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.047450Z digest=sha256:ced51c7b7aac67c9696267b9c75a7e668ebcf2068dbad7da46ad08248384da5e

Observation 32dbaf8f-1567-4bbf-9a52-3ef8a9f551fc · outbound

This paper cites Visual instruction tuning,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.052425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.052425Z digest=sha256:dd9382683d78d8f98246896af5e2004b0a1170dcbe4e9ba8ebffb890640f3c14

Observation cd77a381-8ad5-4801-b6df-7d80726574a8 · outbound

This paper cites Listen, Think, and Understand.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.057432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.057432Z digest=sha256:f5992bb4894027399ac59fca35a5671d467b67703590307867ac4d107a3d4e89

Observation 746508a9-917e-431b-a74b-ea30b275419b · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Sonicvisionlm: Playing sound with vision language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.062234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.062234Z digest=sha256:62408a7aa14f4baf09294563d1e450d29daa1f3f9c9fadd70165171ff315cb8b

Observation 46bd1cec-8f67-4472-80c8-b5a06e2115ba · outbound

This paper cites Grounded language-image pre- training,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Grounded language-image pre- training,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.067115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.067115Z digest=sha256:f91771871c438a8de2a33891be9055b5eeabd3b2406ccac3ef867a5b036faba2

Observation 00734e26-a685-47bb-991a-a66f5039580d · outbound

This paper cites Segment anything,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Segment anything,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.071570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.071570Z digest=sha256:61d221af447aa8f64bac2b0e0771c1dd3cbf6d50c49cd8a9dff531eb320edcc2

Observation 845577d3-c834-4cf7-8a40-9e0b9b155257 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs On evaluating adversarial robustness of large vision-language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.076098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.076098Z digest=sha256:d6832c4140256807375a76b3467780e152d5a3c7260985832377a38d9bd305ac

Observation 067c0329-e16f-4c3e-88ac-a626ce7d55e7 · outbound

This paper cites Mma- diffusion: Multimodal attack on diffusion models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Mma- diffusion: Multimodal attack on diffusion models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.080542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.080542Z digest=sha256:51b8045417a933a4cea55c9277ba1a6d99d26faef43ef2d9408f3e83f71f205f

Observation d508c24d-5592-4712-a758-1cb1b657c72c · outbound

This paper cites Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.084735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.084735Z digest=sha256:92eceb15a73f2385d9996c1391c64eef76129cd317efc831dde1445aefa14ac8

Observation 631adf9a-90c7-43e1-ae22-9610b9e7c1ea · outbound

This paper cites Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.089320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.089320Z digest=sha256:0de52d7ecd966cae40c14541d84ca0df80d36c88f583fc1d43dec8fd2dd63638

Observation 3037f0c1-e04b-4d46-a13a-2a6106a1b05a · outbound

This paper cites Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.093802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.093802Z digest=sha256:e011dac3d02012be64546076f2268b8c72157e1ab12cf78d368bdafc32dd9031

Observation 555c05e6-2912-4c4c-82e6-9d5d054b1f52 · outbound

This paper cites Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.098397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.098397Z digest=sha256:596c1590e64b3f9402974a6c42f4f7d2d3591a524cc93aceee951fd8c2cea272

Observation fd9d8eb5-393c-4e58-a585-61fa9825b70d · outbound

This paper cites Safety of Multimodal Large Language Models on Images and Texts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety of Multimodal Large Language Models on Images and Texts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.103048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.103048Z digest=sha256:c75ccec9cbac14215fa626cb7c3bb89b2119b41071d66bec6f63f0f3587c5aef

Observation 25248a1d-fd28-42f5-be5f-2c1319542a97 · outbound

This paper cites Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.107046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.107046Z digest=sha256:34a9fb8af1fb2496fd7730d210fcc0e9a74c9bdf3f42b45fce380561e702300c

Observation 0db4d2a4-38f5-4cce-ba9f-41abdcbb29b4 · outbound

This paper cites From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.110619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.110619Z digest=sha256:67cede16cac1c505bca5698f83547bb0c9cb86f4d71156ab53872c922ccfae03

Observation 0917c57e-d361-42ac-b959-9fe6f0a3b87c · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.114309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.114309Z digest=sha256:d9ca7abe2725245abee9de3c05d072b92301da82a4bb858b7b1fec4933859e76

Observation dd5065d5-9f21-4841-aeea-08d517887c04 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.117952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.117952Z digest=sha256:778f64af8cd89a9846b924ea8e9f77e94686f1f4593904c5d96b3b0c0b5cc1c3

Observation 1d8946c5-1be5-4d0e-9386-e4773a6471da · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.121586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.121586Z digest=sha256:ddbde574fe41c96fc26cb5ec6e5bc7e7b51aaf2a7a570a52098044389350e31e

Observation 941f2746-7dbf-4864-bf01-cd80d9282a10 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.126010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.126010Z digest=sha256:c2b41c444963549c9d6a646d35cf88876097b5a21791dce50d34cfd064df8579

Observation ef38d53c-ace8-43ce-8583-d49993e6e61f · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.130611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.130611Z digest=sha256:f22092509cc0779c810d6e91d60fdc3acec5b14bfe21b0537499e753bd5a094c

Observation 003f4fc9-c90d-4d73-a018-ee564682fb36 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.134977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.134977Z digest=sha256:1e871d1557269692f2da7e1e429c9e242a627e3ce01e4beb4d949600bb169e2c

Observation 11729ae4-7e6f-4177-b07f-e09cc9c8cd7b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.139189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.139189Z digest=sha256:a34b9d26a80eb4d9de9db1c32151b9521c83b916f37f651d432e50fadb740af4

Observation f32d0cc9-7686-4cc0-aa81-9e02a377b1de · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.143275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.143275Z digest=sha256:4cb61b14b54dd93bdf02ad208d50cd1556cca6b38c58457c5d17a5d085735ecd

Observation e4461dbd-432b-4e2f-b074-9f31f946ceee · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.147602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.147602Z digest=sha256:e502cea3a885cb31cd26a82aee9a56214031415d967542df5bee33655c5b1058

Observation 16c56381-9fb9-4393-9292-6c4bbff720c2 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs OPT: Open Pre-trained Transformer Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.152111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.152111Z digest=sha256:ae98cbaabbf49b2bc2d3d24e50a20c2667e57c1d232be7ae8653140da47aac31

Observation 983b002b-3f83-4f76-a676-5d9ef691c433 · outbound

This paper cites Scaling instruction-finetuned language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Scaling instruction-finetuned language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.156459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.156459Z digest=sha256:ced5be7e0034c632c5496d6dc0b17d8a85aed8ca1bec5d5952f151439caf01e3

Observation edddaa3f-a9ee-4945-bae4-2ae520539233 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.160454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.160454Z digest=sha256:449fb17d9865991192b6cbd1a4c5b63f4903c862c307b91768990b8e2b0ab6c1

Observation 5c14a08b-325f-4be1-aaf3-66564b0eb654 · outbound

This paper cites Introducing mpt-7b: A new standard for open- source, commercially usable llms,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Introducing mpt-7b: A new standard for open- source, commercially usable llms,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.164950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.164950Z digest=sha256:5b880faf1ff576122c05da7aa1ee7bc048af494ac978e85d108cac05c94e64e8

Observation 3416eb19-6004-46f8-980b-2cfd858a6705 · outbound

This paper cites Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned and chat models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned and chat models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.169487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.169487Z digest=sha256:813d93ed1399489d2b3f5c1817a00f62a56011e84d47d42229f351b66a96d654

Observation bdc303a5-d975-4f94-b9d6-8dc809e39f17 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.173697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.173697Z digest=sha256:91fd2726ca5a8493bf166d025370886cb2369a0ab43b617fc10757e497550f8c

Observation 4f146210-a6f6-425d-98f6-66988df75d7c · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.178121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.178121Z digest=sha256:3a55cf428e4199aefd518434de12bf4b87608a1ab8d28ad430f283aeec61cfb5

Observation d114bf67-f0b6-4562-9331-e42f5946f563 · outbound

This paper cites Gpt-4v(ision) system card,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Gpt-4v(ision) system card,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.182522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.182522Z digest=sha256:655e5d76366107e43cba195ccf545a1c25aed02ce6f62230aec3e875f1aa0d2b

Observation 8c7cf713-d876-4f54-a548-41487f2a4a10 · outbound

This paper cites GPT-4 Technical Report.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs GPT-4 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.186790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.186790Z digest=sha256:4ad37a986a0b3aa4c6495df4f2c0c0b76b379561c24b480445f043952997c9a0

Observation 0eb8dd0f-d1cc-4a0c-be49-0ee2c92b8340 · outbound

This paper cites Gemini: A family of highly capable multimodal models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Gemini: A family of highly capable multimodal models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.190925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.190925Z digest=sha256:1fab5ea76af19d2548eab79282fa53a3b124370be1022717004609fb9be1b6f2

Observation a4d978bc-4cfc-4a3a-9fa0-4c829556a078 · outbound

This paper cites AI, “Bard,” 2023.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs AI, “Bard,” 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.194855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.194855Z digest=sha256:b8ddafe5963c94e7812ca893d7b542d0bbe22a15427c7acb3eb6f621e8abccfb

Observation b7fd3fdd-3840-4f90-9b56-b47225b8d540 · outbound

This paper cites Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.198969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.198969Z digest=sha256:6a26477bbecbce79e1672a9aa9eb7343ee2ea800752b6469bcf57ed2afc4c968

Observation 5625fcc6-37d1-415f-80bb-faa10be4830e · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs High- resolution image synthesis with latent diffusion models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.203461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.203461Z digest=sha256:fec0c08c37e987393c99a06329cae42d1337acf1292f1fe2f223baaca8e30a99

Observation a177be31-e27d-4e68-8ddf-e0b80f2dcfb6 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.207019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.207019Z digest=sha256:f8fe041893e948438869d98ba33bd7d893f1c79abace999395691f63e5f8e5de

Observation ed837fdc-ba9a-4819-bc70-5edd775120a1 · outbound

This paper cites Zero-shot text-to-image generation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Zero-shot text-to-image generation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.210902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.210902Z digest=sha256:744f0fb2ed2542542001572786bf75a33076198cee5882202ba99fcc95335e13

Observation 1100fcc0-ad9e-4ec8-9b38-c0bf3599914a · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Imagenet: A large-scale hierarchical image database,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.214490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.214490Z digest=sha256:99806c8727f0eb5e534daeec47bd2fa4eda5e7f4e9510dbd890431b8108621d2

Observation e3f4e440-2293-4a9d-af2e-25122a822280 · outbound

This paper cites Microsoft coco: Common objects in context,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Microsoft coco: Common objects in context,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.218337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.218337Z digest=sha256:1b30b1a60afc2c01bbc500ec9932ea76939e23a56bf5936de8f43f56c2aa6d61

Observation ac85e91c-d502-4609-b080-0ce4f68c9b18 · outbound

This paper cites Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.221963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.221963Z digest=sha256:ef9c4100490751a37d66e1d1e00cc88eea064914af9d7ecd9f5ee96e3c4cb350

Observation 32692a65-3f1f-48d9-a6b6-76a59f1c9991 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.225531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.225531Z digest=sha256:39e0cac8ef6f5ed3cca28f0d91ca40b145f16d385d1300194d802b3abfb0946d

Observation a1c67bea-2577-4d10-a065-eb1c50eb896a · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.229098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.229098Z digest=sha256:f7b935405e22673b03c39db5a8428277754ac1aa5f9b4040b565912561a97a68

Observation 1e710829-82aa-40cb-95e4-a2246126d1d5 · outbound

This paper cites Laion-coco,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Laion-coco,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.233181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.233181Z digest=sha256:2a4b82627b480b7b1488d2db0f52018bba939de3859b538187921001d42d0ab0

Observation bab26b47-e212-4fc4-9b81-845154b749c6 · outbound

This paper cites Stanford alpaca: An instruction- following llama model,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Stanford alpaca: An instruction- following llama model,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.237161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.237161Z digest=sha256:88ca3d7007d2657d6d1af5bdd5aa6badaf7d25a5bd12886a339f0dfae64ac72f

Observation 3aa30c57-54e4-4291-a717-b60c00693bfe · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.241278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.241278Z digest=sha256:69a2f12480149c678557722a14f8514dd92a62d256e58cdf651fb4278fe1e65b

Observation ee3a88df-54f0-4c89-9c74-a5b3a3237360 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.245674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.245674Z digest=sha256:87535942a1889089b66cea4cd9c9b879c09864447e2c6963395006ea15a4af82

Observation 46bc22fa-4bd9-47a4-a302-53cf613a35be · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.250038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.250038Z digest=sha256:23debc8be4b3173b00cc78f529d33d5928127eb48641545bfc1a6296db04a97f

Observation f11ade04-d062-432e-aca9-7a1f0d77267a · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.254478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.254478Z digest=sha256:ba81b2d34683a7a3cdf1adb3cfb5cf8875103fa19c840e52a480bdeda2ba9aae

Observation 521af43c-1569-4d88-9ae6-c5dfa1c93189 · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.258909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.258909Z digest=sha256:1d3e5aa831cf5e13da33498a4c23557b217281b7e12ac40cd8eb7678eb944f14

Observation 41103c7b-89af-4865-9aa5-3ae0ef32b93e · outbound

This paper cites Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.263253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.263253Z digest=sha256:1121c83714e6a0ee73f8bb4bbdeffc07b1e3cd43ac7bf1fb7626820cd245a445

Observation 7ea2d64e-8301-4b41-9f70-5da442bf8d5c · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.268038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.268038Z digest=sha256:0aa27a9a4d170bd2ff2c14e2310983b7724d3f73df5699df30a41e5db33989e5

Observation 8f76bc81-b1b7-4a23-8d96-39574b69da3a · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs White-box multimodal jailbreaks against large vision-language models,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.272280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.272280Z digest=sha256:7cb5278788cc201d2fe6ce69d581992e195e2571c6341c2a823298320cc96fe2

Observation 708e9679-eccd-4a47-a55c-56ee013447f4 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual adversarial examples jailbreak aligned large language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.276363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.276363Z digest=sha256:d68b2a92b9e337b1da326373b1fcf2bfc80ab6f3ba3b0f8d939b7d207cd30f12

Observation 678e795e-b5be-45ba-8444-16542a03c308 · outbound

This paper cites To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.280459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.280459Z digest=sha256:a5c57f4295b6811c77f0fa7cd00c4095b0da0c39f6983725d32506d339ed8b51

Observation d5283b45-37d3-493e-b99f-6cd61d3b3c6c · outbound

This paper cites Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.285032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.285032Z digest=sha256:18759f7820e6ddf7cb4fd8044a310a47f3283541203aefa7828d101fcb42922a

Observation 83fb702f-4757-4424-8500-8ad9bc8d54a9 · outbound

This paper cites Are aligned neural networks adversarially aligned?.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Are aligned neural networks adversarially aligned?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.289387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.289387Z digest=sha256:d73089e5a0cad6200e0173cdfc0f4813a0b7234ef8a1c5b8886f34263ad7fce0

Observation e90ebd9a-795b-42d7-8e1c-cc832f82d4dc · outbound

This paper cites Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.293414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.293414Z digest=sha256:f338e5706f9223724bc968362d3fd7c11837afaecdb5efa0a14de136e3f2f71e

Observation fa808e8d-d9f1-40b2-9d2c-2f32762d6ecd · outbound

This paper cites Extracting training data from large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Extracting training data from large language models,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.297509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.297509Z digest=sha256:e6adf79484a345e14f45b58f5e9265f724e1d46b2fc0bb4ab9a2e1bdb7ac5248

Observation 1ee4c5d8-a577-4edc-ac4d-6a518faf0807 · outbound

This paper cites Va3: Virtually assured am- plification attack on probabilistic copyright protection for text-to- image generative models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Va3: Virtually assured am- plification attack on probabilistic copyright protection for text-to- image generative models,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.301812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.301812Z digest=sha256:fd83918287e9634ab0d76cd4f010e46937e23fbe2b7fcc2a48feb24d2e289dd7

Observation 623ce292-9288-4cca-968f-ae441b468e05 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.305857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.305857Z digest=sha256:aaf52de8711de26f5130e445b7ba0c5c30330416b36938b677b47f09b7144b22

Observation cabab675-5987-49ed-ba20-d59e8a0236ca · outbound

This paper cites Large language models are zero-shot reasoners,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Large language models are zero-shot reasoners,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.309887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.309887Z digest=sha256:5eab4fb66a4270555917d4656438a23950ba7b10e5cd7a050b4133f9c2226b8b

Observation 24ac7ab2-87d9-43c5-881a-87d8a9933c22 · outbound

This paper cites Black-Box Prompt Optimization: Aligning Large Language Models without Model Training.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Black-Box Prompt Optimization: Aligning Large Language Models without Model Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.314113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.314113Z digest=sha256:0b61ccaf6dae17c78f89d8acc0852246da4b2da8b451d48478527bad608b3ffa

Observation fdf0ceed-5041-4a53-84f5-954b56bd6274 · outbound

This paper cites Training language models to follow instructions with human feedback,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Training language models to follow instructions with human feedback,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.318645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.318645Z digest=sha256:9db2f03e68576dff97e23bbda5e3da8c920cfd6bcc742a277990f399342eddc1

Observation aca12052-b038-4c5e-ac9e-45d7fa2bc64b · outbound

This paper cites RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.323020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.323020Z digest=sha256:c2db3cd30b607ef04524b2d4feb92d22837733a37872947643aa0fd3f4fa8eb5

Observation 3020da8a-d357-4371-8332-36ae6a9b8ec9 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural language feedback,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Dress: Instructing large vision-language models to align and interact with humans via natural language feedback,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.327520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.327520Z digest=sha256:550125064407e5d7382dec5a62253c6d2bf3623c5ad60feba00ac88122d4b2e6

Observation cc4414d5-f3bf-409e-84a4-11a15b4ce7c6 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.331678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.331678Z digest=sha256:448a76bfc997ce1dcbc505a4625307c1c515dc62d4583049f6bf844a70b38955

Observation 5fa57646-fe9f-4fef-ae7e-95bd036355e5 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.335773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.335773Z digest=sha256:a7c91cadead102a0777379ce634c959e1075795f60f0a09c72f31e6852ad3663

Observation 1f6b0708-eff5-43d3-b9f1-d5d481deebad · outbound

This paper cites Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.340113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.340113Z digest=sha256:6949c9f5a1a04cad52a66d6a05afc0506ab0bd2cdac782b154c897f07923e999

Observation 99c2107b-9d2b-4ff9-b4ff-020b4345d0ef · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.343991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.343991Z digest=sha256:9cdaa679c183d615dfdb306e1d96d9df4dbfbea57725364ddddbfa4f2a72f4f6

Observation fb549679-a2df-41f3-ba76-183bff6136c9 · outbound

This paper cites AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.348048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.348048Z digest=sha256:e6aab231bf8ea5d7ed789166035f4dca0c4e04a3ae33a14820caa0473f52648d

Observation 347259e9-cd97-49d3-8e8d-41f127fdb9cb · outbound

This paper cites Adversarial illusions in multi-modal embeddings,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial illusions in multi-modal embeddings,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.352864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.352864Z digest=sha256:95e0b3efb9f5b160facb26f81fd5d84218794fe9ccf4be3868af2d4aef252a1f

Observation e21fc647-f769-4fea-99a0-637a2626a322 · outbound

This paper cites Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.357496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.357496Z digest=sha256:560df3b0f9903a56f8f58833d7206347ef6729a5b61ef74aec097550f6f5727d

Observation 6cd72d6f-3c28-4821-a83a-82c1c71a52e9 · outbound

This paper cites Badclip: Trigger- aware prompt learning for backdoor attacks on clip,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Badclip: Trigger- aware prompt learning for backdoor attacks on clip,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.361725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.361725Z digest=sha256:af31cef315c146a42c452237e68af3758a44402c55f54a900b7436d042ab929c

Observation d5a2d880-df5d-4819-a307-813218e66dd1 · outbound

This paper cites Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.365929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.365929Z digest=sha256:95578eae012b7826d7e86ce7b9171ad1aa6c1a158902d24f3afaaf5cd853c4fa

Observation e43542ff-2dbe-4099-94ca-8c6b8de0e925 · outbound

This paper cites An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.370563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.370563Z digest=sha256:e25b3696e418e903222bf275dc4bec849123d4df2f8efd8721cba179edcaa208

Observation b6cfe0d7-418c-4597-842c-c21e74afd408 · outbound

This paper cites Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.374971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.374971Z digest=sha256:dd13ab3c13e7973836f7c18aec02bb97dfe44e2fb042c25e3520af953ad23b53

Observation 715f514c-6c7c-48bf-9423-830233b485fc · outbound

This paper cites Misusing Tools in Large Language Models With Visual Adversarial Examples.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Misusing Tools in Large Language Models With Visual Adversarial Examples

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.379148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.379148Z digest=sha256:d501eef7f70cd972ff7db1e9db630eac3e439ce8d66755f081533cfbdd19be53

Observation da0d21db-ba63-41b2-967b-01947986e400 · outbound

This paper cites Prompt-driven contrastive learning for transferable adversarial attacks,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Prompt-driven contrastive learning for transferable adversarial attacks,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.383721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.383721Z digest=sha256:ed29f8b67f7cb4758233920a39f2ea39dc208133aaccd650d9fb5f61c911baaf

Observation 395a51a3-4556-4e29-8cee-74623948076a · outbound

This paper cites Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.387834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.387834Z digest=sha256:b500e9886323e3656f0a4fe1d932ba0257e3b09f56b08a76134b6d7d65e9930d

Observation 24d757c0-3786-4b74-83e2-fbd6c3eb2232 · outbound

This paper cites Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.392810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.392810Z digest=sha256:2220536d4309ccc49989d0d2959cdf8a2137b47476c0f8cc3a8ace38c022ba5c

Observation 5cc07c5e-937b-4b1a-8724-f3c748fb577e · outbound

This paper cites Slowlidar: Increasing the latency of lidar-based detection using adversarial exam- ples,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Slowlidar: Increasing the latency of lidar-based detection using adversarial exam- ples,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.397009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.397009Z digest=sha256:31825e0d8f34b66d1301da9e497702d4d163bb9a2b146fba60894bf23e59aa2b

Observation 3cc0603d-1d78-488a-b509-d0acdf803e9a · outbound

This paper cites The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.401241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.401241Z digest=sha256:b18a7e34d1d4a83e43ff2c25cc69462b5ee3f458cfa185586a405212c5840309

Observation 27d46af8-06b9-41d9-8eca-990a446d7cb5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Chain-of-thought prompting elicits reasoning in large language models,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.405556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.405556Z digest=sha256:9c4c4fd015ef690c6cad16e2b8a5f7909b744c5bf3e51cec4c0500c34c3b36b7

Observation 04bf9728-5647-4b68-8153-ba90901a86f0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.409426Z digest=sha256:d4a05faac62910b925dfd0b9b128cb8aae81d64f39b757ac8788741b745bf6dd

Observation 842c88c7-17ed-4f21-8dbb-e531802b598b · outbound

This paper cites Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.413950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.413950Z digest=sha256:cc91a6797133c4a0c36d6c9f59308249b481f38948d78fa41c9a105c5517b3c8

Observation d3cb49b6-d9e9-4162-bbb9-7fc7ad644f0e · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Explaining and Harnessing Adversarial Examples

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.418146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.418146Z digest=sha256:a584d91926cad38e0498bb74980f4e9b0d61a1c49e5cff4e256aae10f82573a2

Observation 5e13c29d-db2d-4fc5-b128-4ad787e5a1ba · outbound

This paper cites Adversarial examples in the physical world.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial examples in the physical world

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.422466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.422466Z digest=sha256:05b947839d06b987d9866c083e12a9684d5ed2d6fcdff06f0b9575a8dd6eb94f

Observation 63a5e170-3e13-4c20-9864-c3fefbd01bd3 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Towards deep learning models resistant to adversarial attacks,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.427191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.427191Z digest=sha256:87b198ed2946a962648547410dddcdd7ff91f90e501033a4a5946f42ee5ef3b6

Observation 96a42497-ffce-4860-ae27-4461d962c1b0 · outbound

This paper cites Boosting adversarial attacks with momentum,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Boosting adversarial attacks with momentum,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.431794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.431794Z digest=sha256:609ad757682c5304866fa1248bf44557cf5c56840379cb70116ab39ec535bc7d

Observation 7e134122-f8e5-49f0-8480-91fcb4c8d387 · outbound

This paper cites On the adversarial robustness of multi- modal foundation models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs On the adversarial robustness of multi- modal foundation models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.866336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:38:17.436070Z digest=sha256:a95c59c0c34ecf0c5dbb909dd65654d6131ab88dfcb00b782ea15a103246fe94

Observation 796a9483-6eea-4aed-93bf-0e36a26ff584 · outbound

This paper cites Transferable multimodal attack on vision-language pre- training models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Transferable multimodal attack on vision-language pre- training models,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.852715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:38:17.440408Z digest=sha256:bef3d2a900a65fcdede0be696b48483bac26ea63d6fd7bfc672a1160c58c0e52

Observation 74fc4860-c25d-4864-acb5-a8d646622cc4 · outbound

This paper cites Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.838566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T15:38:17.444773Z digest=sha256:19efe4a94054dee2e70af1c0cbc1bff634f0c864f92599ac96e4e77b5578bf6f

Observation f730f568-73ba-451a-927c-a27f96afe440 · outbound

This paper cites Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.449204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.449204Z digest=sha256:b00cc8f66e8ace4b783141b3aa04bb074034e10ba980e7e82bbf080af6dec610

Observation 28158ef0-3b69-49a8-b5e6-5c1c8098beb6 · outbound

This paper cites Adversarial Patch.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial Patch

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.453227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.453227Z digest=sha256:d207846a6f69557107579f505d944d26f8b14902e8ee09546c508025231f29a4

Pith citing papers

No inbound Pith citation observations are available.