Pith. sign in

Paper Citation Record · LEDGER

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

As of 8 August 2026, this Paper Citation Record lists 100 of 130 outbound references and 0 inbound Pith citation observations for arXiv:2502.06390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06390 v2

Coverage vector

measured 100 of 130 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:38:17.453227Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 130 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 099f1e85-e121-4dd7-82fa-771525c9fef7 · outbound

This paper cites Efficient multimodal large language models: A survey,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Efficient multimodal large language models: A survey,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.028360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.028360Z digest=sha256:921ac19cd63763b5f175a56720107185f2cdecd14b1c29d98e458138b21ea624

Observation b736f5bf-c1a3-4c91-8983-9ffb7500cb1d · outbound

This paper cites Vision-language models for vision tasks: A survey,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision-language models for vision tasks: A survey,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.033499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.033499Z digest=sha256:72c8d310c3ea322a969507df85ec9e07fe2ef65c3bd98a93583d28324cc7a765

Observation 15d7f862-3768-4384-9218-bfcad2b6fdf1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Learning transferable visual models from natural language supervision,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.038157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.038157Z digest=sha256:5e9b7de68a0fa97145af1234e8ec542343a963b38cc19b57d85a4dddf2ec4848

Observation 1f70bd67-0e81-4ffe-9075-4d11bb96dfdc · outbound

This paper cites Vision language models in autonomous driving: A survey and outlook,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision language models in autonomous driving: A survey and outlook,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.042884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.042884Z digest=sha256:a0aa659dfec9788171ecd5a19ce1e237c7630351950403773cafba357aa8a019

Observation d46bf3ae-44a0-48e9-9bb1-4a01bf7215d1 · outbound

This paper cites Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.047450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.047450Z digest=sha256:f4081afeed44b7745243b193fba91a121f8532a45fa62ce9fec6036e1dce4fc3

Observation 32dbaf8f-1567-4bbf-9a52-3ef8a9f551fc · outbound

This paper cites Visual instruction tuning,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.052425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.052425Z digest=sha256:a65f4c4c057be0acbabb976d1ff5c2f9f92948e0803482a0d1d557b9b6af4c92

Observation cd77a381-8ad5-4801-b6df-7d80726574a8 · outbound

This paper cites Listen, Think, and Understand.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Listen, Think, and Understand

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.057432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.057432Z digest=sha256:cd1336e200e2a5e86719386e0154bf8e7e7bea9488daf07a2cbc3b2503a6e9d9

Observation 746508a9-917e-431b-a74b-ea30b275419b · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Sonicvisionlm: Playing sound with vision language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.062234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.062234Z digest=sha256:b3d7bd07ce06358c7c9414c8216c0b27e3acca8e1450b4704ecdef01534154f8

Observation 46bd1cec-8f67-4472-80c8-b5a06e2115ba · outbound

This paper cites Grounded language-image pre- training,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Grounded language-image pre- training,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.067115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.067115Z digest=sha256:1cbc58cc4f5197433f0c4ebd5e84ec920338bc8c3d2d065ecdd91ef547ac3379

Observation 00734e26-a685-47bb-991a-a66f5039580d · outbound

This paper cites Segment anything,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Segment anything,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.071570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.071570Z digest=sha256:5dd4dc2cb26c419926931a4898d1336502a30044c718f04b2390bb2d6cb34a64

Observation 845577d3-c834-4cf7-8a40-9e0b9b155257 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs On evaluating adversarial robustness of large vision-language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.076098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.076098Z digest=sha256:69c2d77b7eb04b8f8dd902526a3ee8972694bdc6388aa281123f20668db9c928

Observation 067c0329-e16f-4c3e-88ac-a626ce7d55e7 · outbound

This paper cites Mma- diffusion: Multimodal attack on diffusion models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Mma- diffusion: Multimodal attack on diffusion models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.080542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.080542Z digest=sha256:cf2b3d557b299ebb2ca79827afd3dd2c4bf2d6bd3014c36fc9a93c1d41972455

Observation d508c24d-5592-4712-a758-1cb1b657c72c · outbound

This paper cites Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.084735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.084735Z digest=sha256:cc8d8a9b4e08bf5d987d0cfa3523fb8b7c4818710b92487a80cc6cdd2026acd3

Observation 631adf9a-90c7-43e1-ae22-9610b9e7c1ea · outbound

This paper cites Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Character

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.089320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.089320Z digest=sha256:625e8fc4ad16e8ee6d19c09605d2030edca550c192f1002737a71f154afdd5ee

Observation 3037f0c1-e04b-4d46-a13a-2a6106a1b05a · outbound

This paper cites Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.093802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.093802Z digest=sha256:db03508d9f9d43c4a79c30cd963af7f77b379ad4f93427f6949af81167a73358

Observation 555c05e6-2912-4c4c-82e6-9d5d054b1f52 · outbound

This paper cites Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.098397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.098397Z digest=sha256:6a165ca14ac7681bdd66f42845914400a811aef78d16581a9ebba2307709c6d8

Observation fd9d8eb5-393c-4e58-a585-61fa9825b70d · outbound

This paper cites Safety of Multimodal Large Language Models on Images and Texts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety of Multimodal Large Language Models on Images and Texts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.103048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.103048Z digest=sha256:e4b50d1b593402a8f36b610b80e875fde40ad1c0d03a8e194d19bf8ac136cc4b

Observation 25248a1d-fd28-42f5-be5f-2c1319542a97 · outbound

This paper cites Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.107046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.107046Z digest=sha256:d6d063ad4234a53912f993bc5956455715afacc17db46466d9dd45fec695fddf

Observation 0db4d2a4-38f5-4cce-ba9f-41abdcbb29b4 · outbound

This paper cites From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.110619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.110619Z digest=sha256:b4003adb74b72fbf371a68354b6285f15f88a7df14cc9c5c90090fb5f17b09e7

Observation 0917c57e-d361-42ac-b959-9fe6f0a3b87c · outbound

This paper cites A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.114309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.114309Z digest=sha256:1789531bb159164ab577b85974f3b91cac15e01ab31bda5901c754a62ab9b628

Observation dd5065d5-9f21-4841-aeea-08d517887c04 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.117952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.117952Z digest=sha256:e0ff37afe49e5b90adfa175386eaa84b5684f2b37daabdb1ca6fe9338a2e72da

Observation 1d8946c5-1be5-4d0e-9386-e4773a6471da · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.121586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.121586Z digest=sha256:a385affe34ff7d3e2bc2f094ddf36238d8348a18bc3dc6a18bba7f6f79d75574

Observation 941f2746-7dbf-4864-bf01-cd80d9282a10 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.126010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.126010Z digest=sha256:5a01863e7ad6cd76193849d15917e2b2402f86c4a2cc6b0f39303726e9aee395

Observation ef38d53c-ace8-43ce-8583-d49993e6e61f · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.130611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.130611Z digest=sha256:34506316c1e7250c779ab59d8a3e1b77f0e676a34101866898bc30aa24e36b21

Observation 003f4fc9-c90d-4d73-a018-ee564682fb36 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs LLaMA: Open and Efficient Foundation Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.134977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.134977Z digest=sha256:a835dc560b060eccc478a45859926420dacbf0c47b4c1567d02a28180ce6d483

Observation 11729ae4-7e6f-4177-b07f-e09cc9c8cd7b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.139189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.139189Z digest=sha256:14a68fe9c2e7f91dbf03a8740a99824ff9c5d4eb08af0d05500484497a838951

Observation f32d0cc9-7686-4cc0-aa81-9e02a377b1de · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.143275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.143275Z digest=sha256:21ea013c3d62df8f9ac30d7c4dc13fea66c6e506128d110834600502698b98c8

Observation e4461dbd-432b-4e2f-b074-9f31f946ceee · outbound

This paper cites JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.147602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.147602Z digest=sha256:c233ccddb21972638e11df74ed7282f11f0d6f1a613aa6e01e67e770fe091411

Observation 16c56381-9fb9-4393-9292-6c4bbff720c2 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs OPT: Open Pre-trained Transformer Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.152111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.152111Z digest=sha256:3bc1eea8c1e3c6c251dce3dfae84b16a458c2ddaf3b38a613e99c4de0582bbe5

Observation 983b002b-3f83-4f76-a676-5d9ef691c433 · outbound

This paper cites Scaling instruction-finetuned language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Scaling instruction-finetuned language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.156459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.156459Z digest=sha256:9cf42f24cc8cea28c8f3003b1afc20f153f9fb141746bff5be24c70345ac3ff0

Observation edddaa3f-a9ee-4945-bae4-2ae520539233 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.160454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.160454Z digest=sha256:43580585a6e0e10bdde0dd0dddb41082e58d4542b77aa4188236a708cd69c424

Observation 5c14a08b-325f-4be1-aaf3-66564b0eb654 · outbound

This paper cites Introducing mpt-7b: A new standard for open- source, commercially usable llms,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Introducing mpt-7b: A new standard for open- source, commercially usable llms,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.164950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.164950Z digest=sha256:b1fbc97190bdbe4cc2f1a9311d02e31fd88ce30f21bdc8222ec2e7f116fa9600

Observation 3416eb19-6004-46f8-980b-2cfd858a6705 · outbound

This paper cites Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned and chat models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Releasing 3b and 7b redpajama-incite family of models including base, instruction-tuned and chat models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.169487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.169487Z digest=sha256:72568ecbe650f612378a1194d01076a50685dc5ba9481e4019bbfa1b594656ed

Observation bdc303a5-d975-4f94-b9d6-8dc809e39f17 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.173697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.173697Z digest=sha256:13fb8b09ae5e733b0466f1cb2f3d58c0d8c8dec4b86d93d52c3d9cf8d8a91cde

Observation 4f146210-a6f6-425d-98f6-66988df75d7c · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.178121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.178121Z digest=sha256:17621ea922b589e00f60c16125e4ea9a26c1f0a9d2e9ebfd1b231e68b4d34d56

Observation d114bf67-f0b6-4562-9331-e42f5946f563 · outbound

This paper cites Gpt-4v(ision) system card,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Gpt-4v(ision) system card,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.182522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.182522Z digest=sha256:b39a41d0f3b14087588399732b5bb59ed2d7355251f62ad1e1d3fd116736ed76

Observation 8c7cf713-d876-4f54-a548-41487f2a4a10 · outbound

This paper cites GPT-4 Technical Report.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs GPT-4 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.186790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.186790Z digest=sha256:23d36ff1fa1eb611f9979f67563562e50d770be25009d1dcd1a98b3411d5ddde

Observation 0eb8dd0f-d1cc-4a0c-be49-0ee2c92b8340 · outbound

This paper cites Gemini: A family of highly capable multimodal models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Gemini: A family of highly capable multimodal models,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.190925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.190925Z digest=sha256:8e301899b45a5518ae318dd9ef70d50d9c946fc6612386e5175620ba45ee7aaf

Observation a4d978bc-4cfc-4a3a-9fa0-4c829556a078 · outbound

This paper cites AI, “Bard,” 2023.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs AI, “Bard,” 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.194855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.194855Z digest=sha256:311ef1bcb24d1cb8c47c4ce77bd2b6c593730b7a20e7b0b2ef3ca72ffef39160

Observation b7fd3fdd-3840-4f90-9b56-b47225b8d540 · outbound

This paper cites Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.198969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.198969Z digest=sha256:079d32d40ba00890282e058c55c1917e078100204072cae3fd0cda90f30d7fd7

Observation 5625fcc6-37d1-415f-80bb-faa10be4830e · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs High- resolution image synthesis with latent diffusion models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.203461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.203461Z digest=sha256:ac1fb622e3e5246d35e298d42369f522c7910fd10665ee594bfbd6f8ff8d5b48

Observation a177be31-e27d-4e68-8ddf-e0b80f2dcfb6 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.207019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.207019Z digest=sha256:5c2e782dbbfdcfd63a0ea3dc9017253f9cb769e6f92f66c06ec174e1bfdcd871

Observation ed837fdc-ba9a-4819-bc70-5edd775120a1 · outbound

This paper cites Zero-shot text-to-image generation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Zero-shot text-to-image generation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.210902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.210902Z digest=sha256:751cfb70b58eb0b984546d0fd2797b567beb9484e4ba2022cd4ddd8bc1229895

Observation 1100fcc0-ad9e-4ec8-9b38-c0bf3599914a · outbound

This paper cites Imagenet: A large-scale hierarchical image database,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Imagenet: A large-scale hierarchical image database,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.214490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.214490Z digest=sha256:170142d0c43982f2ebdc9667ec1dd10ce3e5d96bedbbef1bf5eff6682bfda928

Observation e3f4e440-2293-4a9d-af2e-25122a822280 · outbound

This paper cites Microsoft coco: Common objects in context,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Microsoft coco: Common objects in context,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.218337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.218337Z digest=sha256:d00ffcd6776d6370dfa568d2ca4d4600d31dad4556d7cb6342acd105deff8962

Observation ac85e91c-d502-4609-b080-0ce4f68c9b18 · outbound

This paper cites Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.221963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.221963Z digest=sha256:da03817696fedb70e6ee7d188b84759904895fe7cf1511f4b09aebebc8a9cc46

Observation 32692a65-3f1f-48d9-a6b6-76a59f1c9991 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.225531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.225531Z digest=sha256:1b8de2efed9b5fae38fe3dd9a60da31a0d81ccf5c34b309cc78f2722c11daf02

Observation a1c67bea-2577-4d10-a065-eb1c50eb896a · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.229098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.229098Z digest=sha256:3537769e71d713fdb03bc99c0388f043633f6f681f67cb149f15453610543327

Observation 1e710829-82aa-40cb-95e4-a2246126d1d5 · outbound

This paper cites Laion-coco,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Laion-coco,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.233181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.233181Z digest=sha256:d74a8df7c7ca85e5f4c25f36dffcd93e8adf2da8b45d5942779cee414b5c84d5

Observation bab26b47-e212-4fc4-9b81-845154b749c6 · outbound

This paper cites Stanford alpaca: An instruction- following llama model,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Stanford alpaca: An instruction- following llama model,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.237161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.237161Z digest=sha256:7e735264a216603a8ec398e7d1ed495e1c0820dd61628f1a5c3edff9ec00e3e1

Observation 3aa30c57-54e4-4291-a717-b60c00693bfe · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.241278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.241278Z digest=sha256:b2f78864b5969b219f5ee4f2dd3d524e067f658a7278934ff9b99841fb3a16bd

Observation ee3a88df-54f0-4c89-9c74-a5b3a3237360 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.245674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.245674Z digest=sha256:90769a7c26d34f5f63bb5a6d8e4d669e87de220cebf10955ffc19b5f6a84b256

Observation 46bc22fa-4bd9-47a4-a302-53cf613a35be · outbound

This paper cites B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.250038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.250038Z digest=sha256:cec8838bc1f531d197f427c0ee7b0f20e97ed1f2ac4622a81ab3bf11604133bd

Observation f11ade04-d062-432e-aca9-7a1f0d77267a · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.254478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.254478Z digest=sha256:636680f886181dcce7626c06017bcc1a3281058b33bc010b5a4613a337011cbc

Observation 521af43c-1569-4d88-9ae6-c5dfa1c93189 · outbound

This paper cites Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Arondight: Red teaming large vision language models with auto-generated multi-modal jailbreak prompts,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.258909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.258909Z digest=sha256:1b631742c72e4fef6888766d7e6df0e3e2528a841b569b4e6220a308b1f822d2

Observation 41103c7b-89af-4865-9aa5-3ae0ef32b93e · outbound

This paper cites Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.263253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.263253Z digest=sha256:82114d60d6cdbdaab048660af686f50d5169cbddb7edae27f5dd80ebf6e06ff6

Observation 7ea2d64e-8301-4b41-9f70-5da442bf8d5c · outbound

This paper cites Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.268038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.268038Z digest=sha256:ded19cff0f7bd119b789d53b1aa283a4a710c13d659b45579e416a8e7fdf928c

Observation 8f76bc81-b1b7-4a23-8d96-39574b69da3a · outbound

This paper cites White-box multimodal jailbreaks against large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs White-box multimodal jailbreaks against large vision-language models,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.272280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.272280Z digest=sha256:83ed101e6e7550801052ac072d45e17ced82b3957b97ab3fa06072b992d9508d

Observation 708e9679-eccd-4a47-a55c-56ee013447f4 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Visual adversarial examples jailbreak aligned large language models,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.276363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.276363Z digest=sha256:4fa3c238daded7236ee96160822c36757250528526bff774a5708978690a22e0

Observation 678e795e-b5be-45ba-8444-16542a03c308 · outbound

This paper cites To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.280459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.280459Z digest=sha256:5dd695b26311da53305dd336a3a6dd98bfa22ce3a2b857179f32648b2953f7cf

Observation d5283b45-37d3-493e-b99f-6cd61d3b3c6c · outbound

This paper cites Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.285032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.285032Z digest=sha256:f6674e85be1759eb500c7dc689022dd0f6870b4a86e6c32b01bb6e5bbc85445c

Observation 83fb702f-4757-4424-8500-8ad9bc8d54a9 · outbound

This paper cites Are aligned neural networks adversarially aligned?.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Are aligned neural networks adversarially aligned?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.289387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.289387Z digest=sha256:08e011afc38f2407fab125b302b93ab37e829d0952c4bba43771b93a4a1d92c6

Observation e90ebd9a-795b-42d7-8e1c-cc832f82d4dc · outbound

This paper cites Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.293414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.293414Z digest=sha256:5e7d6347130670ef3c10d7899c27437759ab8ddf921a7bb9dd065a32e40d7cc7

Observation fa808e8d-d9f1-40b2-9d2c-2f32762d6ecd · outbound

This paper cites Extracting training data from large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Extracting training data from large language models,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.297509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.297509Z digest=sha256:479eb4a28db4fa6ebf524f257ffe049842a8b3f594e4df8d716b41dc6ad7a4f6

Observation 1ee4c5d8-a577-4edc-ac4d-6a518faf0807 · outbound

This paper cites Va3: Virtually assured am- plification attack on probabilistic copyright protection for text-to- image generative models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Va3: Virtually assured am- plification attack on probabilistic copyright protection for text-to- image generative models,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.301812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.301812Z digest=sha256:d598b973f33172234b6eea6a2ecc60f50d002d80c5951b7bb7f0244224218c5f

Observation 623ce292-9288-4cca-968f-ae441b468e05 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.305857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.305857Z digest=sha256:acbd593d689e148f4f00674d25426615ad203a0e26ce4f5db021e92050bdcfd8

Observation cabab675-5987-49ed-ba20-d59e8a0236ca · outbound

This paper cites Large language models are zero-shot reasoners,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Large language models are zero-shot reasoners,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.309887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.309887Z digest=sha256:29a7c8ab906efd957a3b267abb12c65d185c45d975f2881091bfe9483ae2af5a

Observation 24ac7ab2-87d9-43c5-881a-87d8a9933c22 · outbound

This paper cites Black-Box Prompt Optimization: Aligning Large Language Models without Model Training.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Black-Box Prompt Optimization: Aligning Large Language Models without Model Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.314113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.314113Z digest=sha256:effb766de1586317a9670ca8d88d21ade6c69909570f1706bf5814220f9eae4a

Observation fdf0ceed-5041-4a53-84f5-954b56bd6274 · outbound

This paper cites Training language models to follow instructions with human feedback,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Training language models to follow instructions with human feedback,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.318645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.318645Z digest=sha256:e82d34f032ddf2c96b7d6f3d6ab2f6780b15beec11d05beb4ea3f8d63661fc58

Observation aca12052-b038-4c5e-ac9e-45d7fa2bc64b · outbound

This paper cites RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.323020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.323020Z digest=sha256:1f400f3823cc91af4f07ca62b9d89530c088b9d64a6559187c360b291dac2c88

Observation 3020da8a-d357-4371-8332-36ae6a9b8ec9 · outbound

This paper cites Dress: Instructing large vision-language models to align and interact with humans via natural language feedback,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Dress: Instructing large vision-language models to align and interact with humans via natural language feedback,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.327520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.327520Z digest=sha256:6c40786495c62a72bd976ba069d8a43a7213157f9135869f0ec796b8f0d22df0

Observation cc4414d5-f3bf-409e-84a4-11a15b4ce7c6 · outbound

This paper cites MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.331678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.331678Z digest=sha256:2128df928fb4cbe3258ccce9914c38185252a30254775175fac2fef0e2f8c13f

Observation 5fa57646-fe9f-4fef-ae7e-95bd036355e5 · outbound

This paper cites Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.335773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.335773Z digest=sha256:eb01add6c84b565ed62c69984a46685375c97ffbec3a76ea076f48b9c15e9695

Observation 1f6b0708-eff5-43d3-b9f1-d5d481deebad · outbound

This paper cites Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Eyes closed, safety on: Protecting multimodal llms via image-to-text transformation,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.340113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.340113Z digest=sha256:32fc71fbd41c394f20f43af560012243f01675b35008e22dd89f4b977f4b293b

Observation 99c2107b-9d2b-4ff9-b4ff-020b4345d0ef · outbound

This paper cites JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.343991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.343991Z digest=sha256:e0d393d03592233452dd41b7e8527a6c46f8882d0db233aaa568e86fd9cc21d2

Observation fb549679-a2df-41f3-ba76-183bff6136c9 · outbound

This paper cites AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.348048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.348048Z digest=sha256:426bc44a5cbd9eedcbe3e5ffe058e91a4dbb1422a116935af7e5b9fbecf34a6c

Observation 347259e9-cd97-49d3-8e8d-41f127fdb9cb · outbound

This paper cites Adversarial illusions in multi-modal embeddings,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial illusions in multi-modal embeddings,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.352864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.352864Z digest=sha256:0538e3ee19b56e221509b03eb2929d5e7cca802262b60791806b7c8fd993240a

Observation e21fc647-f769-4fea-99a0-637a2626a322 · outbound

This paper cites Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.357496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.357496Z digest=sha256:822e4f6bebd41f26720a0578c3efee73f5d6d0d11e764d6ba59eb2e68bdbe7e0

Observation 6cd72d6f-3c28-4821-a83a-82c1c71a52e9 · outbound

This paper cites Badclip: Trigger- aware prompt learning for backdoor attacks on clip,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Badclip: Trigger- aware prompt learning for backdoor attacks on clip,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.361725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.361725Z digest=sha256:97ed21eda90bed44bb11ac9084a76371b87943c0a2407fae705e23ff7ba30997

Observation d5a2d880-df5d-4819-a307-813218e66dd1 · outbound

This paper cites Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.365929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.365929Z digest=sha256:a73140ed154bf5842c55df0414663bc2f55020d9cf9c0bcbe975e5f1a7430997

Observation e43542ff-2dbe-4099-94ca-8c6b8de0e925 · outbound

This paper cites An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.370563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.370563Z digest=sha256:ad7719332c24f12c75d90f0df5626fde1723979627dfdceefb9649408adb2c12

Observation b6cfe0d7-418c-4597-842c-c21e74afd408 · outbound

This paper cites Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Break the visual perception: Adversarial attacks targeting encoded visual tokens of large vision-language models,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.374971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.374971Z digest=sha256:e4fa2e5ddb92c4cacc839341c612b0de13b0dfea8943a52de85aa298728feaab

Observation 715f514c-6c7c-48bf-9423-830233b485fc · outbound

This paper cites Misusing Tools in Large Language Models With Visual Adversarial Examples.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Misusing Tools in Large Language Models With Visual Adversarial Examples

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.379148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.379148Z digest=sha256:1828777ba5c8231c30aaa0ce39ac55c44ab79c236563a1eacaec6357fd63ee82

Observation da0d21db-ba63-41b2-967b-01947986e400 · outbound

This paper cites Prompt-driven contrastive learning for transferable adversarial attacks,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Prompt-driven contrastive learning for transferable adversarial attacks,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.383721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.383721Z digest=sha256:cef26d3bfc3423e68ad0a2ae79bf84c3b58978667b583bd2dfeab63be8a56dbe

Observation 395a51a3-4556-4e29-8cee-74623948076a · outbound

This paper cites Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.387834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.387834Z digest=sha256:cd2b6ae1ec8295e661f2b7b965f1b36811738c3bec8ad8bdb329661fa6ed101d

Observation 24d757c0-3786-4b74-83e2-fbd6c3eb2232 · outbound

This paper cites Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Nicgslowdown: Evaluating the efficiency robustness of neural image caption generation models,

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.392810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.392810Z digest=sha256:a9dbe302074c868f215d1c9b545e61681d04214adc2a2af71af110537f5b036b

Observation 5cc07c5e-937b-4b1a-8724-f3c748fb577e · outbound

This paper cites Slowlidar: Increasing the latency of lidar-based detection using adversarial exam- ples,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Slowlidar: Increasing the latency of lidar-based detection using adversarial exam- ples,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.397009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.397009Z digest=sha256:5c22470edaf70b3e777772ab63e33ae1ec4958bc7685c46cc3f91faec573a103

Observation 3cc0603d-1d78-488a-b509-d0acdf803e9a · outbound

This paper cites The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs The dark side of dynamic routing neural networks: Towards efficiency backdoor injection,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.401241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.401241Z digest=sha256:1fc6edc0060693ef9d910c58f865ec4f2640ae51149b540268e5e5c88033b3c8

Observation 27d46af8-06b9-41d9-8eca-990a446d7cb5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Chain-of-thought prompting elicits reasoning in large language models,

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.405556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.405556Z digest=sha256:acd1ecc0aaa9f71ed0097964a95230a53837af6e151afcf65daa10fb900dd828

Observation 04bf9728-5647-4b68-8153-ba90901a86f0 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.409426Z digest=sha256:1e83ac44c26175e4ec0d669ca5cd9f800124be0831ce1cc9fc46a86e69a45468

Observation 842c88c7-17ed-4f21-8dbb-e531802b598b · outbound

This paper cites Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Stop Reasoning! When Multimodal LLM with Chain-of-Thought Reasoning Meets Adversarial Image

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.413950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.413950Z digest=sha256:a3ce8b7652a8a9b99ac4e51ce7f2170f820fd7a4783ecff8dca3f6402fdce52f

Observation d3cb49b6-d9e9-4162-bbb9-7fc7ad644f0e · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Explaining and Harnessing Adversarial Examples

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.418146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.418146Z digest=sha256:b745cb59f1b972fd591603e42a7f324747225ac4bda6b076d33188bb444029bf

Observation 5e13c29d-db2d-4fc5-b128-4ad787e5a1ba · outbound

This paper cites Adversarial examples in the physical world.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial examples in the physical world

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.422466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.422466Z digest=sha256:1de8d5fc402f23a5cc284f5df474b7f9577980d82b7b82960cea22ce9f15a8c0

Observation 63a5e170-3e13-4c20-9864-c3fefbd01bd3 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Towards deep learning models resistant to adversarial attacks,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.427191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.427191Z digest=sha256:c5abeba7989aea62f3308d61aafd45cf0af55f5f17fa4bacf24e22f8f4b1e047

Observation 96a42497-ffce-4860-ae27-4461d962c1b0 · outbound

This paper cites Boosting adversarial attacks with momentum,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Boosting adversarial attacks with momentum,

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.431794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.431794Z digest=sha256:a88981d78a5aa0a84e46a2763b6c6fc903ee8a4eb1dbcff5748df4039fad4459

Observation 7e134122-f8e5-49f0-8480-91fcb4c8d387 · outbound

This paper cites On the adversarial robustness of multi- modal foundation models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs On the adversarial robustness of multi- modal foundation models,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.866336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:38:17.436070Z digest=sha256:0f5e4fac5a44188099c169a92cb747a25be03df5deee0e7c5146b715b0205515

Observation 796a9483-6eea-4aed-93bf-0e36a26ff584 · outbound

This paper cites Transferable multimodal attack on vision-language pre- training models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Transferable multimodal attack on vision-language pre- training models,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.852715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:38:17.440408Z digest=sha256:429f4ee0e5af6c6c6401fa68c3b80b45191e689900d694f2f6a2530518b7e5de

Observation 74fc4860-c25d-4864-acb5-a8d646622cc4 · outbound

This paper cites Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Set-level guidance attack: Boosting adversarial transferability of vision-language pre-training models,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:38:18.838566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:38:17.444773Z digest=sha256:52f972bfb4af0688c180439c144c199e1f1b229b355ba33cbcad8dd500b20ce3

Observation f730f568-73ba-451a-927c-a27f96afe440 · outbound

This paper cites Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.449204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.449204Z digest=sha256:6eb75fcc00d449bfe3db1807eb54cf1c4b02a273f19fc0acce18ddbf93814a8e

Observation 28158ef0-3b69-49a8-b5e6-5c1c8098beb6 · outbound

This paper cites Adversarial Patch.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Adversarial Patch

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.453227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.453227Z digest=sha256:a8e44bc20984d203ab5baee55e0c6ff47ed941315c8dea14b26f901e580c24f5

Pith citing papers

No inbound Pith citation observations are available.