Pith. sign in

Paper Citation Record · LEDGER

Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2402.12336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.12336 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:58:53.816120Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:59:07.063832Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d90ae248-beb3-4b99-83b0-7ce5358d1cfe · inbound

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models cites this paper.

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:58:53.816120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:58:53.816120Z digest=sha256:af2f6290aee2ac178e5da671741e2b74cfa2d0eba5842d93ae724f5bb1a16081

Observation d2f5efb4-3dbf-442c-b739-0e378715e711 · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.480616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.480616Z digest=sha256:2f9b19304a5e7b32c6edcc731853ab7b0081065dd87f2f9d20fe2f751df39331

Observation 0b0d825d-4486-4c91-b147-d572b13c2511 · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 148

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:19.969645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:19.969645Z digest=sha256:46b9ae1cd11af3cd9532a28bdbb19b7f61a3a0292a52efbd4c5be77e323add70

Observation 60a5a2dd-ed63-4bae-94e8-2a741ef83965 · inbound

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment cites this paper.

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.555198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:31.555198Z digest=sha256:ff31d9c1ec5306e9d2d8bb571fa3f35ce5ae01e8941c5985ec6ad5301314d9f6

Observation ebf584df-d99f-4edd-a28a-dc9ee716d36c · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:14.409174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:14.409174Z digest=sha256:bd1f0ef0b1cde360634f9474a218b349ce35ab1d76f640274d2fcb81a9fd960d

Observation 7bf6b250-1b9c-40ed-92e4-0b3daabdcbca · inbound

Diffusion-based Cumulative Adversarial Purification for Vision Language Models cites this paper.

Diffusion-based Cumulative Adversarial Purification for Vision Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:59:19.962834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:59:19.962834Z digest=sha256:42ad2d567a56dd9621d746348bbae91a60af6baad19a905de5dee63004836f3e

Observation 0f6951c7-d207-492e-a4ee-d99974500a89 · inbound

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents cites this paper.

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:45.196277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:45.196277Z digest=sha256:b89e72b4206e113f40081fd8d392eb4a5481d1bab32bcf2f1a479c2e19731461

Observation 40c52470-0a62-4685-8bf2-293a3fe76fa5 · inbound

Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models cites this paper.

Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:59.332470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:59.332470Z digest=sha256:3f20704ac6a77f9edab61f13330d38184e78b6717050dbf6fcbd49ef8119b323

Observation 44628e35-95c1-46cb-a172-88e77b633620 · inbound

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding cites this paper.

Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T11:54:43.232843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:54:43.232843Z digest=sha256:f01e2b5f565076e63c75d48af7319f34c60886d4d4badc80d0df9bacb4e3c71a

Observation 42f5acdd-0d3a-4206-9f1f-8a3e676722b5 · inbound

Beyond the Textual: Generating Coherent Visual Options for MCQs cites this paper.

Beyond the Textual: Generating Coherent Visual Options for MCQs Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:37.853389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:17:37.853389Z digest=sha256:0707db80582e2db27ee91c54e9b3baeea6ea50bd7794d0658dacd3255fa4ccc9

Observation 9d741b0b-7ea8-47c8-b0a1-15ab48bef964 · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.866893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:23:13.318310Z digest=sha256:1ce4e30fc539da6588ec566b1476922858b349e8f47ed1ae4a2c02add21a095d

Observation 1f6c570e-8d54-4f01-bab1-e703eb2c4c22 · inbound

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models cites this paper.

ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:30:39.404154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T21:28:59.898029Z digest=sha256:447a3429acd90a5b33899b2a5353d9004d41d271f35d739bb36988a37b236b39

Observation 79c270ba-1df5-424b-9319-3061f87632e6 · inbound

Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting cites this paper.

Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:50.528290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:50.528290Z digest=sha256:7e5c15052360b81c60ebbc4ec9933d559a9d51670e4ae0a4417c549c91cca751

Observation 07b7e299-dd12-43fa-9994-0c7e0a87f8b5 · inbound

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings cites this paper.

Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T04:19:00.649555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:15:38.632039Z digest=sha256:75b27c62c0318a571a50b767f932e9e97c2094e48f9a4d8fc0cba028f7a9c885

Observation 47352615-a4e4-4a93-941d-2173c1eb5178 · inbound

Pay Less Attention to Function Words for Free Robustness of Vision-Language Models cites this paper.

Pay Less Attention to Function Words for Free Robustness of Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:33:44.684493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:33:13.977360Z digest=sha256:616f4457d47baaeb8cc5aa62c514ba8a0f85f5c767bbc591486abe3d24e1d7d7

Observation 18c41af6-dc16-4dfd-9daf-0e3ec428f094 · inbound

Hierarchically Robust Zero-shot Vision-language Models cites this paper.

Hierarchically Robust Zero-shot Vision-language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.436342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T04:25:24.008261Z digest=sha256:e8530cec313c7ec1cdf7892e174552f0e7f4caf73d53d889ba62d12741aafbb7

Observation 2de6a4d2-d14f-4c8e-a78b-2f7dec54061f · inbound

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models cites this paper.

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:56:08.096327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T14:24:48.999632Z digest=sha256:260b0962569119d021d8098693fd12b40edd02f8c10a4c9a2e58124870fe745a

Observation bf82e486-a7ba-410e-829d-9b0784ee6381 · inbound

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers cites this paper.

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:26:30.625340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:52:25.317140Z digest=sha256:46fc8c02fe1397591110dbd77b4c873c0fff832e37e6184e77cab19ea9f2912c

Observation 8fad4a87-bd7c-45d5-90b0-8083bf4bd21b · inbound

AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models cites this paper.

AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:38:56.014305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T19:37:09.983024Z digest=sha256:ed7816e6426112a8586d289b3a940f9f1a3586e67081e55fccee5c0331c4459e

Observation e2c317cd-cc9f-4d88-a85a-0070af298475 · inbound

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models cites this paper.

Closed-Loop Bidirectional Prompting for Adversarial Robustness of Vision Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:44:01.424531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:40:26.803098Z digest=sha256:c3155056c7213d5c283928919d707c1a4efffdb7faee0ba40cc42f070b25414e

Observation 88d00781-61dc-4169-be5e-9483b0047586 · inbound

Investigating Adversarial Robustness of Multi-modal Large Language Models cites this paper.

Investigating Adversarial Robustness of Multi-modal Large Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:27.643485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:11:34.152223Z digest=sha256:fe7bf4e2f8c9e47af1f4932ebb324b726d09089a89713fd7f6c8d3cf2f7285ac

Observation 8c1a1912-7c5a-4c20-9366-be2a7de70aca · inbound

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models cites this paper.

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.844167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:04:30.654255Z digest=sha256:786741be904e4562eada69dc98f207dcf5161398ac912dba995ce1e34061a490

Observation 15bcbbc8-c1d1-4e41-b3bb-1cf236b83bef · inbound

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models cites this paper.

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:16:34.811575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T10:10:38.261635Z digest=sha256:e71e2081652bfe37f756c366e3827c30135dcb6e7804308a3ae9bd539e3707f4

Observation 89952cc8-2263-4c1c-b85d-c46a9375ba8c · inbound

Semantic Robustness Certification for Vision-Language Models cites this paper.

Semantic Robustness Certification for Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:07.065828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T21:32:54.522220Z digest=sha256:175da83839e9cb417992c73175c3ea22d2f28ccef8e8fe0f653b879c619ace01

Observation 538d2546-01bd-4239-a8ac-34713d4da38b · inbound

Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness cites this paper.

Rethinking Brain Decoding with CLIP: The Role of Adversarial Robustness Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T04:24:39.660564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:24:39.660564Z digest=sha256:714a10c7463b79a3d9464071174c5743f6806d8264f9a257891ac19c5e94a965

Observation 4ac2d30b-6440-46da-a1cd-4dd36c0c6ef1 · inbound

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning cites this paper.

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T01:16:40.492300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:16:40.492300Z digest=sha256:511ee70b8d657344121dab4a3695fdec6426e22602d840cf4b202b66fceab737

Observation 45b3fb0f-1089-4e95-9caf-ba4113ac01b8 · inbound

Unifying Adversarially Robust Model Experts in Vision-Language Models cites this paper.

Unifying Adversarially Robust Model Experts in Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T23:31:48.797271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:31:48.797271Z digest=sha256:8c143de04b70a402717c05a6c194aac77e55a5140473a4e11595f6bddd99cfa2

Observation 8da3932e-745e-4220-96c4-9f02a0a37b5f · inbound

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models cites this paper.

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:18.042445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:51:18.042445Z digest=sha256:984ba72fcb0a53151b536be9b66e6ec626f412cc04d9f37d5150ee40435d21fd