Pith. sign in

Paper Citation Record · LEDGER

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization

As of 9 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 0 inbound Pith citation observations for arXiv:2605.26501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.26501 v1

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T18:17:52.158294Z

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

5 of 5 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7176d271-9528-4a1d-af55-a273d94059b4 · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization How Robust is Google's Bard to Adversarial Image Attacks?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.907500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:17:52.158294Z digest=sha256:02577fe573cc321078ae47ee0364c373e1b6dcdf113bf27543085da5d953c98a

Observation 24ef5266-77f0-41bf-a008-7888ba977f85 · outbound

This paper cites InCVPR, 6904–6913.

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization InCVPR, 6904–6913

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T18:17:52.158294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T18:17:52.158294Z digest=sha256:eec358867027b688a39e473c4a886f89c4460ad2458ca4f960ecd95cf225f9cf

Observation 50da154a-4226-48d4-a3a9-9e4b0fb06060 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.904906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:17:52.158294Z digest=sha256:791082a6061aa056a0f576aa4f524e3b5251ee5e809309d8bfed5afa35731971

Observation e26aab9f-c653-43bd-9514-be1e631b06f8 · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:50.902580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:17:52.158294Z digest=sha256:7adef68fde2c632c05189c2607042ae3e39753676ad523e1e7ea34c1966f5821

Observation 9afc7fd5-f3b4-4eb9-a074-bc4815ce08c1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.899700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:17:52.158294Z digest=sha256:10f35718241ded1091f8c2da176a7ac9289c1f00f11492c5807ec077c1c92030

Pith citing papers

No inbound Pith citation observations are available.