Pith. sign in

Paper Citation Record · LEDGER

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2412.00114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00114 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:46.283268Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:24:49.439929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:28:10.399987Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959f4d25-889c-4b7e-8222-165a1f45a24e · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.512263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:44.876264Z digest=sha256:bc69b24c11d3d97d4e02f2173ce5ef1199f3d3f907245d4a69f28abfe18ee211

Observation 42e7d51d-2a0b-47b5-a97f-57815e83a9dc · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.941580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.941580Z digest=sha256:3ca02f71ae43734abfc124fce73a32795a9a6d8414a2246a1a4c9d5b074ab17b

Observation 93cd3123-85b5-4ec7-aff2-0fd7fc026c59 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Learning transferable visual models from natural language supervi- sion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.946224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.946224Z digest=sha256:9b3b791650799cb4ddf56fa97c260987ccc2685853f58d9c88e8be8753701d50

Observation 1c5de27f-7e44-42aa-8475-5cb864f2e4c0 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.950726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.950726Z digest=sha256:d9826346752b4f6bc2a487dae26a0c4f4fcc5d793d3a806b68551c2711928fa7

Observation dd7e1a67-32cb-4e9e-8dac-39c719987036 · outbound

This paper cites Visual instruction tuning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.954832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.954832Z digest=sha256:e08827fe6a17df94f8231bc87cfc92b176ad00410dd2564540ab170e0e452acc

Observation 09c59ec9-9fc4-4daf-b530-db04b9f84557 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.958598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.958598Z digest=sha256:f39b8a4dc6af7746efe4e174d2f2b33108a9581bcff3c749284ecddc6a80080a

Observation ce35016b-867c-45e8-851c-c68f624124e4 · outbound

This paper cites Irad: implicit representation-driven image resampling against adversarial attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Irad: implicit representation-driven image resampling against adversarial attacks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.481016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:44.962914Z digest=sha256:4d80a2c52f54330153b332c7f4b3b34735e16227dac751735b0d23e93a86afcb

Observation 2d383b11-c351-4d92-ba72-cd6e3a1d9365 · outbound

This paper cites Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.468554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:44.966374Z digest=sha256:a90561dff0e0e1b7e3f0adc5f08920186f888b3a52d6b7b31faef760b0a1113a

Observation b78f2d46-0d51-47ef-9122-d9ed6bb4c198 · outbound

This paper cites On the Robustness of Segment Anything.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the Robustness of Segment Anything

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.970401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.970401Z digest=sha256:4286af1ba30e12f6e771a207978f89c264f32e573403519f86b9065e8b10b9f1

Observation 3a4e8087-764a-4fde-ab6b-7e2d6d3dade8 · outbound

This paper cites Adversarial relighting against face recognition.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial relighting against face recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.391424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.013666Z digest=sha256:4784adaf8090017427972e5ecd4d1f9dfc38a8d3dbdb6ad0efa7a32cbc3c8603

Observation 145a7d42-1fff-401a-a238-e83ccaf636cf · outbound

This paper cites MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:46.475722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.061056Z digest=sha256:9c5ec2fa481bd7ddef216895e530be14f8d3ad6fa62785d63bb0fd4673d77cfe

Observation 3216e681-342c-400c-96e0-d4848e02de32 · outbound

This paper cites ALA: Naturalness-aware Adversarial Lightness Attack.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments ALA: Naturalness-aware Adversarial Lightness Attack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.066082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.066082Z digest=sha256:01deed1bce85b4af642103ed7929f2bbaee5f3c54f938540713058a164e67d16

Observation ac7130ca-b754-4af9-a1ad-cb6de616d5e7 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On evaluating adversarial robustness of large vision-language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.260042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.070592Z digest=sha256:10baf2d900f8eb01ae28f28205c073608e2430a85579f496cfac449a11d18243

Observation f5256f5e-fe27-4a98-bd75-772143bdbc7a · outbound

This paper cites InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.074404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.074404Z digest=sha256:4108fd43304c00751ba9f56fc578d0fd4d08d2414e98de7311e33f0604821f76

Observation cd03d5ca-c492-46a3-89e3-2166e0b99744 · outbound

This paper cites Transferable multimodal attack on vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Transferable multimodal attack on vision-language pre-training models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.248497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.081051Z digest=sha256:683e8e1d3370de1b51ff0b9725f570d66c6ad187dc48ab5040b1fe51b8742df6

Observation c0cdd27f-7577-49e8-84f5-1c4ae4132fa9 · outbound

This paper cites Towards adversarial at- tack on vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards adversarial at- tack on vision-language pre-training models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.237153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.130705Z digest=sha256:8859cd32a228413408381d5fbab4d0100593174d73f107c0322192861d2b6dbd

Observation 11faa285-7bac-4ba0-873f-ab61e1f4b131 · outbound

This paper cites Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.101647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.206041Z digest=sha256:4e4c7f6a3e60f3b06081f04bbd9db784e25c6177fdb292b40ef4224893667218

Observation fb43b57d-0e3f-43d5-a6a7-51b8d61b2136 · outbound

This paper cites Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.059640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.283076Z digest=sha256:cc177f5d0c13d7a36146be121b74098d80a2be4d78f0d50523237f51ac1dbaa4

Observation ffa708c9-3947-4b29-a55c-9cf780b3067a · outbound

This paper cites Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.287680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.287680Z digest=sha256:5eaaa4a9d01a4e1ff1518be77ac76af8810c828169cc7e954aac5eee46933ee6

Observation 474571cb-8333-43b4-815a-9125db1e13e4 · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.997162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.292824Z digest=sha256:f772ca5be1f517cd34e45578475e2be1c7fba627562f5a97252b34adfc7457fc

Observation 09c6dfb6-7f04-4e80-9e47-3c58018adf70 · outbound

This paper cites An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.296501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.296501Z digest=sha256:716d5ce5353d6886a39a7e6e0aea89d531d01ab5023a303f24df066d86592092

Observation 3c31f85b-ca59-4af3-b03c-cc86b6626988 · outbound

This paper cites On the robustness of large multimodal mod- els against image adversarial attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the robustness of large multimodal mod- els against image adversarial attacks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.950273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.300764Z digest=sha256:551f2e295cd210589615b554a077bca81db30cc8f86f0f2bfdb2db08de5d5c69

Observation e49005a2-0855-4804-b1a8-fc25f761621d · outbound

This paper cites Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.365436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.365436Z digest=sha256:756fef6d38ff4255f07d3479ea9e96e26b2758c3013b9c4f84f9379405ad751a

Observation 7977e5ac-6b8b-49aa-b551-2e6a9470ce12 · outbound

This paper cites Multimodal neurons in artificial neural networks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Multimodal neurons in artificial neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.385682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.385682Z digest=sha256:171f459e01492afb5781b8d8368cb08b2c68625cef6973ce201bd994d41f6b2b

Observation f87c4e92-7e62-42c5-b29a-d313afa8ad4a · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Blended diffusion for text-driven editing of natural images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.389653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.389653Z digest=sha256:08806232db9a1434d33c4b406ed7736010761fa22bc12f341e52db3e17d268f5

Observation f7952bc4-54b3-4cea-802d-1ac0f53be564 · outbound

This paper cites Dis- entangling visual and written concepts in clip.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Dis- entangling visual and written concepts in clip

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.838749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.393835Z digest=sha256:0415734bf4f8dbfc302583edc06629e2115689fff97531bb8c7f5f7fc9875ce5

Observation 6695e2fb-87b9-4fa1-98bf-2e5cb403c4c6 · outbound

This paper cites Patching open-vocabulary models by interpolating weights.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Patching open-vocabulary models by interpolating weights

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.397035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.397035Z digest=sha256:901d696030b5b8318fd3fb00e32b3256b6017f32ccd8c2c0ccf08f3464af662f

Observation 60625371-7448-446f-be21-67bd553096ee · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defense-prefix for pre- venting typographic attacks on clip

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.788280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.400822Z digest=sha256:d69ec3f9e86b44260ef25cda304a1041715a2e78fd530aa1484adbaa56f40c25

Observation 070e2b28-f9a3-4f9a-a4c0-fefd294b99b6 · outbound

This paper cites Defending lvlms against vision attacks through partial-perception supervision, 2024.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defending lvlms against vision attacks through partial-perception supervision, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.681126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.404438Z digest=sha256:f02a69ab5554cca4d5fab0b258819fd9a5a7f055da6807720cd8a0383c61cb3c

Observation cbe3766e-6983-4894-a68a-269518e540d8 · outbound

This paper cites Adversarial Machine Learning at Scale.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial Machine Learning at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.408168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.408168Z digest=sha256:74f3cceef8a0648839cbc4c45c4cf22dc6465b3706fb0901044393495cf5d454

Observation 4e34ba8c-1c18-42cf-af62-868a7192a2f2 · outbound

This paper cites Adver- sarial examples in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adver- sarial examples in the physical world

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.549168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.549168Z digest=sha256:35b2c5781030068ff6853c6645dc62c990627fb71873bae5b198b38898a89fe4

Observation 8c0960cc-a898-421b-ba7d-434aa136a5dd · outbound

This paper cites Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.601997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.580877Z digest=sha256:a66835eea585e145a7c25bddaeeb0f1eca151ce719d82b80919cce93789a160d

Observation acbefc5f-ff81-44d3-b807-6ad022e671fb · outbound

This paper cites Robust physical-world attacks on deep learning visual classification.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Robust physical-world attacks on deep learning visual classification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.588605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.584533Z digest=sha256:5c2cbce22a6bc7e5c5b19a57c0afc06a2a039682d40329bf17abfbfd42098e5b

Observation 2cdecedd-0da2-4719-b7f6-7a826f997784 · outbound

This paper cites Towards transferable targeted 3d adversarial attack in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards transferable targeted 3d adversarial attack in the physical world

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.398716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.588548Z digest=sha256:1a0f26ebb6719fdb920a684efec97e17ab7f6d53f1a1b11b736dc3082f2bc74a

Observation b3b4cd52-3786-4608-ae85-fbe44f906476 · outbound

This paper cites Adversarial t-shirt! evading person detectors in a physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial t-shirt! evading person detectors in a physical world

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.353693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.592280Z digest=sha256:988d7f33457e4564d90e39f43339a9e95635506d7a39cff58cf7861c2efd6042

Observation f3ceabdf-d21f-4450-b350-97bd03b9b2e9 · outbound

This paper cites Fooling thermal infrared pedestrian detectors in real world using small bulbs.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Fooling thermal infrared pedestrian detectors in real world using small bulbs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.343421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.597309Z digest=sha256:12cb87c87785484f1bb1b8843a835332cd169264724292da54816c76a50b7f5b

Observation 2b2d1cd6-6f27-4c9a-9029-48c52e810cef · outbound

This paper cites Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.330137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.636881Z digest=sha256:c5f4833f1db78668efa1563037825d0f984c433e8ec9af14107675e7ec0671d2

Observation af9c87b8-c84c-4b71-a13a-e3dc5b7207f4 · outbound

This paper cites Hotcold block: Fooling thermal infrared detectors with a novel wearable design.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Hotcold block: Fooling thermal infrared detectors with a novel wearable design

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.280549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.682499Z digest=sha256:3d589ba6297bca39f67ef3915eb01142e3b96d498986be7d4cfc30a6a86bbc62

Observation 69b89cfe-1f66-4efa-affc-6774cc2a4bf3 · outbound

This paper cites Adversarial camouflage: Hiding physical- world attacks with natural styles.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial camouflage: Hiding physical- world attacks with natural styles

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.213692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.700581Z digest=sha256:073f32307f4990beb7fe28c233ba77d95ddb13c3b3f94e867e301e69187a0373

Observation dc9c1e9a-f670-481d-9cef-c6846de4cbec · outbound

This paper cites Uni- fied adversarial patch for cross-modal attacks in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Uni- fied adversarial patch for cross-modal attacks in the physical world

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.068983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.704277Z digest=sha256:dfbeda452b0931322dcc83969d778d882852ba7ebeaf87c323886bc9aa6ca668

Observation 6b928e87-84ca-4978-ab13-29b03ccb9fdf · outbound

This paper cites Visual instruction tuning, 2023.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.057281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.708071Z digest=sha256:3cf6e8cfdefb2952c22c4c770938ac5f6ffe890725e973d6317e3ee69d43924c

Observation 1db9c836-f655-43c3-be4d-8e930646beb8 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.044858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.711818Z digest=sha256:96eee60d99baf204a8c8c9cdc1c61fb29a735d21cf20aeec42532b7d395be184

Observation 676dac00-ae12-478b-9c19-07031a83cdb6 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.715928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.715928Z digest=sha256:0eda9aa3ea6063a857102f6dc948becf114e3892622d34d01165f3da9496e1bd

Observation d3e52cf9-54d1-4c00-9fb4-86d53d62740c · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser: Diffusion models as text painters

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.030989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:45.720628Z digest=sha256:1d2da97138b15b0a819f901a82f322141797ba20459b22fbc59473b0e50d3461

Observation 1e92ab44-e181-4564-93c5-4b565c4bc00c · outbound

This paper cites LingoQA: Visual Question Answering for Autonomous Driving.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments LingoQA: Visual Question Answering for Autonomous Driving

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.724257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.724257Z digest=sha256:ba0016a33a2a480be54e736572d61fe55295c96e7c9b535dfa98cdffb89193da

Observation 8b04b758-3b03-4cd7-a2ab-a7f3b704e047 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.728189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.728189Z digest=sha256:2b6085fb8a8c1686959e867f3e1298891ee83a7ddb8105df078c0cb537ef94a6

Observation 3541e0bc-722a-44f4-9506-9c0f60ac74cd · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.868817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.868817Z digest=sha256:31f0ae20c2a31fad1693c0a68f17ea994bfabe977e2e47b0ec954f5e5714cea1

Observation 833ecf8b-6c0a-45c1-b5a0-0b8c86684a08 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.950480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.001670Z digest=sha256:c1dac44be2a19bf6a33a36b57d69b3aaa585d5956eb6b65a9496d6b8ad7623eb

Observation a796c973-ab8b-4611-b8f6-56d235886a28 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.937711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.006249Z digest=sha256:5a25283ab0184a95c938cd17b2cbfe23f5c3e6706a33c5d80d0357caafb6fd77

Observation 21916217-b062-4b09-aa23-d164e0d419e9 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.879841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.010987Z digest=sha256:d83ad524c48d2db6e6151ebc67104de46b18d47da0ce6184d3a5e08b84bd0f32

Observation 3750fd6c-f257-4ab3-917a-4af553051763 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.788254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.015327Z digest=sha256:6d01061ab1d07bba6de3aae70b0a64feef0187f8eee9c2fc7cf2da6f3238a7d1

Observation 52dd5a22-b884-479b-be75-0153a69ef098 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.776382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.119310Z digest=sha256:326a83820003609fb43661bc26dff76bcadb316d921067d64b148d20357253f1

Observation f1d58c31-a94e-4c01-adec-81ef33110274 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.720866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.176161Z digest=sha256:8287935b55c72ff230f50d3dcdf83bf1d48f8b5253ab7cafc34c032caa48ac77

Observation 44c9d7d8-c66f-45b5-a7b4-f06cc5e64c39 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.653487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.271794Z digest=sha256:9f39f7c3070239b247ed91a28c7fadec42c8ae772ba95538c71b349d247f09b2

Observation 844466fc-bf8e-45b0-b5fd-b3452e64edbb · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.642018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.275743Z digest=sha256:c122ad34845a4ef37c908552e3f6ecb53a653e0e56439bc1465309d525de73cd

Observation f80e2030-bbc6-4cf3-81fe-b1914748ed19 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.631034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.279441Z digest=sha256:0505c3652c1f3a8bc276002ea1185cfb99cc96fad18e03eb15a6ec917ecd6d71

Observation 0271bf1b-1787-46e6-8cab-e11bbb31e059 · outbound

This paper cites colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:46.561343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T10:48:46.283268Z digest=sha256:dfeaf3389f0bddd97b1ea19a72ddad9e54ef10641a8bdc2294c61ad66623fd6f

Pith citing papers

Observation 90f37915-c44a-4850-a22b-07dd97cc1639 · inbound

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents cites this paper.

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:49.439929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:24:49.439929Z digest=sha256:0979dad377831dd26266932938356c95988a135fd4d30f9f98c019c8b07390bd

Observation f6e69485-f1b8-4bfb-a5b2-c017f58bba82 · inbound

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision cites this paper.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.423820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.423820Z digest=sha256:6f2a5e3635ecd65542dc141f6a41056a0b5073c05e048622a2021f2830ec1087

Observation 906c3d0e-f8b3-4f13-a6de-6af3f225549f · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:c6a29a0b09be9ba5179fcf02af056a13600f7cdb4be81c9ee01775a951280d67

Observation 82290c1d-6857-4760-b304-ac8dc1d4b0fc · inbound

Not What You Asked For: Typographic Attacks in Household Robot Manipulation cites this paper.

Not What You Asked For: Typographic Attacks in Household Robot Manipulation SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:28:10.401596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T09:26:39.444806Z digest=sha256:42fc293de142344c887d1108cfc410033b1de2153a4a9f3ed3a22a46f3bf11c1