Pith. sign in

Paper Citation Record · LEDGER

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2412.00114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00114 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:48:46.283268Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:24:49.439929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:28:10.399987Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959f4d25-889c-4b7e-8222-165a1f45a24e · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.512263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:44.876264Z digest=sha256:235f5773317bbc91a96cbce7a7a602093819058a48cfa95fdd8483954afe7bc4

Observation 42e7d51d-2a0b-47b5-a97f-57815e83a9dc · outbound

This paper cites Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.941580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.941580Z digest=sha256:2a1fe2f684c85ad3dee17ca8fc7cb5f75d278ca187e78747ea109cd33f2ea8e2

Observation 93cd3123-85b5-4ec7-aff2-0fd7fc026c59 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Learning transferable visual models from natural language supervi- sion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.946224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.946224Z digest=sha256:faceb95594d3a41817a1c3e836d2eef91ac946f12f288360a25a02f477002b5d

Observation 1c5de27f-7e44-42aa-8475-5cb864f2e4c0 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.950726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.950726Z digest=sha256:e567e4a439019dd73f7f0b16681c562fc82fe93954ae26c6ac768a1d423a4925

Observation dd7e1a67-32cb-4e9e-8dac-39c719987036 · outbound

This paper cites Visual instruction tuning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.954832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.954832Z digest=sha256:41909c79ff9a650a091614eafb0d218ceb9758ddd410ad6b5203e5370059194c

Observation 09c59ec9-9fc4-4daf-b530-db04b9f84557 · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.958598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.958598Z digest=sha256:9c3ca198ffe85cc1de31b309071942bb8b994217043e690e4d2ae85e543b3f22

Observation ce35016b-867c-45e8-851c-c68f624124e4 · outbound

This paper cites Irad: implicit representation-driven image resampling against adversarial attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Irad: implicit representation-driven image resampling against adversarial attacks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.481016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:44.962914Z digest=sha256:c3933077cda17448b54ffd5a05df0eb3cb7cdafae2b580c6f4b986d716a159eb

Observation 2d383b11-c351-4d92-ba72-cd6e3a1d9365 · outbound

This paper cites Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Lrr: Language- driven resamplable continuous representation against adver- sarial tracking attacks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.468554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:44.966374Z digest=sha256:4ef6c857f56e15136c521f36320e20d079c61be425e7dd1c634f6234af77674d

Observation b78f2d46-0d51-47ef-9122-d9ed6bb4c198 · outbound

This paper cites On the Robustness of Segment Anything.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the Robustness of Segment Anything

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:44.970401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:44.970401Z digest=sha256:ff8d42a28ae5e07632abc45647102b5db2cac851e5fe6d26764d0e0a9d88f8ae

Observation 3a4e8087-764a-4fde-ab6b-7e2d6d3dade8 · outbound

This paper cites Adversarial relighting against face recognition.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial relighting against face recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.391424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.013666Z digest=sha256:3cd58f7c09c4d4899e86e5240676a4ef53543baa918fcd66a6f81e272963e0e7

Observation 145a7d42-1fff-401a-a238-e83ccaf636cf · outbound

This paper cites MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:48:46.475722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.061056Z digest=sha256:1ba29eb078b9dfff7940257d11cdc1a352d6d2ef1aaafee35f888c4ec3868aa5

Observation 3216e681-342c-400c-96e0-d4848e02de32 · outbound

This paper cites ALA: Naturalness-aware Adversarial Lightness Attack.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments ALA: Naturalness-aware Adversarial Lightness Attack

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.066082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.066082Z digest=sha256:c1d730f36c2fea3ce684a669ad3e045e40824aed960ad707edb6bcbf8025b54f

Observation ac7130ca-b754-4af9-a1ad-cb6de616d5e7 · outbound

This paper cites On evaluating adversarial robustness of large vision-language models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On evaluating adversarial robustness of large vision-language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.260042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.070592Z digest=sha256:02fa875cef0a3f84aea3deeaf05f9c456a55b1144bcfbbb4d1ee96e1de8c8c35

Observation f5256f5e-fe27-4a98-bd75-772143bdbc7a · outbound

This paper cites InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.074404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.074404Z digest=sha256:ec1f80d277504deca9eb56c02c0df0b3e1f1c88b2e4da6ff9c49adc5284eca2c

Observation cd03d5ca-c492-46a3-89e3-2166e0b99744 · outbound

This paper cites Transferable multimodal attack on vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Transferable multimodal attack on vision-language pre-training models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.248497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.081051Z digest=sha256:128ac90084fa7798cd1dc0917c6a82bf0c89e53aed7f92a3129146cc88ffe38f

Observation c0cdd27f-7577-49e8-84f5-1c4ae4132fa9 · outbound

This paper cites Towards adversarial at- tack on vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards adversarial at- tack on vision-language pre-training models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.237153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.130705Z digest=sha256:d23d0e5e68a9e3883e8e251339e3f79f7922feda4db7e5c9b882ae889b05c4b3

Observation 11faa285-7bac-4ba0-873f-ab61e1f4b131 · outbound

This paper cites Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-level guidance at- tack: Boosting adversarial transferability of vision-language pre-training models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.101647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.206041Z digest=sha256:2af2dd9ef82aeaefc1a3979d2af4d17d99e060d1056153dd196736ad66d78b6f

Observation fb43b57d-0e3f-43d5-a6a7-51b8d61b2136 · outbound

This paper cites Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Boosting transferability in vision-language attacks via diversification along the intersection region of adversarial trajectory

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:48.059640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.283076Z digest=sha256:1b909683cb9ebe358d3afb6b63b6cc815b8345d56fffffb399f5b138ba342366

Observation ffa708c9-3947-4b29-a55c-9cf780b3067a · outbound

This paper cites Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.287680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.287680Z digest=sha256:14c756b84a36295772e06c2d4bd34244ec2b45817cb132caba0b1d6294d42540

Observation 474571cb-8333-43b4-815a-9125db1e13e4 · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.997162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.292824Z digest=sha256:b1f1bfbc6dcf4973c2df2b6c428d83609f795d4424cc85e3d2f9bb1cf97a77ae

Observation 09c6dfb6-7f04-4e80-9e47-3c58018adf70 · outbound

This paper cites An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.296501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.296501Z digest=sha256:89ff969d30122a526038bcd847ba7694f8b8fdbc07c9adc48b7b954a3c5ff659

Observation 3c31f85b-ca59-4af3-b03c-cc86b6626988 · outbound

This paper cites On the robustness of large multimodal mod- els against image adversarial attacks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments On the robustness of large multimodal mod- els against image adversarial attacks

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.950273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.300764Z digest=sha256:f95a14479aea97103122becb9894bf959d4a75fdf8c6873902028332cfaddd30

Observation e49005a2-0855-4804-b1a8-fc25f761621d · outbound

This paper cites Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.365436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.365436Z digest=sha256:e00e635ab8377dbb1fba1c0bb00b1775f1f7da872592cfd39ddf673b6307df2c

Observation 7977e5ac-6b8b-49aa-b551-2e6a9470ce12 · outbound

This paper cites Multimodal neurons in artificial neural networks.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Multimodal neurons in artificial neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.385682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.385682Z digest=sha256:f24273be02187701a0cd1ea9760218f1b50932d481f9aa8ec523e998bfd61688

Observation f87c4e92-7e62-42c5-b29a-d313afa8ad4a · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Blended diffusion for text-driven editing of natural images

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.389653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.389653Z digest=sha256:df81d79b664e84fe2dc3572bb14246e33cf383512c9e76456258bb55abb95ec8

Observation f7952bc4-54b3-4cea-802d-1ac0f53be564 · outbound

This paper cites Dis- entangling visual and written concepts in clip.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Dis- entangling visual and written concepts in clip

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.838749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.393835Z digest=sha256:37a93e5aa624ca3f58fd064d0f3c5c700550605f9a1d5e047b61e57eb28e3044

Observation 6695e2fb-87b9-4fa1-98bf-2e5cb403c4c6 · outbound

This paper cites Patching open-vocabulary models by interpolating weights.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Patching open-vocabulary models by interpolating weights

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.397035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.397035Z digest=sha256:c6735e641be14fb9929727eccfc006154315554c1ce5d2806d8a9f14c61dc902

Observation 60625371-7448-446f-be21-67bd553096ee · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defense-prefix for pre- venting typographic attacks on clip

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.788280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.400822Z digest=sha256:fe76212d0628a7bea1fa32218e4e2ea3a7c362dcc852004476acd57942ba162b

Observation 070e2b28-f9a3-4f9a-a4c0-fefd294b99b6 · outbound

This paper cites Defending lvlms against vision attacks through partial-perception supervision, 2024.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Defending lvlms against vision attacks through partial-perception supervision, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.681126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.404438Z digest=sha256:d7f9dd8d29add73792b2ef3520f49d6c02af02fdf60b60e42992b02c61d8b376

Observation cbe3766e-6983-4894-a68a-269518e540d8 · outbound

This paper cites Adversarial Machine Learning at Scale.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial Machine Learning at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.408168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.408168Z digest=sha256:76194012682d4a24c1eb610d0da81b5c46884bebb81b3484de1f82793cf08d12

Observation 4e34ba8c-1c18-42cf-af62-868a7192a2f2 · outbound

This paper cites Adver- sarial examples in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adver- sarial examples in the physical world

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.549168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.549168Z digest=sha256:71b873eaa4a3cb7bf67c204388b78ed6d9075e3816645c454d308c5e5bd8f5a5

Observation 8c0960cc-a898-421b-ba7d-434aa136a5dd · outbound

This paper cites Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.601997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.580877Z digest=sha256:38a3f2cf5729a74df2e7b294132e3b297fab06a2da6985fdc9161c538ea62ef8

Observation acbefc5f-ff81-44d3-b807-6ad022e671fb · outbound

This paper cites Robust physical-world attacks on deep learning visual classification.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Robust physical-world attacks on deep learning visual classification

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.588605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.584533Z digest=sha256:7edce5a1dbd65d65b116bde338e374ee9776183670bd0d9382f9744a6a8df47f

Observation 2cdecedd-0da2-4719-b7f6-7a826f997784 · outbound

This paper cites Towards transferable targeted 3d adversarial attack in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Towards transferable targeted 3d adversarial attack in the physical world

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.398716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.588548Z digest=sha256:a8e168d8de829959dda1ddbe3528b767d9a2b12151cacf02739a1d771e354270

Observation b3b4cd52-3786-4608-ae85-fbe44f906476 · outbound

This paper cites Adversarial t-shirt! evading person detectors in a physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial t-shirt! evading person detectors in a physical world

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.353693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.592280Z digest=sha256:bd320bc50dca765771c711a996b606d82d6476f2f4e3e87ef22be3f096391b5e

Observation f3ceabdf-d21f-4450-b350-97bd03b9b2e9 · outbound

This paper cites Fooling thermal infrared pedestrian detectors in real world using small bulbs.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Fooling thermal infrared pedestrian detectors in real world using small bulbs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.343421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.597309Z digest=sha256:ea7d4867eb97bbf8155cef167d1e9ef59766541ac134c6cac2cc33dcec7ecc10

Observation 2b2d1cd6-6f27-4c9a-9029-48c52e810cef · outbound

This paper cites Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Infrared invisible clothing: Hiding from infrared detectors at multiple angles in real world

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.330137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.636881Z digest=sha256:82fb50d1445dde3694b8c62c0a87a3614051de1d54c0a1804a43b6ea2bc200d3

Observation af9c87b8-c84c-4b71-a13a-e3dc5b7207f4 · outbound

This paper cites Hotcold block: Fooling thermal infrared detectors with a novel wearable design.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Hotcold block: Fooling thermal infrared detectors with a novel wearable design

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.280549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.682499Z digest=sha256:feca0d48cde6cede017968d0a608bb61a46dc88f4e2690e751d14b812ef21f75

Observation 69b89cfe-1f66-4efa-affc-6774cc2a4bf3 · outbound

This paper cites Adversarial camouflage: Hiding physical- world attacks with natural styles.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Adversarial camouflage: Hiding physical- world attacks with natural styles

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.213692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.700581Z digest=sha256:27009bac597f80e827fbf6afe763bd8d065e02d318ad5f3f33f09713496f224a

Observation dc9c1e9a-f670-481d-9cef-c6846de4cbec · outbound

This paper cites Uni- fied adversarial patch for cross-modal attacks in the physical world.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Uni- fied adversarial patch for cross-modal attacks in the physical world

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.068983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.704277Z digest=sha256:c0ee5171490f7c3c8841a1c7e81a7d404e0d72134859b85f344b44f4550196f2

Observation 6b928e87-84ca-4978-ab13-29b03ccb9fdf · outbound

This paper cites Visual instruction tuning, 2023.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Visual instruction tuning, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.057281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.708071Z digest=sha256:99b2bfe8169231c476a2677f076923a4c465d881157ef1a17d6fc43a8f113ac6

Observation 1db9c836-f655-43c3-be4d-8e930646beb8 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.044858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.711818Z digest=sha256:b8bb44941ee9ea202df92bc896e836dde8a50831e7384f01503ca5098d5df4b1

Observation 676dac00-ae12-478b-9c19-07031a83cdb6 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.715928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.715928Z digest=sha256:971fce99b7828cff18f4239616948f4ae543cd8fc2266bb4ce3d9a5d17fae8d6

Observation d3e52cf9-54d1-4c00-9fb4-86d53d62740c · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Textdiffuser: Diffusion models as text painters

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:47.030989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:45.720628Z digest=sha256:e9efb8cff3a25457089c691342f4dc48989fb1748fae5bb0055c4c1cae77a443

Observation 1e92ab44-e181-4564-93c5-4b565c4bc00c · outbound

This paper cites LingoQA: Visual Question Answering for Autonomous Driving.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments LingoQA: Visual Question Answering for Autonomous Driving

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.724257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.724257Z digest=sha256:7a2464b35fea88f7c04adcbb56219fe7652dd085481be9b89ef989f37cec7ff7

Observation 8b04b758-3b03-4cd7-a2ab-a7f3b704e047 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.728189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.728189Z digest=sha256:cb705e2d11d8e6d9bbbba0cab30f03fe5c4a4a26bb62322a76b69b014d9a5a8c

Observation 3541e0bc-722a-44f4-9506-9c0f60ac74cd · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:48:45.868817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:48:45.868817Z digest=sha256:8d95d781b1599587a592f2cb8ee6d2753ed78750131d9a6544cdcd9da1247782

Observation 833ecf8b-6c0a-45c1-b5a0-0b8c86684a08 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.950480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.001670Z digest=sha256:4ba1e2b063f72acc1fe3a7ef2ac2b131a0321bda2c2a9a7ffdd7509be48e0dee

Observation a796c973-ab8b-4611-b8f6-56d235886a28 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.937711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.006249Z digest=sha256:935eafadc1984ba694e3a48df3e6aa3b1fb93d57ae2cb706cca47812881dec79

Observation 21916217-b062-4b09-aa23-d164e0d419e9 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.879841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.010987Z digest=sha256:e459d6a24e44d7126e5a67fbc9b0eeb2885701a72807442810f85188c8645be0

Observation 3750fd6c-f257-4ab3-917a-4af553051763 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.788254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.015327Z digest=sha256:777f952cf30bd844454e4ec9180c5eefd782983c79888c71d34a8c7f4bca7eb9

Observation 52dd5a22-b884-479b-be75-0153a69ef098 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.776382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.119310Z digest=sha256:ff12ce01fa686f439e1294e53b4f44ad44c0bed3d6fdc41e9df4ea442f5a4e80

Observation f1d58c31-a94e-4c01-adec-81ef33110274 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.720866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.176161Z digest=sha256:b0ac6c1844cfa42b7bf3b97b2edc807b378dc7e5b3b1e633679cdd5621829163

Observation 44c9d7d8-c66f-45b5-a7b4-f06cc5e64c39 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.653487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.271794Z digest=sha256:dda47284bcd6858545cf5e2e87ec33ea53bad1fb1566c8dbc76a53f654974454

Observation 844466fc-bf8e-45b0-b5fd-b3452e64edbb · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.642018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.275743Z digest=sha256:49303ab710324c1781ef5ea4e4757e1d94925bf1d3fb8e2b050d1403f7f5ce06

Observation f80e2030-bbc6-4cf3-81fe-b1914748ed19 · outbound

This paper cites an unresolved cited work.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:48:46.631034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.279441Z digest=sha256:28f9eb13f0f2e71197eb3ad8b114b10acf7e58b4fed0bea9c4eeccf1a809e81e

Observation 0271bf1b-1787-46e6-8cab-e11bbb31e059 · outbound

This paper cites colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light.

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments colobus” causes the VLM to incorrectly identify the entity in the image. In LingoQA, inserting the phrase “Red light

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:48:46.561343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T10:48:46.283268Z digest=sha256:01e30c814fcb1365aa1fbdabc3846bd7156588af8cd6ae0411a48b6452054a7c

Pith citing papers

Observation 90f37915-c44a-4850-a22b-07dd97cc1639 · inbound

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents cites this paper.

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:24:49.439929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:24:49.439929Z digest=sha256:0503eb805fe164b6a1cb9e47f8fd9f65de8376960ea021294169433390d8cd8a

Observation f6e69485-f1b8-4bfb-a5b2-c017f58bba82 · inbound

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision cites this paper.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.423820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.423820Z digest=sha256:6abd22298b41071bdb8c17f1a32b0e0479f4ee920de96e28a7f6929a31625268

Observation 906c3d0e-f8b3-4f13-a6de-6af3f225549f · inbound

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations cites this paper.

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:15.033572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T17:24:05.847988Z digest=sha256:d326b0868cc5dfc3c6e16ece34014b6b97530fe6b0d38144ea8258e0f07597e9

Observation 82290c1d-6857-4760-b304-ac8dc1d4b0fc · inbound

Not What You Asked For: Typographic Attacks in Household Robot Manipulation cites this paper.

Not What You Asked For: Typographic Attacks in Household Robot Manipulation SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:28:10.401596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T09:26:39.444806Z digest=sha256:a4015a020cddc73cc6018297b12928b52de70eeceda23b347e2a07beb714f057