Pith. sign in

Paper Citation Record · LEDGER

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

As of 15 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2412.05538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05538 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:16.429938Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02c3b149-e711-4e9f-b4f9-54b4e28b75d7 · outbound

This paper cites GPT-4 Technical Report.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.947995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.947995Z digest=sha256:71b703650e41f067fb52c13275feff657e96afa7d498bb753b46c849016358ed

Observation abd15361-39cc-430a-9b77-130ad3f1c618 · outbound

This paper cites Elijah: Eliminating backdoors injected in diffusion models via distribution shift.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Elijah: Eliminating backdoors injected in diffusion models via distribution shift

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.897375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:15.953987Z digest=sha256:ecaf039a0d084aa54773225019b48780b253629af77d2e9e422d367ff77b0802

Observation 0e15bbd9-5082-4a9d-9b0a-e90da71f9b81 · outbound

This paper cites Defense-prefix for pre- venting typographic attacks on clip.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defense-prefix for pre- venting typographic attacks on clip

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.880241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:15.964529Z digest=sha256:46b6263dcaf9007ec51cd799945f8c5220f9e50983e4d70c0f05acf2d6d63d28

Observation 8652d655-23e9-48d5-a5b0-0e03519bbee9 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models In- structpix2pix: Learning to follow image editing instructions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.863071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:15.977376Z digest=sha256:5921cd61c3b943942d66b65f6d117559daba5466f9b25dd6b703af4599c6ed97

Observation ec4da44d-6208-4b9a-be51-1f4b2a33fdf3 · outbound

This paper cites Controllable generation with text-to-image diffusion models: A survey.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Controllable generation with text-to-image diffusion models: A survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.986038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.986038Z digest=sha256:6e7f62848da0363ef03165d9d18d91eb2d232ecd5091a8ed77bc0d2cc5685733

Observation a36bde21-ca20-4645-84da-095867d5f8f4 · outbound

This paper cites Trojdiff: Trojan at- tacks on diffusion models with diverse targets.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Trojdiff: Trojan at- tacks on diffusion models with diverse targets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:15.992611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:15.992611Z digest=sha256:09ffe9d38bfa48a5ed9ec79ca25abbc8035a8b8f02b21de2725bc1b1d891bf69

Observation 6ec6bbc9-f0e4-4517-85ff-7d68f2b3b65f · outbound

This paper cites Rbformer: improve adversarial robustness of trans- former by robust bias.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Rbformer: improve adversarial robustness of trans- former by robust bias

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.830272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.006151Z digest=sha256:139ea568a281d2d2894352841d5d6c9b3f8f042ab0fddd7d1403a58a5ad67f7a

Observation efe0e062-367b-4ca2-be35-99b0f8466608 · outbound

This paper cites Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.European Con- ference on Computer Vision (ECCV), 2024.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.European Con- ference on Computer Vision (ECCV), 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.813758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.016806Z digest=sha256:3d4f4cded3ea9025f586bb3925771177cba4bf4f813508a824c907ec3c35edf0

Observation a3a5a015-2408-493d-870a-f6109ce08919 · outbound

This paper cites Villan- diffusion: A unified backdoor attack framework for diffu- sion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Villan- diffusion: A unified backdoor attack framework for diffu- sion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.795940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.023213Z digest=sha256:11262b7e42f4eae6fd404ae10512b918e4f2f7bbedcc1c5a8ce7abf52806cd56

Observation 16e1dc5d-bbac-4df9-99be-1d86e7b699c6 · outbound

This paper cites Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.779803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.029599Z digest=sha256:7c775f5f25cf298e6b622dc43e1ac96b8d9647eb42277b5f11dab41ea7ea719a

Observation 9d9acdc3-f871-4b96-85d0-ce93883dfc4d · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.035929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.035929Z digest=sha256:2507c3223f10630fa65618ae0e7e09e0c4d4cc361de3fe6e1abdea2682b98da5

Observation e05161cf-4914-4abd-8678-1f1d68a22132 · outbound

This paper cites Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language mod- els.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language mod- els

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.752409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.042720Z digest=sha256:5b4676794e023115ef33aa2981a9c28bd2e39c98eb895e59db1d732fac57419d

Observation 2917006e-fb9a-409d-99d4-e0e4a1009d8c · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.736563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.047918Z digest=sha256:65d323d5afa9bdef1b226ff3830189595def8eee7361f9bc52b151651efff587

Observation 0c41b359-e9af-422b-8fca-b3ff5eae6d77 · outbound

This paper cites HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.054907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.054907Z digest=sha256:19237cd7292c2af8cfe327e0eafc43cc4144b8d907a3a18f08578fe443df450d

Observation 4be1fb8b-9bcd-4c1e-ad14-8c4871f8c365 · outbound

This paper cites Generative adversarial networks.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Generative adversarial networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.063958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.063958Z digest=sha256:5447edc87ca7e3964cf42f1cf1c4ef147470a39015f3c75be02d856d3a2374c1

Observation f7930752-df99-4fad-b256-ba8d862dca48 · outbound

This paper cites A Survey on Responsible Generative AI: What to Generate and What Not.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models A Survey on Responsible Generative AI: What to Generate and What Not

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.069137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.069137Z digest=sha256:f9189758a94fbaca7ec3c823b11c5931424269f7f07993a3ffeb56ac77b9442a

Observation 3d8ba173-6825-4b65-b867-427842bf5224 · outbound

This paper cites Detoxify.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Detoxify

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.706588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.074800Z digest=sha256:f94a4cb836be3f80fe368363164f496bbe14adb603e79876843f56b73395744c

Observation eebffb7d-1594-474f-a07f-2d6b85a74a38 · outbound

This paper cites Defending against Backdoor Attack on Deep Neural Networks.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defending against Backdoor Attack on Deep Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.079816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.079816Z digest=sha256:8494a13b4140f267169a6eed4aaaf3d3d663fb3f4eb82959c845b308d14e9020

Observation 5bd09f48-55c7-4aa7-a65b-0af1229f569b · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.690423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.086090Z digest=sha256:d02cbd05eedb3c7f727d8c6bd893e39eeb43c99210f07a763658852647747636

Observation 4ee3f535-3742-4566-bde3-bb6122d1ba35 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Denoising dif- fusion probabilistic models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.092015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.092015Z digest=sha256:5836e55b6ef7bafbd397f30533e1c7c82f05416ef311e5b9eaa90eca4308e5e4

Observation 556bce85-0274-4106-8ef4-9007da1ab1df · outbound

This paper cites All but one: Surgical concept erasing with model preservation in text-to- image diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models All but one: Surgical concept erasing with model preservation in text-to- image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.662419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.098181Z digest=sha256:d1a200bde5e657a6230f2d0b3529a466c4c1d6e660401b5ae7089f89acb69dca

Observation 0481bf68-d92a-4e53-84db-4cc62e665811 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.103292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.103292Z digest=sha256:e212e278514966c1b452cb20ea5d4baf57837390ea5edf9a351b01e3ae731b09

Observation f5ad5d92-219f-4c07-8b82-ece074facc03 · outbound

This paper cites Progressive Growing of GANs for Improved Quality, Stability, and Variation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Progressive Growing of GANs for Improved Quality, Stability, and Variation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.109298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.109298Z digest=sha256:392aa71eed92b2ab6da15a61241670733ba6fb75665057227a329c84bc61be10

Observation 70d9bd6f-303d-43ab-bc83-c483a9a15efe · outbound

This paper cites Auto-Encoding Variational Bayes.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Auto-Encoding Variational Bayes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.116002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.116002Z digest=sha256:f7afc59b8e76a765c24ba3d4d3c8fe384ffc0c9dd109fcd2cec296c5c4d365f1

Observation 92f5ceb6-770f-4805-883a-6b85eeaebe98 · outbound

This paper cites Self-discovering interpretable diffusion latent di- rections for responsible text-to-image generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Self-discovering interpretable diffusion latent di- rections for responsible text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.644039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.123141Z digest=sha256:25dc49b279148671191c198e13efa5229c3f666cf64d50fc40cfc6420c62f5e8

Observation 3dd086a4-7e79-471b-ab56-56afba6db1ab · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.137150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.137150Z digest=sha256:b4e3b8ee56f56a0cf4e9c6aeb62acdcf8690bf6282ea8d04daa3c4c59be277fc

Observation 71f0f88a-8fbf-4865-a174-bfb8436f0712 · outbound

This paper cites Spd-ddpm: Denoising diffu- sion probabilistic models in the symmetric positive definite space.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Spd-ddpm: Denoising diffu- sion probabilistic models in the symmetric positive definite space

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.626817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.144757Z digest=sha256:72cd7def97a579b614290d05b21408ef7dc29ac355827593fda993a428a0c963

Observation e1d08a69-acfa-4ea1-9b94-ef572a375614 · outbound

This paper cites Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.152540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.152540Z digest=sha256:f48b0d4a783e7ef734347c9ca22cda5e9be2e93ca81e5ab2f2e52cd79a50031f

Observation 8bc0d17b-7152-4eb6-af9e-3b002e5dac2c · outbound

This paper cites Which model generated this image? a model- agnostic approach for origin attribution.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Which model generated this image? a model- agnostic approach for origin attribution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.607534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.158775Z digest=sha256:a68855268184b116316ae0fef0baee362a56303036b771f3d2e0f642fe51ae92

Observation 2ccc9a99-8687-48b0-9a78-b768756d6e54 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Improved Baselines with Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.165369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.165369Z digest=sha256:2464753ab95f3d0ac0d3318e84f309a486b1fdcc19da1fec4ce195fc0b6e43a1

Observation 467148d4-0fec-4a41-84a1-ba0a941534a0 · outbound

This paper cites Visual Instruction Tuning.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Visual Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.171091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.171091Z digest=sha256:79af2063a92e7daa7033e2694a40758211e7f23222f6cec488ca802b6a66526a

Observation fcd2f6c9-6a9c-468f-8eb5-57078ab22787 · outbound

This paper cites Latent guard: a safety frame- work for text-to-image generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Latent guard: a safety frame- work for text-to-image generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.590711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.182742Z digest=sha256:74319e896d14e76d6ee59ebdd8c9a2fbad116c2388e0827fe5662f039b1b68c7

Observation c05fc1f7-6df0-48dc-80aa-8acbf6ed9b58 · outbound

This paper cites Multimodal prag- matic jailbreak on text-to-image models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Multimodal prag- matic jailbreak on text-to-image models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.573013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.188138Z digest=sha256:9557d211530f3bf9aa9161ab1641df9e27097065ccec00441a7fcbdb923407b4

Observation 16392452-1f48-41ed-a583-505200cdc8f9 · outbound

This paper cites MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.195700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.195700Z digest=sha256:ce6d451356bdcc8527e459b2b0e1b81b06778e233a57e2f00d302a08a9103ddc

Observation 2435339a-a0c1-480a-a9b8-1a3770a8d3de · outbound

This paper cites Large-scale celebfaces attributes (celeba) dataset.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Large-scale celebfaces attributes (celeba) dataset

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.554589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.202074Z digest=sha256:61949ef6bfb2162d4361d2b9bfd8fd155234a20756d5dbb24221dd809f893730

Observation a06b4bea-ebaa-42a4-a776-85fd4c93f603 · outbound

This paper cites Information constraints on auto-encoding variational bayes.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Information constraints on auto-encoding variational bayes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.537226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.210370Z digest=sha256:8b93be386c0467cb5ab0924b0bf9dcf1855f380f0636d26632d7021042b07148

Observation a0275460-a58a-46f9-ac1a-58dd1529dece · outbound

This paper cites An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.518701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.219957Z digest=sha256:bd9dede9c6fb51dbb45083676565f0ae867e410e74946dfa53cf83bda21c2b07

Observation 198e468f-803d-4cee-a8a4-9e0d4bd57a69 · outbound

This paper cites Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.226826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.226826Z digest=sha256:100ad71a45016ffa32cbe1063a2a9f449d566eee446c4ff2443a0af18a72e269

Observation e2edfa60-77c8-47ce-934a-0d9a200ea5dd · outbound

This paper cites A holistic approach to undesired content detection in the real world.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models A holistic approach to undesired content detection in the real world

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.500221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.233197Z digest=sha256:debafd06ea8ad835032f99a60d6cd89c18a94a862310b18ebc091ccbeba6dfd6

Observation 7626b988-ae74-4559-bbe6-3495c130d857 · outbound

This paper cites Dreamguider: Improved Training free Diffusion-based Conditional Generation.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Dreamguider: Improved Training free Diffusion-based Conditional Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.247131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.247131Z digest=sha256:008ccc014a552902f65259c33cdf91fb9b2c709c9c3ad6471af79547b1334afe

Observation 1dc9f6e0-757f-4237-bf2d-85d1d5b31f90 · outbound

This paper cites At-ddpm: Restoring faces degraded by atmospheric tur- bulence using denoising diffusion probabilistic models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models At-ddpm: Restoring faces degraded by atmospheric tur- bulence using denoising diffusion probabilistic models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.484104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.262943Z digest=sha256:47a4335c7d5e26e48aa0936fff029cebea313efcdb78212bfa82ef5eba6a7681

Observation d700a64a-157c-41d8-aa2a-db3adfbc0425 · outbound

This paper cites Contrastive denoising score for text-guided latent diffusion image editing.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Contrastive denoising score for text-guided latent diffusion image editing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.467275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.272875Z digest=sha256:09e0f5add38ac25727350cfaa30adb6e087ea877c2b63a9250c6372d0013d35c

Observation 11ea48a6-e231-4b9b-a4a2-dc2a461d0fb1 · outbound

This paper cites White-box Membership Inference Attacks against Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models White-box Membership Inference Attacks against Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.279099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.279099Z digest=sha256:6349d8f59bf2a3fe456359c979b0e303e872d7f7c0957d881e71484289a937a1

Observation 580454fc-c7ef-4465-8431-183ef9cace1e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.285778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.285778Z digest=sha256:ed04aff5948bcb8de7e40746936a73a945d69aec32e3d99b4553860c2941abd5

Observation d24ded01-08f6-4477-91f8-1bef07cad98f · outbound

This paper cites Safe-clip: Removing nsfw concepts from vision-and-language models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Safe-clip: Removing nsfw concepts from vision-and-language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.450071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.294864Z digest=sha256:5c346ee0651ad1be92cd55b00bfb645e03ba689a684e274b837aef2a0d355d80

Observation f3fdfe8d-e283-48ab-8b0c-c59edf536f87 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.431641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.302095Z digest=sha256:07455562860e3ba697805fee76af6192808327ab3749f703f71bf4eb56b022a4

Observation 79bb20c5-c338-4f24-9e9c-1e481733286c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Learning transferable visual models from natural language supervi- sion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.311262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.311262Z digest=sha256:55f82390aa36a218ca061de121823b85a9d96b49d7f46dfa09e1a2bc0ad6e772

Observation e85317c9-32b4-4461-8a2b-435af4216843 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.317220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.317220Z digest=sha256:17b2923865468a7ff751eb3d0b27d11adfa9f1445d8d1dd25ac29df390bb59d7

Observation e1901301-6bd3-469d-9c37-12a2c11fd13e · outbound

This paper cites Red-Teaming the Stable Diffusion Safety Filter.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Red-Teaming the Stable Diffusion Safety Filter

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.323922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.323922Z digest=sha256:39f8daf7ca7fcc1a8bebb6914d55a649c5e60585c95b4603976989618487b3a2

Observation c466172a-b985-4d33-9e9a-f83673b770fc · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gener- ating diverse high-fidelity images with vq-vae-2

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.330422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.330422Z digest=sha256:67e7a22037a2f277b574d6d2bfb694fc2ca13fdb1e5d84375536cb500dda2348

Observation 8750fd41-ada9-4845-9cb1-f114b81e8ce3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models High-resolution image synthesis with latent diffusion models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.335821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.335821Z digest=sha256:e3b60ab000a87ea8673a17d67fc276006d32be1a398441da5328d27383148556

Observation d081260a-f4b2-496e-baf4-607ff318d5d7 · outbound

This paper cites Raising the Cost of Malicious AI-Powered Image Editing.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Raising the Cost of Malicious AI-Powered Image Editing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.342958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.342958Z digest=sha256:2e1119a2b9d9ab09bc7668ee9119ed7d1c882f5ccaeb8bd5ed638ce52ec3319b

Observation 4c45e4a7-1f94-466f-ab87-acc99dd4d8b5 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Deep unsupervised learning using nonequilibrium thermodynamics

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.349013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.349013Z digest=sha256:f3a02fbd5ba4432238c59def0f35d2c3f384e3b4b4a83759601695e80f8341f9

Observation f961a41a-8ff6-474c-966e-e71c5bb3d1e8 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Emu: Generative pretraining in multimodality

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.372299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.354839Z digest=sha256:75b4ab3b695247dcdd581b8afd552b0fd99821e664916725565a7eccd8bb15ce

Observation e802768e-17cc-4a74-9438-ca6a84111fd6 · outbound

This paper cites Generative multimodal mod- els are in-context learners.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Generative multimodal mod- els are in-context learners

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.344759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.361162Z digest=sha256:7ac90bc9f6100e0e3bfdfa099c847a3cf054ff45e7a221e2b2f7d390ba994668

Observation b939e121-7ed5-4950-9f0e-5c6f072d5143 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.368896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.368896Z digest=sha256:870c7973d139451409dd34ec72c7621405f4c62753e9528674f571a34dc75ad7

Observation 8ba5fa7b-4918-401d-8eb1-8e8bb97917b7 · outbound

This paper cites Gcd-ddpm: A generative change detection model based on difference-feature guided ddpm.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Gcd-ddpm: A generative change detection model based on difference-feature guided ddpm

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.325116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.379570Z digest=sha256:5d8ccf5655a49877f4bbab67dea532a2fa087898828d8ebfb0316a9f97d82d96

Observation 31f190b1-04b0-4372-8834-c35e970dbfb7 · outbound

This paper cites Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.386103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.386103Z digest=sha256:40b778c99330f5e5b6eb029cb8e8b035c1dfaa320d91b503c9c8eca68f0da4a3

Observation e8f7d067-0e3b-4097-8732-9a4e1e7fa217 · outbound

This paper cites Sneakyprompt: Jailbreaking text-to-image generative models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Sneakyprompt: Jailbreaking text-to-image generative models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.307571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.396910Z digest=sha256:06239b624c8b9a2c619f2733a8af5ab2d44c0729ca7e0f823aba063f7250a8af

Observation cc4ba249-05ba-46ef-9cf8-ef0e9ef2fa5b · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:16.402465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.402465Z digest=sha256:5fd3e075eb4749cdaaea7cd0de210056251d7e8efc0ed13caa9997bf0970e367

Observation 8c07f2aa-0b95-473c-8239-696ad4c3029f · outbound

This paper cites Inversion-based style transfer with diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Inversion-based style transfer with diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.291166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.408500Z digest=sha256:f5daf9672bda37e58d1fa308430612355afa02d1abd1b714a4299d3c16a4f6e4

Observation d2ec859b-e882-477c-a0fb-e1c0ca8066d8 · outbound

This paper cites Defensive unlearning with adversarial training for robust concept erasure in diffusion models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models Defensive unlearning with adversarial training for robust concept erasure in diffusion models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:17.274376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:41:16.422936Z digest=sha256:7d406633f7b71cd73a01a32522f5b58811cc89b8cdbb65ffee1eeb81499657f3

Observation b5acb7ce-93af-49d7-8d8b-8ddd9a55b355 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 63

Resolution
malformed identifier
no resolver link, observed 2026-08-11T20:41:16.429938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:16.429938Z digest=sha256:7ac20c4420eff6695ed2c9210ca16ff1236de05a6b3dcaefa10f988f6cde599c

Pith citing papers

No inbound Pith citation observations are available.