Pith. sign in

Paper Citation Record · LEDGER

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2507.04673.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04673 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:46:27.712204Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbe1ae7c-e1cd-4754-b783-12ba81f019ae · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.135856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.135856Z digest=sha256:b84950f19312be83dfe31109d3bcd00676701b575c476f5849dd5c06714a9fe7

Observation 68b78950-31ea-4ae7-a11a-c446900e07a2 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.203611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.203611Z digest=sha256:16e1f755bed61a0ef245e07958e0e927224356a8fa33b207d71bdd8f7855d6fa

Observation 02c6bfce-8387-4b88-961f-51cdbc6e695a · outbound

This paper cites Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.332759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.332759Z digest=sha256:77fa129e41578fb9e9c18b7f3a652f9cad7831175a44e626d23ed5bfd63635c5

Observation 42e7b7f6-a4ff-454f-98b2-b7b0f6a236fd · outbound

This paper cites & Lowe, R.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message & Lowe, R

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:30.312037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:25.420968Z digest=sha256:c79e9c92c7b7a3d35b2a39d179042e6393d2bc09d9d7f3cd47c63ac181f551fb

Observation 39ef1738-b7ea-4156-ba56-a47d600122a4 · outbound

This paper cites Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.503493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.503493Z digest=sha256:29c40f82d7dcbdcd9cfcdb8f456a90685a421ce9076b3056e57f6e616f4fab59

Observation a68bd5e6-34fd-4093-8a56-43f078dbbba9 · outbound

This paper cites Red Teaming Language Models with Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Red Teaming Language Models with Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.612647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.612647Z digest=sha256:ed6a2fc3dc55f04c530458cbd93255b49426ec597d3a26e036a8099d379be68b

Observation 6433a5ab-e4ec-4c88-93f2-acf3754ec031 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.761176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.761176Z digest=sha256:9ccd59b3dc0b53598ed6004016da40a76adb3a07b7a8368921f25ffa3e14e0e3

Observation 3d0799c7-1f3f-4ff7-b50c-591d077226c6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:25.870469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:25.870469Z digest=sha256:90e44c75ba2dcffc9566af093af6c090c0c2ca9ff6e7a2a51cc809c0794397d7

Observation 7fd45238-5c2f-435e-842e-0f93b792de74 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:46:30.169777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:25.984097Z digest=sha256:b5f492cb9ffa5b9d54cfba3398782d8b32c2599c059b645baeb43e1dd7aa665d

Observation c72c4f5c-94a3-4041-b26b-747ee4884a48 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:46:29.946114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:26.136525Z digest=sha256:48d752fb2d34a992b72dec33abf6ab6e8325bf302eb8ab516340cd316d4c04c8

Observation 129cc2a2-71a6-4f6e-9f65-da89169c3590 · outbound

This paper cites Multimodal Pragmatic Jailbreak on Text-to-image Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Multimodal Pragmatic Jailbreak on Text-to-image Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.221727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.221727Z digest=sha256:3a6db2d6168a781fe92a1285bc3f44f9b42e34e9e8471b84a4ed1f2ec687b57a

Observation 86029344-f94f-4cc8-958d-392e967e577e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.337360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.337360Z digest=sha256:5a8cb20400e291bca2fa4bc8f09e4cbda89a0d29860a0a7a462b012e9e4c226b

Observation 5b2c57e0-11c6-4d76-903a-19dec5ccff73 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message F., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:29.747993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:26.401264Z digest=sha256:983494afb1729c14f2c73c1dbaebf0069df777dd3e90812edbdc727a595c113c

Observation 0d2a6696-a20a-4316-9611-5342157ccb23 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Jailbroken: How Does LLM Safety Training Fail?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.491295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.491295Z digest=sha256:f8f4766a41313df98c391305c5349db72cb9ef41f20da749e8c73d4d1b50b039

Observation e6418170-4cf4-49d2-b7f2-4d1e2b9b094d · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.685840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.685840Z digest=sha256:7d053b5a2b1215826b5d9495259f53f165fbb0bca93ca91164e74b5773091170

Observation 2b6982bd-78ac-4b73-aa55-315d166a41e4 · outbound

This paper cites Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:46:28.715500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:26.766578Z digest=sha256:d98b262da712fda57e2eeef356fc849a8735cca8b42d7d99b17db980a587c247

Observation e4dedf08-ce85-4eb0-92da-562fef96fbe0 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:46:29.028272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:26.858023Z digest=sha256:a568e26fd61c0cb7e9e3583a39edffe886bcf0c557c48edfd47d3ea8959f775c

Observation 5128f266-0187-4ed2-974e-3b94f5a72697 · outbound

This paper cites GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.940175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.940175Z digest=sha256:8eefc80568ed8b04f3fd37a35cf56cc331ecf67d10f63a41399e457147012c18

Observation 0a5cbd56-88a6-4c88-8df7-b6e574ac8c03 · outbound

This paper cites Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:26.999796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:26.999796Z digest=sha256:94d2f822addfcd651d2a3ebe0a193437211dd06da591d038653aaddc0f42f2d2

Observation a6b983d6-49ab-4cdd-967c-1d567a4bd97d · outbound

This paper cites HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.054948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.054948Z digest=sha256:e2ffe44ce9c256e12bc8e86749425b12aff5451e0c2bf264331e174991dd1ed6

Observation 3e668489-c586-4ce6-84b2-c3e279662485 · outbound

This paper cites SneakyPrompt: Jailbreaking Text-to-image Generative Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message SneakyPrompt: Jailbreaking Text-to-image Generative Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.138588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.138588Z digest=sha256:bc4456bc13958b4f65e26b12f66ea3005315c08f0beae9e2b2ca38649eef834c

Observation 634b7aab-52c0-4de4-91af-93bce1a07b10 · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:46:28.438341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:27.219717Z digest=sha256:e59034188eca9f7f0f389fce291b5314d5b0d20d79e446e44ce316ee1b855cb9

Observation b2067cff-6d94-47f2-8e03-bcdfb55f1ec1 · outbound

This paper cites Gradient-based Jailbreak Images for Multimodal Fusion Models.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Gradient-based Jailbreak Images for Multimodal Fusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.317130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.317130Z digest=sha256:2782bccf1af51d40dae0a5da4590cfd931a8472d1d0c6ff4042e99833c5fa556

Observation 6a6c5220-9606-4a96-be45-78998596fd50 · outbound

This paper cites Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.368912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.368912Z digest=sha256:46a946a3d2404b488b786b70f0ddffd55fb0847c36b428c49127d991b85d2007

Observation f4e24e59-c187-4a25-b386-c1458f598470 · outbound

This paper cites Red-Teaming LLM Multi-Agent Systems via Communication Attacks.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Red-Teaming LLM Multi-Agent Systems via Communication Attacks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.451948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.451948Z digest=sha256:6329ef49f68f781e19959d2020548add8021ccc15d9d5de6dfecad16c4b523ce

Observation ed12d4c6-a0e6-4378-b560-30a9911f5ae0 · outbound

This paper cites GPT-4 Technical Report.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:27.536570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:46:27.536570Z digest=sha256:7c8f8e8808b0dd4b566851ca58cb5755a78dc72e6c492e8e05bdd641fa73fd00

Observation 3e0cb174-ef09-44f7-a443-6fab71e75112 · outbound

This paper cites Consideration Set Sampling to Analyze Undecided Respondents.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Consideration Set Sampling to Analyze Undecided Respondents

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:46:27.921011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:27.623901Z digest=sha256:bd9ed416de6cac3a9a52f340b04bf977b8f678b06388cf5b0a8d5ad92fc8ebd4

Observation 0cf1581b-227e-41b4-9945-b6c7ee04bc4a · outbound

This paper cites an unresolved cited work.

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:46:29.529563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:46:27.712204Z digest=sha256:9494564053bbe851f92b70c621bbfd84c5338f46cde3f7c3d15c2902c4fc785e

Pith citing papers

No inbound Pith citation observations are available.