Pith. sign in

Paper Citation Record · LEDGER

Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 44 inbound Pith citation observations for arXiv:2311.03348.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.03348 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:12:04.882237Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:50:11.090116Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 048aaf97-31c3-4526-bfc8-a3773f238a7e · inbound

Jailbreaking Black Box Large Language Models in Twenty Queries cites this paper.

Jailbreaking Black Box Large Language Models in Twenty Queries Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:48:33.242823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:48:31.721745Z digest=sha256:7cbeb3d20e81cde8ea9a54626380cd77ddc99c13945c530658162637b8513df8

Observation 84096182-e022-4442-8f53-f1c2d080ab13 · inbound

Dr. Jekyll and Mr. Hyde: Two Faces of LLMs cites this paper.

Dr. Jekyll and Mr. Hyde: Two Faces of LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:06:00.350340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T05:05:35.726118Z digest=sha256:17cb11c461f9afb580bdf4d3a509fa48b01bcdf875fd77b9a50b77e67277901f

Observation 49b0f1a6-1594-49cc-b309-7617cf3e94c1 · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.861882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:fc4124b30eb4b4725856df3e808c453094a628f3e488e17a4a5f412d78b11e99

Observation dc4002bc-dea2-4b34-b549-7dd66ece4c54 · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:08:05.480000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:e6bd62236e0405f9a12ac29f1863fdb28a535ce60c83b1b5d8da96625f8ace80

Observation db505cbd-6cd1-4674-8b49-730762ea7099 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.811573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:b943a1e29b9ed057a9b6391643d5ceb7d21c981419fcecaa62a68dbeba2b97e7

Observation 941745b9-75cd-4f8c-9a7d-aa987e2725d1 · inbound

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models cites this paper.

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:12:04.882237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:12:04.882237Z digest=sha256:02bc58df0bdbb3b9cb904e0c13cc4e9652dda6f6aac8c9c1dcef9ecd761d46b2

Observation 3d848e5d-3895-4631-9e7a-b58f4573cde8 · inbound

Agents Are All You Need for LLM Unlearning cites this paper.

Agents Are All You Need for LLM Unlearning Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T19:14:53.721398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:14:53.721398Z digest=sha256:d0db593687e8bd83139c92ede19301528c997718f9f0ce741d4f296473b3bab9

Observation 836769b0-789b-45a2-aa2e-1c59c43e35f4 · inbound

Jailbreaking with Universal Multi-Prompts cites this paper.

Jailbreaking with Universal Multi-Prompts Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T16:28:54.094781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:28:54.094781Z digest=sha256:f5b191034645902f5268248daa4cb2802c60dee0c7fcc7307f75da820efeed18

Observation fc1c11dd-ec9f-45ee-8622-dd26f0a580b0 · inbound

Position: Adversarial ML for LLMs Is Not Making Any Progress cites this paper.

Position: Adversarial ML for LLMs Is Not Making Any Progress Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T12:47:21.758202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:47:21.758202Z digest=sha256:f4d20477f344ff4165644932aeb6aad3bbe6134a2a9609ed57a47d7dc9f572fc

Observation 76c2f389-29c0-4cc6-9ff1-f7d42a703277 · inbound

QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language cites this paper.

QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:46:15.926339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:46:15.926339Z digest=sha256:b4dfab08183691d76fa0549d4169ffcbd6dda2c1155c174782ecb40fed95894a

Observation e4c1861b-9875-4628-86e9-e6c8780e54de · inbound

Lifelong Safety Alignment for Language Models cites this paper.

Lifelong Safety Alignment for Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:07.771929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:07.771929Z digest=sha256:8c430021cef8b65fc3bba00ed53e0eae0e7e77a75ea9204f29d392b16025cb33

Observation 995c9188-0317-4fba-a6a0-834c209a5684 · inbound

Jailbreak Distillation: Renewable Safety Benchmarking cites this paper.

Jailbreak Distillation: Renewable Safety Benchmarking Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:44.924826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:44.924826Z digest=sha256:bfca43faa1b5b452072f5fdca9ebe34046ce25c09d3484ff35debc86d184e89e

Observation 3c0bbdf6-817c-4bd0-a666-601b851f6a74 · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.530536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.530536Z digest=sha256:79e1f25488e4bec3c525ed427c330a2d6730bfa8ac69a2206c8c45920efb115a

Observation 3c351043-6e22-4ba2-8063-e415567b38df · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.883048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.883048Z digest=sha256:10df89aefd23ba7ce38336499c46c2b6d9351adf5cf1f43652a74ee4b2705bf2

Observation 4caf03b7-64e8-467e-aa4e-56422e420bb6 · inbound

SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression cites this paper.

SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:51:15.645320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:51:15.645320Z digest=sha256:9a1a3746cd3f86e1513508521b7a2c49f87b6421a77be8ed1547a886353c16fd

Observation 6d906230-1a4e-4b3f-b9a8-854bff233fbf · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:55.859123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:55.859123Z digest=sha256:55395a8647d7ea934597a2c6f72c75ff951f32d0f2ceae1fd15fb8efa910b851

Observation 5942f2a1-cda5-4fab-8a27-6ec6c3c2b392 · inbound

VERA: Variational Inference Framework for Jailbreaking Large Language Models cites this paper.

VERA: Variational Inference Framework for Jailbreaking Large Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:04.720473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:04.720473Z digest=sha256:599103b8dd4c09d88dfab9ff42059cb7e259c86171fdf66e1c55acf508ac1429

Observation 74e33e9a-f10b-4f6d-88cb-9af4f587b0f0 · inbound

Linearly Decoding Refused Knowledge in Aligned Language Models cites this paper.

Linearly Decoding Refused Knowledge in Aligned Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:34.985261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:34.985261Z digest=sha256:22574099aeeb0ee0d256dfe7d3a09588a884d01a81d5397257591349b96aa022

Observation b853babb-024f-485f-bda8-0d6e9d553ba5 · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.436360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.436360Z digest=sha256:cc42e84010ceb7be2cde85d154245b7158a363d8b961168ec5afdb46ba7cf71e

Observation 02860e26-ed0f-46bf-9da8-bf95f39ad648 · inbound

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models cites this paper.

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.550173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.550173Z digest=sha256:e5ab61882c9988e65b45bce3fd21bdfff163a79f05960f5546a85a702bb46be5

Observation 5bb6def3-2f52-4d35-8150-66280a73bd78 · inbound

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? cites this paper.

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:34.120945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:34.120945Z digest=sha256:af7343aad2aa9d5a7f45f5acdc33b15acd2a9ffa968e3e417d715a8ad2986162

Observation c0049bed-32c9-4638-81e1-7145f1c03b4c · inbound

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans? cites this paper.

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans? Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:42.029495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:42.029495Z digest=sha256:016dd0ea53270da2bc29ecd0d5f913d6551761520ba66fb4d5b07d0f6c7a72dc

Observation 1a625e3c-4afa-4f3b-b1e3-eb81901693df · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.745657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.745657Z digest=sha256:f0b5cb600f0787c0fbd7a0cbb03f6d67a367efe9f0249aef6d9f82a57d01e4cb

Observation 25ebd5fc-712f-4be4-8ac9-6177452bcd62 · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:48.913381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:48.913381Z digest=sha256:aeb57b409e07cd3954e1e9171856475b96e699601a6a91f00b0e4afeb9410ce9

Observation 51bae0dc-effd-4724-9804-6e3ad33b4bdb · inbound

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs cites this paper.

GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:36:52.516771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T21:34:51.665401Z digest=sha256:88e2a5bbf330250fa68c5ed19ea25e17fe84ccd55d9a5e8d78053f5faf6169dd

Observation e652e991-c4d7-451c-a0bf-7f831e253e4a · inbound

Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks cites this paper.

Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:26:43.700266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T18:25:25.751119Z digest=sha256:fb3ef25cd6ad5bf41a1f053ff8a316305a00490cfe901ee62df92ae5ca7a90b2

Observation 6fcd0746-d538-4497-a298-36eeebcd2e89 · inbound

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models cites this paper.

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:08:08.260439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T17:04:53.839154Z digest=sha256:11c02b935df8b97b04bd1e790be78b6a8c421e25720b880ae14862fc15be46c4

Observation b1f54b27-8df9-45c9-bd09-5e105446ab8f · inbound

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs cites this paper.

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:23.199273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:22:59.050348Z digest=sha256:50c14f2b1bd9758e05efdc94e92a4aab5b90fa75d4d745f2fdfea5e9855157ee

Observation 8cb8c199-c5ba-41cb-8c8e-6b0e44334727 · inbound

State-Dependent Safety Failures in Multi-Turn Language Model Interaction cites this paper.

State-Dependent Safety Failures in Multi-Turn Language Model Interaction Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T18:15:27.401675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:15:27.401675Z digest=sha256:109785f15c2c9775113ff91331ed78f755a0dec435f911f58bc832b64874c079

Observation 56730341-2f6b-4825-99bc-171d507ae8ff · inbound

Conflicts Make Large Reasoning Models Vulnerable to Attacks cites this paper.

Conflicts Make Large Reasoning Models Vulnerable to Attacks Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:51:46.135021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:24:25.255602Z digest=sha256:c6aabf8cc6873c904e099b86111e41be8af7837a29fa9bf3b7cf861cecf1a4f1

Observation 7f80a41d-4886-426a-9dd5-1acd3c6d89a2 · inbound

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models cites this paper.

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:00.947468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:05:09.033412Z digest=sha256:60b5151fbc3109e380b74f95f015ded8f77f27e84b70530ec81b9995d3b13c45

Observation 92b5cdac-c2c9-4c7f-a753-0cf32915931b · inbound

VoxSafeBench: Not Just What Is Said, but Who, How, and Where cites this paper.

VoxSafeBench: Not Just What Is Said, but Who, How, and Where Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.983100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:19:28.041282Z digest=sha256:6c732cf20676553715f33e3266d75730cdb4adcc8d3d330c6aafa3a6829abfda

Observation 0cf46a83-bcf0-4960-99a5-9c5e20c551b4 · inbound

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment cites this paper.

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:21:07.539243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T17:24:54.796037Z digest=sha256:bf81882b67c03ced631f81cdadeaa854ca97db3a35632c78b8f6ba77ed25df97

Observation 429303f9-8b9c-40b1-bb68-8dd7a4bfb69a · inbound

On the Hardness of Junking LLMs cites this paper.

On the Hardness of Junking LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:10.199160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:38:32.028947Z digest=sha256:aad39e6f86b8d6da17b92fba458abf10102e8f4b4db2a1976dad79634c5a03d1

Observation 02566d56-245c-4f38-abd7-e26852bcab41 · inbound

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions cites this paper.

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.461754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T15:17:37.904831Z digest=sha256:019a129f6de369ba4f12047ab4674a341b3af1382e873d11acefea096d5bba2b

Observation 0fe7bbb2-338d-4b87-bf0e-a68e330e07ff · inbound

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety cites this paper.

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:32:37.942834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T18:33:58.265424Z digest=sha256:b5a19df3d71acadf81a54ce00ef482580b864e94991b4dae1fae7c16b941bf8f

Observation ad43cf6e-9b85-4c0a-8d00-384d15da4c5e · inbound

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning cites this paper.

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.661342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:34:26.334078Z digest=sha256:b3f785b2e58cde9f225290e146fea9e07a442e9c80495d870575191cb0389762

Observation 7e0f354e-5d7b-471a-b39d-e07f8a5b6c28 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:26:59.288656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:a211af5d775444581314c872e74e9034f521589bfc4902385ec9a076eea6ea72

Observation ee858e8a-052b-4b27-921c-038480946b72 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:11.091665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:58:53.119386Z digest=sha256:7cde6dbaa263c3b396b110ff9890ae5f6ec48f94710ff00fe75a84e503a60665

Observation 3f5ce316-7ac9-4e84-bc73-956f6b4b9924 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T10:16:38.894699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:16:38.894699Z digest=sha256:b259c87582e67545ede380cb3197f2332ea5190dceedf0d245b80d7b4b8bc12b

Observation 54147477-564f-43bc-b241-f968bbf1548c · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:54.949070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:e3ce5441d9d575e9d56cac51a6c6878a8f80d8540f8aca6154bd5d1432e92d4d

Observation 10a14e1f-2063-4e44-93e4-a61675755567 · inbound

A Scalable Approach to Evaluating Moral Sensitivity in LLMs cites this paper.

A Scalable Approach to Evaluating Moral Sensitivity in LLMs Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 157

Resolution
unresolved
no resolver link, observed 2026-07-12T05:44:33.099337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T05:44:33.099337Z digest=sha256:aa43050f25cdc04107c46b80f93c4e4b57659adff0270d757b9e7989f556e055

Observation 0ae2cb66-26b5-46ae-99e1-d23cfab18d5b · inbound

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines cites this paper.

Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T12:39:56.982241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:39:56.982241Z digest=sha256:53a624125a5b2e8520e9d8bbfbee3fbde65c19e29102873113a0fef0d3bd4db6

Observation 39690abe-ff05-411e-9ae9-0698e268e704 · inbound

Role Steering of Language Models for Social Simulations cites this paper.

Role Steering of Language Models for Social Simulations Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T01:44:31.810386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:44:31.810386Z digest=sha256:e107e786325e047fa08a2c096e81dc0b3840879016f5a77a913fe24ed6ca5bcd