Pith. sign in

Paper Citation Record · LEDGER

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation

As of 22 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2502.00580.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00580 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:31:14.158845Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:18:23.845618Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:51:02.591575Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a46e19c-6fb8-45fa-9d94-1ed5cba42769 · outbound

This paper cites Pushing Boundaries or Crossing Lines? The Complex Ethics of ChatGPT Jailbreaking.SSRN Electronic Journal, 2023.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Pushing Boundaries or Crossing Lines? The Complex Ethics of ChatGPT Jailbreaking.SSRN Electronic Journal, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.503188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.075522Z digest=sha256:834813e6bda8e76a87964b5a77aad792757950df5392695b6563e93d9a13f723

Observation 936fb21b-621b-463d-88cc-c76fbc986007 · outbound

This paper cites Automatic Jailbreaking of the Text-to-Image Generative AI Systems.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Automatic Jailbreaking of the Text-to-Image Generative AI Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.079617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.079617Z digest=sha256:8db4a1dd9fbb1e9a6b6cddd4932fe807727f3cf4b9f2c490dd4f3145d27978ab

Observation 1319f387-e11a-446b-afa5-4aa9bb9c1ef5 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, andhorizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Jailbreakzoo: Survey, landscapes, andhorizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.083521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.083521Z digest=sha256:a29a1127077bb12059ae07a283437365a624206160f6df1fff117f5f85f3f35b

Observation 9286d120-4a50-4efb-b285-aeaf00e32d60 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.087069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.087069Z digest=sha256:259577793cab8255653b8ac262144d0ea88fdad41c08218dba61b4fc534b0608

Observation 61eb7d40-2757-4f15-aed8-30a48687bceb · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.091976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.091976Z digest=sha256:ed217fe129b984e9d81b62075aeb6faacddeaa6cb225cd46dd39196f0667a031

Observation 52f07ae2-6466-4927-92f7-1b219248bd14 · outbound

This paper cites Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.095655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.095655Z digest=sha256:cc7824527cc33ba7fca850ebdf5887c5733555c463941e08f6fbdc500845d6f1

Observation d1313410-f756-460b-b972-a8ea4336bf16 · outbound

This paper cites Jailbreaking Large Language Models with Symbolic Mathematics.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Jailbreaking Large Language Models with Symbolic Mathematics

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.099686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.099686Z digest=sha256:b632eaa7925df5b606e30b8f03640a902e9dd3e07bd1091ae181bde7089580a6

Observation 208c5995-eb14-4156-9167-17de298fe388 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.103240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.103240Z digest=sha256:96fb9d65da62e723615045dc1579412b83aec00e3d8becbd2977c3636fec53ae

Observation 94c39fb8-8774-4ed4-a068-6cc17fce0e1a · outbound

This paper cites Zhang et al.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Zhang et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.491451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.107167Z digest=sha256:f55489ed96d816e356adc75df809eb14580a051ca8b34cd49e56d5a4f832a1d5

Observation 8eafc1a8-8fcd-4740-80ca-e08de5517ad0 · outbound

This paper cites Li and R.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Li and R

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.480233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.110810Z digest=sha256:ae3c14b1bb6455099d9b3d3ffb746a580aad8cbbde20519307e57034ebed33cc

Observation 99cba573-d03e-4bad-a81b-663b460c6b30 · outbound

This paper cites Defending ChatGPT against jailbreak attack via self-reminders.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Defending ChatGPT against jailbreak attack via self-reminders

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.468253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.114240Z digest=sha256:8e2db912bbf3c86cd5e1cb1e3b1f0e8467c3c7d24c67f4830867289adc95e6e5

Observation e84f6c0b-5284-4b79-bec1-c8d787b250b4 · outbound

This paper cites Self-Guard: Empower the LLM to Safeguard Itself.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Self-Guard: Empower the LLM to Safeguard Itself

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.118069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.118069Z digest=sha256:795ce5625d485856f2fda57f38e8a22246c1c5f888bd58c90fba826a8f411783

Observation 3e81fb7a-eae5-43bf-b059-9faa0bdbf9bc · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.122592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.122592Z digest=sha256:873e6649c6ff4a706ecdefcc4005e3498e25ae46c57d94bb9316989545da4374

Observation fe8dacb2-5f74-4406-8f01-e7cb4bdccd0f · outbound

This paper cites Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks.International Conference on Language Resources and Evaluation, 2023.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks.International Conference on Language Resources and Evaluation, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.455148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.126856Z digest=sha256:1c9d23ae650cfa1d9aba55675b9511df17fcce8d23cae9c4cd19c98014ca344d

Observation 654d6091-2f5b-44b0-a4fa-eee35cbe27a0 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.130552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.130552Z digest=sha256:79249491d72bc170eaffb31cb52d2170ca49fbba913b84664e773834536571bd

Observation db77dbce-d16b-4bd4-8f48-30b32cc54638 · outbound

This paper cites Best-of-N Jailbreaking.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Best-of-N Jailbreaking

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.135569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.135569Z digest=sha256:0e7077ec522a471fdf9b8ed9c8a6ce84ce02d57d0b019283eb3e473f8bd1906f

Observation aa908190-0b93-4514-bcf0-da73429615ba · outbound

This paper cites Im- proving alignment and robustness with circuit breakers.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Im- proving alignment and robustness with circuit breakers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.443254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.139808Z digest=sha256:227975b5ffdc6bd3bb6023eb90ec289c9e8968f5ecff7e8c99bbf6f72f1ffc2a

Observation 981c3c2e-53a4-4360-9142-8192753486c3 · outbound

This paper cites Using gpt-eliezer against chatgpt jailbreaking, 2022.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Using gpt-eliezer against chatgpt jailbreaking, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.431110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.143480Z digest=sha256:ade6504c0ad7341124a88f10794aa2d9ae8d6658350d266ee68d1355eed74c1d

Observation 87eb2491-43c1-4e75-92e0-49a52eda2327 · outbound

This paper cites chatgpt-prompt-evaluator on aligned ai’s github, 2022.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation chatgpt-prompt-evaluator on aligned ai’s github, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.419367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.147228Z digest=sha256:e771ea305466c3fa1057d9926ddabf14161914ab9102344e825a3fe916d6f16e

Observation 699a1ea2-a4ef-4991-944b-a2e38d47c0cb · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.150815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.150815Z digest=sha256:03516e26adf95d0df4eb41000b9ef98d1e578ecaca4d8fed5fa65635e4c05205

Observation ea4b7eaf-2ebd-4752-8871-5f912d20ecd2 · outbound

This paper cites The use of confidence or fiducial limits illustrated in the case of the binomial.Biometrika, 26(4):404–413, 1934.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation The use of confidence or fiducial limits illustrated in the case of the binomial.Biometrika, 26(4):404–413, 1934

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.405865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T18:31:14.154989Z digest=sha256:8f210f2b7fefea400c9ddf7e9a7f606fdcaac8e35bae00561265e8ffe9188df9

Observation 56058db1-e979-4b2c-862e-a32b97257ba5 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation AI Control: Improving Safety Despite Intentional Subversion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.158845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.158845Z digest=sha256:f194ac499b6c7a358b0a5b058afe09af038c006085fe54203d0a3b7d695e669e

Pith citing papers

Observation 25b35c97-8634-4097-b9db-4af25988495b · inbound

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection cites this paper.

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:18:23.845618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:18:23.845618Z digest=sha256:27d23d724d2fb4de049868c09bdb7403bfac432bbc83d82fad94909a89d1344a

Observation f37cef00-b916-4826-81d8-595ceaa0cd8a · inbound

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems cites this paper.

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:02.601349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T02:54:35.336251Z digest=sha256:ade065d86d86ad35472cf13d903ec80f62f0a371dd8df3adaddfe6c154bd8dd0