Pith. sign in

Paper Citation Record · LEDGER

JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2312.10766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10766 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:45:18.551103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T18:57:16.788391Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 84898307-2d6a-460b-9b0e-844b61e48c63 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.478385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:c131ec20c12c107f06c2d06df59643c67935f56faa361286a2de7e0a26982bcf

Observation 16d0d785-3c51-4686-beff-511e41aa0e5f · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.065394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:2eaa82116444d6593d8b4a3668e7b64d319c191dd2b236f02d265e965aa56a0b

Observation df94a690-ddda-4a53-80ec-fb328d755a0d · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.551103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.551103Z digest=sha256:d87502ab3142aca0fb5a16907219a9870f9d9ada618a4df90aa2837addaf5935

Observation d2c92e86-d9b9-42da-a6d9-68ff65260dcd · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.890497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:32a3c785984d468bc2eba558165e06940ed6ade918ff6a32b488c2bd93d52e15

Observation ddb5be6a-e6e3-426e-8d88-ce0dd430be2c · inbound

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models cites this paper.

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:24.937849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:24.937849Z digest=sha256:b2b0260065efc696d61d6c1cfdef7af5d2de288fb2176ca88a45276e2fed333d

Observation 12287ed5-b3a7-4724-a9f0-f073f7520f66 · inbound

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders cites this paper.

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.349854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.349854Z digest=sha256:1f110725498df23c1efbbc85100946de7d36f358798d86896d6d39c65af94da6

Observation ff35eece-cd72-4c90-8265-2c4068223693 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:53.745955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:53.745955Z digest=sha256:be3a25ee488a74102776d9d6f3fcbc9a492392eca7ae6f275e5ff73d71c5f1a9

Observation 6d76b6df-9bab-4d18-b895-567e9462008f · inbound

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking cites this paper.

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:37:01.100658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:36:24.013477Z digest=sha256:721985311dc52e5b8a2390f61b3124badef3680e83964e0c41ebd23aa7831114

Observation 78f37935-3f4c-4ef1-a4f3-9a2a0695e737 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.671389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.671389Z digest=sha256:6dc63276169922d57f03df944b64a13c27f206aff4ebc5dabfa2261ad4b54c4d

Observation 5958d44c-33f8-4679-820f-2af6eaebb139 · inbound

Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors cites this paper.

Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T13:28:09.757042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:28:09.757042Z digest=sha256:31f740e8c7670fc6a7b2b45047cc187702fe53a7b2d6dd9d1e8c7d0288d1d46d

Observation 82d3e7b4-c6d6-4e96-8b15-bd50fd26f621 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.307952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.307952Z digest=sha256:ce9b2430156763d1d238d535752adafa6366b3a73c351d29a9ad9c09eaaef658

Observation b29b5132-8deb-4996-9b96-274fe6c9a157 · inbound

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute cites this paper.

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.357656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T20:12:03.715605Z digest=sha256:e76081fef768c1b945819750d6b291feaf31989bdfa2e195622f94d96efc5778

Observation ef969e13-3f7d-458f-b67f-603b00687965 · inbound

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models cites this paper.

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:57:16.789828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T21:47:22.295896Z digest=sha256:c45b5221b658bb82807b353ee1142bc7f38b7d0bfdfec85a105636c4c3306890

Observation 97814ace-f5d1-403f-b93e-e31e5bdf6161 · inbound

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs cites this paper.

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T17:33:49.041193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:33:49.041193Z digest=sha256:6a54b76e6647c2a69086eb026c56cdfb3e7e31372f71e26add1da00232fadb1b

Observation e072f608-1759-4290-9c5b-f3c538d2635e · inbound

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination cites this paper.

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.648221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:41.648221Z digest=sha256:b41c6bb2ad3f3f751adadb23e9617c8b580d245b6ccb7f092a8f82617667c7e9