Pith. sign in

Paper Citation Record · LEDGER

JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2312.10766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10766 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:53:15.076252Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T18:57:16.788391Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 84898307-2d6a-460b-9b0e-844b61e48c63 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.478385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:aefeeae3e901e1fd23122065eb713afddf41164baf8db81fab55097e579038cf

Observation c65117e9-eb84-4ed4-a311-3893bb85bceb · inbound

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey cites this paper.

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-12T20:53:15.076252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:53:15.076252Z digest=sha256:afdf2cb043e4d87492c18f81d62235fed5e29f826c420c1992d487f278c86887

Observation 2841c414-df88-4635-878f-f82f8640645c · inbound

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks cites this paper.

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:24:32.677034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:24:32.677034Z digest=sha256:4493a25ab15479cbe4800e6933e66211528bc9b1d9bc817e0fa7ff636bbe76de

Observation 16e4edef-3d02-41cd-8ff9-60685049dde4 · inbound

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment cites this paper.

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T11:03:01.289178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:03:01.289178Z digest=sha256:90258826eb4d6a28cc5c1e059453553bf1cc3293b767f3c33b7fa69317579f5f

Observation a11d28d9-92be-48aa-a712-ae8ff8efe47e · inbound

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage cites this paper.

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:25:36.223853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:25:36.223853Z digest=sha256:5daa1aeddee12cedc6786d0637c4949df2169a91ce35a25f2ce8f3adb9cf9895

Observation ab662ad0-9046-4ac6-8eff-add245f9eadc · inbound

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision cites this paper.

Defending LVLMs Against Vision Attacks through Partial-Perception Supervision JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:52:06.542812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:52:06.542812Z digest=sha256:e0b7e768a610778cc248ed2d081086b50a8a1d4f0e707d621b6f5b767d15ad89

Observation 9a986c19-45c8-4093-bc7a-6d045c19f6af · inbound

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency cites this paper.

Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:26:00.657668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:26:00.657668Z digest=sha256:7bc60bf81df33f7dc581556e393a1f2cbd1b6d36ed385baa486a6b7ef501fb85

Observation 8fe2da77-eb63-4419-866c-1629db8032a0 · inbound

Topological Signatures of Adversaries in Multimodal Alignments cites this paper.

Topological Signatures of Adversaries in Multimodal Alignments JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T04:32:58.023306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:32:58.023306Z digest=sha256:ccb8914ab036bd23c4c6e8ea936ee3d55e66a54e73a87858d3097205a5718ba7

Observation 16d0d785-3c51-4686-beff-511e41aa0e5f · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:34.065394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:45798c8d264fd05dcc0ed5df50405c87b81003d02b4ceb6015d4dcfe363fd8f5

Observation 99c2107b-9d2b-4ff9-b4ff-020b4345d0ef · inbound

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs cites this paper.

When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T15:38:17.343991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:38:17.343991Z digest=sha256:cfa6179f0123acb2c68edd9c2f098fb5acce09f54bdd96021b9ab47c6e3977b9

Observation df94a690-ddda-4a53-80ec-fb328d755a0d · inbound

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations cites this paper.

A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T19:45:18.551103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:45:18.551103Z digest=sha256:900e999f5da7264bb3b6d51fa382b9dc4c428aa9673cd45b991853881648ef8a

Observation d2c92e86-d9b9-42da-a6d9-68ff65260dcd · inbound

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion cites this paper.

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:15:14.890497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-23T00:13:08.603115Z digest=sha256:2729a1f6ec5deff8e7ebd4c810a184eb0fb6d2748131570456faac6111fe0b50

Observation ddb5be6a-e6e3-426e-8d88-ce0dd430be2c · inbound

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models cites this paper.

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:24.937849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:24.937849Z digest=sha256:f2e00d37887c65457e02f1b15cfeade5a138fc53efd0f045860981e72502316c

Observation 12287ed5-b3a7-4724-a9f0-f073f7520f66 · inbound

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders cites this paper.

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:29.349854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:29.349854Z digest=sha256:9038fdfcb2d500e98c16d54a416a8f54273baffc275b30c9455a16eec06a62a7

Observation ff35eece-cd72-4c90-8265-2c4068223693 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:53.745955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:53.745955Z digest=sha256:0507e99c8d6d9cf07ebbbd6ab7a9c3651cbae33e4c3a2424e12fe08135ee5ded

Observation 6d76b6df-9bab-4d18-b895-567e9462008f · inbound

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking cites this paper.

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:37:01.100658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-19T03:36:24.013477Z digest=sha256:0db64e08dfcb1f55f87bb87f903fe348896879a7b50158714c8150774f3647bf

Observation 78f37935-3f4c-4ef1-a4f3-9a2a0695e737 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:43.671389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:43.671389Z digest=sha256:dee26e7c8bf5960ba1474de71e4f36fd0859d6003dc49ccdb0916fc9e51759a4

Observation 5958d44c-33f8-4679-820f-2af6eaebb139 · inbound

Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors cites this paper.

Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T13:28:09.757042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:28:09.757042Z digest=sha256:4b61de90c88cd18a1f2368b6e4a5b2508796ee708bd5a04f581a0d1178bad7d2

Observation 82d3e7b4-c6d6-4e96-8b15-bd50fd26f621 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 232

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:58.307952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:58.307952Z digest=sha256:44fbffb85501d187f06bffac4ae64d0ddcbd816f699a096871452f523b083975

Observation b29b5132-8deb-4996-9b96-274fe6c9a157 · inbound

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute cites this paper.

Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:13:43.357656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-20T20:12:03.715605Z digest=sha256:96f48eccf3dab2b28168f8c4608cbf26a70d54bc481340eb6cd8e4017d9d15dc

Observation ef969e13-3f7d-458f-b67f-603b00687965 · inbound

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models cites this paper.

MLingualFC: Evaluating Jailbreak Vulnerabilities in Multilingual Vision-Language Models JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T18:57:16.789828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T21:47:22.295896Z digest=sha256:972012ce83dac04d90dc9e2daebb72a52aa6f7eb213d92d2533a4e02893317aa

Observation 97814ace-f5d1-403f-b93e-e31e5bdf6161 · inbound

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs cites this paper.

Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T17:33:49.041193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:33:49.041193Z digest=sha256:02bd7287c1adb96de45dcf88291cb0d35cf5420fc465da238968ffcdef564df3

Observation e072f608-1759-4290-9c5b-f3c538d2635e · inbound

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination cites this paper.

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:41.648221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:41.648221Z digest=sha256:237b6eea1e0c3c8e4111fd4f74806428656e0b4df99635a40827575f64bec704

Observation 7f4db920-1fcb-4cce-bbcf-8f0984e851a9 · inbound

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks cites this paper.

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:41.315199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:57:41.315199Z digest=sha256:98b7b052639a44c2a95581fbf549de8058d6055cab666515b60b5bdf56b0c6cb