Pith. sign in

Paper Citation Record · LEDGER

JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2410.12855.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.12855 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:35:24.532055Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 565eb3a4-3dfe-4625-8f9c-22a310fb75a3 · inbound

Bag of Tricks for Inference-time Computation of LLM Reasoning cites this paper.

Bag of Tricks for Inference-time Computation of LLM Reasoning JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:24.532055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:35:24.532055Z digest=sha256:58311ce14c676f40c90399c8d0e77a845b5ad4be8e95b0bed25012c08527eb5a

Observation 39a21f67-dc69-4ff4-9ce6-46a0f5608084 · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:06:38.110538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:2ac71f4819e5c498165cd34a37f7ce156374bf9656bf2c58f789209d5312dbce

Observation f7634ac1-9eb1-4a9e-bc49-b57e71142404 · inbound

Concealment of Intent: A Game-Theoretic Analysis cites this paper.

Concealment of Intent: A Game-Theoretic Analysis JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:38.624953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:53:38.624953Z digest=sha256:d6fe01b3dbcf11a1f9ab2678afc5fbab4230e8c2a145f2fec3e4bb3c65ab43cb

Observation 39f87ec4-b572-468a-8ca6-49826f53a9bd · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.801585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.801585Z digest=sha256:5929cc6143ae37b5842210a1a10b4d94e50aea1e1d0eef86d4aa7b71592c0c17

Observation 6223093a-3c22-47d2-82ef-f45af762f462 · inbound

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety cites this paper.

FORTRESS: Frontier Risk Evaluation for National Security and Public Safety JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:01.751718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:01.751718Z digest=sha256:30005fb6e88d915c94aecc034dc54b38ceaa828b6341bf9c2dc2538e9b4834c2

Observation ece9b557-aa9d-4cf1-824d-1872b7de1e45 · inbound

Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework cites this paper.

Triadic Fusion of Cognitive, Functional, and Causal Dimensions for Explainable LLMs: The TAXAL Framework JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:33:22.555010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:33:22.555010Z digest=sha256:4fefef04f2c338e7d8e6a7f6f8fe8c2d6371d9c4fc30a4f0ecb3c5a28aeb6428

Observation f995386e-5d23-4b4d-9c7d-da58d65381bd · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:47.497380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:47.497380Z digest=sha256:1a16f9636dce606778ae205e84585699d83b88e6740605d7c7f9114ed53b24ad

Observation 044a9cfa-ca14-46a0-b790-4e35ae7fddd3 · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.079346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:6f0afaa8f15c9faabaa9ccf8388ca207c9c00c2a34805129f820b87b414b27ed

Observation ec0abbdd-1819-4a77-92cb-0f92abe0cdf2 · inbound

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation cites this paper.

RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.206039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:56:24.757369Z digest=sha256:0c65ed442ab3089b6c69cb5e14f03b0274a6b164fa05dc4c2f32925e9653a044

Observation 4af86a28-e59e-4e7c-afa3-03f8cab20f5c · inbound

Adaptive Prompt Embedding Optimization for LLM Jailbreaking cites this paper.

Adaptive Prompt Embedding Optimization for LLM Jailbreaking JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:06:14.615189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:28:32.989179Z digest=sha256:1376071afe27ba309bc312dcc318b828506448a4e5a72b37397754fd2bece805

Observation a6899ed4-b57a-40e6-9c66-51f356d13c40 · inbound

A Theoretical Game of Attacks via Compositional Skills cites this paper.

A Theoretical Game of Attacks via Compositional Skills JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:43.559011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:07:13.287966Z digest=sha256:4f70aba201572b86767c35d4c02ffd02a038fe7a882436f77697a19b1d2d2f60

Observation c275d417-15b6-4620-b06b-d58a7b9fa940 · inbound

SoK: Robustness in Large Language Models against Jailbreak Attacks cites this paper.

SoK: Robustness in Large Language Models against Jailbreak Attacks JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:08.897517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:42:41.137808Z digest=sha256:9c0607dbebd3eec97fe735cae7d31e34b566b8b356543df10fd18de691de8a47

Observation d790ebaa-a486-4fec-a3b9-cac388296db1 · inbound

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring cites this paper.

The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:46:18.605173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:42:08.565972Z digest=sha256:f8de5b7c5714c6eb58490ab3ef667b4feac615b5c4d6d4d25c243757ea87ec69

Observation d1edbe04-c7eb-4098-8df8-d49ba457bb0e · inbound

CoT-Guard: Small Models for Strong Monitoring cites this paper.

CoT-Guard: Small Models for Strong Monitoring JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:39:23.737664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T19:39:16.264272Z digest=sha256:f45bd0da2b2cdb47115070d0ec07f4e5af062d90740b5f0d66eb1724ceb9e23e

Observation f79c2419-5225-4c37-b5d1-a1ce7771b5ec · inbound

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics cites this paper.

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.917676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T21:55:48.561400Z digest=sha256:d03b6055c440a55affa23a915b64b920af5c2426c4efd764e10621eaa26b73db

Observation 3d521efe-a601-4b9e-ac31-b727fefb18a3 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:50:11.060181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:58:53.119386Z digest=sha256:ab323ff647b7d9e3c5eccdfb9d130a0aba518af565bd7b32f60d2d7787b7ec8e

Observation 5487eb26-535a-4c4c-8155-82866cd22886 · inbound

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation cites this paper.

A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T10:16:38.410457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:16:38.410457Z digest=sha256:8d29add4979e451124738c7d6e18277b8ef0fc68ee774904f37858d54c63995f

Observation f96759ab-285c-4f88-8130-3c07ec865a93 · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T22:40:37.839133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:40:37.839133Z digest=sha256:7fd051f8796d231ad93f06d482fc07253d8942b86ecd060d14a117f2576731d1

Observation 92b849b1-c7f8-4bfe-ba88-60e431f2db42 · inbound

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents cites this paper.

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T07:01:49.222325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:01:49.222325Z digest=sha256:f04a76a63abc14ce5fd1a27835c147998f36a4b6d5c939aa38ac77b2323fc093