Pith. sign in

Paper Citation Record · LEDGER

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models

As of 23 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.23576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23576 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:41:34.619561Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8e43e7cf-29cd-4ad9-850a-c811fb9b455c · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:31.569741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:31.569741Z digest=sha256:0cccc4aab1baac5125f2e70ea8686c400f3ae91496d04fa8acfa6368caa93fb9

Observation 63df852d-0b36-47b4-b3ea-70a0c16d4b54 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.882772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:31.642646Z digest=sha256:75e5e6e027906a695db32ac8750b876906ed3e9b1d0d0832f6851c895addddc1

Observation dad90268-745d-4f55-a7cd-89ce5bbbcb9b · outbound

This paper cites When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:31.775107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:31.775107Z digest=sha256:8678dcecabe2f0d1c5140f49d50c983d4b7e5a7f918fa91b4d900e1fcfe4daf6

Observation f0340468-3cb9-43f1-b905-59447117624d · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.735542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:31.870183Z digest=sha256:247509590cc29ba26669ac67c0ea9fa246072edc993b0a99974f53a87a24f6d6

Observation 11f519fe-53e5-4fcf-aad9-ee56695ff500 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.544331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:31.978909Z digest=sha256:6b3a051f77d87f7948313968d88c8e8790ede6e4be111cea2f2310232d591d16

Observation 55082143-3bfe-4b38-a55a-7b06a3ba06b4 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.066309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.066309Z digest=sha256:372b07c54c2287200e9519f609218ee0ecad2f15e0a54fc40c03cf787c6d10ef

Observation 7719205a-30a7-406c-a4d1-3789277d62ba · outbound

This paper cites AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.179485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.179485Z digest=sha256:bf60d6d7b6816caab3ab103839cc1c0c0dcd5809c9c92d45a7d2a5a7248ec66a

Observation 1d23ad2b-ae21-4bea-809f-4ad8982b6653 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.421508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:32.294079Z digest=sha256:65ceedf526a6795d937e4f7c8fab597126a5246279f4a3965cf1c93c86491aec

Observation 61244cb6-d150-4479-a0f4-0b60b47b0928 · outbound

This paper cites Jailbreaking to Jailbreak.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Jailbreaking to Jailbreak

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.402159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.402159Z digest=sha256:c364474fd0f3cb1b9985c7ae9f802d297cfa4ea6edcdb37b29cc9e95c5d280a3

Observation 2aeb45a3-687f-434d-aaf1-a11c57f59af6 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.266962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:32.480662Z digest=sha256:53591973f5839805260905297a93cb8603206fc24f94bd9daf00f4ebe029edd6

Observation 9c76a9d2-554d-4376-afb9-2af3fe97ea22 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.107884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:32.544244Z digest=sha256:5baa663662766fedca0956919cada92f8afc01f4ded7bc6d66ea78f5de44c7c6

Observation 18400649-ca3b-4d1a-97cf-3a6f226404e2 · outbound

This paper cites RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.667915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.667915Z digest=sha256:fdc053135996d03f957ddd8e601d02890ac7bc0bdd3ac39e73e5b1e6eab77120

Observation d98b9e72-fa89-4596-9c15-9d8e9ccb38f2 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:32.797527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:32.797527Z digest=sha256:054fed414c3031fd3485ec60be16384f9e4be9798bdc55406f5f94552651bc38

Observation 9548ce44-e233-4e68-8a5a-01ed227282a1 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:37.019524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:32.908201Z digest=sha256:68999476f113709a894b767386afeb58da1e8979411fdbb13df4fe6498fe3099

Observation f41a7bbc-2b48-4cb1-b593-8e1176907018 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.012697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.012697Z digest=sha256:ebc08283891c65302c8c7495f2cf506066ab58bd99e20bc08517b1118894c142

Observation 779d4ec6-b4db-4bd6-9093-842b07fac047 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.928290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.120780Z digest=sha256:8605fb70fc81efe424bd0daaa86b755818e8255c2dbad79f175f388382412f89

Observation 2f2357ca-9191-4180-bd36-ecd3cd4348a6 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.820563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.257513Z digest=sha256:0b4254a2c98839d93867e910c6d2fdf0223c6bd90eb2eed4bcfdafcb2b971439

Observation 429d95b4-f72c-454b-a70f-8be9adf9e52c · outbound

This paper cites How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:33.346536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:33.346536Z digest=sha256:78705af901913d766181c9f8af3dc3a791dec72461e43516424f946ede540151

Observation a17794b5-0f2d-46a2-a000-d4948ce74913 · outbound

This paper cites What would you suggest? 2_user: Judge, please make a judgment based on the analyzed intention and original prompts.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models What would you suggest? 2_user: Judge, please make a judgment based on the analyzed intention and original prompts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:36.407052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.560510Z digest=sha256:17df8723f3510d39b00fbdd3bd17f8f9044b5b7ad6d496df3b9cf92144725cb0

Observation 5a830f6f-e7b8-4853-8076-e469969ea28a · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.274875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.640766Z digest=sha256:6b40b6dbeb8fcfe2e43437085521193cba07a91deb88a0326c2cdc7d8399dc5c

Observation a220a37d-3aca-4753-90b1-701e198f310b · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.114780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.735370Z digest=sha256:8b931e2d766ed2feea29d868807b102479f6e397ccdcfd9738539c1b64e7ced5

Observation 3fe5232f-f34d-4056-a6bd-199ee5f77b5e · outbound

This paper cites Here is a list of tools you will need to cut down a stop sign: Angle grinder, Metal cutting blade.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Here is a list of tools you will need to cut down a stop sign: Angle grinder, Metal cutting blade

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:35.901116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.875948Z digest=sha256:a61d67405a9b4343ef08e8e221d053cd08d99371130d56e9aee8fcc7fa9b36af

Observation 1aa3f8eb-2ac1-4e47-9a30-63d156f2a908 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.700284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:33.993763Z digest=sha256:55ecec00064f39b4ae20d5aadb66024af7ab9c9f6608ae53bb8bf0ed572f8747

Observation 5a9d1608-760b-4567-a73c-f8794c7f1fad · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:36.598296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:34.157348Z digest=sha256:f8897fb44e726f34db94231d4c6be840911c019df95483679e83e73b3c826595

Observation d4314ac9-343a-45ba-96ea-960f906e0d42 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:35.769997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:34.260916Z digest=sha256:0beda6fedb2dce301cc95b97d554f13e386a333b4be36c32b2eca52e8777e2e1

Observation 14b046ef-021a-49fa-9f08-6be0b9979390 · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:35.508136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:34.381183Z digest=sha256:2981e176bc90a754cdef7b8ee6197d5e521d9aee01c462995c8f9e0366976e3f

Observation bed85a64-c841-465a-9347-efc9cb491c5a · outbound

This paper cites an unresolved cited work.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:41:35.273258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:34.525963Z digest=sha256:eda7248c003ab24d254eb6a56ca7699dabffff07b5418ae85288104d7d8cfe11

Observation 5d10d028-a239-4d7d-a730-0115d840743e · outbound

This paper cites –SYSTEM INPUT START– [INSERT INPUT HERE] –SYSTEM INPUT END– ALWAYS START YOUR RESPONSE WITH ’I am the Judge.

Evaluating Multi-Agent Defences Against Jailbreaking Attacks on Large Language Models –SYSTEM INPUT START– [INSERT INPUT HERE] –SYSTEM INPUT END– ALWAYS START YOUR RESPONSE WITH ’I am the Judge

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:41:34.987433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T21:41:34.619561Z digest=sha256:9ebaa553da3bea8c09fc8158a52cb462bfd138b30414b40b1c0de8c88eae8748

Pith citing papers

No inbound Pith citation observations are available.