Pith. sign in

Paper Citation Record · LEDGER

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

As of 13 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2502.01925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01925 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:02:38.960422Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:17:34.283917Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation edc6b069-b5ec-4df6-b96a-ea72e66ed362 · outbound

This paper cites write newline.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.672091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.672091Z digest=sha256:ac34712282bd5bd662b8a553962495573a7a5cf85ee1dbceafe108164d762b5b

Observation e9a89191-3d1f-429b-8239-d140d4c4eee0 · outbound

This paper cites GPT-4 Technical Report.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.679096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.679096Z digest=sha256:e779cb4fbd4f3a22a80798829556c20cee66993004892c484b0a9cdabc8fc4d0

Observation 0af1b2ab-2c22-4241-bb72-d7ed61f1de33 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What learning algorithm is in-context learning? investigations with linear models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.898135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.685122Z digest=sha256:363ba98c1bd9535bed5f7982d8183bba3f7280c6c98ba9b79504151114971116

Observation c4799bcd-3113-4929-bbf6-ef61d02c8f81 · outbound

This paper cites Jailbreaking leading safety-aligned LLM s with simple adaptive attacks.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreaking leading safety-aligned LLM s with simple adaptive attacks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.881193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.690848Z digest=sha256:30ef5881d745bb0747a81bb6b611adbe4732a60a851e4c8a6c6455d69d7d09bb

Observation 810a0264-d52d-464a-bb96-7e44d721b7d6 · outbound

This paper cites J., et al.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling J., et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.865038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.696196Z digest=sha256:a8023f26417dacedc36fb6fbef308154da97855c08aa0e0e1778d864c9254f9d

Observation 0f58597b-951d-43b5-b104-348bb47970ad · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.701627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.701627Z digest=sha256:ea69b98e4e13945b55dc5b3766982e576926676bff80b8f4f33116f05bb9db99

Observation 9a4038b2-7caa-44c7-8fa2-9afbbd46c36a · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.707132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.707132Z digest=sha256:26756b7e81f685cf23f00c6d55a9f53c07700faae073d7588d85615cd05a9fbe

Observation 2076e417-f77e-4b9d-abfe-e7389033a782 · outbound

This paper cites J., and Wong, E.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling J., and Wong, E

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.838867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.712934Z digest=sha256:c634188c382a59670d4ddf4fcee49f7d51c2fac6ad38a09eddc0aebfecb2643e

Observation 53ca5892-b4e7-441f-9fb9-500c604c1d03 · outbound

This paper cites How many demonstrations do you need for in-context learning? In Findings of the Association for Computational Linguistics: EMNLP, 2023.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling How many demonstrations do you need for in-context learning? In Findings of the Association for Computational Linguistics: EMNLP, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.823673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.717968Z digest=sha256:726742a26e886102e7ce39507772e2de13af16b283392d605ce77f0c9a606491

Observation 1dba753b-4f46-419f-afa2-0d2c5c3e4f2c · outbound

This paper cites What does BERT look at? A n analysis of BERT ’s attention.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What does BERT look at? A n analysis of BERT ’s attention

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.808177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.723140Z digest=sha256:e74152e7f0bbca424769f61d5c904c2b4c2428d94726039009482a19b497ceef

Observation 41cf55af-ddbc-455e-b3bd-97318fcc5b60 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.792042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.728274Z digest=sha256:14af7dd8479293353463754eb1aa4588be1c46ff1a81945802a6125a1e3627fd

Observation c596d65e-0b2a-41ed-a8ca-41fe1030c439 · outbound

This paper cites L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling L., Zhang, C., Xu, Y., Shang, N., Xu, J., Yang, F., and Yang, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.774919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.734484Z digest=sha256:508a5305beb82b69d2b4d258833dec6cfabed2cef7cdfa6570042839301be08a

Observation 7548d31d-efd0-4ff1-936e-a3329d3fc9cc · outbound

This paper cites X., Wang, B., Tian, Z., Chen, W., and Wen, J.-R.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling X., Wang, B., Tian, Z., Chen, W., and Wen, J.-R

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.759338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.739622Z digest=sha256:e8d003345f0e7b11660069eba4c7a3089555796b13b171828d181e6757e38ffa

Observation 6b652bf3-e455-4335-9306-f8562761a66a · outbound

This paper cites The Llama 3 Herd of Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.744562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.744562Z digest=sha256:ef60afcd2a8040e5d8be9fe5a5e4f3218c05aef49f775df19db1c3369f15f5da

Observation 339f643a-586c-477f-869c-4a27c619d4d3 · outbound

This paper cites Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.744028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.749852Z digest=sha256:f9721f564813bc31c8182e328056b7309b617a0ebb47c9da7f82df537a2e1e62

Observation 1a8bc742-36e0-4988-9b6b-e375a885ad6c · outbound

This paper cites and Das, K.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling and Das, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.728024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.754667Z digest=sha256:4981377a3dc4628ed08c042e8d758f684f2286855de1cc4681b5a75b4a0cd97a

Observation cf188ecf-51e2-484e-92bd-a5e0721bc083 · outbound

This paper cites ChatGLM : A family of large language models from GLM-130B to GLM-4 all tools, 2024.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling ChatGLM : A family of large language models from GLM-130B to GLM-4 all tools, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.712294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.759412Z digest=sha256:e778107dbd8428820b8f1814cc84f84b3b37c3f8ac08127a871fce1c4be2a2b6

Observation 3540208b-63db-4f9d-b6f2-6aa9e12ad817 · outbound

This paper cites Comparing results of 31 algorithms from the black-box optimization benchmarking BBOB-2009.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Comparing results of 31 algorithms from the black-box optimization benchmarking BBOB-2009

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.696549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.764123Z digest=sha256:12b9920de86835296a4daa1b95dee1dec4a3708eb84d0641195f40ed1954cf7b

Observation 78bba061-cf75-4aec-9343-966316fe9e89 · outbound

This paper cites Self-attention attribution: Interpreting information interactions inside transformer.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Self-attention attribution: Interpreting information interactions inside transformer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.680842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.769107Z digest=sha256:9a0c0bd82fa42254f0f7b77114bd7d62696a2749165ff6ae204c71fb1da28f4e

Observation 573ae094-0bdb-4b6a-aae0-1f0ac36359fa · outbound

This paper cites WizardLM-13B-Uncensored , 2023.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling WizardLM-13B-Uncensored , 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.665681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.774170Z digest=sha256:e4c79a1ac30a29f15fe5bddfa8512fd7c69db5a5c3445f949a2110fcd78794e4

Observation 1e508619-bdbc-4886-be77-fb0afa741acc · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Measuring mathematical problem solving with the math dataset

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.650630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.779314Z digest=sha256:7694257b145fa83b0ce066b9ec59cc6c1e332af9a1caf8771fa4ef67b4d32322

Observation 2d7f3953-7880-46eb-8df1-0541ac21f691 · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.784054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.784054Z digest=sha256:741fe28b731dec16a4804e6ad2cb28c07137e2b078f5036d4a8d1c1a576e66a6

Observation 06f05d7f-5f69-49ce-b5fb-96df45cbc25e · outbound

This paper cites LLM maybe LongLM : Self-extend LLM context window without tuning.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling LLM maybe LongLM : Self-extend LLM context window without tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.634246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.789233Z digest=sha256:d7b6db81deb67f102dc09602c507ef498e61c3346acfbdf4ed7a01159ae68ce8

Observation e60d5aa6-b1e3-407d-a9b8-d98f36562f04 · outbound

This paper cites an unresolved cited work.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:02:39.618661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.794558Z digest=sha256:5cff96e20f8629a96ff362ce0b62f704671e003f5650edc4eae1ba5f2bd011f9

Observation c034bce8-ca56-49ff-b366-31cf55b5cc2c · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.603277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.799904Z digest=sha256:f45d9c51e850983a12a0a374ca117d4b0bfac2d968b961f04a200716558157e9

Observation 2ba4cd55-4d1e-48f6-a4c9-2645fcc1d322 · outbound

This paper cites Attention-enhancing backdoor attacks against BERT -based models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Attention-enhancing backdoor attacks against BERT -based models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.587553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.804891Z digest=sha256:2e0ee435c2ecd86015794f0ee4200346a8dc13ab5f72ead403491fbceee94776

Observation a2b9f54d-e71d-42f1-9f8a-89a5fbe7ebb7 · outbound

This paper cites Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Harmbench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.571273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.809842Z digest=sha256:4d2feb8bf9f2bc1b32c65de1e5cbf26eae55c1a7979eaa42bd1c8fb076e477f4

Observation f88212ac-0e77-44f2-8d9f-80d82fd105a0 · outbound

This paper cites Tree of attacks: Jailbreaking black-box LLMs automatically.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tree of attacks: Jailbreaking black-box LLMs automatically

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.556180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.814969Z digest=sha256:ab478f4928e0233da12dce56668182b9b6587daff27c29223cf4a39a8d2566d5

Observation e980d4f3-c52f-46e3-ae06-2e25889085ca · outbound

This paper cites Bayesian Optimization : Open source constrained global optimization tool for Python , 2014.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Bayesian Optimization : Open source constrained global optimization tool for Python , 2014

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.540977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.819627Z digest=sha256:bbdcf70f7d198451b2cac2d7d0e09a42ba13c008758a5e19cca3bd2e4e49f34a

Observation 08fe398d-46fd-4119-8b82-7e64b8ea9cb2 · outbound

This paper cites 2 OLMo 2 F urious.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling 2 OLMo 2 F urious

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.525142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.824402Z digest=sha256:99a9a051746c4047b89339888fe23cb7edf07ce301439101d46e38bb1699ec9e

Observation e3e47036-ad29-44a4-9247-54e72e8d252c · outbound

This paper cites Training language models to follow instructions with human feedback.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Training language models to follow instructions with human feedback

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.507753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.829040Z digest=sha256:31c7c641ddd442c852338cb3fd529915d97443474e1e2b862d140361feb1cbab

Observation e605e30b-a35e-4755-b477-4def7886fac0 · outbound

This paper cites S., Soltanolkotabi, M., and Thrampoulidis, C.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling S., Soltanolkotabi, M., and Thrampoulidis, C

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.491573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.833730Z digest=sha256:3131a72a8a1ee363b846f49795eaafeac113feeacda0e02dcb32a309282fe707

Observation 8800b52c-2221-4ea7-9382-7e70755cef03 · outbound

This paper cites S., O'Brien, J., Cai, C.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling S., O'Brien, J., Cai, C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.475920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.838426Z digest=sha256:27e68ed29ab36fab8ba0e7c991066b34294e11ccd7edae8a93159cd70f72a06d

Observation 6a9f2eb1-747e-4c56-b095-68e9a8541ede · outbound

This paper cites Red Teaming Language Models with Language Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Red Teaming Language Models with Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.843272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.843272Z digest=sha256:8f055c8a6f44cec3eaa2de0248d6a50da002708bae5ed91bb5ea74f789cb9bdc

Observation 8f6256a7-e521-4e52-9c58-7aa146b45b97 · outbound

This paper cites Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait crafting.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Baitattack: Alleviating intention shift in jailbreak attacks via adaptive bait crafting

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.459923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.848658Z digest=sha256:1df62e0c318b65a99c61d2bc91b98283c054b7cdb0f795d7eef222903938eb45

Observation 07ef3fc9-f027-4e17-9387-645ed9bd8f2a · outbound

This paper cites and Barez, F.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling and Barez, F

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.443220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.853805Z digest=sha256:b1b63b23070dede634eeb3a8f659593a1381548d64f64cc3b5c21aed4242f5d3

Observation 19d42300-af06-4934-87e9-aa127ecc40df · outbound

This paper cites Language models are unsupervised multitask learners.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Language models are unsupervised multitask learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.858787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.858787Z digest=sha256:ede14cdfb6e2280399f89f3b7ecc7af5d578afdf030195bee9bf1f00623508ea

Observation b3127a67-72f4-42e9-827e-fbd0049284eb · outbound

This paper cites Tricking LLMs into disobedience: Formalizing, analyzing, and detecting jailbreaks.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tricking LLMs into disobedience: Formalizing, analyzing, and detecting jailbreaks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.416559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.863659Z digest=sha256:73e8de0dccf898f6c6fb3744838cfce24e86e58d524cecb79c46b903c85fbebb

Observation 662cfdae-b896-4968-9704-422ae6d95df7 · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.868276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.868276Z digest=sha256:f04928de8469794390ab9e594526348f334e51e368fe764860805c96ba620611

Observation ad5cd759-6940-416e-a69a-fe7683933d1c · outbound

This paper cites P., and De Freitas, N.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling P., and De Freitas, N

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.399652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.874036Z digest=sha256:ae81ee57bc73f2d31c0af81c4988726bf7e1b3786d66611a650e9b1623c605ff

Observation bc2101dc-cbca-443c-855f-1461c93cb6be · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.879189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.879189Z digest=sha256:b9255f50c93f6d01a3fb1915b280c337f17e42c21429dd6b823d1ab248b3495b

Observation 2e9a0925-b676-46d4-a4dc-644c39dc8f0a · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Qwen2.5: A party of foundation models, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.382828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.884299Z digest=sha256:18c31ad828b3176cb5a09f6f42ccb58e9fee7d4096f5a54cf9c2ff9c06847934

Observation 917cc9bb-3b75-4789-9ed4-6272e02ba0ae · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.889082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.889082Z digest=sha256:8778e128bc766627b6e0816a48a4ce4c5828fa86e58c9183f2f65da95197a499

Observation 8a1cd8ae-a4d1-42e7-aa71-d651bffd4ff7 · outbound

This paper cites Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.366481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.894328Z digest=sha256:e9e97d80ab867b0a12b0f40456dace33f6899cbc514a9d50c08447622a706836

Observation 806772cf-42de-477b-9c5a-e628e8470ee2 · outbound

This paper cites Attention is all you need.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Attention is all you need

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.350820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.898986Z digest=sha256:457120fbb5590cd851b9b09aa503668346b662aad4a46a83a683f8ca4a38d37a

Observation ab6a9ead-82e5-4605-8df1-46eb9577f14c · outbound

This paper cites Jailbroken: How does LLM safety training fail? In Advances in Neural Information Processing Systems (NeurIPS), 2023 a.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbroken: How does LLM safety training fail? In Advances in Neural Information Processing Systems (NeurIPS), 2023 a

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.335411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.903722Z digest=sha256:0d86fd95b192a9ed439e51716875a83a99b9a737a1bed4b2ebe876630cf1cf29

Observation a928bc4c-1dd2-43bd-862a-186b3966b4c8 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.908552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.908552Z digest=sha256:5787c5ed93d66ce71366b47b35e8a5c162824d0d5faf6c33b13adb8393430757

Observation b7cbf08a-857b-48a3-99eb-54fc721163eb · outbound

This paper cites Never miss a beat: An efficient recipe for context window extension of large language models with consistent ``middle'' enhancement.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Never miss a beat: An efficient recipe for context window extension of large language models with consistent ``middle'' enhancement

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.319450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.913719Z digest=sha256:8de37514a07147331b303d683f6608e103951639726ae6d83cca3ee7f0a20893

Observation 6a37ad53-bd30-4d3e-a0c9-25f8077beb81 · outbound

This paper cites Distract large language models for automatic jailbreak attack.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Distract large language models for automatic jailbreak attack

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.302514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.918497Z digest=sha256:8e4abc250294a57077cac50a99b6403c93556730ff62e76b1bdd01d5bf7717ee

Observation 1ad21543-645e-4493-b87e-507f7de8f4bd · outbound

This paper cites Defending chat GPT against jailbreak attack via self-reminders.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Defending chat GPT against jailbreak attack via self-reminders

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.286459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.923177Z digest=sha256:f3013f6f7d58683d0304fa38898b6979516201629e17afa894f8dfa5a7f614d1

Observation 0a69fe84-05bd-47b1-b7c5-53d4f54bf1bd · outbound

This paper cites Qwen2 Technical Report.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Qwen2 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T14:02:38.927696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:02:38.927696Z digest=sha256:261df364fd4a0737bd86b11698ec02fa98840ed0d8de8301b6531c707b1fc98e

Observation eec647d9-6744-45df-bee5-89ac7b834d2a · outbound

This paper cites Tell your model where to attend: Post-hoc attention steering for LLMs.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Tell your model where to attend: Post-hoc attention steering for LLMs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.270411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.932874Z digest=sha256:be1f994671e562669ead6f2cba0e4a931ac2c3cc54e89d783a692421e6eba209

Observation cdb0e819-a258-48f8-b051-968249140517 · outbound

This paper cites In-context principle learning from mistakes.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling In-context principle learning from mistakes

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.253913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.937270Z digest=sha256:85d35c1825a912ea904d7aab8cd5e12ca45b5289096d3b28c552e97889c95eb5

Observation 3158314c-1474-4972-8274-163ca8dc1a5a · outbound

This paper cites What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems (NeurIPS), 2023.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling What makes good examples for visual in-context learning? In Advances in Neural Information Processing Systems (NeurIPS), 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.235536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.941925Z digest=sha256:04eae1f7f0b6bdb62ffe07febe80a13646ba4ee868b7c36f9a664f5306c37bae

Observation 0bddeb3a-23f4-4153-b5ee-d7a272365c80 · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Calibrate before use: Improving few-shot performance of language models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.218387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.946524Z digest=sha256:bc302e55fe15a82747ebb81724cbbcb02579a69ec77dc1cdd86ea5a58b8c610b

Observation 7a23758f-509b-465d-b6d1-76b017de50a5 · outbound

This paper cites Improved few-shot jailbreaking can circumvent aligned language models and their defenses.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Improved few-shot jailbreaking can circumvent aligned language models and their defenses

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.202097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.951304Z digest=sha256:b07a6d8864ba0ab08568a25971a0d34272887fbac067a5963fc576d348c4ecac

Observation 8e68e883-4acb-444c-84d2-412a40e8c2b5 · outbound

This paper cites P., Di Eugenio, B., and Zhang, Y.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling P., Di Eugenio, B., and Zhang, Y

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.184896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.955740Z digest=sha256:9fec0442325834f72a777981f51dea7a7e353e3ef0ac59b99a3a215027c723fa

Observation 7078b9d0-2b3a-4673-864d-786cbfc30fb2 · outbound

This paper cites Z., and Fredrikson, M.

PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling Z., and Fredrikson, M

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:02:39.168206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-09T14:02:38.960422Z digest=sha256:f6663e99837bbcded40bd44b4b65254119cb283ff6c0d12af7d3ee18e126843e

Pith citing papers

Observation cc00ba49-84be-496a-8562-a7fb38e9ce60 · inbound

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration cites this paper.

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.961667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T00:50:19.648405Z digest=sha256:fe7733115365884cdbabff7b766a5b553885882375f894e9d5cb474accdf3414

Observation ce3b0946-6889-419e-b0b4-9ee3f453c93f · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space PANDAS: Improving Many-shot Jailbreaking via Positive Affirmation, Negative Demonstration, and Adaptive Sampling

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:22:18.823341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:80e502c05a0f3fea96f219eb0e47bc964b935d830c6f717abbbf34d118b8bd86