Pith. sign in

Paper Citation Record · LEDGER

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

As of 18 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2501.10639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10639 v3

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:08:01.799775Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T15:26:23.290009Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T15:27:20.106151Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0dddf4d-c200-4382-af12-531fb113d44c · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.566687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.566687Z digest=sha256:656d6d9d5e8917985cd1563c0fcb44fef36ccd85dbf2a6ba6f4c954ea9132514

Observation 5189c3aa-57f5-4e1d-ad00-3031b6a7f22f · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.571828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.571828Z digest=sha256:b7b23d7dc05728c8584d117421b5a310bccbb2a81c35b4d92291bcae7cadeacc

Observation 9888649f-64a0-4d39-860d-ab2b2a2291d8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.578805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.578805Z digest=sha256:17ed5147c21e22fad778e77febc6c2da1c53a45c2901414e16460329a058c0e5

Observation b189b13d-7f8c-41dd-9fd9-c0b334bbead5 · outbound

This paper cites , author Ghosh, S.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ghosh, S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.565008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.583668Z digest=sha256:a6b41fc105457abe9dde8e4f972e94e3deaaca51df8fca105a548bec00aa6e66

Observation c67ebe50-2a32-408a-8f48-c3504b14a518 · outbound

This paper cites , author Ye, H.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ye, H

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.551904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.588539Z digest=sha256:95c2595270e2b882952f2425503435e5c7727c6816c62e82a8f71f6fab5a3ff4

Observation c7cf5986-6253-44bc-b21d-cd996d3ce9f1 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.593502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.593502Z digest=sha256:f16847c900e2929683977c35cd55c3831a254367c0238608226aa7289fd95c52

Observation d2ac95ef-8939-4c86-9984-a03b39db9a6f · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.599083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.599083Z digest=sha256:cafc32469c02aecb38fbdb594b274d3955368bf4b8a48143e9c9424cea40051f

Observation 1d04c9aa-2ef6-4740-8ab8-0f33ee4c147d · outbound

This paper cites , author Robey, A.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Robey, A

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.538492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.604071Z digest=sha256:4529d2ba8a61e715b6b4ca239a0ab652c2c47edc77c4be0ad93b118aa2166399

Observation c530ab1e-d492-41e7-8627-266a1c0dfe3b · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.608940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.608940Z digest=sha256:62d0437e60335ef1c4ddb142199d5dece77778968c7c62e431447fbc6f454290

Observation cd124da8-9e9e-4263-b6e3-6b983ac61d1a · outbound

This paper cites , author Ruoss, A.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ruoss, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.524902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.614032Z digest=sha256:d00596d36429c9ee118556a94e36813cfd142f8b865c408a4e26217937d21435

Observation bdfe378a-b3be-4cf4-9931-f706dc698eb1 · outbound

This paper cites , author Chen, Y.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Chen, Y

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.511111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.618703Z digest=sha256:bea2c2607724a271d1fa2d581946a91ac433417fbdcb652f310b5f3b47db1906

Observation 97378996-f677-4d99-ad7e-a4440b3760e1 · outbound

This paper cites MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.623284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.623284Z digest=sha256:2db86c268621fb992be64385e54331de160816e64b438b0d4e8b4289729755e9

Observation 94febe71-c2b2-4e3a-8d95-b055e2ad8e64 · outbound

This paper cites PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.628010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.628010Z digest=sha256:0b6974e65b59b666d2afa15a6b964b13eb2fb4569354317b30a1b0c383961575

Observation dcdc1a55-96d3-4e61-bdcd-9bfbe6f56725 · outbound

This paper cites , author Yu, F.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Yu, F

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.496311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.632772Z digest=sha256:85b1b054ab0c88fb1b8fa7bf23db80400ca8cc72d6d53c1e32441cf0b08c103e

Observation 8d403bd6-9448-4d89-8a6a-d07e32812eae · outbound

This paper cites , author Burns, C.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Burns, C

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.482280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.637090Z digest=sha256:087aa57b352ef0e6549e303f892eda63b7784fd5c3d1cc9d6a88a9787a38e40e

Observation fe449179-5845-4b2c-821b-386d09820afe · outbound

This paper cites Inspecting and Editing Knowledge Representations in Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Inspecting and Editing Knowledge Representations in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.641136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.641136Z digest=sha256:36b51a8dfa2e51dbabaf750346b9b6eaac2986845a18c41171a3fda8af055c4d

Observation 9c367882-5566-4d4f-baab-1155ff1a9ef4 · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.645564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.645564Z digest=sha256:95329b7367dc7a91c6c902fa3262ae6dd91ec3e63976768b012bb84230847b79

Observation f8e218e1-6d2d-4e11-b27d-620525c5b7c6 · outbound

This paper cites , author Choi, E.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Choi, E

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.467635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.649858Z digest=sha256:98a0cca25ae6a56d175162390e7aa6ca74af9b65456d1a1391d98661553b7804

Observation d697efc6-c479-4a4c-b0ff-2a0700cfffb6 · outbound

This paper cites , author Li, X.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Li, X

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.454451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.654313Z digest=sha256:284b22569fe62c3207a062f7f4e153755500c663410f1078f3120ab3820b1b90

Observation a5e4a53d-703d-4811-b94d-203576856598 · outbound

This paper cites , author G \"u rel, N.M.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author G \"u rel, N.M

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.441977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.658208Z digest=sha256:f221f26f96b9fca5ba1fd2d7a498f92cc9364b241779a340a626456e5e884a98

Observation a2e3eaa3-3127-4a17-9a5e-66867126df9b · outbound

This paper cites , author Al-Rfou, R.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Al-Rfou, R

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.426978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.662138Z digest=sha256:388135210b9387e1920a27348c613b6066bdd0a3eb8ccdebf0373738b8f2660f

Observation 6be3978e-c578-4853-b6b2-d25fb739904e · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.666387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.666387Z digest=sha256:323fce3343a9f0b8e1b2476709da0ceaa84258cbbed37542623c85661f48660c

Observation ee780b3c-0826-449c-ab4c-ea88790560ba · outbound

This paper cites AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.671246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.671246Z digest=sha256:d6250302c80cd794db96aa19053f5bc7f3c50e010ed40866ca31b3feeaa575f3

Observation 8214eded-d1f5-4045-a1ef-9b3c8ccda58b · outbound

This paper cites Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.675859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.675859Z digest=sha256:58861d87ad4f3e639af052a9717338a9ab8359f6077c87089e16a2f7179039f5

Observation 8081f179-da4f-4f01-9f50-a433654c7da6 · outbound

This paper cites , author Xu, N.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Xu, N

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.413897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.681800Z digest=sha256:7a47cbc5d6e8fcdf21b0af0d23ad9c9ed73c6aff4f77e2e68e30d96c49a58afd

Observation c1595a31-3a4c-489e-9d57-1e06d74e2ae0 · outbound

This paper cites , author Feng, Z.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Feng, Z

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.400845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.686421Z digest=sha256:d6c676ea42f794dd3428838d0f63dc42895e074f702100af8c4fad475281c586

Observation 63b0eb18-1d74-4ad7-8018-a6f9b920b10e · outbound

This paper cites , author Phan, L.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Phan, L

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.387649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.691155Z digest=sha256:a6d88a965c8c500aeb6a8675a2a3c71eaa4bc7492f369723d77b763c085abec8

Observation 9e20dac0-cba0-46c5-872f-f3e9580ec8f9 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.696363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.696363Z digest=sha256:caa314ee7d0be309a0cc520d05715b621ab2f9c660eb78781f90a1393899f767

Observation c8102a52-8326-4d43-a705-e939f9ecd07a · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Steering Llama 2 via Contrastive Activation Addition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.701320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.701320Z digest=sha256:19cdb7a21f83b621d18b04af35d8b57d3066992cd213c5f2e2127b7750729de4

Observation 6607a1e4-649e-4bdf-9212-e6e75f0cdb88 · outbound

This paper cites , author Wong, E.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Wong, E

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.374326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.706116Z digest=sha256:e6d7a02b93840e8956200eb247925d982af40ff52585676a84a0ce4942da3b4a

Observation 04604d75-cde7-4c1e-a00b-1aff869832e0 · outbound

This paper cites Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.711503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.711503Z digest=sha256:0558c2858ec04bfb6510a2a9a348d01f9205a8287a5730d7aa218892788630b7

Observation 2b769401-6465-427b-8405-f0c8a8162dec · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.716089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.716089Z digest=sha256:5c81a3132a2ced21b7ed607f990b7ac21aaa71f372576a699702d2bf7e55acf0

Observation f0417e44-1297-491b-ab59-55f0d9698a60 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.720455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.720455Z digest=sha256:b4d82c7653e84121cbc70c8c5c7300c31fbbe07b53200f86fa6adac862fe3ac3

Observation 181e5b86-7bd9-43df-8c52-5359bc7c1512 · outbound

This paper cites , author Chen, K.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Chen, K

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.359315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.724684Z digest=sha256:8f553b4e59e5f92671fa1c18355eb1d08b45fad02b4779c528946c163a89ae7f

Observation ca1e0899-d08a-4dda-983b-41757f54af9f · outbound

This paper cites Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.728950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.728950Z digest=sha256:766c8562707722460aa2bf75ea54effa0b6d108405b8082d0a96eeca4b88662d

Observation 049bbf7f-8aca-4712-b114-6937e92cb423 · outbound

This paper cites , author Haghtalab, N.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Haghtalab, N

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.343484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.733123Z digest=sha256:991a39435dd00d04820352dcb56fc9f316f8cfd59a9517bc049820254e2d1372

Observation 19cf3e0d-9de7-48fe-b988-f87b47517819 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.737120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.737120Z digest=sha256:104e1f201c1bdd0bc731064450b1beba4ca01b713bfd4e4ba8d843bc3b88a6ae

Observation 9163df2c-75da-4783-9748-385de29a21c1 · outbound

This paper cites , author Yi, J.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Yi, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.328209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.741382Z digest=sha256:14e2350ae0a364ce637399c820187f620271a7b3cc2f884935e1857ad6417319

Observation 24072a27-516c-470f-9abc-3e840554bcff · outbound

This paper cites , author Huang, R.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Huang, R

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.312988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.745608Z digest=sha256:19e4358bfde09e81b4a147e947ed6d60efd7d30a36941c10458d48e715804522

Observation 0c14d0ef-7e00-433e-b8f6-874c769a2d98 · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.749591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.749591Z digest=sha256:b43e1e317a89a5ce204ccef50f73a2bceb6f3e80a52834c61d07fd29bd50fbf4

Observation b290920e-109b-4cc1-83dc-4f29b1d9f7a4 · outbound

This paper cites , author Ye, R.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ye, R

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.298063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.754352Z digest=sha256:73e997ee16a3981ac9a42c4208998a5942dfc3e1d2d513b0e89a90876dbb99ba

Observation b5d3cf62-d8ef-4013-b15c-4e685fa11192 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.758667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.758667Z digest=sha256:6ec2ecc5207e3915dc2d6cd19bb9a22952944b6a427dda6de28f659d45a2f690

Observation 4fee971d-24aa-41bf-8321-ddf3da8abf56 · outbound

This paper cites Robust LLM safeguarding via refusal feature adversarial training.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Robust LLM safeguarding via refusal feature adversarial training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.764576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.764576Z digest=sha256:efd5aa3e3cf4352cb7bdb6d4878f83ec241ed6d4facefd7895af99c206dfcbbf

Observation f76712e0-a05a-4230-92d8-9c76d534f4a9 · outbound

This paper cites AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.769602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.769602Z digest=sha256:81644533a615f7dcbe51e7af466955a8f3a2cd8b6ef15706f5bc3497b7a6130c

Observation 10e3f891-9a1d-4105-ba86-3c0a9a86e5eb · outbound

This paper cites Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.773900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.773900Z digest=sha256:f7bf3ec30fcc581401537db6f09e61bb683b762d31be83afd915b9e0a6842bc4

Observation 4469a48d-a05d-45a0-a860-e2e544ac5b45 · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks On Prompt-Driven Safeguarding for Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.778736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.778736Z digest=sha256:421a381c6b279ab8ac1d98e3fb5ca769efd039a542a0bea18d2dee7029c78fe9

Observation 38181aea-7c25-495f-89aa-4591513365cf · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Representation Engineering: A Top-Down Approach to AI Transparency

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.783546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.783546Z digest=sha256:41978e3c251512e5495a72b09859460090147021ffa61b78d04f47f081718ff1

Observation 45c68195-138c-429a-ad67-aac2c22e6de5 · outbound

This paper cites , author Phan, L.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Phan, L

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:08:02.282597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T19:08:01.791055Z digest=sha256:d76dd46140284dcf9875d0ef7c374621eccd67121dfbf7875b23680a967c4a3e

Observation 0d93d1b8-eeda-45b8-a0b4-ecfec518d5c7 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.795490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.795490Z digest=sha256:ca249dc15a34e7bf2a74553895e0ff036062aa8ccea2cf3e4ab480951aaf147c

Observation 259f26f1-f9ba-4b42-b94c-8491424e6f70 · outbound

This paper cites write newline.

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks write newline

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T19:08:01.799775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:08:01.799775Z digest=sha256:9f9c43276431ad6a9457cae29ff7612a948a88f24edd2c83e3654fa53f1ccc09

Pith citing papers

Observation ee1a0b98-dd8a-45a9-b9d0-7bf12bcd306d · inbound

Efficient Safety Alignment of Language Models via Latent Personality Traits cites this paper.

Efficient Safety Alignment of Language Models via Latent Personality Traits Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-10T15:27:20.107446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-10T15:26:23.290009Z digest=sha256:e37d9446acf4bc860e5de9138612ae43dac65697e91e0d64aef932ecac077d30