Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:08:01.799775Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2501.10639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:08:01.799775Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T15:26:23.290009Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T15:27:20.106151Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e0dddf4d-c200-4382-af12-531fb113d44c · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5189c3aa-57f5-4e1d-ad00-3031b6a7f22f · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Refusal in Language Models Is Mediated by a Single Direction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9888649f-64a0-4d39-860d-ab2b2a2291d8 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b189b13d-7f8c-41dd-9fd9-c0b334bbead5 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ghosh, S
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c67ebe50-2a32-408a-8f48-c3504b14a518 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ye, H
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7cf5986-6253-44bc-b21d-cd996d3ce9f1 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ac95ef-8939-4c86-9984-a03b39db9a6f · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d04c9aa-2ef6-4740-8ab8-0f33ee4c147d · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Robey, A
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c530ab1e-d492-41e7-8627-266a1c0dfe3b · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd124da8-9e9e-4263-b6e3-6b983ac61d1a · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ruoss, A
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bdfe378a-b3be-4cf4-9931-f706dc698eb1 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Chen, Y
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97378996-f677-4d99-ad7e-a4440b3760e1 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks MoGU: A Framework for Enhancing Safety of Open-Sourced LLMs While Preserving Their Usability
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94febe71-c2b2-4e3a-8d95-b055e2ad8e64 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcdc1a55-96d3-4e61-bdcd-9bfbe6f56725 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Yu, F
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d403bd6-9448-4d89-8a6a-d07e32812eae · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Burns, C
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fe449179-5845-4b2c-821b-386d09820afe · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Inspecting and Editing Knowledge Representations in Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c367882-5566-4d4f-baab-1155ff1a9ef4 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8e218e1-6d2d-4e11-b27d-620525c5b7c6 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Choi, E
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d697efc6-c479-4a4c-b0ff-2a0700cfffb6 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Li, X
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5e4a53d-703d-4811-b94d-203576856598 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author G \"u rel, N.M
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2e3eaa3-3127-4a17-9a5e-66867126df9b · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Al-Rfou, R
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6be3978e-c578-4853-b6b2-d25fb739904e · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee780b3c-0826-449c-ab4c-ea88790560ba · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8214eded-d1f5-4045-a1ef-9b3c8ccda58b · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8081f179-da4f-4f01-9f50-a433654c7da6 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Xu, N
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c1595a31-3a4c-489e-9d57-1e06d74e2ae0 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Feng, Z
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 63b0eb18-1d74-4ad7-8018-a6f9b920b10e · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Phan, L
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9e20dac0-cba0-46c5-872f-f3e9580ec8f9 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8102a52-8326-4d43-a705-e939f9ecd07a · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Steering Llama 2 via Contrastive Activation Addition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6607a1e4-649e-4bdf-9212-e6e75f0cdb88 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Wong, E
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 04604d75-cde7-4c1e-a00b-1aff869832e0 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b769401-6465-427b-8405-f0c8a8162dec · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0417e44-1297-491b-ab59-55f0d9698a60 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181e5b86-7bd9-43df-8c52-5359bc7c1512 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Chen, K
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ca1e0899-d08a-4dda-983b-41757f54af9f · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 049bbf7f-8aca-4712-b114-6937e92cb423 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Haghtalab, N
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 19cf3e0d-9de7-48fe-b988-f87b47517819 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9163df2c-75da-4783-9748-385de29a21c1 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Yi, J
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 24072a27-516c-470f-9abc-3e840554bcff · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Huang, R
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0c14d0ef-7e00-433e-b8f6-874c769a2d98 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Uncovering Safety Risks of Large Language Models through Concept Activation Vector
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b290920e-109b-4cc1-83dc-4f29b1d9f7a4 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Ye, R
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b5d3cf62-d8ef-4013-b15c-4e685fa11192 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fee971d-24aa-41bf-8321-ddf3da8abf56 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Robust LLM safeguarding via refusal feature adversarial training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f76712e0-a05a-4230-92d8-9c76d534f4a9 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e3f891-9a1d-4105-ba86-3c0a9a86e5eb · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Adversarial Representation Engineering: A General Model Editing Framework for Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4469a48d-a05d-45a0-a860-e2e544ac5b45 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks On Prompt-Driven Safeguarding for Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38181aea-7c25-495f-89aa-4591513365cf · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Representation Engineering: A Top-Down Approach to AI Transparency
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c68195-138c-429a-ad67-aac2c22e6de5 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks , author Phan, L
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0d93d1b8-eeda-45b8-a0b4-ecfec518d5c7 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259f26f1-f9ba-4b42-b94c-8491424e6f70 · outbound
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks write newline
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee1a0b98-dd8a-45a9-b9d0-7bf12bcd306d · inbound
Efficient Safety Alignment of Language Models via Latent Personality Traits Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.