Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:45:55.362359Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.12060.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:45:55.362359Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5f5ca2fb-bb12-474c-ac85-7b04b62ded77 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3634ff-3a7f-4799-831b-b9a59b111d77 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b12999-6932-43f8-b9ce-e7d29cc44176 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Detecting Language Model Attacks with Perplexity
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd3e906-c155-48d7-a3c5-abfe4e6e4ba3 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f0dac6-028c-4fa4-8fe6-48587e7016a1 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bd4047-d1a1-4601-8414-9f581ac26c13 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f548a7b-3fc5-4eac-9ff3-3dedfacf0622 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3988a0-6d84-435c-9019-45aa0021fcf2 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Christiano, Jan Leike, Tom B
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a10d4a-ed63-4aa5-937e-582dc1ceeede · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ae1dcd-99fb-4380-b91e-9b055704f53e · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94cd4ae-7843-4f5f-b1cb-2dccc3688720 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db4da0b-733e-4971-a04f-ca46c8ffe92f · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ce8e54-3942-40b9-bb4d-fdd9659524cf · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement The Ethics of Advanced AI Assistants
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e70198-796f-40b0-b3e4-dd8c3fb3dfc5 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a67fc54-b43f-43c3-a5b4-b11b99e9e51e · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Measuring Massive Multitask Language Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24369157-fdf1-4369-9ea4-7edcf20d8af4 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96e56c8-99b4-4f6f-919d-da1aebb16906 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d314247-71de-421a-9805-f6aa364a32b3 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Buckley, Jason Phang, Samuel R
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 051b2a5a-2b35-45c1-8767-fa7a6da36ecd · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d2ec25-105a-494a-b06d-c3993c108f27 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76963d67-9297-4349-8e49-f749002a0979 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 27c125a0-efac-4f97-aca5-bd1986b51276 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689c87a8-a493-457f-a151-30b5f9fab1dc · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226b1c8b-c92b-4e19-a3a8-a5a52ad963c8 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement GPT-4 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8338340-6c13-4c01-9fe8-0cd385686949 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a8464431-916f-4204-8d0d-7c8d97e59d59 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9820e6e7-89f9-41bb-8760-54975fc5fc25 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Qwen2.5 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a222ed-31f7-49ee-8ec7-1168988832d5 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 664885f0-8c54-405a-994f-33af0ee3186e · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement do anything now
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 97929627-fe81-4ea4-bc8b-701da48999f6 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Gemma 2: Improving Open Language Models at a Practical Size
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de251238-8b88-4a7d-a869-d2776c500b30 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb52288-476f-4e6d-b3fd-d833ec2325f7 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 505700fc-08fe-42b2-868a-0df040fe0d5e · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 97e70e77-2b67-495b-9d54-d0d2e56fccd6 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2908be1-0bdb-4385-8425-46098b195243 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a3d297-22ff-4cde-b225-c4f24ce8a503 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c226f6fc-d887-4cdf-826e-39b600e10ddf · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bffd855d-015d-46e8-9e16-cfe41e746466 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83de0397-68ad-4fbd-b876-9211e302c92a · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7eaa9f-2b88-4bdc-afa5-949f2daba906 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 410697ab-5fa0-4466-a3c3-715ecfe5905b · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c72befb0-ed0d-48e0-9948-e8042d70e250 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c1b6c0-72f5-4802-b42b-09673dcd3c4a · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0b270832-d925-4afe-950e-90d540832e67 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4ea130-3938-4c43-89fc-9d0bc8f5936b · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 661e75be-95a1-4712-9f48-5c72250d2c3c · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Can Large Language Models Understand Context?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f86432d-de1c-4f80-95a9-30a9affc5135 · outbound
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.