Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2401.17256.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:06:39.024687Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:29:44.257667Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c586368a-150e-4f55-84a3-29d1350e9d56 · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 22481f73-652b-4e7a-abb2-ec51598fee79 · inbound
Peering Behind the Shield: Guardrail Identification in Large Language Models Weak-to-Strong Jailbreaking on Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2ed9ab0e-3aec-4809-979b-34f74362d5ad · inbound
Confidence Elicitation: A New Attack Vector for Large Language Models Weak-to-Strong Jailbreaking on Large Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e17858c5-6305-4ec4-a4cc-c9d78526c34f · inbound
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Weak-to-Strong Jailbreaking on Large Language Models
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 65d59232-1a7e-4c6b-8b98-3e67cee7e60b · inbound
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Weak-to-Strong Jailbreaking on Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ca2ae1-9574-45a8-a4bf-cc5df38bc296 · inbound
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion Weak-to-Strong Jailbreaking on Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb482b29-547d-4846-bd6b-42e9cf4b000f · inbound
Security Concerns for Large Language Models: A Survey Weak-to-Strong Jailbreaking on Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c35c1c48-c59d-4924-9014-51bde1ac83e3 · inbound
SafeTy Reasoning Elicitation Alignment for Multi-Turn Dialogues Weak-to-Strong Jailbreaking on Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 595452c4-6357-4760-8f52-30d432d8a57b · inbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014a108e-00d5-492c-8dad-fb42b13f5e67 · inbound
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Weak-to-Strong Jailbreaking on Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 31368c05-277e-4bbf-b0e3-6e2c92dff6f8 · inbound
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Weak-to-Strong Jailbreaking on Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2687ee65-1b68-46c1-91c9-905b049d5d40 · inbound
Logit Arithmetic Elicits Long Reasoning Capabilities Without Training Weak-to-Strong Jailbreaking on Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6b8c244-c238-478b-9765-920e0081a3a9 · inbound
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Weak-to-Strong Jailbreaking on Large Language Models
Reference 196
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0431ae-c85a-46f5-878f-6dcce83602e6 · inbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Weak-to-Strong Jailbreaking on Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfedeb4-c79b-4671-9818-36d308a6b432 · inbound
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Weak-to-Strong Jailbreaking on Large Language Models
Reference 241
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e2cef7-bad0-463a-8847-12988eecd8cc · inbound
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure Weak-to-Strong Jailbreaking on Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4cca6e4-e5df-47f5-9d7f-d15684c99b5a · inbound
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification Weak-to-Strong Jailbreaking on Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe3c3664-c480-4d42-b6c2-57b112b72c48 · inbound
Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes Weak-to-Strong Jailbreaking on Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0b49d19-bc3f-44ad-a27c-e72e6c28fd64 · inbound
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Weak-to-Strong Jailbreaking on Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 030fe290-72b6-4030-bff2-7b26a25262c6 · inbound
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 760ed90a-bad4-42fb-85e2-c538bd131da1 · inbound
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs Weak-to-Strong Jailbreaking on Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52fadcb1-083a-450b-ba00-b19f4d4df369 · inbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Weak-to-Strong Jailbreaking on Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47529e3a-b68b-475c-a2c1-45808eaf61f4 · inbound
Weak-to-Strong On-Policy Distillation Weak-to-Strong Jailbreaking on Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.