Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:08:02.426832Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 14 inbound Pith citation observations for arXiv:2502.09674.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:08:02.426832Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:32.812422Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
53 of 53 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 68c228c1-18b4-4fbf-8bd8-d9a9a0851f0c · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47def67-bff9-4b0b-a1d5-7a51bc53b627 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d64c84fb-3bbe-40cd-b08f-19a4c1208da9 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Refusal in Language Models Is Mediated by a Single Direction
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f029c5-4cbe-4ffe-a398-b1589183a55c · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ebf5cf4-3e8d-4fa7-8a88-0ba84ff25d77 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1366a7b7-3958-406c-a0bc-56ad3f7b7479 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4efdb86d-c17f-407e-a93f-312281ef46f5 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538642cd-e89d-4373-82b3-8385d2300887 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions E., Hume, T., Carter, S., Henighan, T., and Olah, C
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7587eac8-b019-45c7-8adf-5c40a3c9eafd · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff6efc6-042c-4596-9565-be56fa075660 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A., Jagielski, M., Gao, I., Koh, P
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3e3daaf-7914-43bb-833d-200f1a153255 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe9b0f2-95fa-46b2-9c3c-c021d0c8a076 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b20750b-9e2c-47e2-ad85-3430e0d371e9 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e0e3a3-3aa5-415e-bf3e-68a6fcb04ac2 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82608076-ae80-4037-8042-964c5bed937a · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da61c799-cd6e-46da-bef2-457e7d73ce2e · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Not All Language Model Features Are One-Dimensionally Linear
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d9a90ac-738b-4e33-8fb4-b77772dd73cf · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5082f84c-cc85-4ff2-9bfe-86d9e9a4b889 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ca78a9-997c-4099-b4ae-02e4c3ff15ca · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b9a9920-c739-4a89-8c69-f234c6012681 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae28165e-f542-45d2-89f6-304d6ab43b01 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Artprompt: Ascii art-based jailbreak attacks against aligned llms
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d041aa4a-cf62-4cd8-a29d-8289e8798809 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e45a7e-76cf-463a-ba7d-93d8e5d1f7e3 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions and Moeller, M
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7646f38f-246d-4cbe-8366-4d53756e1f44 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 348a6b4f-c389-49e8-bae9-04b0b7cfce4b · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Inference-time intervention: Eliciting truthful answers from a language model
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a175d9f-deb9-4e8c-8ceb-07f6bee309d0 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Safety Layers in Aligned Large Language Models: The Key to LLM Security
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92e54623-1853-48b9-b9eb-63bc7ac3864a · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions FlipAttack: Jailbreak LLMs via Flipping
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6119c0-1d99-4d2a-98ad-af587fc43ac6 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a97a2c1-4831-4e8f-8023-7d664b488f88 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Latent space translation via semantic alignment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5cab35a3-1d74-436a-89d8-f36345162ce0 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e786ec-6733-49a6-be99-61c98ac31902 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Tree of attacks: Jailbreaking black-box llms automatically
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 064659f7-bb0b-408b-9e84-4ff66ae2cf6e · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Introducing the world's best edge models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7c76989d-7e2f-45e5-8883-e1d923dfd6fd · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Interpreting gpt: The logit lens
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b99e435c-d37b-4686-b7bd-a8e81b859708 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Training language models to follow instructions with human feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2dc371c-e6eb-4442-9208-e6935ee776fb · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242608d1-ef0b-419a-a143-43eac70ff1bc · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d77954-225c-4b4d-b47a-08898a04088a · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions D., Ermon, S., and Finn, C
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af117371-113a-4dbd-b623-17812b0a1182 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A strongreject for empty jailbreaks, 2024
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8462c8aa-9c0b-4ab3-8872-8cd0db12909b · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4d21228-1af9-4c6f-b676-051099dda54b · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320e0a42-2383-4b22-ace2-29f7b06777cb · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Hermes 3 Technical Report
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c5667f-71a5-480b-8dee-e1ba52d91567 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Detox: Toxic subspace projection for model editing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 737f3945-a6cc-4a1f-b817-56b3c983d795 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b894c94-9f12-4f6c-b447-82eaea163ff6 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions a ger, T., Elstner, J., Geisler, S., Cohen-Addad, V., G \
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434dbf62-dba0-407d-a3db-c61502453e29 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 921adb6e-823d-4c74-8b2d-93922965bd24 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Beyond toxic neurons: A mechanistic analysis of dpo for toxicity reduction
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 18e6a65f-3ff8-458c-a9f6-fd034c45081d · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A safety realignment framework via subspace-oriented model fusion for large language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe7a971-d1ae-4437-b7b3-138e698230a1 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Low-Resource Languages Jailbreak GPT-4
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e620374-a280-466a-895f-1e0c9b06836a · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d896d403-2005-443e-9c5b-fa6408a31bde · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Survey of Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 621f29b2-202f-4cd2-829c-3e80ce477a10 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions How alignment and jailbreak work: Explain llm safety through intermediate hidden states
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 36b35852-d14a-4f02-82c6-ac02953c7d69 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions On the Role of Attention Heads in Large Language Model Safety
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f5f7fef-e0d1-41c1-8fba-b0f34a872051 · outbound
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63ed797c-92c4-44b6-b78d-d3a1b93a6163 · inbound
GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3385d22a-b528-4772-8026-0aa662c4779e · inbound
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b416694e-7921-4807-93b9-12890c8c63dc · inbound
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ffcc5d-2c1c-4592-ac8a-bbbb34600921 · inbound
The Geometry of Harmfulness in LLMs through Subconcept Probing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c16fe6ed-c4e6-4d00-ae74-9673c6369022 · inbound
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a35c1ace-fc4d-4df3-a944-7ca233b559b7 · inbound
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35832151-671c-438c-a24e-3c270f324cf5 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff3a7668-6906-42ca-a745-a1b6dbb5d8c2 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b8a03693-8402-4396-b1c5-4cf60e4ed2c4 · inbound
Before the Last Token: Diagnosing Final-Token Safety Probe Failures The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation de65d733-af68-417c-93a9-723c6581aa4f · inbound
Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 21c0292c-32c3-4fa8-8f9c-d2e454c28328 · inbound
Why Do Safety Guardrails Degrade Across Languages? The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a83a1452-56b9-4d24-b7fb-d5c0ae4a39b0 · inbound
Low-Resource Safety Failures Are Action Failures, Not Representation Failures The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94f4d1d4-0e25-495f-bbea-ddcd37b40db5 · inbound
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66214a24-5bc7-423d-93f9-55adbc683088 · inbound
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.