Pith. sign in

Paper Citation Record · LEDGER

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 14 inbound Pith citation observations for arXiv:2502.09674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09674 v4

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:08:02.426832Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:32.812422Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 68c228c1-18b4-4fbf-8bd8-d9a9a0851f0c · outbound

This paper cites write newline.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.177293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.177293Z digest=sha256:ea08462259e19236dc8aa201dc5e5a1ba9f73ed99d8ba2a6219de7bd45beff78

Observation e47def67-bff9-4b0b-a1d5-7a51bc53b627 · outbound

This paper cites AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.183687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.183687Z digest=sha256:1837c550c9974f06532f1559201cc4f50b5a7ea89d09eb5e80a626e8cf5a1584

Observation d64c84fb-3bbe-40cd-b08f-19a4c1208da9 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.189529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.189529Z digest=sha256:aa71e6288f1ec8b1342957a70fab453789f4793bdae7b4f87cd58298e26e6a93

Observation 83f029c5-4cbe-4ffe-a398-b1589183a55c · outbound

This paper cites On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.194796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.194796Z digest=sha256:d7b22df48b208ac2aebc307194506f22ef80a75c43d6d104845dc22f9483e239

Observation 5ebf5cf4-3e8d-4fa7-8a88-0ba84ff25d77 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.199864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.199864Z digest=sha256:7a1ee0fc4440f28450fa411bce1a14216c96c84e0ff0b7130da3b660f4e11bda

Observation 1366a7b7-3958-406c-a0bc-56ad3f7b7479 · outbound

This paper cites Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:08:02.994811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.205089Z digest=sha256:aaf1f81e36956cee95280ce9de16321558fbd68e98a2161ad5a42394e73ce78f

Observation 4efdb86d-c17f-407e-a93f-312281ef46f5 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.210046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.210046Z digest=sha256:c061aceba2ac8179100054135c9cf3a0bc025c36c68ef5033474e6b4fa89eb6f

Observation 538642cd-e89d-4373-82b3-8385d2300887 · outbound

This paper cites E., Hume, T., Carter, S., Henighan, T., and Olah, C.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions E., Hume, T., Carter, S., Henighan, T., and Olah, C

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.215418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.215418Z digest=sha256:d5ed09310f6b0943c9d1129f107ca5ee8681d80ae31da9835839d7001be55009

Observation 7587eac8-b019-45c7-8adf-5c40a3c9eafd · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.219881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.219881Z digest=sha256:36e64ce9f0fc57ebf82703bf293035ee1f9acd5e3a11edfa03151bf150540a90

Observation 7ff6efc6-042c-4596-9565-be56fa075660 · outbound

This paper cites A., Jagielski, M., Gao, I., Koh, P.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A., Jagielski, M., Gao, I., Koh, P

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.298174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.225773Z digest=sha256:d17147ce35a8341b6d69adf7ed78c56aaef81380f648b71b126d0b6fa659944e

Observation a3e3daaf-7914-43bb-833d-200f1a153255 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.230810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.230810Z digest=sha256:6a8f301c3a498d0fa3f00e67122028de06979759e9170050f013d428c781b4e1

Observation 4fe9b0f2-95fa-46b2-9c3c-c021d0c8a076 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.235641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.235641Z digest=sha256:27f65d49b63c11462df243bc1cc2433d7432199c7764de8af8227a0f798df7aa

Observation 9b20750b-9e2c-47e2-ad85-3430e0d371e9 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.240706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.240706Z digest=sha256:bca9943a5e121ea654ba31dcc1334a370748acfccb0a4318ee919ba4779a8e6b

Observation 42e0e3a3-3aa5-415e-bf3e-68a6fcb04ac2 · outbound

This paper cites A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.245664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.245664Z digest=sha256:c98e1f986a62e69a2ca4984fe4d4cd07e7402ebb674aea1b0227616b09c0530d

Observation 82608076-ae80-4037-8042-964c5bed937a · outbound

This paper cites The Llama 3 Herd of Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.250294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.250294Z digest=sha256:32a193e36f6b541c995d0e2dde293c6dfcf49149a61e1520eaa9ffd11e2b4015

Observation da61c799-cd6e-46da-bef2-457e7d73ce2e · outbound

This paper cites Not All Language Model Features Are One-Dimensionally Linear.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Not All Language Model Features Are One-Dimensionally Linear

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.254996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.254996Z digest=sha256:513934664f59757e006ad70776d986a26f07a4794e9f22d9c7a5c5751360fd8f

Observation 3d9a90ac-738b-4e33-8fb4-b77772dd73cf · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.259789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.259789Z digest=sha256:a2dd6933f1b3eaf28a0a687f84c73904ef91e1558cd879c441e61725ce08180e

Observation 5082f84c-cc85-4ff2-9bfe-86d9e9a4b889 · outbound

This paper cites an unresolved cited work.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.264888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.264888Z digest=sha256:e4a590ddd81b05cfb557392857b79a722d1a506d4eb3b121dce8e92f64db50a0

Observation 23ca78a9-997c-4099-b4ae-02e4c3ff15ca · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.269258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.269258Z digest=sha256:5eaaaf21ede65567d589a6facf3f4bc98806831c43b2d04f78432db14ee3307c

Observation 6b9a9920-c739-4a89-8c69-f234c6012681 · outbound

This paper cites What Makes and Breaks Safety Fine-tuning? A Mechanistic Study.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.274152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.274152Z digest=sha256:b0f8fac0cfab375d59dfd12d0cd14177306d4da742ef648293735f135a612566

Observation ae28165e-f542-45d2-89f6-304d6ab43b01 · outbound

This paper cites Artprompt: Ascii art-based jailbreak attacks against aligned llms.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Artprompt: Ascii art-based jailbreak attacks against aligned llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.251840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.279049Z digest=sha256:6e090810645809526460091ef9d49ba321deba6a72dd938340d3e056cc3d6673

Observation d041aa4a-cf62-4cd8-a29d-8289e8798809 · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.283463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.283463Z digest=sha256:94c73c0be38def49ea30335b9d9550df415ac1f98f63fb651d8618fdefa30f39

Observation 79e45a7e-76cf-463a-ba7d-93d8e5d1f7e3 · outbound

This paper cites and Moeller, M.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions and Moeller, M

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.288290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.288290Z digest=sha256:cae461301968c9045d31b0b5b4cb353efd1390322a667b8786b7c6edf3b55388

Observation 7646f38f-246d-4cbe-8366-4d53756e1f44 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.292674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.292674Z digest=sha256:05e090e71d83797a4ae360fa79388ab8e1310cbbd025176a0878fdb5a4202fb2

Observation 348a6b4f-c389-49e8-bae9-04b0b7cfce4b · outbound

This paper cites Inference-time intervention: Eliciting truthful answers from a language model.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Inference-time intervention: Eliciting truthful answers from a language model

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.216380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.297629Z digest=sha256:14d5a0ba2ea43cb353c6ed59a3242bc0e9649c5091134769b05a87af763f4394

Observation 9a175d9f-deb9-4e8c-8ceb-07f6bee309d0 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.302336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.302336Z digest=sha256:5d3198ec730ef5d75b5962a4cabb717b6f94c3e8731b94589250e0720dd1852a

Observation 92e54623-1853-48b9-b9eb-63bc7ac3864a · outbound

This paper cites FlipAttack: Jailbreak LLMs via Flipping.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions FlipAttack: Jailbreak LLMs via Flipping

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.307047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.307047Z digest=sha256:372fe6e4c0221e61646e9e28887cc49405720f3d9e131398b7a91d69ce4701ad

Observation 1f6119c0-1d99-4d2a-98ad-af587fc43ac6 · outbound

This paper cites CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.311644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.311644Z digest=sha256:402c21e6802b1ab7486d2fd1611cb5e5c44d2fbf86414cacad1bdb7aa132b5ee

Observation 2a97a2c1-4831-4e8f-8023-7d664b488f88 · outbound

This paper cites Latent space translation via semantic alignment.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Latent space translation via semantic alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.199745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.316472Z digest=sha256:d8f70278a1d79c1fe6e272113cfa5f3644f64e36dc3198ec6f27a6aa2fb819ef

Observation 5cab35a3-1d74-436a-89d8-f36345162ce0 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.320868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.320868Z digest=sha256:bc10816a4e07c4b944bff93722478d41fcd13671e3157388b70b26dd8afc1814

Observation 19e786ec-6733-49a6-be99-61c98ac31902 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Tree of attacks: Jailbreaking black-box llms automatically

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.325772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.325772Z digest=sha256:aaa0fb7cd4a7068e399c6ddf9616443a10d4de85610c68d115738851b71764d9

Observation 064659f7-bb0b-408b-9e84-4ff66ae2cf6e · outbound

This paper cites Introducing the world's best edge models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Introducing the world's best edge models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.170884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.329996Z digest=sha256:0351a5fe11ca2b41be872f64da2566590a1e7fa16e36121574f02ea5cbdfcc12

Observation 7c76989d-7e2f-45e5-8883-e1d923dfd6fd · outbound

This paper cites Interpreting gpt: The logit lens.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Interpreting gpt: The logit lens

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.154811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.334399Z digest=sha256:81f6ade9bf31e6c204659879991f04ef973d790517d3c4f80308902dadef63c8

Observation b99e435c-d37b-4686-b7bd-a8e81b859708 · outbound

This paper cites Training language models to follow instructions with human feedback.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.339096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.339096Z digest=sha256:76b90f6b7b7af39e5cb7e6af097e0837a28060e1d2793f6d37d051f33aa6ca8e

Observation b2dc371c-e6eb-4442-9208-e6935ee776fb · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.343341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.343341Z digest=sha256:5de0e0f9f00592e16f62344cb536c4ef7e330d870d777ba4ebc073e728eabb31

Observation 242608d1-ef0b-419a-a143-43eac70ff1bc · outbound

This paper cites Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Multilingual Large Language Model: A Survey of Resources, Taxonomy and Frontiers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.347819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.347819Z digest=sha256:d4c8ecd5715c775bd13307102a4f074a9897ba70060b5c475c2ea2e44d729ce7

Observation a5d77954-225c-4b4d-b47a-08898a04088a · outbound

This paper cites D., Ermon, S., and Finn, C.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions D., Ermon, S., and Finn, C

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.352361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.352361Z digest=sha256:a391a205a963ae33ae45e1cdd073fc0fd3f440a5640d218ea3ec206b54523db6

Observation af117371-113a-4dbd-b623-17812b0a1182 · outbound

This paper cites A strongreject for empty jailbreaks, 2024.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A strongreject for empty jailbreaks, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.114182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.356677Z digest=sha256:06a1b3a56f657ae1aea7b88ecca2771a974260a4f6a834954d2c33b3276597d1

Observation 8462c8aa-9c0b-4ab3-8872-8cd0db12909b · outbound

This paper cites Mission Impossible: A Statistical Perspective on Jailbreaking LLMs.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Mission Impossible: A Statistical Perspective on Jailbreaking LLMs

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:08:02.707111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.360944Z digest=sha256:a78681e1ba98bd959a2bb125206ec500efc3c5c67da7c18a4bfea80e2164fdf1

Observation a4d21228-1af9-4c6f-b676-051099dda54b · outbound

This paper cites an unresolved cited work.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.365325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.365325Z digest=sha256:7c662a5c077f8435f1b663c63e41288132c82a71fbd562987887ae5c7c0d78ac

Observation 320e0a42-2383-4b22-ace2-29f7b06777cb · outbound

This paper cites Hermes 3 Technical Report.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Hermes 3 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.369692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.369692Z digest=sha256:066f46846cc60bd5e4df9d7185105afc6aab96b9f0d6e248c4c2d86d70977046

Observation 59c5667f-71a5-480b-8dee-e1ba52d91567 · outbound

This paper cites Detox: Toxic subspace projection for model editing.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Detox: Toxic subspace projection for model editing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.088650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.374485Z digest=sha256:eb5a466f629608fae2c392055d92836dc89ab591d872cd2f8af8acfc19c215cd

Observation 737f3945-a6cc-4a1f-b817-56b3c983d795 · outbound

This paper cites Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.378812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.378812Z digest=sha256:fa468d2e07274649826e180827b2cc5ca4c4caf0b08cd4540d9fabc35e2d1713

Observation 8b894c94-9f12-4f6c-b447-82eaea163ff6 · outbound

This paper cites a ger, T., Elstner, J., Geisler, S., Cohen-Addad, V., G \.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions a ger, T., Elstner, J., Geisler, S., Cohen-Addad, V., G \

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.383256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.383256Z digest=sha256:c542458c5d58368f9e71a1828e22eec05bf1c04504f334593ec15877364d6132

Observation 434dbf62-dba0-407d-a3db-c61502453e29 · outbound

This paper cites SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.387653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.387653Z digest=sha256:a6f5579a732b1ab327499dccb29042de9750bde1beaab41b61befb21619dfd3e

Observation 921adb6e-823d-4c74-8b2d-93922965bd24 · outbound

This paper cites Beyond toxic neurons: A mechanistic analysis of dpo for toxicity reduction.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Beyond toxic neurons: A mechanistic analysis of dpo for toxicity reduction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.073111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.392401Z digest=sha256:09acd126803d2e20d7e6646150eaec704e3473fe641a00953c39d0c6b18f4732

Observation 18e6a65f-3ff8-458c-a9f6-fd034c45081d · outbound

This paper cites A safety realignment framework via subspace-oriented model fusion for large language models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A safety realignment framework via subspace-oriented model fusion for large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.396873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.396873Z digest=sha256:056f6814a946921e723fcaf43dcdf430b9870f1382ee3e0a36760e2382c5f401

Observation ebe7a971-d1ae-4437-b7b3-138e698230a1 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Low-Resource Languages Jailbreak GPT-4

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.401728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.401728Z digest=sha256:0c0bda7b423eee60c51f52a9bdbe2265a740d91636f46752e1e3a06741f70755

Observation 9e620374-a280-466a-895f-1e0c9b06836a · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.406517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.406517Z digest=sha256:86375096b36ab29a8d67d39d29df88bdc4c4a43a2ad1af04b735cda6689c60b0

Observation d896d403-2005-443e-9c5b-fa6408a31bde · outbound

This paper cites A Survey of Large Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions A Survey of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.411449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.411449Z digest=sha256:3b0d47ad4329be8fccaf07f8151f422b72fefbaf7cf3df3bbb38ec445bb25c90

Observation 621f29b2-202f-4cd2-829c-3e80ce477a10 · outbound

This paper cites How alignment and jailbreak work: Explain llm safety through intermediate hidden states.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions How alignment and jailbreak work: Explain llm safety through intermediate hidden states

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:08:03.057177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T23:08:02.416060Z digest=sha256:8db447e09e3ab35758d384116bb04a12045caf740774fd298e26f98f0d0f0547

Observation 36b35852-d14a-4f02-82c6-ac02953c7d69 · outbound

This paper cites On the Role of Attention Heads in Large Language Model Safety.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions On the Role of Attention Heads in Large Language Model Safety

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.421540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.421540Z digest=sha256:fe50bd7ca173c8a609cab0f5d0a78ce7c985ae599ceea90b5093c63e2fe45c72

Observation 2f5f7fef-e0d1-41c1-8fba-b0f34a872051 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:08:02.426832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:08:02.426832Z digest=sha256:ac20335b436959273cd64a407353c3544a33c83d400e0645310af2f3c46c9204

Pith citing papers

Observation 63ed797c-92c4-44b6-b78d-d3a1b93a6163 · inbound

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace cites this paper.

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.812422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:32.812422Z digest=sha256:7bb1399401b07d1cca021dff2787be476a50937283731bee7005876b849f679d

Observation 3385d22a-b528-4772-8026-0aa662c4779e · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.976764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:f544d53ed9b5cb78e010b916c201c138f30bbdd361c9ae232a2f036994aa45d4

Observation b416694e-7921-4807-93b9-12890c8c63dc · inbound

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs cites this paper.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.398404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.398404Z digest=sha256:91eba43c4eff2a21d44ca4a8fcd5245bdf91104c6ae358d85f636d871494943b

Observation 70ffcc5d-2c1c-4592-ac8a-bbbb34600921 · inbound

The Geometry of Harmfulness in LLMs through Subconcept Probing cites this paper.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.822986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.822986Z digest=sha256:708ad82e3a261395768d4c44a50b414dd9ff723a5f60d744f2b2a7b962660fcf

Observation c16fe6ed-c4e6-4d00-ae74-9673c6369022 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:17:36.674517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:f7f81e0e42315929dfb28ae5503dfc3d1171f4a7c9e5b3fd5a72f13a51e079c4

Observation a35c1ace-fc4d-4df3-a944-7ca233b559b7 · inbound

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models cites this paper.

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:55.132127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:23:15.746795Z digest=sha256:4fa0bbd16fe80620c8bdbb2911a2b7b4c8d75ceffe12d30a4063d17b992b3b22

Observation 35832151-671c-438c-a24e-3c270f324cf5 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.581183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:03:10.263663Z digest=sha256:4e4ac6bb33d2005ab86f9faa0c046fb8dc9cc2e8e41744c3328fc3978f2a915f

Observation ff3a7668-6906-42ca-a745-a1b6dbb5d8c2 · inbound

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion cites this paper.

Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:12:58.880685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:12:06.989077Z digest=sha256:8297148b509c8480a5f3df93216171dc96a5e6c33a5fce061d4ab56fdd0f0e59

Observation b8a03693-8402-4396-b1c5-4cf60e4ed2c4 · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:38:00.101836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:352e52881fbc983699005f7d9225f0759d605688ec820a8d66761c46ebb2416b

Observation de65d733-af68-417c-93a9-723c6581aa4f · inbound

Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models cites this paper.

Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:43:27.385024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T01:43:22.777232Z digest=sha256:01f8227618527b9a69b285e1ad2053980b61ea5329a6bfebd492e8b9d07230cc

Observation 21c0292c-32c3-4fa8-8f9c-d2e454c28328 · inbound

Why Do Safety Guardrails Degrade Across Languages? cites this paper.

Why Do Safety Guardrails Degrade Across Languages? The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:18:21.005340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:17:41.347193Z digest=sha256:6e34bfb0bdb0c9f7f614199aedf008e04c1f51e1a56c7113d08b6523577e2628

Observation a83a1452-56b9-4d24-b7fb-d5c0ae4a39b0 · inbound

Low-Resource Safety Failures Are Action Failures, Not Representation Failures cites this paper.

Low-Resource Safety Failures Are Action Failures, Not Representation Failures The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:02:24.045703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T16:59:51.969896Z digest=sha256:d4a12de97ad62d440133e61291625d9ddbb36eaf558409398f293b5a796e8703

Observation 94f4d1d4-0e25-495f-bbea-ddcd37b40db5 · inbound

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment cites this paper.

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.875530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T14:39:11.178976Z digest=sha256:a0256bf321727f82c0acc72dcfd1d44ece9ca4b473b4d49c64f5b8428743aa7b

Observation 66214a24-5bc7-423d-93f9-55adbc683088 · inbound

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map cites this paper.

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.487293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T10:54:37.039277Z digest=sha256:20c28f7682a027df8a7df04be178ab3f44e946133ca2c87c40123636afb105d0