Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:19.347784Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2506.16078.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:19.347784Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
53 of 53 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 44757a17-c1ea-4f3c-ad97-f022f4781849 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations URL https://api.semanticscholar.org/CorpusID:268232499
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718eee71-0ac9-4984-8b4d-27ede5b53529 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c35c6b-b441-4b2e-a413-138b5172816b · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6cad00-b61a-4956-b815-a592d136e3ce · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Refusal in language models is mediated by a single direction
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff69e5cd-8704-47a4-94e4-9cc84f296787 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 268798a0-aa0b-4957-9ddf-679229567239 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19916c45-3fad-4278-b18f-facb56a1446e · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Defending Against Unforeseen Failure Modes with Latent Adversarial Training
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a8cbe40-1481-4509-90a3-440ba53a2645 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Jailbreaking black box large language models in twenty queries
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8db9b532-ee68-48e0-95c5-0aa7879e17d3 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f4463e-e514-47a3-8380-78bdbb47f8c5 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e31ae40-8794-4330-9ebb-36b06ab76392 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Scaling Laws for Adversarial Attacks on Language Model Activations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dac8ef6-6477-455b-bf29-81049e1406ef · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec9bbdf3-572d-4998-909e-05f08a5733e8 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Explaining and Harnessing Adversarial Examples
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ed4062-b926-4749-9107-c588f9489a4a · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f40f92-7bcb-4bdd-900e-d092364a26b1 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a44650f-67a2-4463-85e3-d59c58693d32 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Flames: Benchmarking Value Alignment of LLMs in Chinese
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95220e44-bdb4-4e09-8c2f-e2c4eba63848 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Arbitrary style transfer in real-time with adaptive instance normalization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf554321-a7f5-4f89-8968-71a2fe8805f5 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Improving Activation Steering in Language Models with Mean-Centring
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e160b3f4-483a-4f70-8c09-475364f81c59 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Large language model unlearning via embedding-corrupted prompts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e486610-471a-4941-b6e0-fd8eec956859 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Autodan: Generating stealthy jailbreak prompts on aligned large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff108060-af99-4009-a9d4-880f6fa53ce0 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Merge to learn: Efficiently adding skills to language models with model merging
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation faeb8a01-9c95-44ae-9428-54e50b471d98 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Training language models to follow instructions with human feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76f0fe94-24e7-4577-a589-7cde1d0a3520 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28760bb7-0e67-4836-bdd1-37e38d88071c · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations In-context unlearning: Language models as few-shot unlearners
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f0ff625d-af38-4182-b385-564958386640 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af8f1517-15e1-4b94-a4e7-822f29993113 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation abb051b1-11fc-4560-94a1-5d911f89b1f6 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e1364f3-4ddd-4670-9358-11f181162778 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Direct preference optimization: Your language model is secretly a reward model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 99a40b3f-4945-4de1-ab35-7ff78bbae000 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Steering llama 2 via contrastive activation addition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f44eefa-c8fa-4b91-9c38-ee6bf5e3db3c · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Proximal Policy Optimization Algorithms
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d5fb51-d013-41a7-8ec0-1afef777e68a · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef3f9bf9-ed9b-4564-8d7e-a651e0a17486 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations C ommonsense QA : A question answering challenge targeting commonsense knowledge
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96943394-bc7e-41fe-ac19-93d292110571 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen2.5: A party of foundation models, September 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a5d664-4a78-4c63-90de-f10015228f14 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6398be82-68be-4b45-92bb-41b031c6e0d0 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a0367ca-accc-42e2-92cc-c852408c0b1b · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Steering Language Models With Activation Engineering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 632f255a-7e6b-4bd0-a1cb-80707deb3485 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations A Language Model's Guide Through Latent Space
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b6adedc-a47d-4cdd-96cd-7504bbb58849 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5491bab7-e1c9-4d1d-8867-b51f0d8c7339 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Finetuned Language Models Are Zero-Shot Learners
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c06365-89a2-47d4-9275-e0a54e7cb143 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Robust fine-tuning of zero-shot models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7fb5c9af-9b3f-40a6-835c-646d0740e914 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Uncovering safety risks of large language models through concept activation vector
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 167be9c9-28bd-4a7b-a9d7-4d51f7b4efb7 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen2 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f6e146-811b-4ee3-bf4c-6b1e275cff48 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen3 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6800cdc0-9bb0-4750-9704-262036744b00 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59debf4-233c-4fbc-9f42-52a40287bbc8 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Yi: Open Foundation Models by 01.AI
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fe87b3-afc9-4e0e-af64-f0dea814d388 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Removing rlhf protections in gpt-4 via fine-tuning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4def3251-ed86-4529-ae70-b1c3d50e2402 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Controlling large language models through concept activation vectors
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6beb801-d181-4c28-9f4d-f4a3f64d9911 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Representation Engineering: A Top-Down Approach to AI Transparency
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241aea52-f36e-4245-b5bf-9d3cdec25e84 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb83993d-61ca-448b-af60-7ad1dffd570c · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations write newline
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 082b78b0-2268-469f-b2df-6d916aae5b3b · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations @esa (Ref
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa129338-5c1c-4bc9-8ebd-ef596cbd1265 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfd156c-ad65-42cc-85e1-1b0b4660c411 · outbound
Probing the Robustness of Large Language Models Safety to Latent Perturbations Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02cf7c9a-d832-48d7-b219-06c3e5ffc297 · inbound
The Impact of Off-Policy Training Data on Probe Generalisation Probing the Robustness of Large Language Models Safety to Latent Perturbations
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 886bd8ff-0cd4-488a-9940-59e25931abd6 · inbound
Do Linear Probes Generalize Better in Persona Coordinates? Probing the Robustness of Large Language Models Safety to Latent Perturbations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a18f57df-b40a-408d-8aba-c97a249e2be6 · inbound
Do Linear Probes Generalize Better in Persona Coordinates? Probing the Robustness of Large Language Models Safety to Latent Perturbations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d59673f-1414-478d-b516-68d1df279595 · inbound
Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Probing the Robustness of Large Language Models Safety to Latent Perturbations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.