Pith. sign in

Paper Citation Record · LEDGER

Probing the Robustness of Large Language Models Safety to Latent Perturbations

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2506.16078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16078 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:19.347784Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 44757a17-c1ea-4f3c-ad97-f022f4781849 · outbound

This paper cites URL https://api.semanticscholar.org/CorpusID:268232499.

Probing the Robustness of Large Language Models Safety to Latent Perturbations URL https://api.semanticscholar.org/CorpusID:268232499

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:12.823540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:12.823540Z digest=sha256:370b29f6d0fa59c0b47e150c04699f905cc4eee1c51b4f56faf79613c7ba0bc5

Observation 718eee71-0ac9-4984-8b4d-27ede5b53529 · outbound

This paper cites GPT-4 Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:12.948150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:12.948150Z digest=sha256:5db1b8632af0332992ede3f77b7f75f9038d22b00f81647eaf8f5ea8c732d37a

Observation 28c35c6b-b441-4b2e-a413-138b5172816b · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.094114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.094114Z digest=sha256:5aa64f71a5dbfba3f9eb40883c9498211b69aa258ca6cc12cde19405fe5d4365

Observation ce6cad00-b61a-4956-b815-a592d136e3ce · outbound

This paper cites Refusal in language models is mediated by a single direction.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Refusal in language models is mediated by a single direction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:22.260682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:13.253737Z digest=sha256:13356e4d32e5dd74d060ceb449b0513897b5c8a8f2a73b0a0ef68d587191bdde

Observation ff69e5cd-8704-47a4-94e4-9cc84f296787 · outbound

This paper cites Qwen Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.400656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.400656Z digest=sha256:cbfa8d5b425809db2563eebebb341bea8969bb1386f0ff9330599553a8aecdda

Observation 268798a0-aa0b-4957-9ddf-679229567239 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.523748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.523748Z digest=sha256:54b5cce6def417b36435784fed888a2b6bd82f8a4ea7db8989f478673b97a14c

Observation 19916c45-3fad-4278-b18f-facb56a1446e · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.652587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.652587Z digest=sha256:e09e809cdbd88e6fa786a19bdc9f0278eb3ef84f848fcbb37e760afa20604c17

Observation 4a8cbe40-1481-4509-90a3-440ba53a2645 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Jailbreaking black box large language models in twenty queries

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.760620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.760620Z digest=sha256:da78a7bc05e0cb68c547ec07f6d8880fdcb8023e66c1483a403602cfa4c99774

Observation 8db9b532-ee68-48e0-95c5-0aa7879e17d3 · outbound

This paper cites Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.842722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.842722Z digest=sha256:a261f93ca75f606b0a94901922fef873729091d3db202b12753277373a50a8b0

Observation 52f4463e-e514-47a3-8380-78bdbb47f8c5 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.943741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.943741Z digest=sha256:ce2b9f7ee0b6e2c6ec249a1ef5ce5734c81a226f5b5ece26461cc590b4951153

Observation 3e31ae40-8794-4330-9ebb-36b06ab76392 · outbound

This paper cites Scaling Laws for Adversarial Attacks on Language Model Activations.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Scaling Laws for Adversarial Attacks on Language Model Activations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.075962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.075962Z digest=sha256:f2fd9f8d8ee72ca5a5e3cc681d244b1cfca0c686dae38a36885eaf1e0bb19093

Observation 0dac8ef6-6477-455b-bf29-81049e1406ef · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.203220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.203220Z digest=sha256:034849c1b3e7109d7de38be86be31bd33df41de475e1213b9ab89d8ad28710a2

Observation ec9bbdf3-572d-4998-909e-05f08a5733e8 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.339131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.339131Z digest=sha256:4ba28f6de1ec9044ffa0b0b99dae67b34ab874310beb142f45f9b30ca2d3678a

Observation 54ed4062-b926-4749-9107-c588f9489a4a · outbound

This paper cites The Llama 3 Herd of Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.419356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.419356Z digest=sha256:23b091e8a03cae341e696cabbccafd3a83126e3101b04ea1d1fc96c594bee66a

Observation 08f40f92-7bcb-4bdd-900e-d092364a26b1 · outbound

This paper cites MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts.

Probing the Robustness of Large Language Models Safety to Latent Perturbations MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.540498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.540498Z digest=sha256:404521935a6b019ecf96877b612d7b74f41147577783691da86627bbfe5b06e6

Observation 4a44650f-67a2-4463-85e3-d59c58693d32 · outbound

This paper cites Flames: Benchmarking Value Alignment of LLMs in Chinese.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Flames: Benchmarking Value Alignment of LLMs in Chinese

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.676174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.676174Z digest=sha256:6763bb64a6f38eb4eec6575da8e8ba7ab5b2adf11f227b65b9b4eeefbfb29980

Observation 95220e44-bdb4-4e09-8c2f-e2c4eba63848 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Arbitrary style transfer in real-time with adaptive instance normalization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:22.138710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:14.798173Z digest=sha256:137c3a2d6a26ce72c6aeab89f8b5f640122f8fe55fb4ecefef1815cc94b9cdbe

Observation cf554321-a7f5-4f89-8968-71a2fe8805f5 · outbound

This paper cites Improving Activation Steering in Language Models with Mean-Centring.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Improving Activation Steering in Language Models with Mean-Centring

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.924023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.924023Z digest=sha256:a7bdf1e5c8009ef419e5869cfeca9965c2918f82085b6aef39992c63705f73d6

Observation e160b3f4-483a-4f70-8c09-475364f81c59 · outbound

This paper cites Large language model unlearning via embedding-corrupted prompts.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Large language model unlearning via embedding-corrupted prompts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:22.036639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.073777Z digest=sha256:ab398066ef23a057109847afce35c3db6fbb61ea2277f16f7784da4d825d220c

Observation 7e486610-471a-4941-b6e0-fd8eec956859 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.808465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.178054Z digest=sha256:fcf16bc6b127087a4195c7df3baf8f8803f30ba6f4434dc33d8c6b59345284f9

Observation ff108060-af99-4009-a9d4-880f6fa53ce0 · outbound

This paper cites Merge to learn: Efficiently adding skills to language models with model merging.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Merge to learn: Efficiently adding skills to language models with model merging

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.570220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.305379Z digest=sha256:347991f0d97003fe0899cb6e76f31075ee6d4fdda9dc35e3d64860d7cc984d5b

Observation faeb8a01-9c95-44ae-9428-54e50b471d98 · outbound

This paper cites Training language models to follow instructions with human feedback.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Training language models to follow instructions with human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:15.426561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:15.426561Z digest=sha256:6b22d345d291eead4baad2c1c7fd27a5ba4d55bff38bd6972ce67b31da2396a1

Observation 76f0fe94-24e7-4577-a589-7cde1d0a3520 · outbound

This paper cites Training language models to follow instructions with human feedback.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Training language models to follow instructions with human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.306450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.590107Z digest=sha256:fe2af7268464f3315191ad8f00fc391fb37318ff607739456ae41737127f0b89

Observation 28760bb7-0e67-4836-bdd1-37e38d88071c · outbound

This paper cites In-context unlearning: Language models as few-shot unlearners.

Probing the Robustness of Large Language Models Safety to Latent Perturbations In-context unlearning: Language models as few-shot unlearners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.149432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.752562Z digest=sha256:3228e1556d44a87f6fef94e4739df08c104c05d3c38aa1e498c7963caf864cc1

Observation f0ff625d-af38-4182-b385-564958386640 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:15.862324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:15.862324Z digest=sha256:d5662332f82572e443a8b76f25ddff1b59a08d413e1a20135e3ae5464ceff9cf

Observation af8f1517-15e1-4b94-a4e7-822f29993113 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.999488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.981565Z digest=sha256:9255b83112669090890bb9d4835f9cd16ca8289e2ad4ac7d4e015bd89a6c99f8

Observation abb051b1-11fc-4560-94a1-5d911f89b1f6 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.090709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.090709Z digest=sha256:b6aa96738c22962bb980fd0e662ab40575917c078aa7fc34842656ce7769716b

Observation 8e1364f3-4ddd-4670-9358-11f181162778 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Direct preference optimization: Your language model is secretly a reward model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.871420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:16.197393Z digest=sha256:d73be87b6d5d72ba02ee215935a689eb35681c3cc65792eb58a83a1dea012f0f

Observation 99a40b3f-4945-4de1-ab35-7ff78bbae000 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Steering llama 2 via contrastive activation addition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.288237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.288237Z digest=sha256:23745171f1ef65b42dbbb58260547716386f63ba4d4686aec97c396756eebe4e

Observation 8f44eefa-c8fa-4b91-9c38-ee6bf5e3db3c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.388217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.388217Z digest=sha256:fae341960592b8d5c6b77fa74572e0e6dd4714fe6f968449690275486d366538

Observation 48d5fb51-d013-41a7-8ec0-1afef777e68a · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.506515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.506515Z digest=sha256:9635ad4578c51b8eddb7cd4df2e79d1a5a1e166b1dfb34c6c8342954ae81760d

Observation ef3f9bf9-ed9b-4564-8d7e-a651e0a17486 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

Probing the Robustness of Large Language Models Safety to Latent Perturbations C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.624391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.624391Z digest=sha256:c6ff2e9bb73fbfc14fbe3171847196111d51aff34608e7684cc604a767dc9e38

Observation 96943394-bc7e-41fe-ac19-93d292110571 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen2.5: A party of foundation models, September 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.769634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.769634Z digest=sha256:a042594afa8e7b14bcdebb31039829dcd6d1f7dd1edffbf93edc271d61a6eb32

Observation 56a5d664-4a78-4c63-90de-f10015228f14 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.926076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.926076Z digest=sha256:11666ec87cc4482c41bc9b2cb21ab8c1e70e63fc74ec79f081007a657304bec7

Observation 6398be82-68be-4b45-92bb-41b031c6e0d0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.044926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.044926Z digest=sha256:be980e57432212989c5466497b7c1a5ada23b495c1784cbb894c2a22e96d6210

Observation 8a0367ca-accc-42e2-92cc-c852408c0b1b · outbound

This paper cites Steering Language Models With Activation Engineering.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Steering Language Models With Activation Engineering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.179313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.179313Z digest=sha256:0e723f3439baba912f6e9381aa83c1ebcc01edf3723a3f8887802385609b8026

Observation 632f255a-7e6b-4bd0-a1cb-80707deb3485 · outbound

This paper cites A Language Model's Guide Through Latent Space.

Probing the Robustness of Large Language Models Safety to Latent Perturbations A Language Model's Guide Through Latent Space

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.308168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.308168Z digest=sha256:270de62efcb1f74c2054eeb49523a62a5994273988986fdf8dc25c17213b5637

Observation 7b6adedc-a47d-4cdd-96cd-7504bbb58849 · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.484246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.484246Z digest=sha256:e92138b5669e0ce00ca4313fed74e32dea93190231bb32bf2e76518b6e8b7616

Observation 5491bab7-e1c9-4d1d-8867-b51f0d8c7339 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Finetuned Language Models Are Zero-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.692465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.692465Z digest=sha256:599105a62553d5c92f23b4d56ef7bfcda02a75c41ab3dfef652d907972651f86

Observation 33c06365-89a2-47d4-9275-e0a54e7cb143 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Robust fine-tuning of zero-shot models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.692151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:17.802675Z digest=sha256:f0f395d9dc024fffc57f746adce6acbfde7f54565a2f75a559b283c19fbb1998

Observation 7fb5c9af-9b3f-40a6-835c-646d0740e914 · outbound

This paper cites Uncovering safety risks of large language models through concept activation vector.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Uncovering safety risks of large language models through concept activation vector

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.511936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:17.928350Z digest=sha256:03129c2dcbf12b429e277869b7cb974725e5bfd0cbbff06a8e14fdc10b6a09ba

Observation 167be9c9-28bd-4a7b-a9d7-4d51f7b4efb7 · outbound

This paper cites Qwen2 Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen2 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.076065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.076065Z digest=sha256:041e4790e79dceb112bc4cd10bbfc0b417fc280eda9f831c795041b549fc689e

Observation 49f6e146-811b-4ee3-bf4c-6b1e275cff48 · outbound

This paper cites Qwen3 Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.222192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.222192Z digest=sha256:ef4399a9a7e1143ce96b0c88edfced15d4948b47759ed3b2fbc04063aa3e62e7

Observation 6800cdc0-9bb0-4750-9704-262036744b00 · outbound

This paper cites A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos.

Probing the Robustness of Large Language Models Safety to Latent Perturbations A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.339304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.339304Z digest=sha256:115e4bf142f9ef426ce53bd3831461c978781c2ffa45f42293a343528894eab8

Observation d59debf4-233c-4fbc-9f42-52a40287bbc8 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Yi: Open Foundation Models by 01.AI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.489339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.489339Z digest=sha256:aafdc34299165dc1274854eaf9f5ae9702ecd96d982d9b1b689de00674bc67c8

Observation c9fe87b3-afc9-4e0e-af64-f0dea814d388 · outbound

This paper cites Removing rlhf protections in gpt-4 via fine-tuning.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Removing rlhf protections in gpt-4 via fine-tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.370575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:18.646068Z digest=sha256:3afaefc1b398b222130cbf3afc3d2e9658889857989ac06fc0a698f5bb9b94c0

Observation 4def3251-ed86-4529-ae70-b1c3d50e2402 · outbound

This paper cites Controlling large language models through concept activation vectors.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Controlling large language models through concept activation vectors

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.227249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T23:49:18.769725Z digest=sha256:7097ab4a4df0e251f0003996cb50f5ef638f65ddbfb5dfe92a78153823078b7e

Observation a6beb801-d181-4c28-9f4d-f4a3f64d9911 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.890265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.890265Z digest=sha256:3a571b1484dc7955562ba55e396b3f794edc338f8cec6deac91ccb765fcf9a5d

Observation 241aea52-f36e-4245-b5bf-9d3cdec25e84 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.031485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.031485Z digest=sha256:af0f58bfb925afad4cf22d1cb8ded1c6a79d29ad5369d8c671d1d2ce89170591

Observation fb83993d-61ca-448b-af60-7ad1dffd570c · outbound

This paper cites write newline.

Probing the Robustness of Large Language Models Safety to Latent Perturbations write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.162518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.162518Z digest=sha256:2ee508b7fc65f72da4942f135a036b8e25f06f7e77be033ed4819bb9dd74a8e2

Observation 082b78b0-2268-469f-b2df-6d916aae5b3b · outbound

This paper cites @esa (Ref.

Probing the Robustness of Large Language Models Safety to Latent Perturbations @esa (Ref

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.235253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.235253Z digest=sha256:b143f9db25949d39350126efbeded0ef17b34e1093b392ff700d5c111595c9bf

Observation aa129338-5c1c-4bc9-8ebd-ef596cbd1265 · outbound

This paper cites an unresolved cited work.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.294952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.294952Z digest=sha256:47acea01a6bde2ff0d3e653f15913f9b9dba4dc9642b7f1a0d94cdd11778a5b2

Observation 4bfd156c-ad65-42cc-85e1-1b0b4660c411 · outbound

This paper cites Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.347784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.347784Z digest=sha256:4f03738f09a0056bb8a75eb90881e727194226b6d945fa65e1e833b6318ae5d0

Pith citing papers

Observation 02cf7c9a-d832-48d7-b219-06c3e5ffc297 · inbound

The Impact of Off-Policy Training Data on Probe Generalisation cites this paper.

The Impact of Off-Policy Training Data on Probe Generalisation Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.710099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:0ca77b1bb00c4c2c5745be4889a196e7573c2ff6893957f3f7a27ad9a4ae713c

Observation 886bd8ff-0cd4-488a-9940-59e25931abd6 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:51:23.072108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T04:47:47.214726Z digest=sha256:5a1ad7aba44e457159e66186341e1236de3bb337546b63c4ace1899268bd11ec

Observation a18f57df-b40a-408d-8aba-c97a249e2be6 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:40.576154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T17:06:22.481341Z digest=sha256:56e2c22ad02c24dbf78f7f3e6892288d4a07445aefb2c9b0cce9c1a6ee756b1c

Observation 9d59673f-1414-478d-b516-68d1df279595 · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.963521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:35cd518767d5944788de6b2b2f463938c6620698ccf8d0cf4075afef50d4e3ee