Pith. sign in

Paper Citation Record · LEDGER

Probing the Robustness of Large Language Models Safety to Latent Perturbations

As of 17 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2506.16078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16078 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:19.347784Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T03:24:24.714121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 44757a17-c1ea-4f3c-ad97-f022f4781849 · outbound

This paper cites URL https://api.semanticscholar.org/CorpusID:268232499.

Probing the Robustness of Large Language Models Safety to Latent Perturbations URL https://api.semanticscholar.org/CorpusID:268232499

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:12.823540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:12.823540Z digest=sha256:3f5828da7103a6509ad20acc1baf6bfb224d7b2423c79a7f48e2e9d41bee9b8a

Observation 718eee71-0ac9-4984-8b4d-27ede5b53529 · outbound

This paper cites GPT-4 Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:12.948150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:12.948150Z digest=sha256:ee44cf1cb38b48377a279b1d9b9601a8bad7d1bac7b9015dc43a666b1be7b778

Observation 28c35c6b-b441-4b2e-a413-138b5172816b · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.094114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.094114Z digest=sha256:b2822cbcca7e25119e500f7bf4a286db521b7fd7a6b664590c70daf8d1f51c1b

Observation ce6cad00-b61a-4956-b815-a592d136e3ce · outbound

This paper cites Refusal in language models is mediated by a single direction.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Refusal in language models is mediated by a single direction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:22.260682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:13.253737Z digest=sha256:516f434c11da5940867322779bcfae0cf821646560a88dbe9c5b0b30d5592dc0

Observation ff69e5cd-8704-47a4-94e4-9cc84f296787 · outbound

This paper cites Qwen Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.400656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.400656Z digest=sha256:7c8d80abb5e5d2d86fef3ac7f5faf04d8344a8d7fe2a74843dab80a1464e7617

Observation 268798a0-aa0b-4957-9ddf-679229567239 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.523748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.523748Z digest=sha256:e1ba3993310b41ae02db7ef27b1157262e85795c027dcda503bf7353b48e4e2d

Observation 19916c45-3fad-4278-b18f-facb56a1446e · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.652587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.652587Z digest=sha256:891240eefbc3a31eaddbbc09d1f5d323f95a5359b554a1abba9334c75063598d

Observation 4a8cbe40-1481-4509-90a3-440ba53a2645 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Jailbreaking black box large language models in twenty queries

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.760620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.760620Z digest=sha256:442ee1cd451d87272f84466941b4cd8f1f445302048d11506b2a2deb4e108ad6

Observation 8db9b532-ee68-48e0-95c5-0aa7879e17d3 · outbound

This paper cites Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.842722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.842722Z digest=sha256:8cc29ad03a7c39b54aed371d19d153a923636114c01058169dce92dcda10bee8

Observation 52f4463e-e514-47a3-8380-78bdbb47f8c5 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:13.943741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:13.943741Z digest=sha256:687fc402f19e90c73e8a5e3e58456a5b518218f71e95919388d8cff9a2bbf4f2

Observation 3e31ae40-8794-4330-9ebb-36b06ab76392 · outbound

This paper cites Scaling Laws for Adversarial Attacks on Language Model Activations.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Scaling Laws for Adversarial Attacks on Language Model Activations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.075962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.075962Z digest=sha256:3ddf0d55e5f35a05349ca26a119819a5f54aefd0c8e771162f5c30a3c7491310

Observation 0dac8ef6-6477-455b-bf29-81049e1406ef · outbound

This paper cites Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.203220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.203220Z digest=sha256:eea712350c53d6ac3d5ff960cefbefc7d3da66b171c457769a837184aa899e31

Observation ec9bbdf3-572d-4998-909e-05f08a5733e8 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.339131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.339131Z digest=sha256:b50ab3f190ea637b2bbf418a4927e0d40a0c991a3fffcbe2f78d3baf987afb18

Observation 54ed4062-b926-4749-9107-c588f9489a4a · outbound

This paper cites The Llama 3 Herd of Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.419356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.419356Z digest=sha256:0eb886a56d7f6e44681e340d310db4eeb47cff07179a53b2bebd9470791e1701

Observation 08f40f92-7bcb-4bdd-900e-d092364a26b1 · outbound

This paper cites MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts.

Probing the Robustness of Large Language Models Safety to Latent Perturbations MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.540498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.540498Z digest=sha256:3da7ab59caa15dc66cb0a64f77605993d90acda1012992a33545b5f96f33e4f3

Observation 4a44650f-67a2-4463-85e3-d59c58693d32 · outbound

This paper cites Flames: Benchmarking Value Alignment of LLMs in Chinese.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Flames: Benchmarking Value Alignment of LLMs in Chinese

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.676174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.676174Z digest=sha256:87e057279f973e5238248a037be9fabd5752ad03b7ee2cc68d5992aa04f0d8dc

Observation 95220e44-bdb4-4e09-8c2f-e2c4eba63848 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Arbitrary style transfer in real-time with adaptive instance normalization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:22.138710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:14.798173Z digest=sha256:db2758ab7d5436046ed9a59d78bf4629ea5d33ae9cb105ccd767d9d3c3ebcdfe

Observation cf554321-a7f5-4f89-8968-71a2fe8805f5 · outbound

This paper cites Improving Activation Steering in Language Models with Mean-Centring.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Improving Activation Steering in Language Models with Mean-Centring

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.924023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:14.924023Z digest=sha256:59787bcf0414202010402d9374ea4fdcf118d43039f13ae55878199aec03ee25

Observation e160b3f4-483a-4f70-8c09-475364f81c59 · outbound

This paper cites Large language model unlearning via embedding-corrupted prompts.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Large language model unlearning via embedding-corrupted prompts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:22.036639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.073777Z digest=sha256:f3cca68adffe0cda36d8f664c31a6e09dbe0de635dbf878105c15644d320ea4b

Observation 7e486610-471a-4941-b6e0-fd8eec956859 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.808465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.178054Z digest=sha256:7fd38c4e9dcdaa9806289cc5cfa5169749d59f80ea957df63f92499646251370

Observation ff108060-af99-4009-a9d4-880f6fa53ce0 · outbound

This paper cites Merge to learn: Efficiently adding skills to language models with model merging.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Merge to learn: Efficiently adding skills to language models with model merging

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.570220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.305379Z digest=sha256:dcf708f0ea900ebd28261ee9e4382582aad7c632606cf1fb6748416c89247ca9

Observation faeb8a01-9c95-44ae-9428-54e50b471d98 · outbound

This paper cites Training language models to follow instructions with human feedback.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Training language models to follow instructions with human feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:15.426561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:15.426561Z digest=sha256:a385f4d7c2792727b0449b31bb12f9a5cd6407b853af44ffcfa26d9fdb7fb4c3

Observation 76f0fe94-24e7-4577-a589-7cde1d0a3520 · outbound

This paper cites Training language models to follow instructions with human feedback.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Training language models to follow instructions with human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.306450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.590107Z digest=sha256:df7b464a2551363f0498c4bf9fbc6ce104480484b825834780c9533501573b7e

Observation 28760bb7-0e67-4836-bdd1-37e38d88071c · outbound

This paper cites In-context unlearning: Language models as few-shot unlearners.

Probing the Robustness of Large Language Models Safety to Latent Perturbations In-context unlearning: Language models as few-shot unlearners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:21.149432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.752562Z digest=sha256:3f5dba7eb090688b944e3703db58d1e743e3305bba084397bf6a1cafb03b7354

Observation f0ff625d-af38-4182-b385-564958386640 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:15.862324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:15.862324Z digest=sha256:b70fdc3159a62cc6f2fd81b0374181ba186b0a2a1bc661f8a72358c1123ea174

Observation af8f1517-15e1-4b94-a4e7-822f29993113 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.999488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:15.981565Z digest=sha256:3fe7e49adb5208def376fea8f0b4dfb4dbde95d07e0c2e920571f5f52961c85d

Observation abb051b1-11fc-4560-94a1-5d911f89b1f6 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.090709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.090709Z digest=sha256:f98d79326cc2ea04d48576887d7a1ff82e3322321746ba7538a1fb6379fa24cc

Observation 8e1364f3-4ddd-4670-9358-11f181162778 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Direct preference optimization: Your language model is secretly a reward model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.871420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:16.197393Z digest=sha256:57c981965d12241d8a1b966c637bf842a4efb3bfdc1b6e912bb6045b815d51cd

Observation 99a40b3f-4945-4de1-ab35-7ff78bbae000 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Steering llama 2 via contrastive activation addition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.288237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.288237Z digest=sha256:b390a52f1bf48605142789b67b23b2fe7ecd030a460c356092fea84b2c3181d3

Observation 8f44eefa-c8fa-4b91-9c38-ee6bf5e3db3c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.388217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.388217Z digest=sha256:352b682b98d3c6c2249e9f09a1b125f8774fb62227e575e9036b0bf262a4b9f3

Observation 48d5fb51-d013-41a7-8ec0-1afef777e68a · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.506515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.506515Z digest=sha256:cbaf6872bf177f9e20825261374d56fcdf9c0427cc71b5e01fdbef414fb6d67d

Observation ef3f9bf9-ed9b-4564-8d7e-a651e0a17486 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

Probing the Robustness of Large Language Models Safety to Latent Perturbations C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.624391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.624391Z digest=sha256:3439cffba1c2211b0793d25772701131ae794291b3b42aeff4276ba1d89c40f1

Observation 96943394-bc7e-41fe-ac19-93d292110571 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen2.5: A party of foundation models, September 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.769634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.769634Z digest=sha256:4c42d5d69e7151f8e198b3ae9828152a42bbddc2440877dd8a05698be3ca1061

Observation 56a5d664-4a78-4c63-90de-f10015228f14 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:16.926076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:16.926076Z digest=sha256:7b8deaa1d9ce8d5d4f41c88a4469abd37baf1acf7c73bd23f0b0c4aa798ba3c7

Observation 6398be82-68be-4b45-92bb-41b031c6e0d0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.044926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.044926Z digest=sha256:52d0d82e01f403e365b0c364408efbc17cc7cbe2ede1664add0c61cd3cf75a25

Observation 8a0367ca-accc-42e2-92cc-c852408c0b1b · outbound

This paper cites Steering Language Models With Activation Engineering.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Steering Language Models With Activation Engineering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.179313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.179313Z digest=sha256:d96e827643e0bf1e90c64d5043798b22a5d909229b649fb6193e9bb1eb67e27d

Observation 632f255a-7e6b-4bd0-a1cb-80707deb3485 · outbound

This paper cites A Language Model's Guide Through Latent Space.

Probing the Robustness of Large Language Models Safety to Latent Perturbations A Language Model's Guide Through Latent Space

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.308168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.308168Z digest=sha256:0b5498f35c7ba0642032a0f1972a73a75768bb6fb63a6e423a3eee822a49a2d2

Observation 7b6adedc-a47d-4cdd-96cd-7504bbb58849 · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.484246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.484246Z digest=sha256:516323755d136d0b764fd55b2abc38acf468d6a619fbf503000d22a921a43119

Observation 5491bab7-e1c9-4d1d-8867-b51f0d8c7339 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Finetuned Language Models Are Zero-Shot Learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:17.692465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:17.692465Z digest=sha256:3332b637449a8cf60733c2f5055c6a2fa3e17f56818e3b12a3fc2f11ef000e99

Observation 33c06365-89a2-47d4-9275-e0a54e7cb143 · outbound

This paper cites Robust fine-tuning of zero-shot models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Robust fine-tuning of zero-shot models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.692151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:17.802675Z digest=sha256:6d88d2bdd16a36589f5f4137cae348b30467aaf0831a299553642b3eab18c9cf

Observation 7fb5c9af-9b3f-40a6-835c-646d0740e914 · outbound

This paper cites Uncovering safety risks of large language models through concept activation vector.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Uncovering safety risks of large language models through concept activation vector

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.511936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:17.928350Z digest=sha256:83039a5aad77a05a67eaae104ed651b522c17add517fe5096fff691b80ee1095

Observation 167be9c9-28bd-4a7b-a9d7-4d51f7b4efb7 · outbound

This paper cites Qwen2 Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen2 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.076065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.076065Z digest=sha256:87535502061bc30aa5e64cc93434e85e17b4516763f66c993cb8764a1fceaaad

Observation 49f6e146-811b-4ee3-bf4c-6b1e275cff48 · outbound

This paper cites Qwen3 Technical Report.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Qwen3 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.222192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.222192Z digest=sha256:15b58252fc2fcdde1e791d91e08afc16d91c5b23a8fa020fc25117bd73eac374

Observation 6800cdc0-9bb0-4750-9704-262036744b00 · outbound

This paper cites A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos.

Probing the Robustness of Large Language Models Safety to Latent Perturbations A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.339304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.339304Z digest=sha256:ae4ad9bbff80ab2293f104b352067a701283ab34b4df2d3374fe16237a0de7b5

Observation d59debf4-233c-4fbc-9f42-52a40287bbc8 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Yi: Open Foundation Models by 01.AI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.489339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.489339Z digest=sha256:efb106b73eb89da83780cd95edd382af181ae21dc9b2ceaa7898c267df482c6a

Observation c9fe87b3-afc9-4e0e-af64-f0dea814d388 · outbound

This paper cites Removing rlhf protections in gpt-4 via fine-tuning.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Removing rlhf protections in gpt-4 via fine-tuning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.370575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:18.646068Z digest=sha256:a36ca57f35686acfacb8823cb0415e7a6026becc9d28b7eda3efaf309fa8c242

Observation 4def3251-ed86-4529-ae70-b1c3d50e2402 · outbound

This paper cites Controlling large language models through concept activation vectors.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Controlling large language models through concept activation vectors

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:20.227249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T23:49:18.769725Z digest=sha256:91e3f50d98f2b955df67370e191b03b8e35b6526a948c049c7e48c66dac5f3b2

Observation a6beb801-d181-4c28-9f4d-f4a3f64d9911 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Representation Engineering: A Top-Down Approach to AI Transparency

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:18.890265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:18.890265Z digest=sha256:039f05d9a846cee1579e1766d1ccfc8eec6a73fda3baf96bf1b136be9521721b

Observation 241aea52-f36e-4245-b5bf-9d3cdec25e84 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.031485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.031485Z digest=sha256:0236bb0523e16572ece6c71fae83f1c6461603059095192659c9d8ec0801add1

Observation fb83993d-61ca-448b-af60-7ad1dffd570c · outbound

This paper cites write newline.

Probing the Robustness of Large Language Models Safety to Latent Perturbations write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.162518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.162518Z digest=sha256:042acc7282620f65a4caee12053ad49ee611380f13f85e9f5dbb5248843ff526

Observation 082b78b0-2268-469f-b2df-6d916aae5b3b · outbound

This paper cites @esa (Ref.

Probing the Robustness of Large Language Models Safety to Latent Perturbations @esa (Ref

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.235253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.235253Z digest=sha256:d12334655968fa1cb2ef483b399211ca76edc825dc363f66cae5d5fca080b657

Observation aa129338-5c1c-4bc9-8ebd-ef596cbd1265 · outbound

This paper cites an unresolved cited work.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.294952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.294952Z digest=sha256:1c45c11fc6383915f2ebfcd6a7397691e81ebed0cdf4a0cd11efda3bf1a52d80

Observation 4bfd156c-ad65-42cc-85e1-1b0b4660c411 · outbound

This paper cites Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models.

Probing the Robustness of Large Language Models Safety to Latent Perturbations Despite significant advances in safety alignment, we observe that minor latent shifts can still trigger unsafe responses in aligned models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:19.347784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:19.347784Z digest=sha256:a77ba128ae42b6b33a44672412a1d27b0b170674913024774599f6f129f93252

Pith citing papers

Observation 02cf7c9a-d832-48d7-b219-06c3e5ffc297 · inbound

The Impact of Off-Policy Training Data on Probe Generalisation cites this paper.

The Impact of Off-Policy Training Data on Probe Generalisation Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.710099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:e77c31a931780e87b9d37597f0041b9f838008506e8a790340493bea66b2ed79

Observation 886bd8ff-0cd4-488a-9940-59e25931abd6 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:51:23.072108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T04:47:47.214726Z digest=sha256:2575b82d33112d1fb951cfbc4e03bc012eef44696a2a47a9b8c62ac3e071293a

Observation a18f57df-b40a-408d-8aba-c97a249e2be6 · inbound

Do Linear Probes Generalize Better in Persona Coordinates? cites this paper.

Do Linear Probes Generalize Better in Persona Coordinates? Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:40.576154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T17:06:22.481341Z digest=sha256:623fcba308f5af59d52eb0c2b8cd1709f957856494ae51e83d97345c99fd42ed

Observation 9d59673f-1414-478d-b516-68d1df279595 · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.963521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:7b21526d1eba02c62d952fbb8df390ad9bb64cd883726015fd5af0d8b51a7164