Pith. sign in

Paper Citation Record · LEDGER

Obfuscated Activations Bypass LLM Latent-Space Defenses

As of 15 August 2026, this Paper Citation Record lists 100 of 114 outbound references and 12 inbound Pith citation observations for arXiv:2412.09565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09565 v2

Coverage vector

measured 100 of 114 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:59:11.757096Z

measured 112 of 112 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:30.275006Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 114 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 68e024db-1330-4e5a-bfcf-67b34c5a4177 · outbound

This paper cites Adversarial Example Detection Using Latent Neighborhood Graph.

Obfuscated Activations Bypass LLM Latent-Space Defenses Adversarial Example Detection Using Latent Neighborhood Graph

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.405016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.405016Z digest=sha256:0a1fea915121440e45551f7e0b44e0c657951f6ad93775b1014123b9c084c6d1

Observation a15f435f-53d3-4836-9e1b-7165bb085139 · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

Obfuscated Activations Bypass LLM Latent-Space Defenses Understanding intermediate layers using linear classifier probes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.409609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.409609Z digest=sha256:4b1b293e873e888a923f35e83222d39b2e73f181c5664239b21016858515a6cb

Observation cf1a5bce-f3b6-451a-ad7b-0dceb4d66add · outbound

This paper cites Trace and Detect Adversarial Attacks on CNNs Using Feature Response Maps.

Obfuscated Activations Bypass LLM Latent-Space Defenses Trace and Detect Adversarial Attacks on CNNs Using Feature Response Maps

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.413504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.413504Z digest=sha256:0919d8edd4fdec29a68c397cebb3e1531f00a890834323f44bf9775b018bd44f

Observation 3240d49e-82f2-419e-bd48-bfa035f25a40 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Obfuscated Activations Bypass LLM Latent-Space Defenses Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.417219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.417219Z digest=sha256:0103f1b299408566f5c34645ac8122af752b9310ba2565c29210db30c35b09a9

Observation 06df9536-17a1-4d1d-b201-d6e04cd5807c · outbound

This paper cites Many-shot jailbreaking.

Obfuscated Activations Bypass LLM Latent-Space Defenses Many-shot jailbreaking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.421310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.421310Z digest=sha256:d9af22abecfaacaf48ae779d0a01f03ebeb77c03e1bc67eea822890c83dfc533

Observation 97645858-2052-4b3a-9109-d47e732b2347 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.424767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.424767Z digest=sha256:ca7a7f2e02dfcc2a4af0f000dc86fb4b7b05efb8572899f0d0af0db53395a33e

Observation 519b6e9c-ebb2-4eb6-a6b1-13471d8a5e85 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Obfuscated Activations Bypass LLM Latent-Space Defenses Refusal in Language Models Is Mediated by a Single Direction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.428975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.428975Z digest=sha256:80d17c0aeaf374154b6dc0f8257c8b4d8f67726fccdcc33b2440b38d52094220

Observation 71d468cf-5ecc-4165-878b-bcb8161445b8 · outbound

This paper cites Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples.

Obfuscated Activations Bypass LLM Latent-Space Defenses Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.432401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.432401Z digest=sha256:20dc3eff6ac74ef901fde9ca0e6e505822f2d45fcbc18da03377201f2823bc61

Observation 64ce84d3-843d-44f8-abed-738d2280498d · outbound

This paper cites sql-create-context Dataset , 2023.

Obfuscated Activations Bypass LLM Latent-Space Defenses sql-create-context Dataset , 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.436270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.436270Z digest=sha256:f11458647065cf369933cb75dd38f452c832777b96ffd249a9c6b83370385eb3

Observation 65eae594-778f-44e5-b6be-372c71ee8da4 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.439328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.439328Z digest=sha256:7c42a152e9db582543cfc3f8bca154aef3a47133af9eb0b69a5052c66afdad1b

Observation c94f3d3e-9e2a-4317-8b26-45639e4690ba · outbound

This paper cites Probing Classifiers: Promises, Shortcomings, and Advances.

Obfuscated Activations Bypass LLM Latent-Space Defenses Probing Classifiers: Promises, Shortcomings, and Advances

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.442625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.442625Z digest=sha256:a712f681d22868be466031140f223c51308cd0b0e91420276a7816d0bfd86c63

Observation 197b6a51-0ff5-4e50-9f2f-05d50839b840 · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

Obfuscated Activations Bypass LLM Latent-Space Defenses Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.445716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.445716Z digest=sha256:18883bdf7a2bf27fc1bd1867e104d3e31958c54b8543152d2cb972e40b033659

Observation 08a2e4f6-395b-4409-bea1-9f34b18d69c4 · outbound

This paper cites Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning.

Obfuscated Activations Bypass LLM Latent-Space Defenses Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.448935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.448935Z digest=sha256:65f33518e950a8f95a83b29645d57f5af4b1483205c0571ae242936b3f6a2342

Observation ffce4e76-104b-434c-a5ef-1ad91f34c22f · outbound

This paper cites A Sober Look at Steering Vectors for LLMs.

Obfuscated Activations Bypass LLM Latent-Space Defenses A Sober Look at Steering Vectors for LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.452487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.452487Z digest=sha256:1690be6b1d1493aa89f277f990539d1f531bcb45836cfec591f5051dc9568489

Observation 45050759-71dc-4317-9a46-bb5668b90c15 · outbound

This paper cites Towards Monosemanticity: Decomposing Language Models With Dictionary Learning.

Obfuscated Activations Bypass LLM Latent-Space Defenses Towards Monosemanticity: Decomposing Language Models With Dictionary Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.455588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.455588Z digest=sha256:76499310faf8004ceda4caeb6dd7c9f5b13aaf318a35850c99b9fb9408b214f3

Observation 8819b7ad-8e30-4e95-a02c-f2b8238a7934 · outbound

This paper cites Using Dictionary Learning Features as Classifiers , October 2024.

Obfuscated Activations Bypass LLM Latent-Space Defenses Using Dictionary Learning Features as Classifiers , October 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.458685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.458685Z digest=sha256:9717dd69805a13103d3435d73f04e07dc73dcb4cda9aded1d0154728411f1c57

Observation cb396376-fd62-4751-b087-2c1cda20bd0c · outbound

This paper cites Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks.

Obfuscated Activations Bypass LLM Latent-Space Defenses Comparing Bottom-Up and Top-Down Steering Approaches on In-Context Learning Tasks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.462084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.462084Z digest=sha256:1793cf496a05410dae048675f590d049c35330074ee3c16fe8cda1eb8538165b

Observation b87dc70c-aa49-4ea9-95b5-505af1aba19a · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Obfuscated Activations Bypass LLM Latent-Space Defenses Discovering Latent Knowledge in Language Models Without Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.465433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.465433Z digest=sha256:db6633ec504f0f7a90d1aaee08b1f811df6a1ba742a4369e60a74c55784363f0

Observation 94d62af0-14e1-4c89-8257-1b6bede1170e · outbound

This paper cites Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods.

Obfuscated Activations Bypass LLM Latent-Space Defenses Adversarial Examples Are Not Easily Detected: Bypassing Ten Detection Methods

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.468864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.468864Z digest=sha256:92903c38457caef3b96c93f1a80d8202431ac5bc23722b7c102ceeefde1454d2

Observation b6214e8e-c1ab-455f-93f3-baca9257afdd · outbound

This paper cites Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024.

Obfuscated Activations Bypass LLM Latent-Space Defenses Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.472009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.472009Z digest=sha256:89e9ede5f11ac6dd76c2c01a17cd831a44a79804c3b916097e5627e3642c07d7

Observation a2c190ff-a895-4917-b1a8-5f0dec926f57 · outbound

This paper cites Defending Against Unforeseen Failure Modes with Latent Adversarial Training.

Obfuscated Activations Bypass LLM Latent-Space Defenses Defending Against Unforeseen Failure Modes with Latent Adversarial Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.475833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.475833Z digest=sha256:9ecb915b48f1c1acf54b89bf7c3cc2a1fda12320a9900999af85eb58462cf92e

Observation f58b91f1-5a00-4ec5-b371-89dcf355a42e · outbound

This paper cites A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders.

Obfuscated Activations Bypass LLM Latent-Space Defenses A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.479701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.479701Z digest=sha256:8bc881cdea96bfc19630160bcbed92132f5a0124de0cc1c538a73f2620afd2a0

Observation 9ce3b649-da44-426e-b3d4-a55a84d87665 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Obfuscated Activations Bypass LLM Latent-Space Defenses Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.482921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.482921Z digest=sha256:e639d14b4636ae4e7621a43fcf5623254e70699e6add7b667453672060620ba0

Observation 5847f4b0-f741-4eb6-92f2-120231a1d30e · outbound

This paper cites Code Alpaca: An Instruction-following LLaMA model for code generation.

Obfuscated Activations Bypass LLM Latent-Space Defenses Code Alpaca: An Instruction-following LLaMA model for code generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.486569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.486569Z digest=sha256:5bad2f2359ea4927162131ea797983c635d059d3e4522d0ceb7c106d07ccf8cc

Observation e02dce06-f80f-453a-817e-c42acaef0390 · outbound

This paper cites Srivastava.

Obfuscated Activations Bypass LLM Latent-Space Defenses Srivastava

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.490133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.490133Z digest=sha256:0f95d5909f6c1f93df135e0f34547db44718afcaf49afd037aa183c8854e61ff

Observation 3920e612-80c2-46a9-b7e9-6b9fc7469e2d · outbound

This paper cites Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks.

Obfuscated Activations Bypass LLM Latent-Space Defenses Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.494756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.494756Z digest=sha256:b4a1506e947ec466d87a224fbcaedcd37c7bee93970ceb027a2640240425c187

Observation c8d70fba-391f-483a-b470-b4af80013364 · outbound

This paper cites Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning.

Obfuscated Activations Bypass LLM Latent-Space Defenses Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.498355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.498355Z digest=sha256:a6973c4fc1df433062fa10dbd2a9b869ef59e2819e9b319043963c4187accb44

Observation 4b24aece-0821-4cc7-8831-a68c3773d40b · outbound

This paper cites Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals.

Obfuscated Activations Bypass LLM Latent-Space Defenses Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.501870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.501870Z digest=sha256:86a65d309492644031ce56d5b108768518e83035736e1b10061b1326fb9366d5

Observation 4e8e58a9-7bc7-42f5-9dd1-d59622dfbfc5 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.505152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.505152Z digest=sha256:698d01132661ecd66d705500f20fc26535963ba8e6067483026ca6e2e20dfcd5

Observation 0eaf2520-952d-4e82-bd08-1b03eda895bb · outbound

This paper cites Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting.

Obfuscated Activations Bypass LLM Latent-Space Defenses Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.508683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.508683Z digest=sha256:38b673fbb9f62b6f73f8f8e5ceff06d9eef722835611a6cc5460006be0fc0182

Observation f29247ee-e7fb-4180-a7eb-b8ad1cdfddde · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Obfuscated Activations Bypass LLM Latent-Space Defenses Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.511560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.511560Z digest=sha256:2ff651aad7c861ef664f8e601af41af62adc1b6dcf984002cb65942abf870416

Observation ca21f569-dbf6-48ec-aead-d1d372431b20 · outbound

This paper cites Backdoor Attack with Imperceptible Input and Latent Modification.

Obfuscated Activations Bypass LLM Latent-Space Defenses Backdoor Attack with Imperceptible Input and Latent Modification

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.515812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.515812Z digest=sha256:1423b34de16dabb0fc12d57348e6669e3bed9250b699d722625a8052d65c557f

Observation a7848329-05a0-4ac2-94bb-a317b014d96d · outbound

This paper cites Detecting Adversarial Samples from Artifacts.

Obfuscated Activations Bypass LLM Latent-Space Defenses Detecting Adversarial Samples from Artifacts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.519305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.519305Z digest=sha256:87ef4ac750eb028415d5458351e19fe46e7e96eaa2a079348fdb8fdb6407079b

Observation 7a2707ae-d2d4-476b-ae45-71f39a7017a6 · outbound

This paper cites Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness.

Obfuscated Activations Bypass LLM Latent-Space Defenses Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.522790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.522790Z digest=sha256:bd2abcf1ca9623497eae8e7052d3f0b6c48276f5834769a1434d17e15e5eccf7

Observation 6116d3e7-b668-47a1-b5ab-62ef0ec4e3bb · outbound

This paper cites Interpretability Illusions in the Generalization of Simplified Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Interpretability Illusions in the Generalization of Simplified Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.526314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.526314Z digest=sha256:c5d325faf8b0a61f5a1f33163c763f4d00c26bd71687689d8764433edccdb423

Observation 98bd1a90-addd-46e8-8b3e-c924b1990fa4 · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Erasing Conceptual Knowledge from Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.529629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.529629Z digest=sha256:19dc2a7537bc6f364977bf69d37332c69e2a974c95828a125c0ce8e8ae104d1b

Observation 860f0c13-9fde-4212-9dd7-df8718b74178 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Obfuscated Activations Bypass LLM Latent-Space Defenses Scaling and evaluating sparse autoencoders

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.533084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.533084Z digest=sha256:6fd9f14e7445f13339db11ba27594bc328f878c5f0ea9872eabda015ed4945b9

Observation dde99bc3-fe29-495b-80c5-0f79fb43780c · outbound

This paper cites STRIP: a defence against trojan attacks on deep neural networks.

Obfuscated Activations Bypass LLM Latent-Space Defenses STRIP: a defence against trojan attacks on deep neural networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.536844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.536844Z digest=sha256:0cd91cfa55ac067d01742a374756be4943cfba54f4730a369581b3a0a255cec3

Observation 71325723-ed92-47ed-b556-ae14f38b4bc1 · outbound

This paper cites Coercing LLMs to do and reveal (almost) anything.

Obfuscated Activations Bypass LLM Latent-Space Defenses Coercing LLMs to do and reveal (almost) anything

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.540064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.540064Z digest=sha256:16bb305d0aa30f9e3fa23ce7c04f989602708f3a306952949f4ed7a90373250f

Observation aadfe69c-ffee-4088-8934-c09dd1de9cb6 · outbound

This paper cites Planting Undetectable Backdoors in Machine Learning Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Planting Undetectable Backdoors in Machine Learning Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.543468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.543468Z digest=sha256:ae1838024064da1cd0509bc100deeaa4c76d6a6dd917e73cc24a4011eee0a24e

Observation 6bb9200b-ef2d-423a-8ae8-a97217837e3c · outbound

This paper cites On the (Statistical) Detection of Adversarial Examples.

Obfuscated Activations Bypass LLM Latent-Space Defenses On the (Statistical) Detection of Adversarial Examples

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.547667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.547667Z digest=sha256:fc0fb9be38e65fe957148ed01bc55048ba3370d04200a4943ddec097869badfa

Observation e1aebca0-e941-4ae0-ad46-6c4b341e1fbb · outbound

This paper cites Automated Multi-Turn Red-Teaming with Cascade , October 2024.

Obfuscated Activations Bypass LLM Latent-Space Defenses Automated Multi-Turn Red-Teaming with Cascade , October 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.552170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.552170Z digest=sha256:b39eb3eae02df6b37ef3515d975c9f968b68fd6d2ed1e017713e1c5be6093947

Observation 72dc18e0-40be-43c5-b389-02f6512f8397 · outbound

This paper cites SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics.

Obfuscated Activations Bypass LLM Latent-Space Defenses SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.555495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.555495Z digest=sha256:baa40b19d2392da371dc3a4897934b95c35c23cf58ece43bc8de47304cee539f

Observation c74993d0-9e8e-482f-926f-6cdf3cc75bc3 · outbound

This paper cites Backdoors as an analogy for deceptive alignment , 2024.

Obfuscated Activations Bypass LLM Latent-Space Defenses Backdoors as an analogy for deceptive alignment , 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.558780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.558780Z digest=sha256:7b8fb42ce52ef58772592d7847eb8ed2b332f4be6b5c43b2906805c5c47b852c

Observation fe9d3e11-6cee-4b89-a887-085a6ae0cc40 · outbound

This paper cites Are Odds Really Odd? Bypassing Statistical Detection of Adversarial Examples.

Obfuscated Activations Bypass LLM Latent-Space Defenses Are Odds Really Odd? Bypassing Statistical Detection of Adversarial Examples

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:59:12.252118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T16:59:11.562219Z digest=sha256:f05193aa507feb028d4e099e2afadefa220c60dae50e6fd9a82ba9d9be14dd27

Observation 2c63efb2-9cdd-4a50-8b2d-3043cbbbd4ae · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses LoRA: Low-Rank Adaptation of Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.565993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.565993Z digest=sha256:4e322be06e1d1fdd05fb88a5cb2df20b1ffd229b5e39cab8a493393a0ff23573

Observation 094286f2-0ad5-47d3-b92f-8a6c53e2d4f5 · outbound

This paper cites Gradient hacking.

Obfuscated Activations Bypass LLM Latent-Space Defenses Gradient hacking

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.569320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.569320Z digest=sha256:7ae9a12ada9c7a5c50aec27e6648b08442e8af633839d34134d7e12654c1c36c

Observation f9a2a7a4-ce46-4277-9f75-39cec7a8164d · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Obfuscated Activations Bypass LLM Latent-Space Defenses Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.572303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.572303Z digest=sha256:6e869c5c6758c2edf9b2c35b7c2e959d11002ad7bb85aae290b87be170664d2b

Observation c093c635-2b6c-4879-8aa7-3b2d805c3683 · outbound

This paper cites What Makes and Breaks Safety Fine-tuning? A Mechanistic Study.

Obfuscated Activations Bypass LLM Latent-Space Defenses What Makes and Breaks Safety Fine-tuning? A Mechanistic Study

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.575679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.575679Z digest=sha256:45c3d00f5c77b5eff1fc47b16d7bdeeaa4d5fa67e686f54732c33f712e029d08

Observation 2de280a7-46c0-4a80-9330-481447e77ecb · outbound

This paper cites BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning.

Obfuscated Activations Bypass LLM Latent-Space Defenses BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.579221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.579221Z digest=sha256:c6b38980eab555b6ff82c41cc863a42d5216965e2552c2458dd4c2c5b2d775d2

Observation 160d0496-f335-4607-bffa-7be83ac0c2ef · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.582197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.582197Z digest=sha256:03a7fe205162f775a1531edf80a88957d4143cf3fc7dbd50e7296f7a60792fa5

Observation f21a9e4f-f192-4b56-9ac5-3b90093d910f · outbound

This paper cites Generating Distributional Adversarial Examples to Evade Statistical Detectors.

Obfuscated Activations Bypass LLM Latent-Space Defenses Generating Distributional Adversarial Examples to Evade Statistical Detectors

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.585314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.585314Z digest=sha256:2242affa2fde2717ecb073048be62389c009918d3e66d908f38b546b75b2f3e8

Observation e79d9321-d92d-4459-9f02-ce14416434a9 · outbound

This paper cites Auto-Encoding Variational Bayes.

Obfuscated Activations Bypass LLM Latent-Space Defenses Auto-Encoding Variational Bayes

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.589133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.589133Z digest=sha256:2f26e9b2e444bc91ac6aac71320c239810c2834b7640bd2b3278afb3f108ea1f

Observation ee330ea4-a785-4e82-af08-da58f53f3718 · outbound

This paper cites What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks.

Obfuscated Activations Bypass LLM Latent-Space Defenses What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.592515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.592515Z digest=sha256:89001b4ecb293d954c2c0e0ed3d1fb09c0517bff2d035dc80c1ac61e7246d338

Observation 36ff46fd-7f63-4936-a35c-3f3f48fdfdeb · outbound

This paper cites SAEs (usually) Transfer Between Base and Chat Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses SAEs (usually) Transfer Between Base and Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.595573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.595573Z digest=sha256:8e346dcfc79ad1cc25737d34526c1728068abedc5e726be34d7a2d08928786b2

Observation b01c08d5-4388-4694-9a2a-29a141fc57bf · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Obfuscated Activations Bypass LLM Latent-Space Defenses The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.599567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.599567Z digest=sha256:e9377012672adcbc0cdb3d09c86ac6f361553cc29202e9962a2f9b29b77de885

Observation 5115691c-caaf-4891-9db6-e8b3d5780a03 · outbound

This paper cites LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet.

Obfuscated Activations Bypass LLM Latent-Space Defenses LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.603469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.603469Z digest=sha256:daf0777ca899e43eb68f0e7789f908b95df7ce6a2874161a8fb10d3d77bc49c1

Observation 24e45150-6e19-416c-b93d-3936fc5af50b · outbound

This paper cites Adversarial Examples Detection in Deep Networks with Convolutional Filter Statistics.

Obfuscated Activations Bypass LLM Latent-Space Defenses Adversarial Examples Detection in Deep Networks with Convolutional Filter Statistics

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.607066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.607066Z digest=sha256:2b835b64784813d7d33d80099bd5431ba19f17ceb77726e1416e55f413bf6b4d

Observation cc526e77-b464-4c21-b3d7-569ca5bde8fa · outbound

This paper cites BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning.

Obfuscated Activations Bypass LLM Latent-Space Defenses BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.610276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.610276Z digest=sha256:75b10dcf2a62ec6c39ee556bbc86eb1914d93aca0656aa47661f63304c81ac1d

Observation 5a353bad-f017-47ec-beb5-637d5233e086 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Obfuscated Activations Bypass LLM Latent-Space Defenses Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.613283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.613283Z digest=sha256:d96b3b6fcd6b560b03579970d492a7d451c1f382751912616489a51c0ee38a92

Observation 84d0be1d-efd5-47e1-80cf-6b0177c47886 · outbound

This paper cites Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks , 2023.

Obfuscated Activations Bypass LLM Latent-Space Defenses Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks , 2023

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.616565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.616565Z digest=sha256:6b0061c0b62acba720148760fe99d603d9af01ae86e416b3ca173d3ba739eb75

Observation 8fb9a391-e95c-462a-963e-a525e7c51870 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.619662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.619662Z digest=sha256:63a2f09788f2abc29ea54cb3d46a257a268d366ad54498e54652e16eeef2311f

Observation 972924e2-e936-4619-bc0f-b3e58303615e · outbound

This paper cites Piccolo: Exposing Complex Backdoors in NLP Transformer Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Piccolo: Exposing Complex Backdoors in NLP Transformer Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.623124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.623124Z digest=sha256:2bf89b407f855e62ca9741b4b2c041046a35329e53709474885ba49bcf1ba2e1

Observation 9f89346a-2820-491a-92b0-d56708a8f09e · outbound

This paper cites The "Beatrix'' Resurrections: Robust Backdoor Detection via Gram Matrices.

Obfuscated Activations Bypass LLM Latent-Space Defenses The "Beatrix'' Resurrections: Robust Backdoor Detection via Gram Matrices

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.626531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.626531Z digest=sha256:1a6918e61e60698c39b10cf3be2b3b73b505e108d945b133d98218696163b5ec

Observation eb21202b-0928-4cdb-9708-79ae06d21e29 · outbound

This paper cites Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality.

Obfuscated Activations Bypass LLM Latent-Space Defenses Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.630217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.630217Z digest=sha256:b9099a7dab96c0bbd6a9ba2af1a4e79efe4637754e8bce9ae304632a48f086dc

Observation 0af454d7-7518-439a-b9d9-1c1ee6580c89 · outbound

This paper cites Simple probes can catch sleeper agents.

Obfuscated Activations Bypass LLM Latent-Space Defenses Simple probes can catch sleeper agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.633396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.633396Z digest=sha256:a4fab8511871572f8be60aa89b39fcc099f053129474fa602563c9bad715a8d6

Observation 1ac81dae-25c1-4374-a378-f22950205d7d · outbound

This paper cites Deep Causal Transcoding: A Framework for Mechanistically Eliciting Latent Behaviors in Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Deep Causal Transcoding: A Framework for Mechanistically Eliciting Latent Behaviors in Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.636341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.636341Z digest=sha256:7dde341fe157ab67c9ef2a18cc60a0eb4e20c1b1107de6c0d7000b5ca464d835

Observation d3f132ba-7d39-4524-a8da-f8b6a80a6b02 · outbound

This paper cites an unresolved cited work.

Obfuscated Activations Bypass LLM Latent-Space Defenses Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.639563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.639563Z digest=sha256:5fc5ad28864ce62841f1aabf38f947f6e773a0bc3cc0b9a20936d63b2a572652

Observation 060896e8-446b-4d6d-95a2-555d291a6863 · outbound

This paper cites Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching.

Obfuscated Activations Bypass LLM Latent-Space Defenses Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.642814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.642814Z digest=sha256:8d204ed940d046bfa9907501e669832821eb3c0d446134b3182b7dff6f512058

Observation 6381687a-f8ba-4fe1-98c4-bfc07fdeca29 · outbound

This paper cites Eliciting Latent Knowledge from Quirky Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Eliciting Latent Knowledge from Quirky Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.645830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.645830Z digest=sha256:5b27b99607ddcb16e966321295fe4d1256f9a4f12bc5a8c805a36ff347c3a3ab

Observation 8f9b3ec4-de69-4a3d-8543-91cd9cea82ca · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Obfuscated Activations Bypass LLM Latent-Space Defenses The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.649420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.649420Z digest=sha256:4568cf6668a9d12e43a0ff5d0ac9ecce12663e42114bd59db08a991bdee93f10

Observation a5381cfd-d3ce-4df9-8b8c-c0397fdf949b · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.652846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.652846Z digest=sha256:71661c489d4b805a9afce70a599f5ff0d5b8fa63a208d65ce031bbb6c021cdfd

Observation 60eaeec0-91c2-469c-b1cb-9fb6d79a4b1b · outbound

This paper cites On Detecting Adversarial Perturbations.

Obfuscated Activations Bypass LLM Latent-Space Defenses On Detecting Adversarial Perturbations

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.656581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.656581Z digest=sha256:befac0551c94386b13b1e41f16b704444c1c2772785fa6d7b78feb5ea7486ce5

Observation 926200f6-1a52-47bf-acbe-95fadef3d622 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Obfuscated Activations Bypass LLM Latent-Space Defenses Steering Llama 2 via Contrastive Activation Addition

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.663606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.663606Z digest=sha256:d17d0e3cd7b9315b4c133117dfb73c15c19a6e719d209fccb997760dab682dbe

Observation 6d81f312-e153-4cf1-b700-1b9e621a2301 · outbound

This paper cites Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection.

Obfuscated Activations Bypass LLM Latent-Space Defenses Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection

Reference 76

Resolution
verified exact
doi, observed 2026-08-11T16:59:12.106814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T16:59:11.667314Z digest=sha256:bba23dc60d15bcf0da1b486e6c18e6f85fd6c935cdf345ad69b963b86451c69d

Observation f26cb0aa-e5c2-403d-bdb1-00ccba7e67ed · outbound

This paper cites Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs.

Obfuscated Activations Bypass LLM Latent-Space Defenses Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.670440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.670440Z digest=sha256:86cb032646716a2fe324f3d0820fb2286b0e3a19f2e506ed029898112e53f2a0

Observation d011d69b-c5e1-42d3-b97b-182d9f2f4046 · outbound

This paper cites Circumventing Backdoor Defenses That Are Based on Latent Separability.

Obfuscated Activations Bypass LLM Latent-Space Defenses Circumventing Backdoor Defenses That Are Based on Latent Separability

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:59:12.084155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T16:59:11.673715Z digest=sha256:b40a717a6bf588e935e609d0e794b975ddc99268f4d6db0dd1414541bb54fba6

Observation 0d515b3b-34d8-4400-af3f-a065dbdc9047 · outbound

This paper cites A General Framework For Detecting Anomalous Inputs to DNN Classifiers.

Obfuscated Activations Bypass LLM Latent-Space Defenses A General Framework For Detecting Anomalous Inputs to DNN Classifiers

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.677266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.677266Z digest=sha256:57e562ae399a0120705e07a8d6c4343a056bd7d82b9ccafe30c3faf70bfff02e

Observation 94794b8a-9fe5-48aa-937b-9906fb19ad23 · outbound

This paper cites Representation Noising: A Defence Mechanism Against Harmful Finetuning.

Obfuscated Activations Bypass LLM Latent-Space Defenses Representation Noising: A Defence Mechanism Against Harmful Finetuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.680413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.680413Z digest=sha256:ed873423b7bd4ffaf339ba7f89e5fb7663ca8750da65f8df952a89f64ae72ef6

Observation 7b96c86c-c1e5-4224-944f-0a0c50d0ed41 · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.685411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.685411Z digest=sha256:ebd26971c54485ce97139657f161b15e5e8c202b20c348c36687c00c188ac74c

Observation b57e636d-449d-4f02-b514-2a7b3bb47cda · outbound

This paper cites Public comment: Robustness evaluation seems invalid , 2024.

Obfuscated Activations Bypass LLM Latent-Space Defenses Public comment: Robustness evaluation seems invalid , 2024

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.689653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.689653Z digest=sha256:e6dc5100a91f6cb9cdf9afc993a0dad56826f549cd1d0d34e12878d433a16c29

Observation 07c7a426-0484-4dd1-883e-d3af608ed365 · outbound

This paper cites Revisiting the Robust Alignment of Circuit Breakers.

Obfuscated Activations Bypass LLM Latent-Space Defenses Revisiting the Robust Alignment of Circuit Breakers

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.693516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.693516Z digest=sha256:5455539d18a2ae9a016240f665eb2d33c56923c087302ae394f8acbb6b7e1fa6

Observation b559814d-bf1f-4f1b-ba82-4beb643b4ba5 · outbound

This paper cites Circumventing interpretability: How to defeat mind-readers.

Obfuscated Activations Bypass LLM Latent-Space Defenses Circumventing interpretability: How to defeat mind-readers

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.697137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.697137Z digest=sha256:761eaf63b361b9247ad99d71380f35bca60d80ca8ce064d9f5fe41e262970f47

Observation ac3eb96e-f440-4bf5-9fec-23fd00fcebe3 · outbound

This paper cites Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks.

Obfuscated Activations Bypass LLM Latent-Space Defenses Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.700722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.700722Z digest=sha256:60c1f4956613f240e0f0a2a9044ef86bfdcfeb2f77ac28e8bdfa279d20624a78

Observation 379e0092-c038-46bc-ba94-95b0b9927df2 · outbound

This paper cites A Survey on Backdoor Attack and Defense in Natural Language Processing.

Obfuscated Activations Bypass LLM Latent-Space Defenses A Survey on Backdoor Attack and Defense in Natural Language Processing

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:59:12.032937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T16:59:11.704056Z digest=sha256:1640a4fa178a0652c29fdc47acf3fc0708d6e4623d1c6246aa870f27384507df

Observation bb6e60a6-6a20-40f9-9315-c36b3ce69e79 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Obfuscated Activations Bypass LLM Latent-Space Defenses Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.707549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.707549Z digest=sha256:453780b8deaf12ff3a60df2d7b5e49780a290d9c91c09516b4b35a7a649f6c7d

Observation ca3f3d39-4dfc-4eba-bcab-548c2544b76f · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Obfuscated Activations Bypass LLM Latent-Space Defenses A StrongREJECT for Empty Jailbreaks

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.711141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.711141Z digest=sha256:182f3bcd09db85c661b9a14800024c44a66d81ac43bb1f242a4b7b12684ae761

Observation 01d15feb-b7a8-4d0b-94b8-629a9cb20710 · outbound

This paper cites Few-shot out of domain intent detection with covariance corrected Mahalanobis distance , 2023.

Obfuscated Activations Bypass LLM Latent-Space Defenses Few-shot out of domain intent detection with covariance corrected Mahalanobis distance , 2023

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.714736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.714736Z digest=sha256:efecdaf8b956ea4c66ae22087c2abb1d23dca70815689f76480ef785da3fea1f

Observation 008ecee1-61b6-4461-b843-4c969bf2b182 · outbound

This paper cites Codebook Features: Sparse and Discrete Interpretability for Neural Networks.

Obfuscated Activations Bypass LLM Latent-Space Defenses Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.718227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.718227Z digest=sha256:aeab6982caccd23b59595774b5125862972f9f88ca9e3afd5eb6cafaeb6dfe2d

Observation 7e3f88a3-6668-48dc-babd-833c6eddda01 · outbound

This paper cites Analyzing the Generalization and Reliability of Steering Vectors.

Obfuscated Activations Bypass LLM Latent-Space Defenses Analyzing the Generalization and Reliability of Steering Vectors

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.722044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.722044Z digest=sha256:23267b55ee5d90706320bf058996507fc63f13f62948ebc87ac71d349ca3f7e8

Observation 37e91562-d27f-418d-bee3-8df1dd0879e6 · outbound

This paper cites Bypassing Backdoor Detection Algorithms in Deep Learning.

Obfuscated Activations Bypass LLM Latent-Space Defenses Bypassing Backdoor Detection Algorithms in Deep Learning

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.726332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.726332Z digest=sha256:6df8098d7a637f7b117b9e41b763b089cd2e7094d8262b7bfca60f1b0471d10c

Observation 70a508bc-3cac-44d9-94b9-162770d43f53 · outbound

This paper cites Demon in the Variant: Statistical Analysis of DNNs for Robust Backdoor Contamination Detection.

Obfuscated Activations Bypass LLM Latent-Space Defenses Demon in the Variant: Statistical Analysis of DNNs for Robust Backdoor Contamination Detection

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.729476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.729476Z digest=sha256:2a80e0b58ecb0539a3a807f8be815c769742d7a3d62b4af34b32cf80d53ff369

Observation f1c385ed-26c8-44cb-b215-2c25cbc6cb32 · outbound

This paper cites Distribution Preserving Backdoor Attack in Self-supervised Learning.

Obfuscated Activations Bypass LLM Latent-Space Defenses Distribution Preserving Backdoor Attack in Self-supervised Learning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.732772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.732772Z digest=sha256:23463cd7c9d505cf911ea681d2ea91c3b8392a1b926b5ea2531a22b4452f4820

Observation 0eb0af25-1452-455b-a4b0-0cf231e08ac6 · outbound

This paper cites Hashimoto.

Obfuscated Activations Bypass LLM Latent-Space Defenses Hashimoto

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.735899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.735899Z digest=sha256:edddf376137604dd884d6d744bd98b640ed7a2b73d427742bb940ae2be88b522

Observation be33ab39-7ac7-4f30-a755-e106263caf2a · outbound

This paper cites Daniel Freeman, Theodore R.

Obfuscated Activations Bypass LLM Latent-Space Defenses Daniel Freeman, Theodore R

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.739009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.739009Z digest=sha256:10e0312ceb69bb18e5c2d1408b9e9526c32875febbfaf45099daa5ad2c3703bc

Observation 3dc1bfcd-25ce-4af1-a9aa-a1de9df8383a · outbound

This paper cites FLRT: Fluent Student-Teacher Redteaming.

Obfuscated Activations Bypass LLM Latent-Space Defenses FLRT: Fluent Student-Teacher Redteaming

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.742385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.742385Z digest=sha256:9923c1cbe6f3407ed6e7d0789d8a3e73236ab84a39812723ef9be0877bc247e4

Observation 3b7fdaa0-9504-4ad8-8883-50997517749a · outbound

This paper cites Function Vectors in Large Language Models.

Obfuscated Activations Bypass LLM Latent-Space Defenses Function Vectors in Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.745884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.745884Z digest=sha256:d7ab7354f1b4bab7024092cec7033ba54f758b0b360431810ee9b4275bb83708

Observation 0f68f8ff-ba12-4d2a-91b2-c2f0322f7082 · outbound

This paper cites Spectral Signatures in Backdoor Attacks.

Obfuscated Activations Bypass LLM Latent-Space Defenses Spectral Signatures in Backdoor Attacks

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.750028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.750028Z digest=sha256:e1fabfb6cf29427c5a4d44c4c426589b9c6e019f1ab0d8a1c54b93c6cb0542a8

Observation 2da1de87-7f6e-44c0-8648-16b45ab1c789 · outbound

This paper cites Steering Language Models With Activation Engineering.

Obfuscated Activations Bypass LLM Latent-Space Defenses Steering Language Models With Activation Engineering

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.753512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.753512Z digest=sha256:ff94b3ec35cbc2820c9bc4b8ce815825e08d93b9edc5d6124ecf19045b20189c

Observation 7b69fb0a-f2a8-47f9-94ef-8ac776967154 · outbound

This paper cites an unresolved cited work.

Obfuscated Activations Bypass LLM Latent-Space Defenses Unresolved cited work

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:11.757096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:11.757096Z digest=sha256:4218628b634c06e29b8818ddab5813a6abfdcf7fe49de4d7be4995dc1d79a37c

Pith citing papers

Observation 5149911b-abee-4c46-9cbc-ff6d294fcf87 · inbound

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors cites this paper.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.275006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.275006Z digest=sha256:d257ff4b07b15e765dbadb1c554117447e4e820e92256c06839b211ec0114143

Observation f01be2eb-e74d-4b8f-9da1-cc63708d255d · inbound

Benchmarking Misuse Mitigation Against Covert Adversaries cites this paper.

Benchmarking Misuse Mitigation Against Covert Adversaries Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:32:14.724366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T10:29:05.104520Z digest=sha256:b92c11e635f730011dca282a8eeedab39f3b5806ca42390e92810cc386fe1949

Observation d7993807-29b9-48f0-8d70-8d292f805944 · inbound

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems cites this paper.

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:12.058846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:12.058846Z digest=sha256:57d019f983a2d6448b6f27980b596526b027831a2277e7e3cb021f2cee787223

Observation 24a9af66-5191-471a-93b5-341c475e8be1 · inbound

Adaptively Robust LLM Monitoring via Activation Watermarking cites this paper.

Adaptively Robust LLM Monitoring via Activation Watermarking Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:41.404497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:41.404497Z digest=sha256:5cb75299fa70f7a85d3672504def16ccedd5e54c42abb226eba887ef5ef08c40

Observation 0f76f3ef-2629-44f6-9793-bb977860a5e1 · inbound

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework cites this paper.

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:20.974832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-09T19:45:36.021741Z digest=sha256:78c5aa45e3ca5c0292083a59face07ff77cd90c24a455a622909b4059157a603

Observation 1055259d-3caf-4889-b05b-f6b58a07d7bb · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:17:56.463644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:7425e8b56202e6ba806c01692ce612ed8168605c5a502294c59bd24ef276517f

Observation 4279c4b0-7e2d-4e2c-890f-fa7e9480e83a · inbound

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations cites this paper.

Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.519189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T12:40:13.092199Z digest=sha256:ed79d982cde1ccdd818bac6fc1064eab8ce27abe21022598e4ed15e252b62928

Observation 194e3e3a-075a-4ea0-aec6-f7a04bf0650a · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.093705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:48a383b6b5bf9317813e326883d32fca31c2017239bcb22768cc9182641d21e6

Observation 91e4960d-93a5-4b8a-babc-a4acafe21e27 · inbound

PRISM: Recovering Instruction Sets from Language Model Activations cites this paper.

PRISM: Recovering Instruction Sets from Language Model Activations Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:57:30.654273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T16:52:02.948457Z digest=sha256:887a2154e47e2cf429c56dbfe9095ea63b50ab79b820888544a51a08cdf05771

Observation a1c880b8-3b52-4ca2-a2af-03007f1531cf · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 170

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.924841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:49464d34f3e2c84a90157d42f9ef32ca9e60204670b4364bb40d512f1cc04972

Observation 48df94a6-2358-4964-b02f-611e4753c21f · inbound

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier cites this paper.

Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:57:48.255219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:36:09.211639Z digest=sha256:d4303215075a4e07212e6bcfc5d3b40b779ff566ca71540e46587c33d59c7e83

Observation 3ecf09f9-6aba-4b0c-aa42-7b87e27aba39 · inbound

GDM AI Control Roadmap cites this paper.

GDM AI Control Roadmap Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:53:50.342415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:53:50.342415Z digest=sha256:8e367b0fe1211d0cf63f113ca6206c5eec1a9e0605b2ccc492e9c89e85b617ef