Pith. sign in

Paper Citation Record · LEDGER

Detecting Safety Training Modification in Language Models via Activation Analysis

As of 20 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.05578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05578 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1145ef7f-047a-4ef1-8ea8-a333b10b3283 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp.

Detecting Safety Training Modification in Language Models via Activation Analysis Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.220640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.637323Z digest=sha256:46de26259b529f9aeda7e54d55ff3b383d79adc03ff680f5f7393124472c4704

Observation 85033a13-9694-4ca9-b036-b4c508dc651b · outbound

This paper cites Dolphin: An uncensored, unbiased language model.

Detecting Safety Training Modification in Language Models via Activation Analysis Dolphin: An uncensored, unbiased language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.209167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.641526Z digest=sha256:a9dee6199a3503a3fb77ebb753990f022a3783c8ca9b6ad0d807b51d556c2943

Observation a8ccaf23-60e3-4135-81da-5273afaaff09 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Detecting Safety Training Modification in Language Models via Activation Analysis Representation Engineering: A Top-Down Approach to AI Transparency

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.645097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.645097Z digest=sha256:a726c381d8b46ec439f2ae8b27aa9ad27aa8572117bf4c5dbbd51b83dafc80ec

Observation ed92bf25-8632-4f6f-acfa-18a076b6f4ce · outbound

This paper cites Steering Language Models With Activation Engineering.

Detecting Safety Training Modification in Language Models via Activation Analysis Steering Language Models With Activation Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.649072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.649072Z digest=sha256:e0adafbff39e59685a8886b84d8de66de08bdb9fa5ad376522331a4b47dc70c9

Observation d430b2db-35fb-49e0-89b2-2b967986b074 · outbound

This paper cites Instructional Fingerprinting of Large Language Models.

Detecting Safety Training Modification in Language Models via Activation Analysis Instructional Fingerprinting of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.652802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.652802Z digest=sha256:6dcce51c1a4da3db7acb98bfa1e40fcfcc0aecc0de81b80ea339e43e2b369325

Observation d7dba1c4-a689-4636-a7e1-22455d71b97e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Detecting Safety Training Modification in Language Models via Activation Analysis HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.656374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.656374Z digest=sha256:be03c5a05fecb5842e71c70ecce85e6223cbc4465c501660c673f20de3817915

Observation 912bbfdd-f564-45b6-b6f1-af2451286584 · outbound

This paper cites ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Detecting Safety Training Modification in Language Models via Activation Analysis ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.198859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.660030Z digest=sha256:3f8c9040dc526244881e1612a7fde2256cd7c2853c42c6a1bb5c6172398e99e6

Observation 3f998b65-7b57-40f4-8e23-58d45a7e8509 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Detecting Safety Training Modification in Language Models via Activation Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.663609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.663609Z digest=sha256:6a9d89bb176fd9e26fce3c06a26e23644b23f8b39f140b7b86d27a9473de2c5a

Observation 5d4c3a31-a753-4cdb-8bf0-9a98a4617f0f · outbound

This paper cites Sokhansanj.

Detecting Safety Training Modification in Language Models via Activation Analysis Sokhansanj

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.188435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.667184Z digest=sha256:f118d1fad915a971d1f4049021b8cadb5080ab3a8feb57235eb7c4fcc3f82f11

Observation 36a9fd31-55df-475f-a3f3-8fa8785b160b · outbound

This paper cites XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025.

Detecting Safety Training Modification in Language Models via Activation Analysis XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.670840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.670840Z digest=sha256:c7080d241e954c91db218b6d5cfa0ccacb7b424369870d768d6b1f2820cd47ac

Observation 9c0e3c76-2e48-4f09-8a88-634e2dbaacd8 · outbound

This paper cites Defending Large Language Models Against Attacks With Residual Stream Activation Analysis.

Detecting Safety Training Modification in Language Models via Activation Analysis Defending Large Language Models Against Attacks With Residual Stream Activation Analysis

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:13:51.949039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.673991Z digest=sha256:d00901c66c3c915466a075000a4009820d5534ad9f1787ae571dbdd1c01997b9

Observation dd234149-1813-4a52-9f04-abdb222f3955 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Detecting Safety Training Modification in Language Models via Activation Analysis Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.677511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.677511Z digest=sha256:2785efae8a4a282242bcdf681c77dcacb1712f549dd6ba2c85854c4f80c5a3ac

Observation 8b6de393-72fc-4fe3-8aea-7a496d7ddcde · outbound

This paper cites Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022.

Detecting Safety Training Modification in Language Models via Activation Analysis Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.178320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.680835Z digest=sha256:a10438259b5a0d9ab44a814c8366a679e919c9bf4171233b953dfc979b7b6037

Observation a315e742-ee52-403d-b15e-bc24df6b0682 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Detecting Safety Training Modification in Language Models via Activation Analysis Gonzalez, Hao Zhang, and Ion Stoica

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.683805Z digest=sha256:67b84b817f7e2b2171bc9e0f99b8d3bfb583b03236e3be3b0aae244f1e8989e6

Observation 9fbef769-486b-4882-acad-c3121e68a00b · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Detecting Safety Training Modification in Language Models via Activation Analysis The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.686798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.686798Z digest=sha256:4632de045116fd130ad54547ca874533d61fceba86b8346c7066a45e4d4a2a08

Observation 706e3b47-05cf-4bee-b850-6f0c68768f80 · outbound

This paper cites Lo- cating and Editing Factual Associations in GPT.

Detecting Safety Training Modification in Language Models via Activation Analysis Lo- cating and Editing Factual Associations in GPT

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.159556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.690142Z digest=sha256:a0655b98f56ed6e97d88314d00af2939bb8a5c4b665a2c6fcd831a7755a5038f

Observation 87cae2ed-8127-4b6c-a6c9-d1effc757772 · outbound

This paper cites ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation.

Detecting Safety Training Modification in Language Models via Activation Analysis ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:13:51.917259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.693027Z digest=sha256:1b2865daf97260d14e0c8d52f9ae1b3870be48c4fff7ff95be710a538e1e7427

Observation 11f04dd4-9a9c-42ba-86b3-792f5256a686 · outbound

This paper cites VLLM_ALLOW_INSECURE_SERIALIZATION.

Detecting Safety Training Modification in Language Models via Activation Analysis VLLM_ALLOW_INSECURE_SERIALIZATION

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.148643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-08T10:13:51.695862Z digest=sha256:50362afa4ef10e909110e1f54c1881b4ef7d7201ce32446996dbc568bdff94b4

Pith citing papers

No inbound Pith citation observations are available.