Pith. sign in

Paper Citation Record · LEDGER

Detecting Safety Training Modification in Language Models via Activation Analysis

As of 9 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.05578.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05578 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T10:13:51.695862Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1145ef7f-047a-4ef1-8ea8-a333b10b3283 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp.

Detecting Safety Training Modification in Language Models via Activation Analysis Refusal in Language Models Is Mediated by a Single Direction.Advances in Neural Information Processing Systems 37 (NeurIPS 2024), pp

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.220640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.637323Z digest=sha256:8f2cf7dda77a93405f3c887431833ec67eae10514d8bbd54923bd580701dccea

Observation 85033a13-9694-4ca9-b036-b4c508dc651b · outbound

This paper cites Dolphin: An uncensored, unbiased language model.

Detecting Safety Training Modification in Language Models via Activation Analysis Dolphin: An uncensored, unbiased language model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.209167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.641526Z digest=sha256:6bc9254a6fd7ace892e963813f44829533e9595487e03935b9776a2a3462268a

Observation a8ccaf23-60e3-4135-81da-5273afaaff09 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Detecting Safety Training Modification in Language Models via Activation Analysis Representation Engineering: A Top-Down Approach to AI Transparency

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.645097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.645097Z digest=sha256:9ebf32a5146d4742bbd7416375eead3a60509729459490eaae31e0fa8ffc73d6

Observation ed92bf25-8632-4f6f-acfa-18a076b6f4ce · outbound

This paper cites Steering Language Models With Activation Engineering.

Detecting Safety Training Modification in Language Models via Activation Analysis Steering Language Models With Activation Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.649072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.649072Z digest=sha256:e1a9bda1d32f0e6af3499ce4a238874e193d8f63da0dfa2e9517ead5aebb7fa6

Observation d430b2db-35fb-49e0-89b2-2b967986b074 · outbound

This paper cites Instructional Fingerprinting of Large Language Models.

Detecting Safety Training Modification in Language Models via Activation Analysis Instructional Fingerprinting of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.652802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.652802Z digest=sha256:be41091fa77c59196b9e75aa2e17d3d84b0f4a98ee6c1b6d488645e150c19d92

Observation d7dba1c4-a689-4636-a7e1-22455d71b97e · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Detecting Safety Training Modification in Language Models via Activation Analysis HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.656374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.656374Z digest=sha256:6d0749857a6841117d9860f600f248faeb5655bc74d389271caf58b8794bf428

Observation 912bbfdd-f564-45b6-b6f1-af2451286584 · outbound

This paper cites ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Detecting Safety Training Modification in Language Models via Activation Analysis ToxiGen: A Large-Scale Machine- Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.198859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.660030Z digest=sha256:c87aaf06e40c2c386afa9aff69a6e2f195d1789ce221acd687151db61f4748eb

Observation 3f998b65-7b57-40f4-8e23-58d45a7e8509 · outbound

This paper cites WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs.

Detecting Safety Training Modification in Language Models via Activation Analysis WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.663609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.663609Z digest=sha256:0b2172cd6d09708222bf39dec57706d81d584a69650ce66003d4892634978360

Observation 5d4c3a31-a753-4cdb-8bf0-9a98a4617f0f · outbound

This paper cites Sokhansanj.

Detecting Safety Training Modification in Language Models via Activation Analysis Sokhansanj

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.188435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.667184Z digest=sha256:ffe711336328f26c148eda1d3b3382d649bdbbcd61fba69927311cc9cf5dbb4d

Observation 36a9fd31-55df-475f-a3f3-8fa8785b160b · outbound

This paper cites XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025.

Detecting Safety Training Modification in Language Models via Activation Analysis XBreaking: Understanding how LLMs security alignment can be broken.arXiv preprint arXiv:2504.21700, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.670840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.670840Z digest=sha256:625bb2f2aa5ec011521d30ba2dd0601bfe952a44f4321747af40cc666014313a

Observation 9c0e3c76-2e48-4f09-8a88-634e2dbaacd8 · outbound

This paper cites Defending Large Language Models Against Attacks With Residual Stream Activation Analysis.

Detecting Safety Training Modification in Language Models via Activation Analysis Defending Large Language Models Against Attacks With Residual Stream Activation Analysis

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T10:13:51.949039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.673991Z digest=sha256:4cb910b36fea79d280bc484f73018bdda8841d877d0ff78935d544b97077d4ac

Observation dd234149-1813-4a52-9f04-abdb222f3955 · outbound

This paper cites Safety Layers in Aligned Large Language Models: The Key to LLM Security.

Detecting Safety Training Modification in Language Models via Activation Analysis Safety Layers in Aligned Large Language Models: The Key to LLM Security

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.677511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.677511Z digest=sha256:4e2cfbf812e803a8058afd4d75f64d4f5d669c953f0a802499ed74993f3235a3

Observation 8b6de393-72fc-4fe3-8aea-7a496d7ddcde · outbound

This paper cites Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022.

Detecting Safety Training Modification in Language Models via Activation Analysis Probing Classifiers: Promises, Shortcomings, and Advances.Computational Linguistics, 48(1):207–219, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.178320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.680835Z digest=sha256:bec8d59b2d1f874bfcf647b73f7b311aa318d437417b4ee729ab8ebafde66131

Observation a315e742-ee52-403d-b15e-bc24df6b0682 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Detecting Safety Training Modification in Language Models via Activation Analysis Gonzalez, Hao Zhang, and Ion Stoica

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.683805Z digest=sha256:6b550c41265c6032fe12e97cb214a2703fc098e7825c05969228b0773e45b64f

Observation 9fbef769-486b-4882-acad-c3121e68a00b · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Detecting Safety Training Modification in Language Models via Activation Analysis The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T10:13:51.686798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:13:51.686798Z digest=sha256:651a5aee2888548f6c8441e6dd51c4cab432fff8798bd6d1490fbfd9cd2ca598

Observation 706e3b47-05cf-4bee-b850-6f0c68768f80 · outbound

This paper cites Lo- cating and Editing Factual Associations in GPT.

Detecting Safety Training Modification in Language Models via Activation Analysis Lo- cating and Editing Factual Associations in GPT

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.159556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.690142Z digest=sha256:af01961fac9d9fec060435891ad719f66c646c570675b262da90863b0ca59865

Observation 87cae2ed-8127-4b6c-a6c9-d1effc757772 · outbound

This paper cites ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation.

Detecting Safety Training Modification in Language Models via Activation Analysis ABS: Scanning Neural Networks for Back-doors by Artificial Brain Stimulation

Reference 17

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T10:13:51.917259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.693027Z digest=sha256:09a5cf762baaf7b731ab2370cfe116a53ba27c5afac3bb616b055ad1a646606e

Observation 11f04dd4-9a9c-42ba-86b3-792f5256a686 · outbound

This paper cites VLLM_ALLOW_INSECURE_SERIALIZATION.

Detecting Safety Training Modification in Language Models via Activation Analysis VLLM_ALLOW_INSECURE_SERIALIZATION

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T10:13:52.148643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T10:13:51.695862Z digest=sha256:713da6fa9cffc22efe6d2f6d02d70a6727710dc8daf485a85cfd8ec20c331348

Pith citing papers

No inbound Pith citation observations are available.