Pith. sign in

Paper Citation Record · LEDGER

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2607.10226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10226 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T13:23:22.708604Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf4ded39-27c8-4b13-ac80-12b0a360004d · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Refusal in Language Models Is Mediated by a Single Direction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:b78cb9cc11763d3f7901df00bc5cd605d62f0f08da150afdb041a8154ec93b7e

Observation 8ec33f83-6c48-449f-8bf7-0c651fe993d9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:956cd8f906c14023edcdc6926b0ad31248c4adcef6185abb9595a013696b3329

Observation 97537d91-5005-47c6-9585-ee8f53865476 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:9688e914daa4ee938fb6c2b057946f04108c6bc9cfba9258740b53a99f296654

Observation 538b963b-2a7d-41fe-ac9e-80a08fa3bca6 · outbound

This paper cites The Llama 3 Herd of Models.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:939857ba1b546763433c7df8fbf26f84f8421419c882a9275c9a175c00fe620f

Observation 794b1101-7735-42a0-a3d8-a53e13c7148a · outbound

This paper cites Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:3d2187a201cbb876e1e0f3afa66420b67e71617080b968baa22f314572067b7c

Observation 9854b652-f25d-4651-8f55-1cc62a74f677 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Gemma 2: Improving Open Language Models at a Practical Size

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:600ebf28b5e34688f9f99d92f948a19573c5270db29b58636a7bcaa3e920616c

Observation 60b03b10-a014-49a2-907e-824206c2bffa · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:012240a0cf94845d6b92986951c005abf90a5a3298e9d3826affca6c8d7cf0c0

Observation bd962976-c7e0-4b71-9f83-010c868e279c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:ec1f7eaddad16e4803601f8a941cf4b337ced92fce7d8153d81b3ad09ba065a5

Observation 2e1d3879-bf63-4e4d-b305-87c584e031db · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:b1c1d86f38e44ddee62db5361eb819a3ec71c9417ec3a53d1684e339bc4cdc03

Observation 81b2608a-47c8-4ec4-9b14-1a2c108bec51 · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:9b687e43cda4e1f833583803634b51cdefef2b998807d530aa5d90846ceec634

Observation f57db782-a748-44c2-89dc-ecb498631e99 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:39c3e2a606d85af42bfef03a756c4cda473f4782271aa9563d65f9d121c635ed

Observation cecc21b7-45be-4dc0-8944-977367f56a2b · outbound

This paper cites XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:441e43152ea8bb39af2c4baefa78f1dd262fc74d69fc83cc7a6d4d805f97e102

Observation 69d632b0-843a-4e42-adc5-1abec4b73da1 · outbound

This paper cites Steering Language Models With Activation Engineering.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Steering Language Models With Activation Engineering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:c95a3ee1e180a5a24500bdedfb6d39e924f10146f6ad49cf7ca4230123b6a709

Observation dba96bb6-5e97-4caf-91c9-d146da754562 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control Representation Engineering: A Top-Down Approach to AI Transparency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T13:23:22.708604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:23:22.708604Z digest=sha256:abf7254f7c5064d1459ca8fefad27dcad5e4dd5cc45ef0dd1ed4c331f0b0cbe2

Pith citing papers

No inbound Pith citation observations are available.