Pith. sign in

Paper Citation Record · LEDGER

The Geometry of Harmfulness in LLMs through Subconcept Probing

As of 9 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2507.21141.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21141 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:47.854894Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:09:32.736304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T21:38:00.071469Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ecb6aa75-1eaf-4e6c-89a3-701ad29a7dff · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.765176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.765176Z digest=sha256:69c69526aef5857d4f6b785a57aff19bac7d2638a51770e0412e14b0ecddc0fc

Observation 82c5c459-c5e7-4430-96b4-4da7071c1999 · outbound

This paper cites Safety-Aware Fine-Tuning of Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Safety-Aware Fine-Tuning of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.772974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.772974Z digest=sha256:df8b61f0d6390e4a2b2f80f6662afd30443c5aedea46541fdcef676f28770310

Observation f360be14-326c-43cc-8947-f019ae1cba39 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

The Geometry of Harmfulness in LLMs through Subconcept Probing Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.776571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.776571Z digest=sha256:e19ce820df6ade02d9a4ff0235a435ca44fadd43472aa16c90c9de284e089787

Observation 7fb6beaf-c30a-4a6e-a4a4-0d5bf79bceec · outbound

This paper cites The Llama 3 Herd of Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.791537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.791537Z digest=sha256:ac78b0d5b1feb4fa385eff7b2fac79cbbc2167f4e72b48aac4cb57b7aade3653

Observation a528a80c-c0b7-4f4d-9440-a7b03bb7d901 · outbound

This paper cites Safedpo: A simple approach to direct preference optimization with enhanced safety.

The Geometry of Harmfulness in LLMs through Subconcept Probing Safedpo: A simple approach to direct preference optimization with enhanced safety

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.801834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.801834Z digest=sha256:3848713eb6ff011758f177689fcb2f26b2e352e4a8dfc672ced2a98103ed019d

Observation 6b0343f1-d2dc-41e9-9564-a4d601b82724 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.805088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.805088Z digest=sha256:946198709369571ed2cc944548fa8f0d4362c934eb8123fc008499fec07165ff

Observation a1939382-08e1-4005-bcf6-ae0c11ce47c8 · outbound

This paper cites Enhancing LLM Safety via Constrained Direct Preference Optimization.

The Geometry of Harmfulness in LLMs through Subconcept Probing Enhancing LLM Safety via Constrained Direct Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.808584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.808584Z digest=sha256:8fa11fe1ab21b1e43eb771235f33f74dc234deb38bf2e528b5bde5da48f93c33

Observation ff0ef252-7b16-4373-8618-3decc99bb851 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.812006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.812006Z digest=sha256:c54a767b00d69edc6956336ec3e8f67c3cdd1f08ec7cec566ade35848eb772c2

Observation a681fd9b-fb00-472c-b52d-078489649203 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.815767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.815767Z digest=sha256:1cc5f589240822cad4d1584d1e0f1bc0479caa7d00fd92778abcf887decf736d

Observation 509a0729-c1e0-46ab-a923-9a4e81bc055d · outbound

This paper cites Emergent Linear Representations in World Models of Self-Supervised Sequence Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Emergent Linear Representations in World Models of Self-Supervised Sequence Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.819681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.819681Z digest=sha256:4678366ef803035dce5db301a97b49268d3496424ed492b3612dc3701bd2211b

Observation 70ffcc5d-2c1c-4592-ac8a-bbbb34600921 · outbound

This paper cites The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions.

The Geometry of Harmfulness in LLMs through Subconcept Probing The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.822986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.822986Z digest=sha256:708ad82e3a261395768d4c44a50b414dd9ff723a5f60d744f2b2a7b962660fcf

Observation 0a8fb2e9-746c-4bd9-808e-5dc127b0f233 · outbound

This paper cites Interpretable steering of large language models with feature guided activation additions.

The Geometry of Harmfulness in LLMs through Subconcept Probing Interpretable steering of large language models with feature guided activation additions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.208051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:47.826803Z digest=sha256:cf6580581094dda76943a6df2660fffd987bd2489387e3cdb22b73c42a2f512f

Observation a850eefb-8758-4c6a-934d-26aa0304363f · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing Linear Representations of Sentiment in Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.830115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.830115Z digest=sha256:7e11157722684c7c0ac8cabe841e65aa28ef533f56cf5cf8678d308b53633e6b

Observation 4ad54f81-b217-4a3e-b17c-67d477fd1619 · outbound

This paper cites Steering Language Models With Activation Engineering.

The Geometry of Harmfulness in LLMs through Subconcept Probing Steering Language Models With Activation Engineering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.833410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.833410Z digest=sha256:f3243f4789021d94aaaa139e47d1d8fcddf5f19fb380f1615416fc24cb1ebf0e

Observation 47b0c430-ee7b-49a7-acd6-18f093ce0bce · outbound

This paper cites pyvene: A library for understanding and improving PyTorch models via interventions.

The Geometry of Harmfulness in LLMs through Subconcept Probing pyvene: A library for understanding and improving PyTorch models via interventions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.198757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:47.836938Z digest=sha256:acb16854bb1d120c698982390a9fc5c85a6a322a6787a8e6a5346bdda4e960c7

Observation 3407a27e-cdb7-4591-9266-9b563e4e3180 · outbound

This paper cites URL https://aclanthology.org/2024.naacl-demo.

The Geometry of Harmfulness in LLMs through Subconcept Probing URL https://aclanthology.org/2024.naacl-demo

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.189182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:47.840771Z digest=sha256:c4371deb671ec2d37b7e423513b1682cf65613eff4ea49e9f8f5ed27b19c2725

Observation ea24d6d9-d573-408f-b00e-5409aa37c140 · outbound

This paper cites Qwen2 Technical Report.

The Geometry of Harmfulness in LLMs through Subconcept Probing Qwen2 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.843887Z digest=sha256:19e4fed96ec0f05de5ac58ed81273a773c69f6755c032c79312dc9665867d6a0

Observation f5c4d94f-9caa-4c5e-8b5b-e722938ad976 · outbound

This paper cites From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs.

The Geometry of Harmfulness in LLMs through Subconcept Probing From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.847228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.847228Z digest=sha256:e6b91cd44fe287b826e68bee0a097f175fe0b625728fa59277db3acab90f144a

Observation e59f07d6-563b-4f03-b42d-8bf5f66c7b67 · outbound

This paper cites Under review.

The Geometry of Harmfulness in LLMs through Subconcept Probing Under review

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:48.177822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:47.850568Z digest=sha256:f47c056df9fc89101d5ced968f59b620db8604fd8cc030a04ea7609337211843

Observation 323d869c-b384-4e58-b0b8-df3e41f18161 · outbound

This paper cites an unresolved cited work.

The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:48.167501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:47.854894Z digest=sha256:cb2114bad35f9f847d1bbcc5c5b4bd5089667abf19f0e0e5bfaea3921fb6848d

Observation 4a5a077b-c308-4fea-820a-662d0f93c565 · outbound

This paper cites doi: https:// doi.org/10.1016/S0031-3203(96)00142-2.

The Geometry of Harmfulness in LLMs through Subconcept Probing doi: https:// doi.org/10.1016/S0031-3203(96)00142-2

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.769206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.769206Z digest=sha256:69714605b28293cef78e428d495f722edcff18cec8ae84d02f208f427cec3156

Observation 580723f0-2f20-47b0-897a-9660e8d9a8e0 · outbound

This paper cites Refusal Behavior in Large Language Models: A Nonlinear Perspective.

The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal Behavior in Large Language Models: A Nonlinear Perspective

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.794774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.794774Z digest=sha256:e87f3f63e80bcc5dbd12d201cd5a81ffa1b9b44ec669f1b66a1b94eb66d78cfc

Observation 15050118-aeec-4901-9351-987148f61d04 · outbound

This paper cites emnlp-main.273.

The Geometry of Harmfulness in LLMs through Subconcept Probing emnlp-main.273

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.787746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.787746Z digest=sha256:140a369253c7be536ac82f8b93c5871fe9d19dda9fc197afd9e3c820b455060b

Observation 2ad56729-0f6b-43cc-88de-f5f7a594dbb2 · outbound

This paper cites Toy Models of Superposition.

The Geometry of Harmfulness in LLMs through Subconcept Probing Toy Models of Superposition

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.780339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.780339Z digest=sha256:713f55c1c7484a925d7a0425e2eebd3d5dec2452defcee4e39e15965a7058649

Observation 266f749f-500e-42c1-9e67-1ad2a6323e5f · outbound

This paper cites an unresolved cited work.

The Geometry of Harmfulness in LLMs through Subconcept Probing Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:48.218134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:47.784263Z digest=sha256:ca1a139288edb3a2b2a36266ebd692056fc8cc49d95eb07acbccd93a4aade9a9

Observation 2b54dfec-7589-47ba-96ab-80ad776a304e · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

The Geometry of Harmfulness in LLMs through Subconcept Probing Refusal in Language Models Is Mediated by a Single Direction

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.757141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.757141Z digest=sha256:cbf8ef96686ffe61fb7cde40cd1a7c2cccf5947b12ce154bf2aa66bc2e035246

Observation 9ab45db4-4393-4a33-b972-58749b0b829a · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pp

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.761453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.761453Z digest=sha256:fbd1453723ac28ee2b3a39ad360857cca0117dfb5fa7fc339d7f897e48468211

Observation 89355c50-affb-4354-b6b5-88a10c87a02d · outbound

This paper cites On the Origins of Linear Representations in Large Language Models.

The Geometry of Harmfulness in LLMs through Subconcept Probing On the Origins of Linear Representations in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:47.798462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:47.798462Z digest=sha256:b23f5ee1878f2c2b085b4da6f449f9cd9f5a044af43db70d557be517cb40f206

Pith citing papers

Observation 05b043eb-10a1-4ebf-9169-085a654c8545 · inbound

Before the Last Token: Diagnosing Final-Token Safety Probe Failures cites this paper.

Before the Last Token: Diagnosing Final-Token Safety Probe Failures The Geometry of Harmfulness in LLMs through Subconcept Probing

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.073909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:37:02.439454Z digest=sha256:616ce3b674eb9245bad694ba4fa8b37fa8655576021d022901bd8179a93ca527

Observation 28a2ee74-b0c1-4dc9-a21e-8d16199d1ba7 · inbound

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators cites this paper.

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators The Geometry of Harmfulness in LLMs through Subconcept Probing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:32.736304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:32.736304Z digest=sha256:0b74d7434f225e9de5d306c96c9c5206e2eb76da8959dba13874b046a6525d04