Pith. sign in

Paper Citation Record · LEDGER

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 1 inbound Pith citation observation for arXiv:2602.19101.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.19101 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:49:10.161724Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:32:57.295159Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T09:07:47.890063Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved11
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a02c9c72-c23a-4358-9bcb-d17f97eaefa6 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models The Ethics of Advanced AI Assistants

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.436079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.436079Z digest=sha256:6bfe78435a23cc00d3b692d4524659207bd24163900279cea2c04e15a1d82ae0

Observation 476eb455-e5a0-4b8e-8197-863197103191 · outbound

This paper cites Unsolved Problems in ML Safety.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Unsolved Problems in ML Safety

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.522544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.522544Z digest=sha256:fa1bafcda4bd7479dd0270e7b8ab24f225404464d82e5676579ee3492e514055

Observation 6b90d727-52ab-4803-b6fb-de0f44b4b028 · outbound

This paper cites Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.679130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.679130Z digest=sha256:1e72658a360cf999b7c73da31412b4b1c54c85f37083197d8b7dd2f27ca1d7eb

Observation c4bd6c0d-3735-457e-a197-5f642f9620dd · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Steering Llama 2 via Contrastive Activation Addition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.785134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.785134Z digest=sha256:c0b91440f8ee04f664265c97fcbecb3ea1d0b8ac4284df3084c467092a646603

Observation bb8a2115-5feb-4a1f-994e-13e4f4bd8724 · outbound

This paper cites Qwen2.5 Technical Report.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Qwen2.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.913148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.913148Z digest=sha256:47d5bbff77aa007162e4ca3021e1f719a417eea8fd8c6afc54f3f19cc44849ec

Observation 62bc3130-4331-42b5-8127-627d1046e46c · outbound

This paper cites Convergent Linear Representations of Emergent Misalignment.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Convergent Linear Representations of Emergent Misalignment

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.959136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.959136Z digest=sha256:135d3730b8fc2dfc693a197cb2ef5cc105dc4ed13865608afe502109379ee300

Observation 33404c56-3a97-45f9-a987-b6b99de887eb · outbound

This paper cites Model Organisms for Emergent Misalignment.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Model Organisms for Emergent Misalignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:10.039369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:10.039369Z digest=sha256:c203695e816932b832fd6983ecbaac8bdb1f832f45366537c91e20945cc526da

Observation d8b0c6ed-3a59-4f74-be19-a4b844c105a5 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:10.093299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:10.093299Z digest=sha256:7d6cf8babf89fb12e281282b34fe508ff0fa3cd706577ac5f2643d60e1703753

Observation 6d1233c1-472e-4409-8198-6282c031fdd5 · outbound

This paper cites OLEDinstead of going to the optional work event. Neutral $$$$ I chose to watch TV on myLG 65.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models OLEDinstead of going to the optional work event. Neutral $$$$ I chose to watch TV on myLG 65

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-02T21:49:10.161724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:10.161724Z digest=sha256:5e81761afc1a8131379fa7de18519767d8e356e2d66861b1d49dc463e74b277a

Observation bef8f37c-cd31-42e8-a4b3-3b54a546f5e8 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.271327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.271327Z digest=sha256:ad47f37fd67a907619e20716cb97e0bb1ff14ce71667531a5a19c93467447199

Observation d72d5ca9-f660-4242-a164-048a0c90f11d · outbound

This paper cites doi: 10.1016/j.tics.2023.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models doi: 10.1016/j.tics.2023

Reference 2023

Resolution
verified exact
doi, observed 2026-08-02T21:54:29.151885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-02T21:49:09.350529Z digest=sha256:6712272a55ba29e6452a7976eafd6820972c1b4a7f0e326e465367f3e7752e63

Observation ae547ce6-8ead-4da0-982f-2005f7b6f4b6 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Refusal in Language Models Is Mediated by a Single Direction

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.092182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.092182Z digest=sha256:e32904d3f64de569c37fb685ff3e627c25ccd1b0b2d9367c1b4d533aa9381e3b

Observation e1d58311-be77-4314-af1c-edb9ff1d7380 · outbound

This paper cites an unresolved cited work.

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T21:49:09.167520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:49:09.167520Z digest=sha256:86d863dfc434733724772c6c0aeec5f8bd1e051699d7530e219ed1863af7f1da

Pith citing papers

Observation b6ef70d9-560b-4770-94b7-d55c7f28f91a · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

Reference 292

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:07:47.891206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:9baf74847256de06980b5989dfd42c6f66185baee887e7a118d065316d20c728