Pith. sign in

Paper Citation Record · LEDGER

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 6 inbound Pith citation observations for arXiv:2507.03662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03662 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:12:32.717267Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:46:16.678202Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:47:29.934082Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62f4367f-cd43-4297-967f-69d6ae87b075 · outbound

This paper cites URL: " 'urlintro :=.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:30.497834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:30.497834Z digest=sha256:3904c2c89a27d0d3e13033020a80bfe0aaec70696cfa5cc794d9f338fb25dc89

Observation 25dbe0fd-3487-41a7-b685-04a8f11c4335 · outbound

This paper cites write newline.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:30.565244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:30.565244Z digest=sha256:ac40dedfdf038d289c0a608ee640279b8e0f46c5c2621c46a5ae91970a0a3d21

Observation a5c470c2-9592-4749-a3eb-23587dfd7d9d · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.140998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:30.799717Z digest=sha256:bf83df098fda4f7c80beeb8a29015b708b6643cc6f43ea57c05f7b85fc48c22b

Observation ca7bb3f9-e6e4-4ac1-8ab9-ecfc7afcde91 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:30.963700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:30.963700Z digest=sha256:7f06585e4be19b20466d92b00782cb6271bdf6f902ec31d555fb2b9d62d86709

Observation 97a054a0-bc05-4b2c-a4df-343e9ba49a7e · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.133118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:31.086504Z digest=sha256:18441cda0b2653f8aca9ef98b1d59896f57dd8ebac2b6b08a30c0b955464f745

Observation b299773b-7daa-495b-bb1a-9afbb6bba5c0 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:31.194346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:31.194346Z digest=sha256:df6222ef4e90c4e8736f3ac2c4faaee90f7aaa2f4c8df03610f87943ae8e6a84

Observation 625618f7-5876-4f5c-9d46-a9ca13d9fed9 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.125055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:31.342729Z digest=sha256:e0b6f786aae0a12b9a930572b4a3cf70a2f046ee67aa0fce6907366438a61668

Observation 3cc5fc6c-864d-44a4-b643-d10f4d2acc4a · outbound

This paper cites Alignment faking in large language models.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Alignment faking in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:31.417382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:31.417382Z digest=sha256:2902775598aa57a4f1c39e38dfe369e560d7b268627338acb24fd7bb81b7cc37

Observation cd9b383d-6907-4e68-a577-33a83da7a7c5 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Qwen2.5-Coder Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:31.510017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:31.510017Z digest=sha256:b0dce5555e4c95e018226fb36e9b778e8c4807477708644ff0b3be643536f5a4

Observation 6ef3bc1d-de03-4454-8877-edee4a792cf4 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.116840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:31.619729Z digest=sha256:9ae826bcbdfc06cdf05e7336abef9bfc30296e6e934d3b1ef3bf69c4eb91e890

Observation e5f17dd1-7cd4-439e-a439-a55953373809 · outbound

This paper cites What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:31.713901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:31.713901Z digest=sha256:072fbaabe21160b4d54d0df305fae7d7a234c75f1cff4e493fffdc70b3dcb9e4

Observation c72ec5de-4211-420b-96be-6e2b12e39de5 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:31.818876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:31.818876Z digest=sha256:e11b28ac510d1e780a1abbb3c2b4b52ee8a4740813f8fa920af2a9d7899e0627

Observation 4352885e-eeab-4b84-bd8f-08c00e0f0c34 · outbound

This paper cites Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:31.934646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:31.934646Z digest=sha256:ca99fce61f61793f7759a5191b6d0d4d8ecf174800aa8e0f77ca2d597b55cbdc

Observation 9295086e-065f-4692-bae0-5645d92b2092 · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.082534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.082534Z digest=sha256:ede8116abe6533a48533a4e5952111c9cbefc4bebaee7c162bbfac78826da9a0

Observation 89996c97-33f7-4b19-af30-445a6f932ffb · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Frontier Models are Capable of In-context Scheming

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.170069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.170069Z digest=sha256:af96b3166f2d898d2180f27d923728cf7c7779de7ac7d74bdeb5cb3b249a8fa1

Observation 7b2a8d3a-8b9f-4bc4-a50d-dc6304d93ee6 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.103446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:32.292132Z digest=sha256:164c98d8478ed398da3638fc312c7249a7d6abc382ce0b24b84df750715f9a8f

Observation b416694e-7921-4807-93b9-12890c8c63dc · outbound

This paper cites The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.398404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.398404Z digest=sha256:6ad9c5e978aba01bc8240f3d9cf17ab8133e85a8a7a6845b672fea5f5e6029a3

Observation 939f9b4c-eefc-4e53-a938-b2c1c323f77b · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.095322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:32.567511Z digest=sha256:3119d2b0e98c84bba85ce263ea1454268f76d2aa1f999307d6c891a19e836ab5

Observation bafb2894-8f73-459c-8955-986e3df6e39a · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.643003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.643003Z digest=sha256:566c0f4dfded9b6063eccdfcd8a1c2c1e2b39250aca601a4a552372a1decb412

Observation 65189b11-ca14-4d88-a6e2-8eccd78c7420 · outbound

This paper cites Large Language Model Alignment: A Survey.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Large Language Model Alignment: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.681144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.681144Z digest=sha256:68d39c5625bd01cb8377a39a7ac8a8cd6dccc5845048ffbf7a1424c936b0be5c

Observation 9f9a9da4-00a7-4c63-99f1-b22f0e2864bb · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.080992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:32.699969Z digest=sha256:d69f2719a03b469ccd08f76189c8b0e1ed7d7f97c11eba14446bd18cc58a3277

Observation 977ae365-45ac-4b9a-a33d-300974733ea9 · outbound

This paper cites Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:12:33.072891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:32.703259Z digest=sha256:d3167c4a64760aa89260c11a994fcd11ed4315b96ce500d9a4f61c774faf371e

Observation 197cdb0f-11bf-4a8f-b2f3-cdaac7a31d69 · outbound

This paper cites Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Tradeoffs Between Alignment and Helpfulness in Language Models with Steering Methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.706277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.706277Z digest=sha256:b29ffc9abe9c69812a361f67eb51a0b50d821499471f42405c4423f16b6dc66f

Observation bdee2e24-4579-4b38-a4bc-fb759751952f · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.709279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.709279Z digest=sha256:a73cc9390e223ad6c263d62694c25e13aaf7bed17a65c1dd6a0cfbad20de191a

Observation 21face3e-8415-4880-af1a-fe573448aef9 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.711845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.711845Z digest=sha256:1512d82033dbeb0a68593cb2a7888735154bdfe1de43125a960930732fc1ce7f

Observation 11ef7507-24b9-4b9b-8132-d4782471ece7 · outbound

This paper cites an unresolved cited work.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:12:33.063688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T20:12:32.714663Z digest=sha256:63dfd849797ca7dfd7b829e5f78d5e7b201f5cb0f95a1552f06ac33717e0b887

Observation 3798a0a1-33f5-473c-9ba6-a532c6038529 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Representation Engineering: A Top-Down Approach to AI Transparency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:12:32.717267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:12:32.717267Z digest=sha256:ff00f0eb5fd36f175641b998d9501acc796bc8d4ed521b94aaf73b3d05b1798e

Pith citing papers

Observation bb166d1b-0ba0-4538-9556-df44c085b1e4 · inbound

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer cites this paper.

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:39:26.793422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:36:44.104197Z digest=sha256:de87be4e846927c1bd5e3a7c4b8642fbb57a9a8071899e251fc4b48c0f8c8ff0

Observation a9e7daba-9750-43ab-aaa4-3a5d18d69f6c · inbound

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence cites this paper.

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T20:22:37.266421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T20:13:53.972585Z digest=sha256:16226aeb409dde9966c777b43914550e7d3cd9582bb858d156e12154cf2a738e

Observation 54b7d6bf-2f22-4b85-a8b1-d71a6906e547 · inbound

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective cites this paper.

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:23.094369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T20:04:17.744876Z digest=sha256:b3d22401bac54024ceac8be937d339ed0ce4918c6d7cbfc3b8b471c80d544d22

Observation 9f2ba424-287a-47fe-808c-60bbae68247a · inbound

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating cites this paper.

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:29.935594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T17:03:33.199645Z digest=sha256:385a70633e0b402d75ecaacad2803ad8b6d4f41a4002ad0341716d7cfd5210b5

Observation dae3c0ef-8390-43fa-9503-c5e75a1078e8 · inbound

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 cites this paper.

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5 Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T18:20:40.287006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T18:20:40.287006Z digest=sha256:9ec1096872fd8833867eccd15c2c91d977d8b7d4e9e3703514d70f2c395ef411

Observation 71991987-2695-4185-8692-2fddc292ff96 · inbound

Emergent Misalignment Recruits a Pre-existing Persona Subspace cites this paper.

Emergent Misalignment Recruits a Pre-existing Persona Subspace Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-01T07:46:16.678202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:46:16.678202Z digest=sha256:0f306ca26809b6139dca578c9e9d7fb228aeb2ad2efeb7dda255ad1254bade3a