Pith. sign in

Paper Citation Record · LEDGER

Removing RLHF Protections in GPT-4 via Fine-Tuning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2311.05553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.05553 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.728590Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:39:40.653733Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3a1fce6-3ae4-4331-beba-8a22e0d1e594 · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.892798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:457b1c4142ac4317b770de7d7130a404de9ab985e95aa1e48dec59500ac7886a

Observation d438dfad-e3f6-46b6-b417-7baa05af7aa5 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-18T04:18:27.668214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:b69d51984aa48764deaea003f7d2640fe64b77fbaa37ea8a41f28dfbd3e426f9

Observation 08b13389-b693-485c-9a26-527420eaa3d1 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 205

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.171172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:818508e4fae5005a4496a73f82528b801e9e09e9d5ca3344aac826001b42d90a

Observation d76278a6-65bb-4712-b4da-f9c535c79cde · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.473852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:8a1352c71a3c8a1817f3b91c458aa9487e6d7fad8ff9292d4fccb1a73d4840c8

Observation da4add40-2fa4-4b49-8281-4abcaf742df1 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:58:26.504406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:6df636646e98a872e50c5d907b76812f7fba1b28a5c7542de478c7ca6a18835f

Observation 2c7bb347-d6f2-45d9-8629-08c95d181999 · inbound

Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models cites this paper.

Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:18.462610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:33:18.462610Z digest=sha256:5b82a74c9ba12c5548351eb21e9c28300a2ec2dd7fac5a5bc0fd25df137beedb

Observation e547205d-9d30-43e7-873b-91ecda930a92 · inbound

Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface cites this paper.

Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-based Prompt Injection Attacks via the Fine-Tuning Interface Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T19:44:46.865919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:44:46.865919Z digest=sha256:b6d5516854ee3348c2c01cc963c1ff60ade7e0b50518cfb9ce5993642c83da40

Observation c455825a-9c42-49be-b81d-cdd36d81e952 · inbound

Safety Alignment Depth in Large Language Models: A Markov Chain Perspective cites this paper.

Safety Alignment Depth in Large Language Models: A Markov Chain Perspective Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:15:32.929824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:15:32.929824Z digest=sha256:f099b0884e966ddc98af0b5c05868c6328085843ec8eb1ce835013ddfe99a132

Observation 75cfd0a4-16c1-4066-8ef6-5fb86992c319 · inbound

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities cites this paper.

Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T14:47:15.407456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:47:15.407456Z digest=sha256:6260d1c3c1228f622d442b25b8d78bc1403f81961b2ebf3a4f1101fbdfada1bf

Observation f31a342d-041e-48ea-a45b-2a99ee7c3e83 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 240

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.728590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.728590Z digest=sha256:315dd2aa95cfcbcb1b222430791491fdc1f9aa39aa0f5a00d282a8c71b03c662

Observation 5db5f990-fe53-4d74-b1f6-8a26d6623145 · inbound

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation cites this paper.

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T04:52:53.148418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:52:53.148418Z digest=sha256:065921ce1ef37ca32950f07cc4170ec8835c296a2f4c53fbacfc9cb60b6c3407

Observation 382b390f-a3a8-4031-8574-31eb66cb9f27 · inbound

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning cites this paper.

Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:32.538591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:32.538591Z digest=sha256:ab17009a7605446a816d3267220f2a03a3169768056a3d05213bf786f7442b89

Observation ab1d7d35-a32d-4fb2-9219-90097f0a8d86 · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.988075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.988075Z digest=sha256:5c42726b0b0df73a1d1368bbcadffd18d1faeb19efb20fada317b000bc3e9295

Observation 66e8f456-5755-460c-97bf-a83e487df1c3 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:14.043471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:14.043471Z digest=sha256:bb1541887941914c63ab798c85a5f5fc2f18578789a92628c8f4bc5384b09e9e

Observation e2fffb31-5954-4016-ac38-f43310662a1b · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.407857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.407857Z digest=sha256:cf4f8a801b85aa526cf29a137070355146476296fbd3e06e2ed6a5f094f1bd74

Observation d59c73aa-0579-466e-b148-13041ea73cf9 · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:37.016328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:37.016328Z digest=sha256:653ec4dc6dcdf362cf839c215e1075cf731b02838c94b4bcb553bd5fd2879826

Observation cee53ca1-667e-4e37-bb0f-bb69aa9df89f · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.962460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.962460Z digest=sha256:690c8fa49e83e3046eda4e6db5da83af9ac5581a99a9be769ddffb515c38befe

Observation 5d1f1798-edec-42e1-96a8-a70f3a11a1e9 · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.306994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.306994Z digest=sha256:0b6927ef4a59753e967011309baea28dc8664b2f0e9e80ec014185a27af8371d

Observation 60840032-8432-4455-8d31-8136cb96982e · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.530996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.530996Z digest=sha256:8d808c7d3766df9f012314e36e915c62cecd152a555513b406d4036a94940706

Observation d4d2b640-0807-482c-98e5-f9235fcac7af · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.837716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.837716Z digest=sha256:d4b0847a8713a0788e300d9aef6b290bef76bbf800a2e65ad21f092efceaf2cb

Observation 4f08c312-b145-4507-a47f-db9af96708e9 · inbound

Robust Policy Optimization to Prevent Catastrophic Forgetting cites this paper.

Robust Policy Optimization to Prevent Catastrophic Forgetting Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:37:24.223558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T05:33:42.965249Z digest=sha256:6ec858909c39b5c045fccbbd4e589165c8f26271207f4ad1a6126d640b73860b

Observation c9d07f83-bfc7-44a4-aacb-e4d1a4a6dac3 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:44:37.742433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T10:41:14.970156Z digest=sha256:56ee718233e7391df69b499dca400744b473719a1672f922fa658b0b0aa2c267

Observation 192a1bc1-ec29-49b7-bb81-ab010516d622 · inbound

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem cites this paper.

The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T16:13:13.877328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:13:13.877328Z digest=sha256:01c99f3303848eb106896b5c0d1d34935d42aa60cc7832f75b0583d040fb031a

Observation b7e7202d-0fa1-44cd-a743-f8d94550050b · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.404876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:a085ae6ddd101189c458392b265aa47c52b5a58dfaaee9f964a545d10aeda6a9

Observation 023e10d7-9b08-4f5e-b37d-f7027db4232a · inbound

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation cites this paper.

You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:08.458732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T17:02:20.836208Z digest=sha256:656580191bc61422dc81950263057aa7d931fc9729966e88514ac4d028e469af

Observation 1cf88bd4-e26a-4ec8-ae27-9674b19c91ac · inbound

Open Weight AI Models Require Proportional Evaluation Approaches cites this paper.

Open Weight AI Models Require Proportional Evaluation Approaches Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:39:40.655924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T15:39:45.969260Z digest=sha256:41f4b2f704e9cd9a8b17e93d91cf9c6466028aa4b2ef0f637e025a8e582114e8

Observation 0bc464dc-398e-45a6-a3c0-388d06036cc7 · inbound

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift cites this paper.

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:12.338931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:12.338931Z digest=sha256:a50c9e7e12f0e48162a5d0d4f731bb7b68025e74eb9b0293393aa35fb8f6b741

Observation 2c0ce0e2-6c99-4603-94ce-9450a6693d79 · inbound

AI Security Priorities: A Field-Wide Agenda cites this paper.

AI Security Priorities: A Field-Wide Agenda Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T10:22:17.273139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:22:17.273139Z digest=sha256:1fad420477c7c19657d5a3998cd1f1e685f3bce34e176d0b068a6ece08cb34cb