Pith. sign in

Paper Citation Record · LEDGER

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.10236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10236 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:08.568150Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6224a955-ca34-442b-aeb5-51011391d676 · outbound

This paper cites Who's Harry Potter? Approximate Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Who's Harry Potter? Approximate Unlearning in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.784382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.784382Z digest=sha256:f08c88da7f8e20c96cfb0d39fb69b47244e4251a5fd6914a10ee6b836499128b

Observation 784e5666-a270-4ef5-a7d0-1fb1b4a0cedc · outbound

This paper cites 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.905230Z digest=sha256:a4bcdee9fae19c7438c4d7141860ff9b49da12820d5b56e769235f3540ca1d4e

Observation 063ad588-eda8-45ea-9ed0-4369c984792e · outbound

This paper cites Mistral 7B.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Mistral 7B

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.100886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.100886Z digest=sha256:1dfda8d226cabf452f81aa109043b0fa62498668276c157544c823e69b072ddb

Observation fa825a17-919c-4dcf-a4ee-b15251b88615 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.250365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.250365Z digest=sha256:ece1f0c57cb9c09924ed16c942f24df05424bc28306ba9d87c65b6fba0d47030

Observation 7c4baf1d-82b2-43f2-ad2d-084173759af0 · outbound

This paper cites TOFU: A Task of Fictitious Unlearning for LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods TOFU: A Task of Fictitious Unlearning for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.358259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.358259Z digest=sha256:1976247d88c08defc5e9ea80c14f6cb9f1a885b2520ca5c96bdc74820de0163f

Observation 81467daf-364b-48ba-afc0-c08fe4e21be5 · outbound

This paper cites McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.534487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:37:07.411854Z digest=sha256:99efe97d6e7b6e92319ec00b201ae66b98118621501a401aeeed407c3b667032

Observation 3f40819a-a57e-4375-ae14-5ec259a731e8 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.507769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.507769Z digest=sha256:b7f65d51de91f703c513f844eb650e3180665fc9b415034a1f0f092182203a55

Observation 21c52f53-8e11-467b-b0bb-db5c958e57cc · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods tinyBenchmarks: evaluating LLMs with fewer examples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.670301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.670301Z digest=sha256:38eda6dcb9e32a3f758af35ddd90e7e0c87b7f5aa80161cd457fb001aaea9bda

Observation 7e31cf72-4903-4ae0-b01f-ba26f6d10686 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.725074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.725074Z digest=sha256:271d9f3281611b47f7409b80bc7ba4a811b18eb9cb7d8c46174fe127a6aa5fa4

Observation 7f7bb65d-affd-40f1-97b4-2e08bda272ab · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.816508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.816508Z digest=sha256:a4e13ff7751776cce49c90ef680f70654bc900fa0406a669ed393d2fd865b29b

Observation 72cbd017-c19c-46bb-ae8f-632a1a95f874 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.881113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.881113Z digest=sha256:7a767a8f6e317d0f60a5d31869d344b5545061e6b3b89a981a255b5bef6693be

Observation dd246f8c-77b8-44ad-b04e-a47d6ebc2c5e · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.975951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.975951Z digest=sha256:c1f69feb854c158d099ebb16874291f22a78f019e8aed4b2d4e9e2d250d5c098

Observation d61c81e3-d491-44ad-b92a-caff49b54915 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Jailbroken: How Does LLM Safety Training Fail?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.055115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.055115Z digest=sha256:85b7f12c14539d5d0b89a60855830dc30d78db5d0276ed84a6a050d8f5197cd6

Observation 2497b4e0-902a-4fe1-87aa-e909ac647df9 · outbound

This paper cites In-Context Learning Can Re-learn Forbidden Tasks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods In-Context Learning Can Re-learn Forbidden Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:37:08.751340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:37:08.141780Z digest=sha256:22b667f5efb74e070d41b38c0ca922c7498b202ca4df3864287ac74b70ae3a53

Observation e7b22fde-4c56-45b1-8e2e-19c8d5f09556 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Low-Resource Languages Jailbreak GPT-4

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.222926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.222926Z digest=sha256:76b3106bd7fee8afda14feae15e58654eca7d52f114103e98a99b8c1cd4b7381

Observation e9b446f5-9593-4ab4-8bb8-172c1d9c5ae5 · outbound

This paper cites Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.295456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.295456Z digest=sha256:43b15f9a7f25899d1e13cc8a4ddb2831d2501f211d12208180d9ac469a0852fc

Observation daf4a539-e064-49db-90ab-d9a500267018 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Improving Alignment and Robustness with Circuit Breakers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.378893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.378893Z digest=sha256:aa657ffe3e35f3e13a05a8c290b005489484dab298612206785d33f3b5cbd822

Observation ff311a48-1010-47e9-b739-5526929b2600 · outbound

This paper cites [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.299362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:37:08.441861Z digest=sha256:a9093054b437e872e1734d98ddda4ba14abe16188fa40ea5da353ee7c8772438

Observation f746d274-9a40-4499-b45c-c5bba14a7e2d · outbound

This paper cites Right Format.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Right Format

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.131821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:37:08.568150Z digest=sha256:cb75ae303fe0afe5d70f323ea66f4c62f9806adf310b814cecca8df88ee4c6b3

Observation 381b5830-0b54-4e86-b56f-edf92e64f8bd · outbound

This paper cites The Elicitation Game: Evaluating Capability Elicitation Techniques.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods The Elicitation Game: Evaluating Capability Elicitation Techniques

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.012174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.012174Z digest=sha256:c69ddf39e75e8b8e6a705263a82987aa7c00623b587ffaf96e0633f278dc0d5a

Observation 546dad7a-a2f3-4619-85ee-905b22e762e5 · outbound

This paper cites Continual Learning and Private Unlearning.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Continual Learning and Private Unlearning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.178037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.178037Z digest=sha256:96753d7c3c1a16833d4d7363c7be46e50ad5c36212df827e896df9b8ed0337a7

Observation 6e0fc9f6-87cb-4b06-91ca-970599e27cbd · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Erasing Conceptual Knowledge from Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.829821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.829821Z digest=sha256:8e01f58f57358220e5fabca247e360f7aec74e80fdb20d75cb58ee8b8d67b719

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:f7e1eefa60783c714dddf66fa9cab798cbee32a820df7c7c5ea5e5f5e0ce3c08

Observation 9c1bd043-7b6f-478f-bcbb-b698ff529c13 · outbound

This paper cites Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.722626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.722626Z digest=sha256:f3cdd07dfb70714a8737350d463904bba47eeb85b164d548e27be5123e97cf9a

Pith citing papers

No inbound Pith citation observations are available.