Pith. sign in

Paper Citation Record · LEDGER

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

As of 18 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.10236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10236 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:08.568150Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6224a955-ca34-442b-aeb5-51011391d676 · outbound

This paper cites Who's Harry Potter? Approximate Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Who's Harry Potter? Approximate Unlearning in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.784382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.784382Z digest=sha256:79711968f9c1a9b10c39d72178adba9ca7270dc2585f6b71d2c63ce062d21cc9

Observation 784e5666-a270-4ef5-a7d0-1fb1b4a0cedc · outbound

This paper cites 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods 6 Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.905230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.905230Z digest=sha256:a4bcdee9fae19c7438c4d7141860ff9b49da12820d5b56e769235f3540ca1d4e

Observation 063ad588-eda8-45ea-9ed0-4369c984792e · outbound

This paper cites Mistral 7B.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Mistral 7B

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.100886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.100886Z digest=sha256:3c992453207585e4e12b913460926abe6b7ce9778e07e6349c96b598fe407e07

Observation fa825a17-919c-4dcf-a4ee-b15251b88615 · outbound

This paper cites Eight Methods to Evaluate Robust Unlearning in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Eight Methods to Evaluate Robust Unlearning in LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.250365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.250365Z digest=sha256:cda0c655fdaf5f2f3eed0eefae24d168f68f6dcd1bc94b9b4bdce7f53072e3bb

Observation 7c4baf1d-82b2-43f2-ad2d-084173759af0 · outbound

This paper cites TOFU: A Task of Fictitious Unlearning for LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods TOFU: A Task of Fictitious Unlearning for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.358259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.358259Z digest=sha256:5b1740a22bc9465abccc17b8789c3f6746987af06f9bec5fafc02d3853a0e8da

Observation 81467daf-364b-48ba-afc0-c08fe4e21be5 · outbound

This paper cites McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods McKinney, Anvith Thudi, Juhan Bae, Tara Rezaei Kheirkhah, Nicolas Papernot, Sheila A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.534487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:37:07.411854Z digest=sha256:f0aaddae56b8cfdd7b7394217d54865bf5779a94d0185ac393c76e4088f1de49

Observation 3f40819a-a57e-4375-ae14-5ec259a731e8 · outbound

This paper cites Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.507769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.507769Z digest=sha256:6d28acdd7f2c2b340c95f673790a3bfe0c09e4ffcb5cdec4b68ee8213198c390

Observation 21c52f53-8e11-467b-b0bb-db5c958e57cc · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods tinyBenchmarks: evaluating LLMs with fewer examples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.670301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.670301Z digest=sha256:e5e64f7311689c297378c3bffa311aab9b43a97d344d1a6e4667962029db5b39

Observation 7e31cf72-4903-4ae0-b01f-ba26f6d10686 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.725074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.725074Z digest=sha256:5d4d4a0fb11b007fd4ac68a3faa6488fbf85e89ee578ede7f7535cb581a09fe5

Observation 7f7bb65d-affd-40f1-97b4-2e08bda272ab · outbound

This paper cites UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.816508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.816508Z digest=sha256:89c64d9a20091c178cbb8a944ad9b6b6cf33d1cfae5ba6964eeb708c3b7fb8ba

Observation 72cbd017-c19c-46bb-ae8f-632a1a95f874 · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.881113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.881113Z digest=sha256:29fc8ad2327cac12b9443e221f1778306cc65b71c3df90b4e3537375831c020b

Observation dd246f8c-77b8-44ad-b04e-a47d6ebc2c5e · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.975951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.975951Z digest=sha256:6beb4fc8324dd61db8f8e191856c11b9ff10809ad9583adf2524085db57a5428

Observation d61c81e3-d491-44ad-b92a-caff49b54915 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Jailbroken: How Does LLM Safety Training Fail?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.055115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.055115Z digest=sha256:9c8224ef492cfc2a6afa2826be2ba20b6554821457a59adbd84eb30e26731858

Observation 2497b4e0-902a-4fe1-87aa-e909ac647df9 · outbound

This paper cites In-Context Learning Can Re-learn Forbidden Tasks.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods In-Context Learning Can Re-learn Forbidden Tasks

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:37:08.751340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:37:08.141780Z digest=sha256:271fe0d4d8a96f5518030890091dcc88ac02b9c61cf862fea45ffc76857d5187

Observation e7b22fde-4c56-45b1-8e2e-19c8d5f09556 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Low-Resource Languages Jailbreak GPT-4

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.222926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.222926Z digest=sha256:ba4113b6cc88e4828d8c804df135b55e4490eac83f3f2d56a8d26678614e6083

Observation e9b446f5-9593-4ab4-8bb8-172c1d9c5ae5 · outbound

This paper cites Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.295456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.295456Z digest=sha256:c047afc2451e5edc66172d013827cae81be3658032d629c4b4de3fc6ae3f9592

Observation daf4a539-e064-49db-90ab-d9a500267018 · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Improving Alignment and Robustness with Circuit Breakers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:08.378893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:08.378893Z digest=sha256:f199bffb57d6fbe727ddbfc56d05fab0069ce3a2eca0e722fc57231bd0b18c48

Observation ff311a48-1010-47e9-b739-5526929b2600 · outbound

This paper cites [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods [2024], a subset of 100 data points selected from MMLU (Massive Multitask Language Understanding) Hendrycks et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.299362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:37:08.441861Z digest=sha256:2e213ac40144ab5f1bf85255a79fb4ce452325112a67e8ee86b7bd81ed8ae20a

Observation f746d274-9a40-4499-b45c-c5bba14a7e2d · outbound

This paper cites Right Format.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Right Format

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:09.131821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:37:08.568150Z digest=sha256:0c84e852423a4e018d291ae283f4e891512422afefe65f05945a0350b39b621b

Observation 381b5830-0b54-4e86-b56f-edf92e64f8bd · outbound

This paper cites The Elicitation Game: Evaluating Capability Elicitation Techniques.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods The Elicitation Game: Evaluating Capability Elicitation Techniques

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.012174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.012174Z digest=sha256:25ce9865b6c04d52ba6d55edbf216905bd842d050201e8a07db6967b2700add1

Observation 546dad7a-a2f3-4619-85ee-905b22e762e5 · outbound

This paper cites Continual Learning and Private Unlearning.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Continual Learning and Private Unlearning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:07.178037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:07.178037Z digest=sha256:f517971f6fdd738ab5a5f894a43cfec242ab9e10ba30d2580b58a0647206cce2

Observation 6e0fc9f6-87cb-4b06-91ca-970599e27cbd · outbound

This paper cites Erasing Conceptual Knowledge from Language Models.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Erasing Conceptual Knowledge from Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.829821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.829821Z digest=sha256:c738ee63a6251a567043cd15837a9acb7dc5a82ddd3b147ae36d5172e45e791a

Observation 9ac562e3-c5b1-46c0-8e04-7102890168a9 · outbound

This paper cites Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.683217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.683217Z digest=sha256:efbab4339174daac74575ba0bebfb2279ebde91763d4e01767954c31a75c3b19

Observation 9c1bd043-7b6f-478f-bcbb-b698ff529c13 · outbound

This paper cites Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods.

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:06.722626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:06.722626Z digest=sha256:ba08ed71b375fe47a2a4695520890e1ef72abac5cc704d8502c15d8d321681b5

Pith citing papers

No inbound Pith citation observations are available.