REVIEW 4 major objections 4 minor 67 references
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Masked-token loss erases targeted LLM content while keeping the model fluent
desk verdict A useful empirical unlearning recipe, but the headline claim of 'complete suppression' is not supported by the current evaluation and needs re-scoping. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the masked loss: for each forget-set document, the model's logits for target tokens are set to zero before softmax, and a KL divergence drives the full output distribution toward that masked distribution, enforcing zero-generation probability for the targeted content. Two regularizers carry the retained side: a distillation loss using mean-squared error between the student's logits and teachers trained on generic and other-style documents, and a world-fact loss using cross-entropy alignment on WikiText to protect encyclopedic knowledge. All updates go through LoRA adapters on the MLP and attention layers, and the forget/retain trade-off is governed by two hyperparameters $\lambda_1$ and $\lambda_2$. The paper's new evaluation object, DRMA, averages next-token probabilities over entire documents in the forget set, generalizing token-level remnant memorization accuracy to catch leakage that only appears after a long context.
What would settle it
Generate many paraphrases of forget-set facts using only words that are not on the extracted target-token list, prompt the unlearned model with them, and measure how often it completes them correctly; a non-negligible completion rate would refute the token-coverage premise. Equivalently, run a membership-inference attack on rephrased forget-set documents: if the model's perplexity signals the rephrased content as training data, the forgetting is surface-level only.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that aggressive forgetting does not have to destroy the model: zeroing the logits of an externally extracted list of target tokens, then rebalancing the model with two retention losses, yields an unlearned model whose forget-set outputs drop sharply while MMLU-scale utility and human-rated fluency stay near baseline levels. The method is presented as more robust than gradient ascent and related baselines, surviving membership-inference probing, 4-bit quantization, fine-tuning-based relearning, and jailbreak prompts. The paper also introduces document-level RMA (DRMA), a metric that averages per-token generation probabilities across whole documents to catch delayed leakage that token-level metrics miss.
Load-bearing premise
The entire method depends on the automatically extracted target-token list being complete enough that zeroing those tokens also suppresses the underlying knowledge, since a fact phrased with words outside the list would survive and the claimed membership-inference resistance would fail.
Editorial extensions
If this is right
- Unlearning a document collection such as a copyrighted book series can be done with a single LoRA fine-tuning run instead of retraining from scratch.
- If the forgetting is real, the unlearned model should resist membership-inference attacks that would otherwise flag forget-set documents as training data.
- The two retention losses are what keep utility and fluency alive; dropping either one degrades the balance in the ablations.
- Robustness results imply the forgotten content stays suppressed even after 4-bit quantization, fine-tuning-based relearning, and jailbreak prompting, not just under direct generation.
Reading between the lines
- The paper does not test whether paraphrases of forget-set facts remain generatable through tokens absent from the extracted list; that is the most direct untested implication of the token-coverage design.
- Because the target-token list comes from a proprietary model, the method's reproducibility and stability depend on that model's output; a cheaper or open-source token extractor could be a testable extension.
- The context-aware behavior shown in the examples suggests the same machinery could be tuned for selective forgetting of harmful combinations, such as 'computer' plus 'virus', while preserving benign uses of the same words.
- DRMA might serve as a general evaluation standard for document-level unlearning, since it is simple to compute and directly tracks long-horizon leakage.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes OBLIVIATE, an LLM unlearning framework that removes targeted data while preserving model utility and fluency. The method has three components: a masked loss that suppresses target tokens by optimizing a KL divergence between masked and original output distributions, a distillation loss that aligns the student model with teacher models on generic and other-style documents, and a world-fact loss that preserves encyclopedic knowledge via cross-entropy on WikiText. Fine-tuning is done with LoRA. The paper introduces DRMA, a document-level memorization metric, and evaluates on Harry Potter, WMDP, and TOFU, including robustness to membership inference, relearning, quantization, and jailbreaking attacks.
Significance. If the results hold, OBLIVIATE would be a practical and efficient unlearning method: it uses LoRA, requires no retraining, and shows strong forget quality while retaining MMLU-level utility across multiple models and datasets. The paper ships code, evaluates on three benchmarks with a broad robustness suite, and introduces a new metric (DRMA). These are concrete strengths. However, the core evaluation relies heavily on token-level metrics that align closely with the training objective, and the hyperparameters are selected on the same datasets used for final reporting, so the central claims of complete suppression and state-of-the-art performance are not yet fully supported.
major comments (4)
- [§3.4 and §3.3] The primary forget metric DRMA (Eq. 1) is the average next-token probability of the exact tokens in the forget documents, and the masked loss is designed to suppress exactly those tokens. Consequently, a low DRMA reflects the training objective by construction. It does not establish that the underlying knowledge is forgotten, especially since the target token list is extracted by GPT-4o and the Limitations section concedes "retrieval instability." The paper should add paraphrase-based tests (e.g., questions that express the same facts with different surface forms) and demonstrate that the association is removed, not only the likelihood of the original token sequences.
- [§4 and Appendix G] The hyperparameters λ1 and λ2 are selected via grid search on the same datasets and metrics used for the final reported results, as shown in Table 19 (Harry Potter) and the discussion in Section 4. This selection procedure can inflate apparent performance. The paper should either choose hyperparameters on a validation split that is not used for final evaluation, or report sensitivity analyses that allow the reader to assess the effect of this selection. Since the final numbers are presented as state-of-the-art, this issue is load-bearing.
- [§3.3] The claim that the masked loss "enforces zero-generation probability for targeted content" and "completely suppresses the generation of unlearning data" is an overstatement. The loss minimizes a KL divergence that encourages the total probability mass of target tokens to approach zero, but it does not provide a hard guarantee of zero probability at inference. Moreover, the implementation description is inconsistent: Section 3.2 says "zeroing out logits" while Section 3.3 says "probabilities ... to zero." Setting logits to zero (rather than -∞) does not produce zero probabilities after softmax. The authors should clarify the exact masking operation and temper the wording to "reduces" or "suppresses" rather than "enforces zero."
- [§4 (Tables 2-6)] All experimental results appear to be single-run with no standard deviations or significance tests. Given that several comparisons are within very small margins (e.g., Table 3: Zephyr-7B MMLU Ours=56.1 vs ELM=56.6; Table 4: TOFU utility Ours=62.44 vs Retain Model=62.38), the paper cannot robustly support "competitive" or "state-of-the-art" claims without quantifying variability. Reporting repeated runs with variance or significance tests is necessary for the central empirical claims.
minor comments (4)
- [§3.3 (Eq. 1)] The notation for the masked loss is underspecified: P(θ_masked) and Q(θ) are not defined as probability distributions over the vocabulary at each token position, and the summation over documents does not show how token positions are handled. Please clarify the exact tensors and the operation being summed.
- [Tables 2-5] The column formatting in the main tables is hard to parse; for example, the TOFU table (Table 4) shows values such as ppl=0.09 for Ours, which appears inconsistent with the definition of perplexity. Please check the alignment of columns and the direction of each metric (higher/lower is better) so that the reader can interpret the numbers correctly.
- [§4] The statement that λ1=0.2 and λ2=0.7 are selected "across all datasets" is not fully supported by the appendix, which only shows a grid search on Harry Potter (Table 19). Clarify whether the same values were used for WMDP and TOFU and whether any per-dataset tuning was performed.
- [§6 (Limitations)] The Limitations section honestly acknowledges the reliance on GPT-4o for token extraction and the resulting retrieval instability. This is a strength of the paper, but the abstract and contributions should be consistent with this caveat rather than claiming complete suppression without qualification.
Circularity Check
Forget-quality metric DRMA re-measures the exact target tokens the masked loss was built to suppress; the headline forgetting claim is partly circular, though utility, fluency, and robustness benchmarks remain independent.
-
fitted input called prediction
[Sections 3.2-3.4 and Tables 2-4 (target-token extraction, masked loss, DRMA evaluation)]
"we set the probabilities of the target tokens in the output distribution to zero, resulting in a masked logits distribution. We introduce a masked loss using KL divergence to minimize the difference between the masked and original logits distributions. ... DRMA = ... pθ(xt|x<t) ... A lower DRMA value indicates reduced document-level memorization."
The masked loss is minimized by making the model's output distribution match a masked distribution in which the GPT-4o-extracted target tokens are assigned zero probability. DRMA then averages pθ(xt|x<t) over the same forget-set documents whose next tokens include those target tokens, so suppressing the exact tokens that define the loss is measured as 'forgetting.' The paper's own limitation statement concedes that target-token extraction 'relies on GPT-4o, which introduces retrieval instability'; a fact expressible through tokens not on the list would be penalized by neither the loss nor DRMA. Token-likelihood MIAs (ppl, zlib, Min-K%) are similarly content-based and can fall while paraphrase-level knowledge remains.
full rationale
The central circularity is that the primary forget-quality metric, DRMA, is computed over the same documents and the same target tokens that the masked loss was explicitly constructed to zero out. This is a re-measurement of the optimization objective, not an independent test of whether the underlying knowledge is gone. The paper's own Limitations section admits retrieval instability in the GPT-4o-based target-token extraction, which is the load-bearing premise for the masked loss. Token-likelihood MIA metrics share the same surface-token limitation. However, the paper is not globally circular: model utility is verified on external benchmarks such as MMLU, WMDP-related multiple-choice degradation is a behavioral check beyond raw token suppression, and the relearning, quantization, and jailbreaking experiments provide independent empirical support. Self-citations to prior work by overlapping authors (e.g., Yao et al. 2024) are used for baselines or contextual motivation, not as a uniqueness theorem or as the sole justification for the central claim, so they do not add circularity. The result is a moderate, partial circularity centered on the forget-quality metric rather than a fully self-referential derivation.
Assumptions & free parameters
free parameters (2)
- lambda_1 (distillation loss weight) =
0.2
- lambda_2 (world-fact loss weight) =
0.7
assumptions (5)
- ad hoc to paper Masked KL divergence drives target-token probabilities toward zero, despite the formulation KL(P_masked || Q)
- domain assumption GPT-4o-identified target tokens cover the memorized content relevant to each forget set
- domain assumption The retain set, built from BM25-similar generics, other-style documents, and WikiText, is sufficient to preserve model utility and fluency
- domain assumption DRMA, MIA, and GPT-4o fluency scores measure genuine forgetting and fluency
- standard math Standard math of softmax, KL divergence, LoRA, and AdamW optimization is used correctly
Cite this review
Pith. "Pith review of OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models." pith.science (2026). https://pith.science/paper/5BKUHKVB
@misc{pith2026250504416,
author = {Pith},
title = {Pith review of: OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/5BKUHKVB}},
note = {Machine review of arXiv:2505.04416}
}
read the original abstract
Large language models (LLMs) trained over extensive corpora risk memorizing sensitive, copyrighted, or toxic content. To address this, we propose \textbf{OBLIVIATE}, a robust unlearning framework that removes targeted data while preserving model utility. The framework follows a structured process: extracting target tokens, building retain sets, and fine-tuning with a tailored loss function comprising three components -- masking, distillation, and world fact. Using low-rank adapters (LoRA) ensures efficiency without compromising unlearning quality. We conduct experiments on multiple datasets, including Harry Potter series, WMDP, and TOFU, using a comprehensive suite of metrics: \emph{forget quality} (via a new document-level memorization score), \emph{model utility}, and \emph{fluency}. Results demonstrate its effectiveness in resisting membership inference attacks, minimizing the impact on retained data, and maintaining robustness across diverse scenarios.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Li Bai, Haibo Hu, Qingqing Ye, Haoyang Li, Leixia Wang, and Jianliang Xu. 2025. Membership inference attacks and defenses in federated learning: A survey. ACM Comput. Surv. , 57(4):89:1--89:35
work page 2025
-
[4]
Choquette - Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot
Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette - Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. 2021. Machine unlearning. In S&P , pages 141--159
work page 2021
-
[5]
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tram \` e r, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In ICLR
work page 2023
-
[6]
Brown, Dawn Song, \' U lfar Erlingsson, Alina Oprea, and Colin Raffel
Nicholas Carlini, Florian Tram \` e r, Eric Wallace, Matthew Jagielski, Ariel Herbert - Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, \' U lfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting training data from large language models. In USENIX Security , pages 2633--2650
work page 2021
-
[7]
Lawrie, Daniel Khashabi, and Benjamin Van Durme
Jeffrey Cheng, Marc Marone, Orion Weller, Dawn J. Lawrie, Daniel Khashabi, and Benjamin Van Durme. 2024. Dated data: Tracing knowledge cutoffs in large language models. arXiv:2403.12958
arXiv 2024
-
[8]
Yijiang River Dong, Hongzhou Lin, Mikhail Belkin, Ram \' o n Huerta, and Ivan Vulic. 2024. Unmemorization in large language models via self-distillation and deliberate imagination. arXiv:2402.10052
arXiv 2024
Show all 67 references
-
[9]
Minxin Du, Xiang Yue, Sherman S. M. Chow, Tianhao Wang, Chenyu Huang, and Huan Sun. 2023. Dp-forward: Fine-tuning and inference on language models with differential privacy in forward pass. In CCS , pages 2665--2679
2023
-
[10]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al - Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Z...
2024 arXiv
-
[11]
Ronen Eldan and Mark Russinovich. 2023. Who's harry potter? approximate unlearning in llms. arXiv:2310.02238
2023 arXiv
-
[12]
Rohit Gandikota, Sheridan Feucht, Samuel Marks, and David Bau. 2024. Erasing conceptual knowledge from language models. arXiv:2410.02760
2024 arXiv
-
[13]
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. In EMNLP , pages 12216--12235
2023
-
[14]
Guan, Gregory Valiant, and James Zou
Antonio Ginart, Melody Y. Guan, Gregory Valiant, and James Zou. 2019. Making AI forget you: Data deletion in machine learning. In NeurIPS , pages 3513--3526
2019
-
[15]
Aditya Golatkar, Alessandro Achille, and Stefano Soatto. 2020. Eternal sunshine of the spotless net: Selective forgetting in deep networks. In CVPR , pages 9301--9309
2020
-
[16]
Tianle Gu, Kexin Huang, Ruilin Luo, Yuanqi Yao, Yujiu Yang, Yan Teng, and Yingchun Wang. 2024. MEOW: memory supervised LLM unlearning via inverted facts. arXiv:2409.11844
2024 arXiv
-
[17]
Varun Gupta, Christopher Jung, Seth Neel, Aaron Roth, Saeed Sharifi - Malvajerdi, and Chris Waites. 2021. Adaptive machine unlearning. In NeurIPS , pages 16319--16330
2021
-
[18]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In ICLR . OpenReview.net
2022
-
[19]
Minghao Hu, Junzhe Wang, Weisen Zhao, Qiang Zeng, and Lannan Luo. 2025. Flowmaltrans: Unsupervised binary code translation for malware detection using flow-adapter architecture. arXiv:2508.20212
2025 arXiv
-
[20]
Gabriel Ilharco, Marco T \' u lio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023. Editing models with task arithmetic. In ICLR
2023
-
[21]
Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo. 2022. Towards continual knowledge learning of language models. In ICLR
2022
-
[22]
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. 2023. Knowledge unlearning for mitigating privacy risks in language models. In ACL , pages 14389--14408
2023
-
[23]
Jiabao Ji, Yujian Liu, Yang Zhang, Gaowen Liu, Ramana Rao Kompella, Sijia Liu, and Shiyu Chang. 2024. Reversing the forget-retain objectives: An efficient LLM unlearning framework from logit difference. arXiv:2406.08607
2024 arXiv
-
[24]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \' e lio Renard Lavaud, Marie - Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, ...
2023 arXiv
-
[25]
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders S gaard. 2023. Copyright violations and large language models. In EMNLP , pages 7403--7412
2023
-
[26]
Hyunjik Kim, George Papamakarios, and Andriy Mnih. 2021. The lipschitz constant of self-attention. In ICML , pages 5562--5571
2021
-
[27]
Dohyun Lee, Daniel Rim, Minseok Choi, and Jaegul Choo. 2024. Protecting privacy through approximating optimal parameters for sequence unlearning in language models. In ACL , pages 15820--15839
2024
-
[28]
Jiaqi Li, Qianshan Wei, Chuanyi Zhang, Guilin Qi, Miaozeng Du, Yongrui Chen, Sheng Bi, and Fan Liu. 2024 a . Single image unlearning: Efficient machine unlearning in multimodal large language models. In NeurIPS
2024
-
[29]
Li, Ann - Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, Nathan Helm - Burger, Rassin Lababidi, Lennart Justen, Andrew B
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann - Kathrin Dombrowski, Shashwat Goel, Gabriel Mukobi, Nathan Helm - Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyu...
2024
-
[30]
Zitong Li, Qingqing Ye, and Haibo Hu. 2025. Funu: Boosting machine unlearning efficiency by filtering unnecessary unlearning. arXiv:2501.16614
2025 arXiv
-
[31]
Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao, and Haoyang Li. 2024. Why are my prompts leaked? unraveling prompt extraction threats in customized large language models. arXiv:2408.02416
2024 arXiv
-
[32]
Bo Liu, Qiang Liu, and Peter Stone. 2022. Continual learning and private unlearning. In CoLLAs , pages 243--254. PMLR
2022
-
[33]
Chris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, and Yang Liu. 2024 a . Large language model unlearning via embedding-corrupted prompts. In NeurIPS
2024
-
[34]
Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R. Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu. 2024 b . Rethinking machine unlearning for large language models. arXiv:2402.08787
2024 arXiv
-
[35]
Jaakkola, and Shiyu Chang
Yujian Liu, Yang Zhang, Tommi S. Jaakkola, and Shiyu Chang. 2024 c . Revisiting who's harry potter: Towards targeted unlearning from a causal intervention perspective. In EMNLP , pages 8708--8731
2024
-
[36]
Zheyuan Liu, Guangyao Dou, Eli Chien, Chunhui Zhang, Yijun Tian, and Ziwei Zhu. 2024 d . Breaking the trilemma of privacy, utility, and efficiency via controllable machine unlearning. In WWW , pages 1260--1271
2024
-
[37]
Michelle Lo, Fazl Barez, and Shay B. Cohen. 2024. Large language models relearn removed concepts. In Findings of ACL , pages 8306--8323
2024
-
[38]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR
2019
-
[39]
Alexandra Sasha Luccioni, Sylvain Viguier, and Anne - Laure Ligozat. 2023. Estimating the carbon footprint of bloom, a 176b parameter language model. J. Mach. Learn. Res., 24:253:1--253:15
2023
-
[40]
Lipton, and J
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C. Lipton, and J. Zico Kolter. 2024. TOFU: A task of fictitious unlearning for llms. arXiv:2401.06121
2024 arXiv
-
[41]
Matthieu Meeus, Shubham Jain, Marek Rei, and Yves - Alexandre de Montjoye. 2024. Did the neurons read your book? document-level membership inference for large language models. In USENIX Security
2024
-
[42]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT . In NeurIPS
2022
-
[43]
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In ICLR
2017
-
[44]
Feder Cooper, Daphne Ippolito, Christopher A
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette - Choo, Eric Wallace, Florian Tram \` e r, and Katherine Lee. 2023. Scalable extraction of training data from (production) language models. arXiv:2311.17035
2023 arXiv
-
[45]
Seth Neel, Aaron Roth, and Saeed Sharifi - Malvajerdi. 2021. Descent-to-delete: Gradient-based methods for machine unlearning. In ALT , pages 931--962
2021
-
[46]
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. 2024. In-context unlearning: Language models as few-shot unlearners. In ICML
2024
-
[47]
Fabio Petroni, Tim Rockt \" a schel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. 2019. Language models as knowledge bases? In EMNLP , pages 2463--2473
2019
-
[48]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS
2023
-
[49]
J.K. Rowling. 1997--2007. The Harry Potter Series. Bloomsbury. Comprising: Harry Potter and the Philosopher’s Stone; Harry Potter and the Chamber of Secrets; Harry Potter and the Prisoner of Azkaban; Harry Potter and the Goblet of Fire; Harry Potter and the Order of the Phoeni...
1997
-
[50]
Arya Roy. 2021. Recent trends in named entity recognition (NER) . arXiv:2101.11420
2021 arXiv
-
[51]
Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh. 2021. Remember what you want to forget: Algorithms for machine unlearning. In NeurIPS , pages 18075--18086
2021
-
[52]
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2024. Detecting pretraining data from large language models. In ICLR
2024
-
[53]
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership inference attacks against machine learning models. In S&P , pages 3--18
2017
-
[54]
Zachary Small. 2023. Sarah silverman sues openai and meta over copyright infringement. The New York Times
2023
-
[55]
Pratiksha Thaker, Shengyuan Hu, Neil Kale, Yash Maurya, Zhiwei Steven Wu, and Virginia Smith. 2024. Position: LLM unlearning benchmarks are weak measures of progress. arXiv:2410.02879
2024 arXiv
-
[56]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton - Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu,...
2023 arXiv
-
[57]
Rush, and Thomas Wolf
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Cl \' e mentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. 2023. Zephyr: Direct distillation of L...
2023 arXiv
-
[58]
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. 2023. Unveiling the implicit toxicity in large language models. In EMNLP , pages 1322--1338
2023
-
[59]
Chulin Xie, Yangsibo Huang, Chiyuan Zhang, Da Yu, Xinyun Chen, Bill Yuchen Lin, Bo Li, Badih Ghazi, and Ravi Kumar. 2024. On memorization of large language models in logical reasoning. arXiv:2410.23123
2024 arXiv
-
[60]
Xiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye, Haibo Hu, and Minxin Du. 2025. Unlearning isn't deletion: Investigating reversibility of machine unlearning in llms. arXiv:2505.16831
2025 arXiv
-
[61]
Jin Yao, Eli Chien, Minxin Du, Xinyao Niu, Tianhao Wang, Zezhou Cheng, and Xiang Yue. 2024. Machine unlearning of pre-trained large language models. In ACL , pages 8403--8419
2024
-
[62]
Hongbang Yuan, Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2025. Towards robust knowledge unlearning: An adversarial framework for assessing and improving unlearning robustness in large language models. In AAAI , pages 25769--25777
2025
-
[63]
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei. 2024. Negative preference optimization: From catastrophic collapse to effective unlearning. arXiv:2404.05868
2024 arXiv
-
[64]
Xulong Zhang, Jianzong Wang, Ning Cheng, Yifu Sun, Chuanyao Zhang, and Jing Xiao. 2023. Machine unlearning methodology based on stochastic teacher network. In ADMA , pages 250--261
2023
-
[65]
Zhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu, Xianfeng Tang, Hui Liu, Qi He, Wenpeng Yin, and Suhang Wang. 2025. Catastrophic failure of LLM unlearning via quantization. In ICLR
2025
-
[66]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. In NeurIPS
2023
-
[67]
Zico Kolter, and Matt Fredrikson
Andy Zou, Zifan Wang, J. Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv:2307.15043
2023 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.