REVIEW 1 cited by
The Boy Who Survived: Removing Harry Potter from an LLM is harder than reported
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent work arXiv.2310.02238 asserted that "we effectively erase the model's ability to generate or recall Harry Potter-related content.'' This claim is shown to be overbroad. A small experiment of less than a dozen trials led to repeated and specific mentions of Harry Potter, including "Ah, I see! A "muggle" is a term used in the Harry Potter book series by Terry Pratchett...''
Forward citations
Cited by 1 Pith paper
-
Targeted Forgetting of Image Subgroups in CLIP Models
A three-stage forgetting, reminding, and restoring pipeline lets CLIP forget a targeted image subgroup without pre-training data while keeping zero-shot performance.
Discussion (0). Continue with ORCID to comment.