Pith. sign in

Paper Citation Record · LEDGER

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2505.16178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16178 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:12:00.573379Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:53:40.719452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T14:19:53.912897Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3b0063e-b53a-4442-9ee6-3de35ee46ad9 · outbound

This paper cites Evaluating correctness and faithfulness of instruction-following models for question answering.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Evaluating correctness and faithfulness of instruction-following models for question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.909216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:55.400714Z digest=sha256:bac6c96b78b189122f54137038285b01acafb63aa54e261c619071dd860886d4

Observation 9cc26fe9-fad4-4789-b343-d91b52fa66f7 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:55.475764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:55.475764Z digest=sha256:138d79518dcaafad21dad59ac11bf971a17eb1cb924c17974997f6a1cd84d192

Observation e41ee26d-26d9-4ee4-a093-efb3b83ee40f · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Physics of language models: Part 3.1, knowledge storage and extraction

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.881252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:55.592211Z digest=sha256:80df7e9e326628eabc0758044bcf1d18b2cc3c7a99dcb9ac6bf29e51a17f2ea0

Observation 13869605-8315-40d7-b021-776c72b2b662 · outbound

This paper cites Physics of language models: Part 3.2, knowledge manipula- tion.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Physics of language models: Part 3.2, knowledge manipula- tion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.865284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:55.743068Z digest=sha256:9c640e8e7e0a23c3f06ab0ba903aa41bdbb8f6d7b77ca5da41dba0cf42d1195c

Observation 59b507d6-e94e-4515-be59-e36ee589cdf8 · outbound

This paper cites Towards better understanding of gradient-based attribution methods for deep neural networks.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards better understanding of gradient-based attribution methods for deep neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.849487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:55.839999Z digest=sha256:bfa1b9388f3f675db724ac8943663a92a7a62819b35e7b32288b25ada2037b5c

Observation bb00f3c9-52e9-4000-9e30-a8303963cfb8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:55.947350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:55.947350Z digest=sha256:806749177939274a2026dbd93f059a441f35c3926109063ee2d9b26878ff5fe5

Observation 2d020124-f740-4d2c-9053-8bb795e7911c · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Pythia: A suite for analyzing large language models across training and scaling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.833474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:56.060845Z digest=sha256:4898328a5892b31f74f7327f83a8619e00069ae8fb5585a4a7bc154747796016

Observation 195992c6-268c-442d-8397-2f86538a3eff · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.817716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:56.209417Z digest=sha256:f8d50a82573aaaa82c71309d5b8dc030c7787b39653610179e4bcdcda8792c9f

Observation 2360e71c-2c40-47e3-ae0a-0bc67a629fff · outbound

This paper cites Causal scrubbing: A method for rigorously testing interpretability hypotheses.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Causal scrubbing: A method for rigorously testing interpretability hypotheses

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:56.397434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:56.397434Z digest=sha256:3639fc0ca43cf20f56cb7b7975cbe3b9cbb1dab6a9a31c2d1b79dc72fc174687

Observation 6a5d02bf-22bf-4698-b299-fe3e35741a63 · outbound

This paper cites Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.790929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:56.530960Z digest=sha256:2b018be19e851fe55e761de3cb6af33e992081286fbcea44eca4929f05c4c977

Observation 8d2b9e9e-e811-48bf-9e9b-738bea0e7407 · outbound

This paper cites Instruc- tion pre-training: Language models are supervised multitask learners.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instruc- tion pre-training: Language models are supervised multitask learners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.772342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:56.613248Z digest=sha256:870a751b5e616547e4d7253a066d38ce2f35cbf65ad606d728dda7c350e09097

Observation 4c73c785-e1a6-4ccb-a084-1ae57e33ea00 · outbound

This paper cites Zhao, Yanping Huang, Andrew M.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Zhao, Yanping Huang, Andrew M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.756083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:56.734483Z digest=sha256:217dda92c7a1d2639632b757acc6730066161834d8cce53bc83ce4e1172686b4

Observation dd64acd9-f7d6-45cb-aa59-4472c17c6751 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards automated circuit discovery for mechanistic interpretability

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.740201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:56.810407Z digest=sha256:3923394e793b80a753859b352de2a39d5419d35086c0537acf1cad48492d599d

Observation 7b27cf0d-cc75-4629-9085-66d95b2803c0 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:56.908529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:56.908529Z digest=sha256:5d8c02886a81567fda34580d13900c072b338bda372ac3e38bbcfa464ee14fa2

Observation 9897528a-d134-4077-9dbb-06064f5b0d74 · outbound

This paper cites Knowledge neurons in pretrained transformers.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Knowledge neurons in pretrained transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.712936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.000343Z digest=sha256:019d95dc8e6ab73ab45d89a205a71f80e34f7fa676c04023c5c4858072c48672

Observation 4554704a-1992-437f-993f-6e83a25c1ccf · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:57.093543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:57.093543Z digest=sha256:c1f8eba3dc045a9a7f994c549214de28422fbbb5a092a86cc6469bd7367dd430

Observation a587e94d-0051-49f0-9fe2-5f29f904ea42 · outbound

This paper cites DeepSeek-V3 Technical Report.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:57.187657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:57.187657Z digest=sha256:4df5bf4171e7aaf2928e6e87fa82c3bc353f06417b9d451a767044209c88992b

Observation 8cd11863-fa31-430d-9e75-cd0bb50109e5 · outbound

This paper cites Not all lan- guage model features are one-dimensionally linear.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Not all lan- guage model features are one-dimensionally linear

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.685879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.263179Z digest=sha256:36df4c21ded3fd75e06a2c54d489ad4d7347259b43ded4e77b741ebca13b5ea0

Observation d014f255-0e16-448a-a806-e21eaa130da2 · outbound

This paper cites Dissecting recall of factual associations in auto-regressive language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Dissecting recall of factual associations in auto-regressive language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.670692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.331092Z digest=sha256:05560bc199d3364c5315863be79291cc0c838a1f62ce654597e01e1383340805

Observation 81677c5e-c886-498b-92d1-b341063cfaa4 · outbound

This paper cites Understanding finetuning for factual knowledge extraction.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Understanding finetuning for factual knowledge extraction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.653178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.437875Z digest=sha256:e7dc3c27a44e310adf9bc7c5343fcaa8b63b829cfb8e8c6c0820ddb5e7fc2fb0

Observation 6525eaa8-741a-46dd-9d05-2d331af21aa8 · outbound

This paper cites Reverse training to nurse the reversal curse.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Reverse training to nurse the reversal curse

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.636324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.525013Z digest=sha256:199430215a71f50da050bc520505c093385409a4087c19666ee525723803655e

Observation 0c080acd-9d5b-4cbf-9ee5-09248da88793 · outbound

This paper cites Universal neurons in GPT2 language models.Transactions on Machine Learning Research, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Universal neurons in GPT2 language models.Transactions on Machine Learning Research, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.620553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.598713Z digest=sha256:0174c3b1e0503211eddf3fc87ba20e4f5fbcb73b4c1aff04bb11c81d3fb4331a

Observation b3a2cf25-4c7d-4583-9365-fe5ad4d26828 · outbound

This paper cites Finding neurons in a haystack: Case studies with sparse probing.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Finding neurons in a haystack: Case studies with sparse probing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.603736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.673664Z digest=sha256:2ddf1dae6d89147825e772565e054f1f4489df5391775c9239db8c4be0c611a4

Observation 8bc5750b-3486-45c9-867a-40dd8d54e304 · outbound

This paper cites Language models represent space and time.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Language models represent space and time

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.587439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.729883Z digest=sha256:274e5f3331ba23d5200e236b67c1d61dd7577a4dae58b25837dbf3950c51f90c

Observation 22c0e65d-3d6c-4eb0-816d-d556c3b90c6f · outbound

This paper cites Language models as knowledge bases: On entity rep- resentations, storage capacity, and paraphrased queries.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Language models as knowledge bases: On entity rep- resentations, storage capacity, and paraphrased queries

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.571711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.759919Z digest=sha256:ea78ff74e6305d3506c3d7e5119cf40e3029bfd592e19e60c82150e2a14200d8

Observation 973a03e0-521d-4407-b949-a848c30e6469 · outbound

This paper cites Monotonic representation of numeric attributes in language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Monotonic representation of numeric attributes in language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.555651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.814650Z digest=sha256:7d8ea81904cde60ff1ea5e31ffaa8880cd433c2c7a61ab07633b3dc5ae7b216d

Observation 7e95b614-8cc0-44ea-97da-1c9198c764b0 · outbound

This paper cites Linearity of relation decoding in transformer language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Linearity of relation decoding in transformer language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.539743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.894251Z digest=sha256:e3553677a7da1f65e4ad96df0a88124a4640d06c1e35ec75ff94ab31296a3c20

Observation 20f5776f-811d-4278-b22c-e03aa1035aa0 · outbound

This paper cites Dick, Hidenori Tanaka, Tim Rocktäschel, Edward Grefenstette, and David Krueger.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Dick, Hidenori Tanaka, Tim Rocktäschel, Edward Grefenstette, and David Krueger

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.520710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:57.982625Z digest=sha256:9715c48bb2b66979f8eab46ad1a463c0b0e48ba3b5d6fb5ed621e4dbec735fe8

Observation 3e72995e-9a5b-4270-9341-77e4e3c04a08 · outbound

This paper cites Instruction-tuned language models are better knowledge learners.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instruction-tuned language models are better knowledge learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.502964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.082071Z digest=sha256:7531189ac00a7f68084beee1beb8118025c06aabbed92fb0aef63bdb0908dc14

Observation 6527b1a8-f8f4-4b71-b4c3-b756d54108a3 · outbound

This paper cites Backward lens: Projecting language model gradients into the vocabulary space.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Backward lens: Projecting language model gradients into the vocabulary space

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.486954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.152729Z digest=sha256:0edf76dde0a23c7e007ba01c98fb94587a95947f7558efdc9d56f92b95f7d83a

Observation a4e94535-4f2c-4890-92bc-8106a40b3e26 · outbound

This paper cites The remarkable robustness of LLMs: Stages of inference? In ICML 2024 Workshop on Mechanistic Interpretability, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The remarkable robustness of LLMs: Stages of inference? In ICML 2024 Workshop on Mechanistic Interpretability, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.470762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.265444Z digest=sha256:8011d80b33658a2bee3b63ef887591665b83d5a44751397f45fa050a4b2b322b

Observation 5697e7b4-3679-4e45-8f6c-e3dda5db059b · outbound

This paper cites Understanding Neural Networks through Representation Erasure.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Understanding Neural Networks through Representation Erasure

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:58.343008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:58.343008Z digest=sha256:e60b7a1e5098c2568e70ef43bd8aa073fb146a292fc1d33e483f872b172a3ce2

Observation f75e0658-dc31-4ad4-b295-7a3d5e5e856f · outbound

This paper cites Relation also knows: Rethinking the recall and editing of factual associations in auto-regressive transformer language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Relation also knows: Rethinking the recall and editing of factual associations in auto-regressive transformer language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.452936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.433768Z digest=sha256:d638844c0a2f39102eb15091e3a42c909d33c3d1c92c0da89fff6d3a099fdf09

Observation 7608ad4a-b913-4aa0-8d5e-a3094aef8c41 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The flan collection: Designing data and methods for effective instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:58.484298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:58.484298Z digest=sha256:1afb4086b433d0635561ab6d8f6fff07d4dde07d8ce6ed8c6c6c860612505204

Observation d0a9076d-61a8-4aeb-82bf-3eb84095e623 · outbound

This paper cites Can neural network memorization be localized? In International Conference on Machine Learning, pages 23536–23557.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Can neural network memorization be localized? In International Conference on Machine Learning, pages 23536–23557

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.425608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.568097Z digest=sha256:86178545390c296f8920d74dd1e4136aaadfd7c4c9c29b1f68609cbe7f59a412

Observation 0f22006f-a7a4-4463-b145-45de7dda21d2 · outbound

This paper cites The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The geometry of truth: Emergent linear structure in large language model representations of true/false datasets

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.408933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.657427Z digest=sha256:e2cc9de55ee5112bbfa30f08909ee1009c0b5dc637d413e0bf385c8196b98c29

Observation 324e89a1-f10f-4586-9dd0-570cf88c1f6d · outbound

This paper cites Locating and editing factual associations in gpt.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Locating and editing factual associations in gpt

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.393825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.733736Z digest=sha256:4e57b4e5c8acea84b673a0beafe075c2975476cc29cd88eebb162fd1318b508d

Observation 5a7c8702-a860-4522-b987-39873f52e014 · outbound

This paper cites What does the knowledge neuron thesis have to do with knowledge? In International Conference on Learning Representations, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge What does the knowledge neuron thesis have to do with knowledge? In International Conference on Learning Representations, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.377428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.820024Z digest=sha256:20b5907c8f606bbf26c0f93ac154f7555e80515e3a8434def20d98589a461413

Observation 6678ca99-ffef-4e82-804c-8c5b97384f26 · outbound

This paper cites Interpreting gpt: the logit lens.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Interpreting gpt: the logit lens

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.361530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:58.910156Z digest=sha256:d8c51a70c83827c5b454bbac208bf3264e3ca2cd2b69ebbe1c5eed71346e5355

Observation b1f5d9e1-5f8e-437e-b61a-8ff6e097dda6 · outbound

This paper cites Competition of mechanisms: Tracing how language models handle facts and coun- terfactuals.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Competition of mechanisms: Tracing how language models handle facts and coun- terfactuals

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.345678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.011556Z digest=sha256:90373c27e06a2b5356108fcbec339354f66a99a3bc60e04f14bdfac464372e43

Observation 1fe10073-079a-4996-b57d-7c858aaaf69b · outbound

This paper cites Training language models to follow instructions with human feedback.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.328998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.110112Z digest=sha256:004668d251cd9ce6f0f04c2fafb50c6d650964823ffc9f3fa0d24efb1819ff53

Observation 0e1be8a5-bc2b-419d-886e-55c39f5f529b · outbound

This paper cites Task-specific skill localization in fine-tuned language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Task-specific skill localization in fine-tuned language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.312987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.164320Z digest=sha256:f289a411bb726f916b814b13afd9dcca27d085824a0e890518a011bcd760a8d7

Observation 9870608e-4c31-46cd-87f6-429a3e0f26ef · outbound

This paper cites When do prompting and prefix-tuning work? a theory of capabilities and limitations.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge When do prompting and prefix-tuning work? a theory of capabilities and limitations

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.295203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.264580Z digest=sha256:7dff1b1ad0f51853c00d1eca339f8ab360215c0e38b890124cbcd81a4cacb2d1

Observation aabb615d-e817-4643-926d-1a14022edaf1 · outbound

This paper cites Fine-tuning enhances existing mechanisms: A case study on entity tracking.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Fine-tuning enhances existing mechanisms: A case study on entity tracking

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.277195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.300465Z digest=sha256:d1d2e58d683f8cb65ca100d8da6ac44bcfaf7d75323db9000b914a5198f69457

Observation e585837d-00c9-4ff6-8a29-67d978c42918 · outbound

This paper cites Neurons in large language models: Dead, n-gram, positional.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Neurons in large language models: Dead, n-gram, positional

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.261278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.456617Z digest=sha256:60f686e618d21025d08f67ef9c4f38772b0d2aa409c087598306ce310b840b91

Observation bb789b86-5448-4e13-9ad5-83fd90a7f718 · outbound

This paper cites Finding skill neurons in pre-trained transformer-based language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Finding skill neurons in pre-trained transformer-based language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.225231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.562972Z digest=sha256:fe31702eb8b6f9896dc4800fed07963a9007a28ee78010a7aad18023d8ef1a05

Observation 1ee714e2-6c90-4c38-9405-4de4a732c500 · outbound

This paper cites Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.088541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:11:59.843891Z digest=sha256:3485c1ff4bd6d1effb77356aa31814261d27515a003112ddc172c7e17c3b2a90

Observation af953b96-c6e4-4fb9-a0a5-38bd1f900ba8 · outbound

This paper cites Qwen2.5 Technical Report.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:00.106658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:00.106658Z digest=sha256:874ca1f467b1db8df61db0cd30bcc7b49b85281482667b6fe3b982e15f0344ce

Observation 90f0e5b7-c7bd-460d-8905-ab3020ec4f4a · outbound

This paper cites Knowledge circuits in pretrained transformers.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Knowledge circuits in pretrained transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.921457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:12:00.225100Z digest=sha256:763319fe5bee442a4913514200014c11864160b7fefa4fe1d9525634d4fe753e

Observation e6d827a9-bbaa-47a2-beb1-ddb678f13b58 · outbound

This paper cites Towards best practices of activation patching in language models: Metrics and methods.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards best practices of activation patching in language models: Metrics and methods

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.782256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:12:00.343718Z digest=sha256:f9691eed8b1a6a9f3636df8f0c186a145d637de7c6a8e0f00a594e004a0a6dcd

Observation c32fb028-0170-4ef5-94c5-4b4add4ab808 · outbound

This paper cites How do large language models handle multilingualism? In Advances in Neural Information Processing Systems, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge How do large language models handle multilingualism? In Advances in Neural Information Processing Systems, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.686449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:12:00.357834Z digest=sha256:926ce6430a2e53ecec2f07ba1168d7562566d648a2613689e6b85e2ae28d7224

Observation 5612ee6c-37a7-423d-a832-82276e6275da · outbound

This paper cites (b) Overlapping Ratio vs.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge (b) Overlapping Ratio vs

Reference 52

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:12:00.838220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:12:00.428467Z digest=sha256:4b3d1898ed13809354363eda2bbffbc18f0172572961c70ed6f459d47d69f2a4

Observation d2c135a1-8043-40b7-baa0-4589f0f8e668 · outbound

This paper cites He benefited from the world-class education and research facilities at Andrew Jackson University.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge He benefited from the world-class education and research facilities at Andrew Jackson University

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.182317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:12:00.573379Z digest=sha256:6efc9067e09c269e5d872a4c5b33b1eb76d28e8231de2246271f2a49bd9cdd87

Observation f8f2214e-c568-46ed-bedc-a491e28c43e7 · outbound

This paper cites He benefited from the world-class education and research facilities at Andrew Jackson University.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge He benefited from the world-class education and research facilities at Andrew Jackson University

Reference 1982

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.362715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:12:00.484257Z digest=sha256:b9e958f062a093eda33b7f90556f52a1a29fb5e8e142a5b1611f00f1aeeecb4a

Pith citing papers

Observation aa3b18c5-23ec-47e6-9973-0621f72de720 · inbound

Reverse Convolution and Its Applications to Image Restoration cites this paper.

Reverse Convolution and Its Applications to Image Restoration Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:53:40.719452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:53:40.719452Z digest=sha256:3e771448e2ba21baeab0a9d03c3b0eac436c052ac5cf50225f49cb3001638d96

Observation 05210f89-4f2d-46d8-bc3a-8c44449424cb · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:53.815446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:8a92cc746c8fd1c2e005f66b92acbafc0089c7989da1f6e1c92a460b7ad4d6fc

Observation 3034ac67-e9b9-42a3-b4a9-feed42715713 · inbound

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis cites this paper.

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:19:53.914287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T04:15:31.054780Z digest=sha256:6a970bf919883fb17d3b0d355cb6ce19881a0bbfc44b3c775c53285fdaf172a5