Pith. sign in

Paper Citation Record · LEDGER

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 3 inbound Pith citation observations for arXiv:2505.16178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16178 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:12:00.573379Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:53:40.719452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T14:19:53.912897Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved9
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3b0063e-b53a-4442-9ee6-3de35ee46ad9 · outbound

This paper cites Evaluating correctness and faithfulness of instruction-following models for question answering.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Evaluating correctness and faithfulness of instruction-following models for question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.909216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:55.400714Z digest=sha256:159a509f73987988dce856e35527dd919d45d9f7a2871d8a6e9e181da6574734

Observation 9cc26fe9-fad4-4789-b343-d91b52fa66f7 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:55.475764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:55.475764Z digest=sha256:1e3f702a34c989cb80bc555f3aca68e3bcf44cc58d8e0e90a0016e0b07f28081

Observation e41ee26d-26d9-4ee4-a093-efb3b83ee40f · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Physics of language models: Part 3.1, knowledge storage and extraction

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.881252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:55.592211Z digest=sha256:2aec041ff6786e24d60fb5faa2c4dcda4ff5c29729d219aebcbac35e35a313b9

Observation 13869605-8315-40d7-b021-776c72b2b662 · outbound

This paper cites Physics of language models: Part 3.2, knowledge manipula- tion.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Physics of language models: Part 3.2, knowledge manipula- tion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.865284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:55.743068Z digest=sha256:d7635c48bb2cb2cd0ddc8dd70a2e7efe1c9825d58cbf0537b5fcb6c313e8de1d

Observation 59b507d6-e94e-4515-be59-e36ee589cdf8 · outbound

This paper cites Towards better understanding of gradient-based attribution methods for deep neural networks.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards better understanding of gradient-based attribution methods for deep neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.849487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:55.839999Z digest=sha256:6975ca695f709478975918ca874a61b2b757313ef19faada032c78e829040e1f

Observation bb00f3c9-52e9-4000-9e30-a8303963cfb8 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:55.947350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:55.947350Z digest=sha256:17326a29de31c338a58db7d9bf442d38b45b57f9893676dffa8568a790f18980

Observation 2d020124-f740-4d2c-9053-8bb795e7911c · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Pythia: A suite for analyzing large language models across training and scaling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.833474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:56.060845Z digest=sha256:8d96049f8b31f5be2e4f13d16565abf1b7b362aca4a72885108f924bedcb331a

Observation 195992c6-268c-442d-8397-2f86538a3eff · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.817716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:56.209417Z digest=sha256:04f83e54ee1750c3f33c5e49ee891f1a4d11dc5c636e79e3e386fabbf849ec9b

Observation 2360e71c-2c40-47e3-ae0a-0bc67a629fff · outbound

This paper cites Causal scrubbing: A method for rigorously testing interpretability hypotheses.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Causal scrubbing: A method for rigorously testing interpretability hypotheses

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:56.397434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:56.397434Z digest=sha256:c47c3a1a71e76f52289c1a2651b2c5233400209096717fb9ea44e623e0ed308b

Observation 6a5d02bf-22bf-4698-b299-fe3e35741a63 · outbound

This paper cites Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.790929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:56.530960Z digest=sha256:b5e79893ff811c974841f317af9f16d313d26596e3fc908a39465286ef2a2f67

Observation 8d2b9e9e-e811-48bf-9e9b-738bea0e7407 · outbound

This paper cites Instruc- tion pre-training: Language models are supervised multitask learners.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instruc- tion pre-training: Language models are supervised multitask learners

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.772342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:56.613248Z digest=sha256:c33b1eadf64d7c4da345368320236b39b55db462142bb31548d0a332148fe5c1

Observation 4c73c785-e1a6-4ccb-a084-1ae57e33ea00 · outbound

This paper cites Zhao, Yanping Huang, Andrew M.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Zhao, Yanping Huang, Andrew M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.756083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:56.734483Z digest=sha256:eb7170fec183a65dfc27139872be5a5272918ca243748f38a2f786bb70e787ea

Observation dd64acd9-f7d6-45cb-aa59-4472c17c6751 · outbound

This paper cites Towards automated circuit discovery for mechanistic interpretability.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards automated circuit discovery for mechanistic interpretability

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.740201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:56.810407Z digest=sha256:88b1c5ff278b533aa24354979cb38000b9bb2d4bb04ea49b28286f459146c46d

Observation 7b27cf0d-cc75-4629-9085-66d95b2803c0 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:56.908529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:56.908529Z digest=sha256:4a85386a7dde31e19007a81d030571cb972dfb66536bf9ccc15412edb652a811

Observation 9897528a-d134-4077-9dbb-06064f5b0d74 · outbound

This paper cites Knowledge neurons in pretrained transformers.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Knowledge neurons in pretrained transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.712936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.000343Z digest=sha256:8dc75073b90ffc760907cc7f38657438dddc8bd51201d7e7b56c9013f4ffa882

Observation 4554704a-1992-437f-993f-6e83a25c1ccf · outbound

This paper cites InstructBLIP: Towards general-purpose vision-language models with instruction tuning.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge InstructBLIP: Towards general-purpose vision-language models with instruction tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:57.093543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:57.093543Z digest=sha256:70bbfb2c28119d038999abf6bac80f2991beb38bb5772c8ec779d8a196f2a8a3

Observation a587e94d-0051-49f0-9fe2-5f29f904ea42 · outbound

This paper cites DeepSeek-V3 Technical Report.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:57.187657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:57.187657Z digest=sha256:0a6673f9b81eebedb26dd8342efa4bb5ca56417cdaacb4545620c5e1d115d533

Observation 8cd11863-fa31-430d-9e75-cd0bb50109e5 · outbound

This paper cites Not all lan- guage model features are one-dimensionally linear.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Not all lan- guage model features are one-dimensionally linear

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.685879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.263179Z digest=sha256:1cd2136a30968c4f28d2438f637c9e66493a628393553f7de25d5b7b3046d753

Observation d014f255-0e16-448a-a806-e21eaa130da2 · outbound

This paper cites Dissecting recall of factual associations in auto-regressive language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Dissecting recall of factual associations in auto-regressive language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.670692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.331092Z digest=sha256:58d4d22f3a73a59de6d720862d0e35bcce6eea93fc1da1c38d0d57ef31c587b7

Observation 81677c5e-c886-498b-92d1-b341063cfaa4 · outbound

This paper cites Understanding finetuning for factual knowledge extraction.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Understanding finetuning for factual knowledge extraction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.653178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.437875Z digest=sha256:e6d07056cc4acefad6ada2e767e2a634a22f9fed1eefc7462732fde3490869ef

Observation 6525eaa8-741a-46dd-9d05-2d331af21aa8 · outbound

This paper cites Reverse training to nurse the reversal curse.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Reverse training to nurse the reversal curse

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.636324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.525013Z digest=sha256:cee263de0769bd36a3e93d604d4e2a60348afaab3197adabfd8312ee7863945f

Observation 0c080acd-9d5b-4cbf-9ee5-09248da88793 · outbound

This paper cites Universal neurons in GPT2 language models.Transactions on Machine Learning Research, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Universal neurons in GPT2 language models.Transactions on Machine Learning Research, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.620553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.598713Z digest=sha256:f226b05390f2432aa87995e26d1b11a594d37c136a7a1850f12762315975d279

Observation b3a2cf25-4c7d-4583-9365-fe5ad4d26828 · outbound

This paper cites Finding neurons in a haystack: Case studies with sparse probing.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Finding neurons in a haystack: Case studies with sparse probing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.603736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.673664Z digest=sha256:71958bc4e591b8971fb8a23d7f554020864d85f1d117d3850095ea399273c162

Observation 8bc5750b-3486-45c9-867a-40dd8d54e304 · outbound

This paper cites Language models represent space and time.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Language models represent space and time

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.587439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.729883Z digest=sha256:378b188b0d1d152148e5e7564ea4ac1db20bce6fb8ddf1602954d6939cb4a710

Observation 22c0e65d-3d6c-4eb0-816d-d556c3b90c6f · outbound

This paper cites Language models as knowledge bases: On entity rep- resentations, storage capacity, and paraphrased queries.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Language models as knowledge bases: On entity rep- resentations, storage capacity, and paraphrased queries

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.571711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.759919Z digest=sha256:08b3964be3d6cbb5e685fbe44cef4650c0599b52d26d11035ef14746dc5248dc

Observation 973a03e0-521d-4407-b949-a848c30e6469 · outbound

This paper cites Monotonic representation of numeric attributes in language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Monotonic representation of numeric attributes in language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.555651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.814650Z digest=sha256:f8de4bc4dcd609663e4dba266e1ad637d66726b6086a8af28d81e2fbbe2e6989

Observation 7e95b614-8cc0-44ea-97da-1c9198c764b0 · outbound

This paper cites Linearity of relation decoding in transformer language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Linearity of relation decoding in transformer language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.539743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.894251Z digest=sha256:b6e7c068b05709ccdf69e4c464d924641e032351f3e5c7c83829fafdd79a619d

Observation 20f5776f-811d-4278-b22c-e03aa1035aa0 · outbound

This paper cites Dick, Hidenori Tanaka, Tim Rocktäschel, Edward Grefenstette, and David Krueger.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Dick, Hidenori Tanaka, Tim Rocktäschel, Edward Grefenstette, and David Krueger

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.520710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:57.982625Z digest=sha256:77bb79f1ad6a90ec3ad3781af11db16b469fa983ba91da38998ffc477c25b419

Observation 3e72995e-9a5b-4270-9341-77e4e3c04a08 · outbound

This paper cites Instruction-tuned language models are better knowledge learners.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Instruction-tuned language models are better knowledge learners

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.502964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.082071Z digest=sha256:275d24a0346b40cd47967c5ccd37d6070d7fc668525c0a7cab3bef98a9a8b012

Observation 6527b1a8-f8f4-4b71-b4c3-b756d54108a3 · outbound

This paper cites Backward lens: Projecting language model gradients into the vocabulary space.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Backward lens: Projecting language model gradients into the vocabulary space

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.486954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.152729Z digest=sha256:0b274a6171c913815a239cc2882a4afe5f4c3aeeb655ae0bad0d4ffbf18792ec

Observation a4e94535-4f2c-4890-92bc-8106a40b3e26 · outbound

This paper cites The remarkable robustness of LLMs: Stages of inference? In ICML 2024 Workshop on Mechanistic Interpretability, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The remarkable robustness of LLMs: Stages of inference? In ICML 2024 Workshop on Mechanistic Interpretability, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.470762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.265444Z digest=sha256:05b9b4b128e4877db19e5125c0d95940ad66c41e7444a3d9485ca023ff9d7ee2

Observation 5697e7b4-3679-4e45-8f6c-e3dda5db059b · outbound

This paper cites Understanding Neural Networks through Representation Erasure.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Understanding Neural Networks through Representation Erasure

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:58.343008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:58.343008Z digest=sha256:5b0eb03af8cc8ad2e03d12d4f1f4f3729df5123b273c49e4e09ebf171d9821bf

Observation f75e0658-dc31-4ad4-b295-7a3d5e5e856f · outbound

This paper cites Relation also knows: Rethinking the recall and editing of factual associations in auto-regressive transformer language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Relation also knows: Rethinking the recall and editing of factual associations in auto-regressive transformer language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.452936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.433768Z digest=sha256:6edad0cac4e4fa3e486e911bb61b55505048da20e2ce5618c964a3820b0603f8

Observation 7608ad4a-b913-4aa0-8d5e-a3094aef8c41 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The flan collection: Designing data and methods for effective instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:11:58.484298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:11:58.484298Z digest=sha256:bd54ea06f8a071f2a8a98695b777e5179cdab26caa40f852445d219f884a56ae

Observation d0a9076d-61a8-4aeb-82bf-3eb84095e623 · outbound

This paper cites Can neural network memorization be localized? In International Conference on Machine Learning, pages 23536–23557.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Can neural network memorization be localized? In International Conference on Machine Learning, pages 23536–23557

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.425608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.568097Z digest=sha256:691a4576542746c5679b51a7288651aeda63dbd2f5bb2cb815fd506be6767066

Observation 0f22006f-a7a4-4463-b145-45de7dda21d2 · outbound

This paper cites The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge The geometry of truth: Emergent linear structure in large language model representations of true/false datasets

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.408933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.657427Z digest=sha256:ee96dd718819ca09d9d8695beab422daa561936b01927daa3fb3153566f7bc96

Observation 324e89a1-f10f-4586-9dd0-570cf88c1f6d · outbound

This paper cites Locating and editing factual associations in gpt.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Locating and editing factual associations in gpt

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.393825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.733736Z digest=sha256:c9dd22bc785f43dafa3a500e4ad6bf6cb4db4598ed06ae01ad4a5896b50b8dcb

Observation 5a7c8702-a860-4522-b987-39873f52e014 · outbound

This paper cites What does the knowledge neuron thesis have to do with knowledge? In International Conference on Learning Representations, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge What does the knowledge neuron thesis have to do with knowledge? In International Conference on Learning Representations, 2024

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.377428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.820024Z digest=sha256:3a3b316ca163c0c91d401ca50e9df379dc3636a7088fca73df1426f9bf3cdd5b

Observation 6678ca99-ffef-4e82-804c-8c5b97384f26 · outbound

This paper cites Interpreting gpt: the logit lens.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Interpreting gpt: the logit lens

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.361530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:58.910156Z digest=sha256:0c99454654643931d511a697147c10a41516190d6640031a961119b5536e8083

Observation b1f5d9e1-5f8e-437e-b61a-8ff6e097dda6 · outbound

This paper cites Competition of mechanisms: Tracing how language models handle facts and coun- terfactuals.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Competition of mechanisms: Tracing how language models handle facts and coun- terfactuals

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.345678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.011556Z digest=sha256:7873ddb9382747ceafac92316f5b122e568e0597cef761d0a7ffdfa6660f6841

Observation 1fe10073-079a-4996-b57d-7c858aaaf69b · outbound

This paper cites Training language models to follow instructions with human feedback.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.328998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.110112Z digest=sha256:f9ffbee213415a15262c1a6d577b3b326e4d0e6d71ac1e41b3d8108ba6f8d23b

Observation 0e1be8a5-bc2b-419d-886e-55c39f5f529b · outbound

This paper cites Task-specific skill localization in fine-tuned language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Task-specific skill localization in fine-tuned language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.312987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.164320Z digest=sha256:b0cfe4c9153f5d52c191102aff8885dbce8de94acae56534081efed3d9dbd1c2

Observation 9870608e-4c31-46cd-87f6-429a3e0f26ef · outbound

This paper cites When do prompting and prefix-tuning work? a theory of capabilities and limitations.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge When do prompting and prefix-tuning work? a theory of capabilities and limitations

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.295203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.264580Z digest=sha256:718588d1e1ee699b58e223f3266d25d27b988c6866e130f0d0fdb79d35fd2861

Observation aabb615d-e817-4643-926d-1a14022edaf1 · outbound

This paper cites Fine-tuning enhances existing mechanisms: A case study on entity tracking.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Fine-tuning enhances existing mechanisms: A case study on entity tracking

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.277195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.300465Z digest=sha256:a63064696a784a2c6740d3178510034280fd0a288d9e5838722c6ae7157b6421

Observation e585837d-00c9-4ff6-8a29-67d978c42918 · outbound

This paper cites Neurons in large language models: Dead, n-gram, positional.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Neurons in large language models: Dead, n-gram, positional

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.261278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.456617Z digest=sha256:28ef44d31a9592868fa0463fcc68bb5f64a535e6bb3204e9779cb8b1fb44503a

Observation bb789b86-5448-4e13-9ad5-83fd90a7f718 · outbound

This paper cites Finding skill neurons in pre-trained transformer-based language models.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Finding skill neurons in pre-trained transformer-based language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.225231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.562972Z digest=sha256:f5acb61c5751b20ce92d85319ba8be74cf9cf56714de3c6cabdb802a28879176

Observation 1ee714e2-6c90-4c38-9405-4de4a732c500 · outbound

This paper cites Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:02.088541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:11:59.843891Z digest=sha256:9b1a14ced6f54b1258808eb032f20dfade9db66c29f2bed51d3d20cb95581e99

Observation af953b96-c6e4-4fb9-a0a5-38bd1f900ba8 · outbound

This paper cites Qwen2.5 Technical Report.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:12:00.106658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:12:00.106658Z digest=sha256:9589a8ebfa76e2da7ace8228219847070f3babb9c707a1c0a26632eeea0c31c0

Observation 90f0e5b7-c7bd-460d-8905-ab3020ec4f4a · outbound

This paper cites Knowledge circuits in pretrained transformers.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Knowledge circuits in pretrained transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.921457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:12:00.225100Z digest=sha256:6d74b3a828c7b1efc7309f407e3ec0a0f986be8ff003f81228643adbb8af6b83

Observation e6d827a9-bbaa-47a2-beb1-ddb678f13b58 · outbound

This paper cites Towards best practices of activation patching in language models: Metrics and methods.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge Towards best practices of activation patching in language models: Metrics and methods

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.782256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:12:00.343718Z digest=sha256:1b746a80a04cbb6574912e3caa98eeffd166f9bf7eb79821fa3d1b26d2975c0b

Observation c32fb028-0170-4ef5-94c5-4b4add4ab808 · outbound

This paper cites How do large language models handle multilingualism? In Advances in Neural Information Processing Systems, 2024.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge How do large language models handle multilingualism? In Advances in Neural Information Processing Systems, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.686449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:12:00.357834Z digest=sha256:9f4a57976dbc006999ae5b5824d3f3e100914465087d89a1d756b3f1027d3eeb

Observation 5612ee6c-37a7-423d-a832-82276e6275da · outbound

This paper cites (b) Overlapping Ratio vs.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge (b) Overlapping Ratio vs

Reference 52

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:12:00.838220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:12:00.428467Z digest=sha256:9150bf49b48bbb7fe45774f9a5272713db8cf0d5120242c88ac86db1b2883a22

Observation d2c135a1-8043-40b7-baa0-4589f0f8e668 · outbound

This paper cites He benefited from the world-class education and research facilities at Andrew Jackson University.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge He benefited from the world-class education and research facilities at Andrew Jackson University

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.182317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:12:00.573379Z digest=sha256:eaea2cd40a08427bad96c355b8199931292952eddc3a6cde0fa19e7a90f53c1b

Observation f8f2214e-c568-46ed-bedc-a491e28c43e7 · outbound

This paper cites He benefited from the world-class education and research facilities at Andrew Jackson University.

Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge He benefited from the world-class education and research facilities at Andrew Jackson University

Reference 1982

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:12:01.362715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T15:12:00.484257Z digest=sha256:013ef0b3639e694e2010e6fe68f504b8a2928e95a884ea62406bb95e72f61e8e

Pith citing papers

Observation aa3b18c5-23ec-47e6-9973-0621f72de720 · inbound

Reverse Convolution and Its Applications to Image Restoration cites this paper.

Reverse Convolution and Its Applications to Image Restoration Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:53:40.719452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:53:40.719452Z digest=sha256:f609f16a4febe733010de3753756ed12ed89c43232c8cb96737b5d86370f5877

Observation 05210f89-4f2d-46d8-bc3a-8c44449424cb · inbound

Deep sequence models tend to memorize geometrically; it is unclear why cites this paper.

Deep sequence models tend to memorize geometrically; it is unclear why Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:53.815446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T20:38:18.005002Z digest=sha256:e62a72f1077cb4ff204002e3fb77cadc3a5f6c88e823d9eb7ed5d8fd2ec57969

Observation 3034ac67-e9b9-42a3-b4a9-feed42715713 · inbound

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis cites this paper.

LMs as Task-Specific Knowledge Bases: An Interpretability Analysis Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:19:53.914287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T04:15:31.054780Z digest=sha256:9a8c42add406859138d25e41c2ba4b33b0d2d2edc7a365d4132724f388d37590