Pith. sign in

Paper Citation Record · LEDGER

Efficient Knowledge Injection in LLMs via Self-Distillation

As of 14 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2412.14964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14964 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:52:37.270748Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:04:31.647548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.730179Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03e56360-c8f8-4595-8e27-5336625d58f7 · outbound

This paper cites GPT-4 Technical Report.

Efficient Knowledge Injection in LLMs via Self-Distillation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.086452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.086452Z digest=sha256:d120ec0b04d1ea17ba1bd7a82734be5414b5d2f1e92f67d55085e6a8f92ade39

Observation 47d91a7d-ffe6-4e50-8395-36040de64166 · outbound

This paper cites Adapting Language Models to Compress Contexts.

Efficient Knowledge Injection in LLMs via Self-Distillation Adapting Language Models to Compress Contexts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.104311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.104311Z digest=sha256:6f73296d8cf802bf5df6483ebc1ab1ea4f3d8ee3afb728c4cb62fa8e87c35154

Observation 56089bf7-56f8-46e6-a1b0-93035819f9a5 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Knowledge Injection in LLMs via Self-Distillation The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.112713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.112713Z digest=sha256:204071c92ab1818854bcd69116bb7989e44e70b47af4ebebed7b7de2e3b6b2ed

Observation 7470d0fb-9905-4fb6-9e1e-93189fe3665b · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Efficient Knowledge Injection in LLMs via Self-Distillation The False Promise of Imitating Proprietary LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.116782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.116782Z digest=sha256:5f74854409147a7fe518ce5fe136913b0e5d5d1e8fba3a89b3d9b1b87eed4518

Observation 7796b37f-3803-4045-92a5-8329e84c5a90 · outbound

This paper cites Don't Stop Pretraining: Adapt Language Models to Domains and Tasks.

Efficient Knowledge Injection in LLMs via Self-Distillation Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.125237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.125237Z digest=sha256:0bd9d1eb46f444932a66f31ee0f40355f4f4be72bc37e640f401ff2014f77614

Observation 9578ef48-9612-46c7-8956-972139b897bf · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Efficient Knowledge Injection in LLMs via Self-Distillation Distilling the Knowledge in a Neural Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.134081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.134081Z digest=sha256:dcecefaf6ac4dc9729cfb1b9c286dc2c64a4cba8a92cd80150ec0715e329505a

Observation 6126f122-a95f-4b8c-982d-f151609a3ea6 · outbound

This paper cites RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.143760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.143760Z digest=sha256:4169e7a72481254001e5a999c902226cf56247feb9c07022d8de688dd5ba4afd

Observation 2d269c3c-e99d-432c-9a0b-87bbf3a0aaa8 · outbound

This paper cites Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering.

Efficient Knowledge Injection in LLMs via Self-Distillation Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.148846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.148846Z digest=sha256:68e02f1f90bb669bf94c7b9be4bf949b08e9857e3cbcf3d3207a3c162a465760

Observation 805153a0-8015-4b22-91e7-c5dee28c07b2 · outbound

This paper cites Generalization through Memorization: Nearest Neighbor Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation Generalization through Memorization: Nearest Neighbor Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.153376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.153376Z digest=sha256:62642ef583b6955f69085edcee00ef7a4f2fbf7915ccb6b4d145b2849e424f34

Observation 763a6aab-fa9c-41fb-ac48-d46c6e206bb6 · outbound

This paper cites RA-DIT: Retrieval-Augmented Dual Instruction Tuning.

Efficient Knowledge Injection in LLMs via Self-Distillation RA-DIT: Retrieval-Augmented Dual Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.162226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.162226Z digest=sha256:c092e87fdd6caa4695714dda10b65df5cee39aca395c41e866c227be1be05c3b

Observation e457e219-f634-44d5-a729-f0ae9a5cce2f · outbound

This paper cites Structure-aware Domain Knowledge Injection for Large Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation Structure-aware Domain Knowledge Injection for Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:52:37.522227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.167306Z digest=sha256:21aec00e7799bbf540150869ae4224aea544f54dd73e94163d62fdaffbd44328

Observation ef842316-cf51-4bb9-a7ce-5973f06ce800 · outbound

This paper cites ChatQA: Surpassing GPT-4 on Conversational QA and RAG.

Efficient Knowledge Injection in LLMs via Self-Distillation ChatQA: Surpassing GPT-4 on Conversational QA and RAG

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.171818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.171818Z digest=sha256:614a1fb9ce3c170e6b31a30304dc6a99f63bf0f50ddb59f58952659710299143

Observation d9b8178e-5f41-4a8a-ab2b-ab4c75d8a82d · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

Efficient Knowledge Injection in LLMs via Self-Distillation When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.176676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.176676Z digest=sha256:0e32ed9902b8081cf97d74282609e433545b79ebe9b9c4e4c4d94da18c335286

Observation 03a92b77-3952-4e84-b964-e741a3d76cf3 · outbound

This paper cites Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning.

Efficient Knowledge Injection in LLMs via Self-Distillation Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.181478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.181478Z digest=sha256:5c1093ac1a64896a7db2a434ba0478871fc14874a8eef6c7ba635297cd693d21

Observation e55a76c8-fa4c-4235-b91f-3c9b16d18ec2 · outbound

This paper cites Orca 2: Teaching Small Language Models How to Reason.

Efficient Knowledge Injection in LLMs via Self-Distillation Orca 2: Teaching Small Language Models How to Reason

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.186069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.186069Z digest=sha256:317fa1c62c2c9a3bddb33ef215e381d130d39d78672124386de21b3ad1c1bed2

Observation 26f47e3f-fd91-49cd-95c9-4cc7954909a3 · outbound

This paper cites XtremeDistil: Multi-stage Distillation for Massive Multilingual Models.

Efficient Knowledge Injection in LLMs via Self-Distillation XtremeDistil: Multi-stage Distillation for Massive Multilingual Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.191047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.191047Z digest=sha256:ddaa374e2e663a95fa8735e03fc05d5836d00f0817509f879c225ebfb6075fad

Observation 7fe94a89-d997-4c74-b92f-321de45756cd · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Efficient Knowledge Injection in LLMs via Self-Distillation Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.195772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.195772Z digest=sha256:9fad249088b01d0fd4e724e5ac5370f8b426fbcb5b11212f90c6d58eb68a1756

Observation a180c0bd-0333-46ef-aed6-f19a61b16814 · outbound

This paper cites Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation.

Efficient Knowledge Injection in LLMs via Self-Distillation Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.200189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.200189Z digest=sha256:7f1f2842507d1c0c26cc971494c16ef26ce0e428132a3f84bf0f96c79198b065

Observation 46ae1da7-bb1f-4efc-9158-6a63d3e710b9 · outbound

This paper cites Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs.

Efficient Knowledge Injection in LLMs via Self-Distillation Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.204758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.204758Z digest=sha256:c169c6003f04acd1f78f48660fd68f4eda05cbd46164345bf825c8c7629afc79

Observation a7e5019a-b4ba-410f-b011-5f50298570a3 · outbound

This paper cites Instruction Tuning with GPT-4.

Efficient Knowledge Injection in LLMs via Self-Distillation Instruction Tuning with GPT-4

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.209209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.209209Z digest=sha256:23a27dbba15167cc3380bc9393db09d5f40f4c97b11354dc78628c812d0dc8e7

Observation 1154c72e-e39c-4a94-93eb-bcfe8809d26c · outbound

This paper cites In-Context Editing: Learning Knowledge from Self-Induced Distributions.

Efficient Knowledge Injection in LLMs via Self-Distillation In-Context Editing: Learning Knowledge from Self-Induced Distributions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.213565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.213565Z digest=sha256:c9c162cceab61ec26553d935b5a3631dff72630b386211d2c30cb2987f55e467

Observation 9f8816d2-bda9-4d46-8d19-94c6926739ec · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

Efficient Knowledge Injection in LLMs via Self-Distillation Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.222089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.222089Z digest=sha256:1793386228367c33c19756b774a71d64a1ac26a326fadbfc08dd523fd71d7812

Observation a716a0cc-f600-487b-afc5-c2031fd2c167 · outbound

This paper cites REPLUG: Retrieval-Augmented Black-Box Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation REPLUG: Retrieval-Augmented Black-Box Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.226208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.226208Z digest=sha256:abcef7282cf14c0527482aa2c07c4b4c44fe0ed4a1847ab7de59feac0236e7fe

Observation ed94ebae-474a-4acb-a0d6-1e8c7e463793 · outbound

This paper cites InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining.

Efficient Knowledge Injection in LLMs via Self-Distillation InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.230237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.230237Z digest=sha256:cc266c916e834639167c23d82d3b401778ffb20faed81e6a1d6e505ce50cf07b

Observation 6a66c2b8-fb06-43a0-8531-eda96ea69d6c · outbound

This paper cites In-Context Former: Lightning-fast Compressing Context for Large Language Model.

Efficient Knowledge Injection in LLMs via Self-Distillation In-Context Former: Lightning-fast Compressing Context for Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.234248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.234248Z digest=sha256:1a1e4d1238a7a951185cc3f921e3cdcedddf002837506ddb391cbdd0f89f6893

Observation 923c68cc-4294-4dd4-a31e-c97c1eb74b92 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Efficient Knowledge Injection in LLMs via Self-Distillation HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.238593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.238593Z digest=sha256:c46da4a1f65d2808ae79204fed0ed321a0f90d150cd3d4e302042e8c1c7ad423

Observation dbb9e362-91b0-48c8-ab4d-527e837e71b0 · outbound

This paper cites RAFT: Adapting Language Model to Domain Specific RAG.

Efficient Knowledge Injection in LLMs via Self-Distillation RAFT: Adapting Language Model to Domain Specific RAG

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.246602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.246602Z digest=sha256:ceeae8425129ea2549317e518960697aef082005634ac14481d70432656b82cf

Observation 41bedcf0-75ff-4798-ae30-f499cb4282e0 · outbound

This paper cites an unresolved cited work.

Efficient Knowledge Injection in LLMs via Self-Distillation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:52:37.826001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.250250Z digest=sha256:3e2e27c95b92782df287acbb72e6e354faf8178de9ad140dfece0ba6cd00db65

Observation 04e9f70e-b46c-466b-8655-729caef3457b · outbound

This paper cites C Related Work: More Detailed Review of Context Distillation In prior work, context distillation has been used for in-context learning and qualitatively modifying LLM behavior.

Efficient Knowledge Injection in LLMs via Self-Distillation C Related Work: More Detailed Review of Context Distillation In prior work, context distillation has been used for in-context learning and qualitatively modifying LLM behavior

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.811775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.254877Z digest=sha256:5dc8f8cd45ac56e6de9fbec3e1f205e88d553717182e9d9313ce8a4b736cbee4

Observation 9cbae46e-8cd4-4a99-91cb-f11120cdc24f · outbound

This paper cites We found Bonito capable of generating competitive questions for the New York Times dataset.

Efficient Knowledge Injection in LLMs via Self-Distillation We found Bonito capable of generating competitive questions for the New York Times dataset

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.796868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.258972Z digest=sha256:1caecd62f35ea3cce3f3250df0b8710cda2bba68bc6b5861e817a9a2a8b8a497

Observation 9b9bc418-12e8-4212-bca3-54286de8d69b · outbound

This paper cites To explore potential factors underlying this phenomenon, we examine two key statistical properties of the teacher model’s outputs:.

Efficient Knowledge Injection in LLMs via Self-Distillation To explore potential factors underlying this phenomenon, we examine two key statistical properties of the teacher model’s outputs:

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.783268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.262850Z digest=sha256:41af023f8a99be20f7680ff367e0cac2670732487372c0552c5a6d3ba8873d68

Observation 84748132-158c-4def-b593-fbe1bb9e97ad · outbound

This paper cites This characteristic may help explain why Llama-3- 8B-Instruct demonstrates superior performance as an expert compared to Qwen2.5-72B-Instruct (Table 3).

Efficient Knowledge Injection in LLMs via Self-Distillation This characteristic may help explain why Llama-3- 8B-Instruct demonstrates superior performance as an expert compared to Qwen2.5-72B-Instruct (Table 3)

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T11:52:37.766198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.266686Z digest=sha256:33de9248d6519b32699d25ced7819568192ed00b55d8023e0e9dc5b47a0a0c86

Observation 241a108e-0ec2-4009-a715-e53238617d53 · outbound

This paper cites Splendid Cities.

Efficient Knowledge Injection in LLMs via Self-Distillation Splendid Cities

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.752900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T11:52:37.270748Z digest=sha256:124553f52b762cbea707426a1dcc0355a1120ad36b1b147af324521fc81dbc8c

Observation 123ed824-d287-4346-b5ab-1066c915c06f · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Efficient Knowledge Injection in LLMs via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.217777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.217777Z digest=sha256:7ad0e44e0747936d7a29670481d027aa2f9819126f812258b520b78b6285bd32

Observation f5475e12-fc41-4d5d-a90e-3e6d776adc2f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation LoRA: Low-Rank Adaptation of Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.138767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.138767Z digest=sha256:dc6687d4fc4df0d72a393a19b6de308f503a11542319ebc4768667e67fdfc3ed

Observation 8b57b03b-a881-4fed-b327-2bbbb52cb362 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Efficient Knowledge Injection in LLMs via Self-Distillation Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.157917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.157917Z digest=sha256:379939e0208048a2505a91bb30198118e3d678e894ce7f5195ca3221af8bc620

Observation 54ab17d3-19ec-4103-9bdb-3bbfe2983a89 · outbound

This paper cites Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model.

Efficient Knowledge Injection in LLMs via Self-Distillation Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.242707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.242707Z digest=sha256:844e5e00c18b4f7d45dade196e62ec404404573c1bf8ae2af79421c604264122

Observation 072dd78b-6d67-4811-a065-c09d41319785 · outbound

This paper cites Prompt Injection: Parameterization of Fixed Inputs.

Efficient Knowledge Injection in LLMs via Self-Distillation Prompt Injection: Parameterization of Fixed Inputs

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.108513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.108513Z digest=sha256:43a08dcbd29273513742b2729553073c0ec37b56149a8afd4094b97816f8becc

Observation 992cede1-a7a6-4848-b9c0-527f3f93f629 · outbound

This paper cites Revisiting Self-Training for Neural Sequence Generation.

Efficient Knowledge Injection in LLMs via Self-Distillation Revisiting Self-Training for Neural Sequence Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.129494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.129494Z digest=sha256:dd44f7e3571af90933a1fed672c6878fc8ef6b100d4f131adb6cb6ee66ff9597

Observation 3973a647-d213-43d5-a2d9-69ab10f3e25e · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation Quantifying Memorization Across Neural Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.099890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.099890Z digest=sha256:eee07f991468b80b028cf7432eec78aaed531c0e07b5ea4b955f160ac8e5c669

Observation 5740aedf-cc3a-4b95-bee8-5e89c618310e · outbound

This paper cites Language models are few-shot learners.

Efficient Knowledge Injection in LLMs via Self-Distillation Language models are few-shot learners

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.095907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.095907Z digest=sha256:e58392143cc3ba03c8ac2a35f10e6f46a215b01f2bfcd55c2ab1952f4a5ed94b

Observation 2a09967e-9457-49e0-a6fd-06aa0687e991 · outbound

This paper cites RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture.

Efficient Knowledge Injection in LLMs via Self-Distillation RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.121350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.121350Z digest=sha256:06980289c9f0dbc9c96fd6d29b10712cf3be3fc3ab9341dff2a493f0cc5a6c54

Observation 28f18322-0f4e-4095-974c-ca52efd91650 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Efficient Knowledge Injection in LLMs via Self-Distillation A General Language Assistant as a Laboratory for Alignment

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.091278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.091278Z digest=sha256:52c543559f09f3cad8f71d9ade04d8a2acbeb2f0e686986da5c5928de7aa827e

Pith citing papers

Observation f90dd867-10d9-4bb4-b121-883f62cffc77 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.647548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.647548Z digest=sha256:89439216bda359ad95a5741dd81f2d364f005be2bc4191a0fab68e42043001bf

Observation e20039dd-2798-4308-8d3c-9870e2fd9dc7 · inbound

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe cites this paper.

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:26:16.201002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T17:00:49.448352Z digest=sha256:8b6bd5264add688924c0cad473d854a63921940cb7020c8b46454acf0afeb338

Observation fa696f57-f44a-4620-86a8-feb61f88b25c · inbound

Context Memorization for Efficient Long Context Generation cites this paper.

Context Memorization for Efficient Long Context Generation Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.658948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T10:39:09.720412Z digest=sha256:f030b0a3845bc1852acb1d925a3deaae82f178bf58e4f6573907642ced5aa96b

Observation 30f0f7d0-80a4-4a18-96fd-14e643e083d5 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.803207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T10:56:13.058872Z digest=sha256:25b91188c48d9c6f8fe10cc3322e7d15cefc94c0fba4e7ccec3380359c7bb090

Observation 9db19840-b7a8-45e5-9367-6f811e4c40c4 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T07:44:25.325808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:44:25.325808Z digest=sha256:e63300b2a3b6bbd54900584ad11dd342ee78e844506632a0927c5aff9ee3c714

Observation 2cbc327f-02d1-47a7-90a5-93bd6eb5c31b · inbound

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents cites this paper.

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.731556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T06:29:11.398007Z digest=sha256:7610b0dbaf578e3fe9bcc3a5b6655a3ef025c6a6de91ff49bd873a24608c587b

Observation 0d92bd82-ed75-4d1e-914f-a3c13bc959a5 · inbound

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA cites this paper.

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.802808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T05:07:58.441326Z digest=sha256:c392c5a43f7db73170d089edcd68f3d077843f62e894e48ef7084e2946b89367