Pith. sign in

Paper Citation Record · LEDGER

Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2012.09816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.09816 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:55:13.736803Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.312964Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa908253-a6f3-49aa-9a8a-f735aef5c6eb · inbound

Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition cites this paper.

Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:24:24.068765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-24T11:20:49.217671Z digest=sha256:787ca6662cec9a5c0638ff1b39106b62bcb5c37856fd03069079d57b482c3ad4

Observation 9d7127de-539c-4440-b771-13a68a10e26f · inbound

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? cites this paper.

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:36:55.226039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-25T07:36:55.087443Z digest=sha256:5159c2862931a4fa05c79c84c365777d41bc6053d06eac670e23cca8ffa52572

Observation e26a1e47-2346-4b40-8a38-178e9a2f7570 · inbound

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models cites this paper.

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:00:20.988097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-14T23:00:20.720030Z digest=sha256:d8e8884a00b9f901b65150a94795e9d668242160ed705b2c33f57e35e182e0ff

Observation b33b96a4-3cf4-41af-84b2-d8c087597fa8 · inbound

Self-Improvement in Language Models: The Sharpening Mechanism cites this paper.

Self-Improvement in Language Models: The Sharpening Mechanism Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.614276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.614276Z digest=sha256:ee91ca21e690246ab2c774773f814d3b95351409f4c7bada385b5db01e22fc5a

Observation 2b69744c-abff-4a85-a9ab-52589e056b54 · inbound

On Local Overfitting and Forgetting in Deep Neural Networks cites this paper.

On Local Overfitting and Forgetting in Deep Neural Networks Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:35.424919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:40:35.424919Z digest=sha256:fc98fdb00fc107cf03877f3d20b93d8452e28a294a4881cc9716c050c5f752be

Observation 0db7fd84-b5f2-4bac-9003-2cc1f562b098 · inbound

Multi-Branch Mutual-Distillation Transformer for EEG-Based Seizure Subtype Classification cites this paper.

Multi-Branch Mutual-Distillation Transformer for EEG-Based Seizure Subtype Classification Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:01.534949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:01.534949Z digest=sha256:f6f6e367d50b6f93de290068fb14c38589e34fe6449b31b79c5128fc0d967d92

Observation 78188984-e501-4903-ae45-27ab1ba7616b · inbound

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment cites this paper.

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T13:57:21.164110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:57:21.164110Z digest=sha256:4957b1acf93f64b4695bf33563d3d830b067beb6b298661448f97251b8d42ec1

Observation 01614683-ebbf-4344-8e6f-8b0787e9de1b · inbound

sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging cites this paper.

sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:03.502948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:03.502948Z digest=sha256:66c6f9d0cafd16cfbfe7cb848e04e576880de915d96e8673613ade542644090e

Observation 81093135-1811-4529-b77c-43b1a131e698 · inbound

Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions cites this paper.

Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:24:40.048724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:24:40.048724Z digest=sha256:02b53957a62b682edbb4047281dc93b7ca0bfcce0f61ec4082a9940900f905ca

Observation ff5ccdee-d2c8-4cda-a074-b70e877f76b2 · inbound

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? cites this paper.

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T21:50:24.493688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:50:24.493688Z digest=sha256:db87243f8e9507cd4b642ba6c44e1fb8ae86cb119ce4c814c92f7b7ed76220dd

Observation 028bb0b8-5f8f-4062-8df1-89ad6f74f7bb · inbound

Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation cites this paper.

Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T16:31:36.475961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:31:36.475961Z digest=sha256:f999e4b5cdabe10cc0b3eec259bca97d8804f5ecb8f226d9ebbb627115375ff1

Observation 2b59b83d-946c-452b-a971-f18dc250df86 · inbound

Model Fusion via Neuron Transplantation cites this paper.

Model Fusion via Neuron Transplantation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T21:40:12.790902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:40:12.790902Z digest=sha256:0a3a7bf19efceae27b16076f4a04f89bf5e5e6944758c13da013365b6b32a6cc

Observation fbe9a7a8-7549-4797-90cc-c253e3ea3f93 · inbound

Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods cites this paper.

Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.726608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:39:51.726608Z digest=sha256:08e175863c4d00a62af52fddacae8ddc8895ddcefbd709583f5a7d65ef68933c

Observation 9ac6405e-c393-4fe7-be5e-44e9302a72b6 · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.168645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.168645Z digest=sha256:c1e1a92f5d7d36f708e427058e2aa77a14136955481a78e7105de055e9d612ec

Observation 187ee80c-5b8b-4d2e-bde7-363c5159e6eb · inbound

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning cites this paper.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.328935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.328935Z digest=sha256:7216c3ca61b93150e1105abcce5fc03fbeb80a5deb313713895069c38ee7f220

Observation fa11f9c9-8ab7-4336-8494-efbd30b1a47b · inbound

Distillation of atomistic foundation models across architectures and chemical domains cites this paper.

Distillation of atomistic foundation models across architectures and chemical domains Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:21.874305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:21.874305Z digest=sha256:a4ee37ef59188572d5459eb999975f8fa39833f6450b00b04724b86e7824610d

Observation b0c73fea-4d4e-4224-a0ba-7902fca8efe4 · inbound

Learning Causality for Modern Machine Learning cites this paper.

Learning Causality for Modern Machine Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:46.623707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:46.623707Z digest=sha256:15e41c4f790fadbc013932acf3c41dadeae6d3a09dd265224cfee73286760f53

Observation 14de3fd5-95fe-467c-8942-efefb6e14fdb · inbound

Machine Understanding of Scientific Language cites this paper.

Machine Understanding of Scientific Language Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:35.399144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:35.399144Z digest=sha256:5aa59dc27c3ff7879fd3064fac27eff9897870867cd16e21b9a58647151c2548

Observation c799125e-ab5f-436e-9539-60df4ff2888d · inbound

Forget Me Not: Fighting Local Overfitting with Knowledge Fusion and Distillation cites this paper.

Forget Me Not: Fighting Local Overfitting with Knowledge Fusion and Distillation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.213224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.213224Z digest=sha256:5e65720a33c3baffc87aea25eefa4aebbfecb6c579de5c693f2f1233830ae900

Observation 70307df7-99c3-4c1e-8a12-2fd6a0d553f4 · inbound

Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning cites this paper.

Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:25:20.610099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:25:20.610099Z digest=sha256:50efc496c47f1adc12fefc03578fc34ad2b87641136c14005da68d02f7b434a9

Observation 5a7d25de-d0b7-4356-9b7a-ba51c792bef9 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.736803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.736803Z digest=sha256:4cff9ac3c18de39559932fbb326da4f5d26a1456b622f1e04a51cdddccf2f7d2

Observation 9263cb36-40c2-47f6-b769-3111ece9d798 · inbound

Provable Knowledge Acquisition and Extraction in One-Layer Transformers cites this paper.

Provable Knowledge Acquisition and Extraction in One-Layer Transformers Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:10:44.837889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T23:05:56.687644Z digest=sha256:61bfec2d4c098c44f559c8cec807157f430b28dfe125304b72256e136bf12f56

Observation e24103e9-513a-4050-903b-797f72bea4ea · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:11.004696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:11.004696Z digest=sha256:8dcd7ec221eb5448fff38f20dbfecf609737077fb03027343bdfa6c947020fca

Observation 33cf71be-8187-4fb2-a120-fce88488509f · inbound

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws cites this paper.

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:14:12.821429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:14:12.821429Z digest=sha256:6a0b11b2424edc8a8392699a004bb52d709dfad6b143200cc6a79392ab3cd196

Observation 5eceff6e-92a1-49d2-8345-7c95fc5731ec · inbound

Aggregate Models, Not Explanations: Improving Feature Importance Estimation cites this paper.

Aggregate Models, Not Explanations: Improving Feature Importance Estimation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:09:32.249407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:09:32.249407Z digest=sha256:9f68784ca23cc716c9bb5e46b3b011afed13df8f4eac3e1d116f70a2c6ebe1d1

Observation 8064f121-f4e9-43bd-8976-c31d09da6d65 · inbound

FLAME: Condensing Ensemble Diversity into a Single Network for Efficient Sequential Recommendation cites this paper.

FLAME: Condensing Ensemble Diversity into a Single Network for Efficient Sequential Recommendation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.560618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T17:29:06.884078Z digest=sha256:90be9aec46b831f7d47d6577a936d196956e7fa9e2fa815c78897e6d1956f52b

Observation 9f31e33e-2a2b-4eb1-9e76-b8836b0e9461 · inbound

Benign Overfitting in Adversarial Training for Vision Transformers cites this paper.

Benign Overfitting in Adversarial Training for Vision Transformers Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:18.193794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T02:58:29.672338Z digest=sha256:5a5fb2eac55a600f15da53b89788188c14fbd56caac3d67d29ed650a03d7ac65

Observation c18835db-2651-443d-8d1d-4373d5fe908c · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.250639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:6bb0c2eb5c553200eab6c1270920938403abc615678be00852a18c4e6d00655f

Observation 0c46d6dd-4716-4c9c-9a79-b60780bfae17 · inbound

Hierarchical Mixture-of-Experts with Two-Stage Optimization cites this paper.

Hierarchical Mixture-of-Experts with Two-Stage Optimization Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:26.781487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:58:04.218139Z digest=sha256:198ccab477ab5e8e812eeb4cd0fd3fc350919e35b20c570046cd6279d0e63b44

Observation 5355f14c-18dd-466e-8b83-980736b0a9f6 · inbound

Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer cites this paper.

Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.243409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T02:30:57.473569Z digest=sha256:2f6d6b012fbce3564e7bb167cf8d357e3b80028497a51580f7ff496c89c8de1a

Observation fbc414cf-db0b-482c-ae5f-d4d49302a57f · inbound

Quantifying and Defending against the Privacy Risk in Logit-based Federated Learning cites this paper.

Quantifying and Defending against the Privacy Risk in Logit-based Federated Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.932947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T19:17:23.711817Z digest=sha256:e436430ea1823779cf9624e36cdd85991926919646238b041a2e7834d5bf1a7e

Observation 312f287d-60f6-453f-89f0-e5449d3153e0 · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:30.199822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:5590e381761214eb50b0ee2b6cfa631d7d93ec4210f1263b735f8200cfeb2d24

Observation f4fe0ad6-68e4-4bad-a815-703f5f318db0 · inbound

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime cites this paper.

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:29:44.314748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T08:53:46.285233Z digest=sha256:f4506713784df5b11fc8be845214285106ba70f0fde873eccf20c93310c9465f

Observation bf22cc0a-7578-4111-a104-01848e4034b1 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:74bf4f8fabf2b00b92c5f424190827db501ae803d7c167a89ec39aa25dc3f80f

Observation 5a86c0ce-2077-42e8-9bfc-5fe801ea1dec · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.488598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.488598Z digest=sha256:e352eb68a1fe7c5f462ab608491ebe457ce689419c566a0479ab7158637f33bf

Observation cbb8f483-f147-468b-89b7-a361c91ac52d · inbound

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions cites this paper.

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T08:04:29.580811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:04:29.580811Z digest=sha256:be7280f3f582c544c51913b40edad981b904b8d109c9a26d267d66d10b32588d

Observation 45eca94a-1728-4cda-9c6c-3face8a7e4c1 · inbound

SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation cites this paper.

SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T23:44:20.615202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:44:20.615202Z digest=sha256:8c229932d096b3efa1a271af8baa93ae7055dbe0257fd279aed4bdaaa32368ce

Observation 037584ed-01d1-49c4-9ca5-695efc6b310c · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:45.240771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:45.240771Z digest=sha256:7bc8c5bdd084750e09c4239f23ed71cc1808ea5c76bf0704299099ec1baffe43