Pith. sign in

Paper Citation Record · LEDGER

Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2012.09816.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2012.09816 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:55:13.736803Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.312964Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa908253-a6f3-49aa-9a8a-f735aef5c6eb · inbound

Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition cites this paper.

Easy Ensemble: Simple Deep Ensemble Learning for Sensor-Based Human Activity Recognition Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:24:24.068765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-24T11:20:49.217671Z digest=sha256:9363d45cfb9b868e3967e4e2e1662fa2d2d23fec0ed4ce0657bbb1d871b8f002

Observation 9d7127de-539c-4440-b771-13a68a10e26f · inbound

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? cites this paper.

TinyStories: How Small Can Language Models Be and Still Speak Coherent English? Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:36:55.226039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T07:36:55.087443Z digest=sha256:fae2cbc9efe7dbcb514665714b7ef7f64c56ccce8b983b9af8e32020a22ccb23

Observation e26a1e47-2346-4b40-8a38-178e9a2f7570 · inbound

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models cites this paper.

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:00:20.988097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T23:00:20.720030Z digest=sha256:ceaca887a89cacfff5a44cd2c6f174419c7965672f744b4488672f676ceabfba

Observation b33b96a4-3cf4-41af-84b2-d8c087597fa8 · inbound

Self-Improvement in Language Models: The Sharpening Mechanism cites this paper.

Self-Improvement in Language Models: The Sharpening Mechanism Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.614276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.614276Z digest=sha256:ee91ca21e690246ab2c774773f814d3b95351409f4c7bada385b5db01e22fc5a

Observation 2b69744c-abff-4a85-a9ab-52589e056b54 · inbound

On Local Overfitting and Forgetting in Deep Neural Networks cites this paper.

On Local Overfitting and Forgetting in Deep Neural Networks Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:40:35.424919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:40:35.424919Z digest=sha256:fc98fdb00fc107cf03877f3d20b93d8452e28a294a4881cc9716c050c5f752be

Observation 0db7fd84-b5f2-4bac-9003-2cc1f562b098 · inbound

Multi-Branch Mutual-Distillation Transformer for EEG-Based Seizure Subtype Classification cites this paper.

Multi-Branch Mutual-Distillation Transformer for EEG-Based Seizure Subtype Classification Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:01.534949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:01.534949Z digest=sha256:f6f6e367d50b6f93de290068fb14c38589e34fe6449b31b79c5128fc0d967d92

Observation 78188984-e501-4903-ae45-27ab1ba7616b · inbound

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment cites this paper.

Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep Deployment Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T13:57:21.164110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:57:21.164110Z digest=sha256:4957b1acf93f64b4695bf33563d3d830b067beb6b298661448f97251b8d42ec1

Observation 01614683-ebbf-4344-8e6f-8b0787e9de1b · inbound

sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging cites this paper.

sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:03.502948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:03.502948Z digest=sha256:66c6f9d0cafd16cfbfe7cb848e04e576880de915d96e8673613ade542644090e

Observation 81093135-1811-4529-b77c-43b1a131e698 · inbound

Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions cites this paper.

Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:24:40.048724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:24:40.048724Z digest=sha256:02b53957a62b682edbb4047281dc93b7ca0bfcce0f61ec4082a9940900f905ca

Observation ff5ccdee-d2c8-4cda-a074-b70e877f76b2 · inbound

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? cites this paper.

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T21:50:24.493688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:50:24.493688Z digest=sha256:db87243f8e9507cd4b642ba6c44e1fb8ae86cb119ce4c814c92f7b7ed76220dd

Observation 028bb0b8-5f8f-4062-8df1-89ad6f74f7bb · inbound

Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation cites this paper.

Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T16:31:36.475961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:31:36.475961Z digest=sha256:f999e4b5cdabe10cc0b3eec259bca97d8804f5ecb8f226d9ebbb627115375ff1

Observation 2b59b83d-946c-452b-a971-f18dc250df86 · inbound

Model Fusion via Neuron Transplantation cites this paper.

Model Fusion via Neuron Transplantation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T21:40:12.790902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:40:12.790902Z digest=sha256:0a3a7bf19efceae27b16076f4a04f89bf5e5e6944758c13da013365b6b32a6cc

Observation fbe9a7a8-7549-4797-90cc-c253e3ea3f93 · inbound

Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods cites this paper.

Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.726608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:39:51.726608Z digest=sha256:08e175863c4d00a62af52fddacae8ddc8895ddcefbd709583f5a7d65ef68933c

Observation 9ac6405e-c393-4fe7-be5e-44e9302a72b6 · inbound

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models cites this paper.

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:13.168645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:13.168645Z digest=sha256:c1e1a92f5d7d36f708e427058e2aa77a14136955481a78e7105de055e9d612ec

Observation 187ee80c-5b8b-4d2e-bde7-363c5159e6eb · inbound

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning cites this paper.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.328935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.328935Z digest=sha256:7216c3ca61b93150e1105abcce5fc03fbeb80a5deb313713895069c38ee7f220

Observation fa11f9c9-8ab7-4336-8494-efbd30b1a47b · inbound

Distillation of atomistic foundation models across architectures and chemical domains cites this paper.

Distillation of atomistic foundation models across architectures and chemical domains Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:21.874305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:21.874305Z digest=sha256:a4ee37ef59188572d5459eb999975f8fa39833f6450b00b04724b86e7824610d

Observation b0c73fea-4d4e-4224-a0ba-7902fca8efe4 · inbound

Learning Causality for Modern Machine Learning cites this paper.

Learning Causality for Modern Machine Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:46.623707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:46.623707Z digest=sha256:15e41c4f790fadbc013932acf3c41dadeae6d3a09dd265224cfee73286760f53

Observation 14de3fd5-95fe-467c-8942-efefb6e14fdb · inbound

Machine Understanding of Scientific Language cites this paper.

Machine Understanding of Scientific Language Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:31:35.399144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:31:35.399144Z digest=sha256:5aa59dc27c3ff7879fd3064fac27eff9897870867cd16e21b9a58647151c2548

Observation c799125e-ab5f-436e-9539-60df4ff2888d · inbound

Forget Me Not: Fighting Local Overfitting with Knowledge Fusion and Distillation cites this paper.

Forget Me Not: Fighting Local Overfitting with Knowledge Fusion and Distillation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.213224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.213224Z digest=sha256:5e65720a33c3baffc87aea25eefa4aebbfecb6c579de5c693f2f1233830ae900

Observation 70307df7-99c3-4c1e-8a12-2fd6a0d553f4 · inbound

Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning cites this paper.

Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:25:20.610099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:25:20.610099Z digest=sha256:50efc496c47f1adc12fefc03578fc34ad2b87641136c14005da68d02f7b434a9

Observation 5a7d25de-d0b7-4356-9b7a-ba51c792bef9 · inbound

SDD: Self-Degraded Defense against Malicious Fine-tuning cites this paper.

SDD: Self-Degraded Defense against Malicious Fine-tuning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:55:13.736803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:55:13.736803Z digest=sha256:4cff9ac3c18de39559932fbb326da4f5d26a1456b622f1e04a51cdddccf2f7d2

Observation 9263cb36-40c2-47f6-b769-3111ece9d798 · inbound

Provable Knowledge Acquisition and Extraction in One-Layer Transformers cites this paper.

Provable Knowledge Acquisition and Extraction in One-Layer Transformers Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:10:44.837889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T23:05:56.687644Z digest=sha256:92361608525f34cce9596ca72c273663c30ae1f1eeee6e4907b162b1a0e76bca

Observation e24103e9-513a-4050-903b-797f72bea4ea · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:11.004696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:11.004696Z digest=sha256:8dcd7ec221eb5448fff38f20dbfecf609737077fb03027343bdfa6c947020fca

Observation 33cf71be-8187-4fb2-a120-fce88488509f · inbound

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws cites this paper.

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:14:12.821429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:14:12.821429Z digest=sha256:6a0b11b2424edc8a8392699a004bb52d709dfad6b143200cc6a79392ab3cd196

Observation 5eceff6e-92a1-49d2-8345-7c95fc5731ec · inbound

Aggregate Models, Not Explanations: Improving Feature Importance Estimation cites this paper.

Aggregate Models, Not Explanations: Improving Feature Importance Estimation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:09:32.249407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:09:32.249407Z digest=sha256:68235721e124c7bdcd9c6f6488ae9a95c50bfc01d5ced9b2a6831ded6f454a90

Observation 8064f121-f4e9-43bd-8976-c31d09da6d65 · inbound

FLAME: Condensing Ensemble Diversity into a Single Network for Efficient Sequential Recommendation cites this paper.

FLAME: Condensing Ensemble Diversity into a Single Network for Efficient Sequential Recommendation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.560618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T17:29:06.884078Z digest=sha256:9d7e720c82c6cc7afa5d874a397af677a242ef3e4f014e248ca70ec2d0cec710

Observation 9f31e33e-2a2b-4eb1-9e76-b8836b0e9461 · inbound

Benign Overfitting in Adversarial Training for Vision Transformers cites this paper.

Benign Overfitting in Adversarial Training for Vision Transformers Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:46:18.193794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T02:58:29.672338Z digest=sha256:59f8305d210fd1a4caf9c1be0fc0eedb69d96b019eddbd889ef07b53c0c3be0d

Observation c18835db-2651-443d-8d1d-4373d5fe908c · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.250639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:7853a6fc9e1aa0e87c70f36996394595650be0c90a21ff5024e7d21500073ad5

Observation 0c46d6dd-4716-4c9c-9a79-b60780bfae17 · inbound

Hierarchical Mixture-of-Experts with Two-Stage Optimization cites this paper.

Hierarchical Mixture-of-Experts with Two-Stage Optimization Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:26.781487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:58:04.218139Z digest=sha256:2d4d0e7b6e8e686ddd4473a3977e7d7fbeb25ece03edfd41eb4aa2b0ac57c174

Observation 5355f14c-18dd-466e-8b83-980736b0a9f6 · inbound

Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer cites this paper.

Generative Diffusion Prior Distillation for Long-Context Knowledge Transfer Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.243409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T02:30:57.473569Z digest=sha256:c05feb3aacaa47c58dc008334f81779f3ed4158243b5a71ac7bbed641d5b11c0

Observation fbc414cf-db0b-482c-ae5f-d4d49302a57f · inbound

Quantifying and Defending against the Privacy Risk in Logit-based Federated Learning cites this paper.

Quantifying and Defending against the Privacy Risk in Logit-based Federated Learning Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.932947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T19:17:23.711817Z digest=sha256:eb204681ca0f4d4748021fd2622900df47d2036d9fd39cd275d5bd44d6662ab3

Observation 312f287d-60f6-453f-89f0-e5449d3153e0 · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:30.199822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:b5f4093fe4a487fc453257882d4533b3403622fe4bd5f5086c82dbeb3f2d50a9

Observation f4fe0ad6-68e4-4bad-a815-703f5f318db0 · inbound

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime cites this paper.

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:29:44.314748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T08:53:46.285233Z digest=sha256:7632f13a315cafbfebd135bf01fa951abc0343a47d0aad53962c55cbff2ae4b9

Observation bf22cc0a-7578-4111-a104-01848e4034b1 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:74bf4f8fabf2b00b92c5f424190827db501ae803d7c167a89ec39aa25dc3f80f

Observation 5a86c0ce-2077-42e8-9bfc-5fe801ea1dec · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:36.488598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:36.488598Z digest=sha256:e352eb68a1fe7c5f462ab608491ebe457ce689419c566a0479ab7158637f33bf

Observation cbb8f483-f147-468b-89b7-a361c91ac52d · inbound

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions cites this paper.

A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T08:04:29.580811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:04:29.580811Z digest=sha256:be7280f3f582c544c51913b40edad981b904b8d109c9a26d267d66d10b32588d

Observation 45eca94a-1728-4cda-9c6c-3face8a7e4c1 · inbound

SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation cites this paper.

SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T23:44:20.615202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:44:20.615202Z digest=sha256:8c229932d096b3efa1a271af8baa93ae7055dbe0257fd279aed4bdaaa32368ce

Observation 037584ed-01d1-49c4-9ca5-695efc6b310c · inbound

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control cites this paper.

V-Simba: Unleashing the Architectural Potential of RL in Visual Continuous Control Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:45.240771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:45.240771Z digest=sha256:7bc8c5bdd084750e09c4239f23ed71cc1808ea5c76bf0704299099ec1baffe43