Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:50:14.759452Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2502.07436.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:50:14.759452Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T23:29:02.457697Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
20 of 20 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation ab335bab-23d2-4e4f-ad30-f0bae640aea8 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f6d786-4436-4c22-983f-bab6d53ca99f · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9068f8ac-0ddf-4406-b958-8d7da0fa1248 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78b8bee-6c67-4fe5-ad22-3bce9a60a19f · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers For language pretraining tasks, we trained LLaMA models on the BabyLM dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64ee1a55-faa9-42fc-9d83-6d4ad651d957 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Distilling the Knowledge in a Neural Network
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4387d8a-7692-4398-be93-7dd99ac5b25b · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers 11 Submission and Formatting Instructions for ICML 2025 Figure
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4c4921f-a0c1-45ed-be5e-db6e6fb75890 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Improved Precision and Recall Metric for Assessing Generative Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83aac50c-a978-43d1-abf3-aa3bbc9611b5 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7de572-72d7-4a51-80b2-fa93d0779bc4 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Patient Knowledge Distillation for BERT Model Compression
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a4deb7-9465-4e06-8cf2-3d522917c86e · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c86c0533-4ef1-46a9-b04b-f9bc814c9412 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1e988b-8ccc-4f1b-872a-b7e935942d91 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad6542bb-8d4f-49ac-9436-755170fbc72c · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers ViTKD: Practical Guidelines for ViT feature knowledge distillation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b7c2b1-06b6-4e3b-9cef-0ebb3902576c · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce1c0edc-2323-4c8f-8410-c35cf4074639 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers TinyBERT: Distilling BERT for Natural Language Understanding
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6afcf6-867a-4516-9a2e-ebf2d6212a11 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f8a83dc-2af1-42c6-8018-1a4e84847cda · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Training data-efficient image transform- ers & distillation through attention
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1812a03-25c2-4ecc-b1cf-253c2d026f34 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 71d1e575-f4af-4a1e-8231-19ee730fd4bc · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af8114e-bc56-40e5-9dba-d72d80ca3b10 · outbound
Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66acbb21-a70f-4f7b-a051-e73e82b6bed4 · inbound
Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.