Pith. sign in

Paper Citation Record · LEDGER

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2502.07436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07436 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:50:14.759452Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:29:02.457697Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ab335bab-23d2-4e4f-ad30-f0bae640aea8 · outbound

This paper cites GPT-4 Technical Report.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.653697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.653697Z digest=sha256:9c0122263855ef6018b281a3f9fbf8d1d34b0c01590564d471d19de3f508a4d9

Observation e7f6d786-4436-4c22-983f-bab6d53ca99f · outbound

This paper cites Language Models are Few-Shot Learners.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.674642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.674642Z digest=sha256:3fae3395c54919df0278f2aa1f1a12513c9f3f2e03d68594917f7ba7bd5ad343

Observation 9068f8ac-0ddf-4406-b958-8d7da0fa1248 · outbound

This paper cites The Llama 3 Herd of Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.679833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.679833Z digest=sha256:20f1b8ed3ba3bc2cf71c3c624df1d2fc012a66f84866a0e4612dc56336f154e8

Observation f78b8bee-6c67-4fe5-ad22-3bce9a60a19f · outbound

This paper cites For language pretraining tasks, we trained LLaMA models on the BabyLM dataset.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers For language pretraining tasks, we trained LLaMA models on the BabyLM dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.062540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:50:14.754909Z digest=sha256:fa2e9bf55b6a35b223eb6f8240915a53764b53104bb565b76dc731ed3c5fadc7

Observation 64ee1a55-faa9-42fc-9d83-6d4ad651d957 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Distilling the Knowledge in a Neural Network

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.689647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.689647Z digest=sha256:8320ed03b14aec7b999df92d5588b7c95e62aa67f86dd0993fae36c2f2254a6c

Observation e4387d8a-7692-4398-be93-7dd99ac5b25b · outbound

This paper cites 11 Submission and Formatting Instructions for ICML 2025 Figure.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers 11 Submission and Formatting Instructions for ICML 2025 Figure

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.046119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:50:14.759452Z digest=sha256:faf5577615e7efa7cfeafccbaea5754b79624956afe84a0e524e4970feb58392

Observation c4c4921f-a0c1-45ed-be5e-db6e6fb75890 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Improved Precision and Recall Metric for Assessing Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.704900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.704900Z digest=sha256:301cf3a9e5dc9a5b1ddea4a78ec82e1967472c3b02d014f4303f8242e111be2a

Observation 83aac50c-a978-43d1-abf3-aa3bbc9611b5 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.720671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.720671Z digest=sha256:b913af287de0bd09699557d1defa1dcbb38fc70190f33287439221310167cc07

Observation 3a7de572-72d7-4a51-80b2-fa93d0779bc4 · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Patient Knowledge Distillation for BERT Model Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.725290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.725290Z digest=sha256:b71204ab97ea5b0b7a74227278b5a72116a997feaff9f3b38fc539d3bc715f13

Observation a1a4deb7-9465-4e06-8cf2-3d522917c86e · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.730479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.730479Z digest=sha256:224b87e2860912c511d2f1b4eaa6374885781fa3ffa8e23624e3e0d5c83e98f2

Observation c86c0533-4ef1-46a9-b04b-f9bc814c9412 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.739767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.739767Z digest=sha256:8f4a38a0acace8f8e7bd0a0e159a6db464a7c4b2c24bc23fd7917bca80cab8b9

Observation 6f1e988b-8ccc-4f1b-872a-b7e935942d91 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.744931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.744931Z digest=sha256:1484d66c4723814cec173ed4ff44a1b405d38788cb8581d68b82b42acf8e810b

Observation ad6542bb-8d4f-49ac-9436-755170fbc72c · outbound

This paper cites ViTKD: Practical Guidelines for ViT feature knowledge distillation.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers ViTKD: Practical Guidelines for ViT feature knowledge distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.749967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.749967Z digest=sha256:05a9ed85635d1810b6dbdc1bfbb22cd77c3e50f22ee16344b1913808f3345d40

Observation c1b7c2b1-06b6-4e3b-9cef-0ebb3902576c · outbound

This paper cites Like What You Like: Knowledge Distill via Neuron Selectivity Transfer.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Like What You Like: Knowledge Distill via Neuron Selectivity Transfer

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.694727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.694727Z digest=sha256:4d4986b0d91d4b50228f8da58f0227a4b051327cfb96f964e6e7fbb213bd42fa

Observation ce1c0edc-2323-4c8f-8410-c35cf4074639 · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers TinyBERT: Distilling BERT for Natural Language Understanding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.699745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.699745Z digest=sha256:804804957258eed2936c23e97cc9e837a7e6856d1ca890ba608f6f8fb9217161

Observation 0f6afcf6-867a-4516-9a2e-ebf2d6212a11 · outbound

This paper cites Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:50:14.996941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:50:14.669993Z digest=sha256:1dce6b79e0cd3c6cb318e8be42981a5c626793562ae91d263ca9321ea8dd8623

Observation 4f8a83dc-2af1-42c6-8018-1a4e84847cda · outbound

This paper cites Training data-efficient image transform- ers & distillation through attention.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Training data-efficient image transform- ers & distillation through attention

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.077756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:50:14.735333Z digest=sha256:532b295e879a43825197343fa2585ff64a8ca43e18119f20cfc2db0bd12ebe4d

Observation e1812a03-25c2-4ecc-b1cf-253c2d026f34 · outbound

This paper cites Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.093065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:50:14.684755Z digest=sha256:649ff3a7a34027affe484d809245d4dd30067f17a8725606ef2ddd3d62c78644

Observation 71d1e575-f4af-4a1e-8231-19ee730fd4bc · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.659381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.659381Z digest=sha256:e669425e8ce686408ca1ea1cc241a57788227ffd2c8ca896d32b5ee24b431505

Observation 8af8114e-bc56-40e5-9dba-d72d80ca3b10 · outbound

This paper cites $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.710201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.710201Z digest=sha256:1a0ae6187aedf2ae3637d52cdcfb1f8e6d8f660593ef8290793169574b2dcdd0

Pith citing papers

Observation 66acbb21-a70f-4f7b-a051-e73e82b6bed4 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.667666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:b0096e091976b2154ecb7109db1adae01741d79c2c3d3eed4789204da8feca89