Pith. sign in

Paper Citation Record · LEDGER

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation

As of 16 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15281 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:52.024717Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3d4b98e-6203-4ea3-84f8-481944c31916 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.774537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.774537Z digest=sha256:775499be3c7694a19eb346b9f217016a6752fd152bc60f81ddb19a0571f91354

Observation 6e161dd6-471d-400d-a50f-9a5dba66d90e · outbound

This paper cites Fluctuation-based adaptive structured pruning for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Fluctuation-based adaptive structured pruning for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.778052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.778052Z digest=sha256:30dde750ba36e8702e82e3d5919340cdec01b9a2666734d8ad82bb3f9dd5324f

Observation 02aab24a-4e03-4b51-b3af-43518215b7cd · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.780823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.780823Z digest=sha256:a88bfc12b5e73371d85b0ba0f6f9d4f1c44120a0fa6aa9241924e1223d759cb1

Observation b49bea46-440c-4585-bf77-5a4d907f760d · outbound

This paper cites A general language assistant as a laboratory for alignment, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A general language assistant as a laboratory for alignment, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.783833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.783833Z digest=sha256:d0ec44d7fe40ba82e595c923944831bb507f6e4e244c0aa57ba43db3a6012b42

Observation d4aa84aa-226c-415f-94ac-a90b796e9a9c · outbound

This paper cites Mitigat- ing open-vocabulary caption hallucinations, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mitigat- ing open-vocabulary caption hallucinations, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.787009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.787009Z digest=sha256:a1f46b7e74990462d5e5bbf1f92a6028333ca536c54466aa69dcc96fdbd3cf7b

Observation bac38b26-ef54-4f65-9184-f54dbebf4816 · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.790132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.790132Z digest=sha256:d7eb627fce46540e0fdaddd05c293a1b19eb6de07d0c24016a7e06145da1218e

Observation dfd730a8-8e37-4c5e-9891-b058b4c1b3cf · outbound

This paper cites Token Merging: Your ViT But Faster.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Token Merging: Your ViT But Faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.793216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.793216Z digest=sha256:33110527a478dc9136c5038e26adfff4535174e35dfb5140fb2d9e287a2b0764

Observation b91a4d79-f0b6-4bf1-a0bb-ffe534159e3a · outbound

This paper cites Emerging properties in self-supervised vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Emerging properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.796341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.796341Z digest=sha256:2077ace9ad893693600273246a63474f2c03f019ac9fb8d6d7c20a0389616f15

Observation 92479429-200b-4edf-9fc0-11b80eed0654 · outbound

This paper cites Vision transformer slimming: Multi-dimension searching in continuous optimization space.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Vision transformer slimming: Multi-dimension searching in continuous optimization space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.799238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.799238Z digest=sha256:d87658ac0d799c5a366e6c9b4772492c79d87cc051ebd6d031e6d3645dbc10c7

Observation b13b73c1-4e1b-42e8-b792-8ebfafee7847 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.802481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.802481Z digest=sha256:0ddade242854f23aefd74de27b2470673b517d90e148b4b4322636afdee9f840

Observation 01cc1d15-c44a-4cdf-a841-8966e90133a5 · outbound

This paper cites The lottery ticket hypothesis for pre-trained bert networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The lottery ticket hypothesis for pre-trained bert networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.698529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.805518Z digest=sha256:7e7ec2bfbc9aa0916eee1567053541ea623b239e7e0f561c619279f2456dee17

Observation f86d788e-745b-4e07-9c79-257d048ac7ca · outbound

This paper cites The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.689473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.808327Z digest=sha256:def9a1c188f8c2e481d092d1c1a565b75a29bc05eea58aad34ecdcb5a15afc7a

Observation fee15ab5-e6a7-492a-9566-7a05fb5d132d · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A toy model of universality: Reverse engineering how networks learn group operations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.680942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.811213Z digest=sha256:d32c8bdb65abb00f62c7c3926ec3c7339b6d494311c861e449ae3f71d90a1444

Observation dd75c708-4a6f-4750-a7f1-62719a01e8af · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.813816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.813816Z digest=sha256:dc5836d6b8bde9e91cf5a50f767cd91462e85e1926a041277b240ad8d49c1d52

Observation 3996ee05-cad3-4b1d-bff1-23af3959df71 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.816837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.816837Z digest=sha256:2ccaa3b4b01b3cae11d9c40f1178c227d6ce37a3b8b5b88b1240f4f8ded0fc61

Observation fbb301ab-1bab-4673-8ae0-54f5dd6d3740 · outbound

This paper cites Analyzing Redundancy in Pretrained Transformer Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Analyzing Redundancy in Pretrained Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.820203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.820203Z digest=sha256:6c21dc77b9ee7484b903d60b2e2540dfe555d92f0ea9081d360524aaa6d4d755

Observation 5538669b-0403-47af-876a-931014f27133 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Imagenet: A large- scale hierarchical image database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.822860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.822860Z digest=sha256:f5c603722112ac9b7a2f5e1ae6954bbab6b9f91f784d10943c4cad43a354cec2

Observation 650b0c88-dbfe-41b6-b3cb-405e0417c8d8 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Qlora: Efficient finetuning of quantized llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.825418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.825418Z digest=sha256:3e9b26e3ce69630dd14eac37f7086327b918287604d8db874a0b48b167f818d6

Observation f9609917-7f82-4be9-a255-a1b9eef0f188 · outbound

This paper cites Eventful transformers: leveraging temporal redundancy in vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Eventful transformers: leveraging temporal redundancy in vision transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.663161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.828153Z digest=sha256:27cd409347b8af7ba0d2ff73729e4eab46a805fbbda9dc3e34c3f5132790b626

Observation a1f76e72-8720-422c-a01e-9072164333cf · outbound

This paper cites A mathematical framework for transformer circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A mathematical framework for transformer circuits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.830907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.830907Z digest=sha256:b5cb3b63321ad1315a92a0e346847f46cc7f49f5374f7a9fdc60ff7be305d2cc

Observation 199df78c-ee4c-4a08-bc03-183021e11d32 · outbound

This paper cites Depgraph: Towards any structural pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Depgraph: Towards any structural pruning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.833616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.833616Z digest=sha256:1c2726fdf6ea8543ab7b0f8eb84b2ac29a4c266d6c9db158449c5079b6e481cb

Observation 4907c096-c9f4-4529-82a1-5959b9676cdd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.836416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.836416Z digest=sha256:cc4bf6f0947301ac213c1a11ea935755d087f77f7a45b1f731de3adeee9532b0

Observation 7e1c5090-53ce-408f-834f-467a84021cfa · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Transformer Feed-Forward Layers Are Key-Value Memories

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.839216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.839216Z digest=sha256:b5b6283e3a7b3cb3fc9c99a2e9fad1612d22c5a6d3f20844db42b44342e3083a

Observation a0ba1185-ab48-45e3-9583-0794cbcd0c7e · outbound

This paper cites Successor Heads: Recurring, Interpretable Attention Heads In The Wild.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Successor Heads: Recurring, Interpretable Attention Heads In The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.842513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.842513Z digest=sha256:0549c51eed7665d1d11402025fccd3220ebe87e1ee8a464f8f0f35c0e2a87f50

Observation 26169097-f8ff-4db5-af4b-4befc0b14ab3 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MiniLLM: Knowledge distillation of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.845580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.845580Z digest=sha256:ce26e04dea2b50dcd8ace7259cd4a9ec1b2dfbe0800897bf9948c58bcb09b395

Observation 414e3cc2-5e12-4c2a-9203-cec92ff814e8 · outbound

This paper cites Learning efficient vision transformers via fine-grained manifold distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Learning efficient vision transformers via fine-grained manifold distillation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.634244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.848506Z digest=sha256:d4343062bb7a87cb17338abebce97544746dca08f660b420ad8be38f0a57da81

Observation 131feb79-a177-4bfb-a74d-01897e54cf75 · outbound

This paper cites Masked autoencoders are scalable vision learners, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Masked autoencoders are scalable vision learners, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.850924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.850924Z digest=sha256:c937115d88ef43fb98cd2d1b2e7e1e948a06b717d2e996be0e695cce799991f8

Observation 7485a9dc-058c-473a-8750-49079323c695 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation What Matters in Transformers? Not All Attention is Needed

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.853118Z digest=sha256:e64f13689434c28cd73ee3e8e44922537ed7f0a148d4e7fe545cf4790d600ca1

Observation eb07f159-3abb-4775-ac42-e7edb1f28aaf · outbound

This paper cites Distilling the Knowledge in a Neural Network.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Distilling the Knowledge in a Neural Network

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.855575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.855575Z digest=sha256:b5d6034758b46a62c009c1be24f24de100b9f55d4c042b2d3de277c5a91d4f13

Observation 390edfd7-ede8-4410-9e30-cc7a3240712d · outbound

This paper cites Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:38:52.316704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.858131Z digest=sha256:0c2e04e17caac02a15d7e7a68de8527e4e829a8745ba14ea526acf17e06661f5

Observation ad95b5a0-f5e3-493d-a44e-54b6f12f8fb1 · outbound

This paper cites Mixture of Nested Experts: Adaptive Processing of Visual Tokens.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.860966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.860966Z digest=sha256:a14dc04efa8bb718e739c08e4075ef1120a8c6d3f9729608de1d0dc6be07bfdc

Observation a3be0e74-894d-484d-9fa7-405a2e6d81d2 · outbound

This paper cites Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.619430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.864184Z digest=sha256:2d016fa81653f945b2b77d431664ed8feb6a7e7bd2a26a545382b9a7853ba42f

Observation 55eab72c-802c-498f-82b7-411ff9c34e9a · outbound

This paper cites Expedited training of visual conditioned language generation via redundancy reduction.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Expedited training of visual conditioned language generation via redundancy reduction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.610412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.867019Z digest=sha256:4c9e629825e4785c4ee2a62ede88921ae45a98b3166836e66ec23737bd4fdfd8

Observation 084e9234-9b0f-4e80-addb-0417bb2a341c · outbound

This paper cites Mixtral of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.869956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.869956Z digest=sha256:67b368971db677aafe6c7401c77aecfde3a9644379d38a6105f5a6f0cb223fd6

Observation 14383e2d-b807-48c9-b653-95ee6e53d585 · outbound

This paper cites Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit).

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.602038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.872978Z digest=sha256:89932e6a54c5da3d46f298058ab4d29880d96c04fcc2d61f5e9b1c1557833cb1

Observation 682ee29e-5d31-4b4a-9497-a16b0031564e · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation TinyBERT: Distilling BERT for Natural Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.875902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.875902Z digest=sha256:e1cbeec497ce3c63f3aeebf9f71b1582815eb28da48791f2e31bae5506a4bbb5

Observation 74b67fb5-c03e-424c-8183-147344f8fe73 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.879159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.879159Z digest=sha256:cf505b037a7bd26f47158e68760426c97ee46942d3e967bd04d4f680e9f60418

Observation e73ced7e-ee60-4cf1-998e-531e43d0c297 · outbound

This paper cites Self-Distillation for Further Pre-training of Transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-Distillation for Further Pre-training of Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.882149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.882149Z digest=sha256:b71c2abb8527de06f5c7d36fdf0fd91812abe7514fb2627d38c9875ae8ab725a

Observation 4113c3f4-3216-4e76-a6e1-48836cd0a659 · outbound

This paper cites Clustered imagenet labels for training production-friendly image classifier.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Clustered imagenet labels for training production-friendly image classifier

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.594211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.884764Z digest=sha256:bde0462d16e125c6d430259c9b13297c62ae299b3cacf472d37fcf88ae18dce6

Observation 5593dda4-a275-4e98-a5f3-4d9a0ef3351a · outbound

This paper cites Knowledge distillation via the target-aware transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation via the target-aware transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.585753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.887144Z digest=sha256:4e1ac578aea99c390bc2f0076a8caa5595b69bd626b751bd44ab3b56a8290b54

Observation 8b84504f-67ac-4341-814b-58b247eef34d · outbound

This paper cites FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.889606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.889606Z digest=sha256:1df0829bebcefe96dece7d805b904767b05ee74ccf67be9f0bc600ccc5dc2801

Observation 50ac9757-c5a3-4dfb-8f42-4646693131ec · outbound

This paper cites Visual instruction tuning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.892356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.892356Z digest=sha256:e367eb71eef102cbf0f29757f4c1ead6a16cc9477f636707348314a3e0255ae4

Observation 3aa03e23-d79f-447a-8594-7510f5fce33e · outbound

This paper cites Oscillation-free quantization for low-bit vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Oscillation-free quantization for low-bit vision transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.572215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.894734Z digest=sha256:38399628851c2cf48dc76e0a731308af8e6147067470b0a99c50920fcc54df91

Observation a0cf4fec-e8a2-451e-b519-75c6bbb1ef38 · outbound

This paper cites Post-training quantization for vision transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Post-training quantization for vision transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.896973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.896973Z digest=sha256:db4452aec360e3201e1a9c5e5b93df0dc6fad05ec5800044eb5c4b3f496af575

Observation 106afdf3-ea68-44a8-a341-4635adc29ec7 · outbound

This paper cites Anytime Dense Prediction with Confidence Adaptivity.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Anytime Dense Prediction with Confidence Adaptivity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.899751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.899751Z digest=sha256:6e55ef01c1e2277b54ff57a92e9a36777a01bd3973d9b143c37d1fb621e09862

Observation 7688bba1-7d21-40ad-8772-053af3629bdb · outbound

This paper cites A transformer-based model with self-distillation for multimodal emotion recognition in conversations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A transformer-based model with self-distillation for multimodal emotion recognition in conversations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.558394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.902822Z digest=sha256:2f5634efed3d7b0342070d86230ea93123c00dad2a5debe94cad16ee19486961

Observation 89d96134-36af-4bb5-a78c-36ab88e259ad · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Llm-pruner: On the structural pruning of large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.905570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.905570Z digest=sha256:afc8d12b10dfe0eb9484fa26479197d5f4ed21f9eb5a5cad181cfc3cdd4f10f4

Observation 8cc78494-02f8-4e58-ba88-8ed448c12c59 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Copy Suppression: Comprehensively Understanding an Attention Head

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.908556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.908556Z digest=sha256:27fd648bf05a539285ce82a9c5e0e6b7880f5de629b89e40e102ed0c58bfa7c6

Observation 0bd13a09-5670-464f-a3aa-340dc8418f03 · outbound

This paper cites Locating and editing factual associations in gpt.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Locating and editing factual associations in gpt

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.911656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.911656Z digest=sha256:b59f69b370e3dac1ecc6f4a556b8696fa7161444bfe04a89a54eaeedafd18c8c

Observation caa4206b-3f0a-453a-ae07-76d0bcfe2ca4 · outbound

This paper cites Circuit Component Reuse Across Tasks in Transformer Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Circuit Component Reuse Across Tasks in Transformer Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.914561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.914561Z digest=sha256:590aa6322a74774d284312a00155b92e871739ae8b1cee9645c32285eea76e1f

Observation 2baf1a4b-320a-45e0-82ef-de20738605bb · outbound

This paper cites Zoom in: An introduction to circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Zoom in: An introduction to circuits

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.917590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.917590Z digest=sha256:fe5c90a9e2cecca140d0577c1d9c0bb7e0b8b88a948d674ac416ecdce4729d58

Observation 0ad33b7e-34c7-416e-887b-3928d987a573 · outbound

This paper cites In-context Learning and Induction Heads.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation In-context Learning and Induction Heads

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.919899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.919899Z digest=sha256:c87c83bba5bc3af0a68bb71ee39753d74aea1b2f02043a5f233a1f9e589ad417

Observation f652ae6a-d2fb-4fc5-abb2-6ba6a042452c · outbound

This paper cites Ia-red 2: Interpretability-aware redundancy reduction for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Ia-red 2: Interpretability-aware redundancy reduction for vision transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.534655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.922482Z digest=sha256:ef0f1c9edef47830ea54dab84dd98444f9c79a74d73a3d68da443b4fa1fc50ca

Observation 5add9e11-2b4b-45b8-be8d-75586f889ba1 · outbound

This paper cites Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.526272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.924867Z digest=sha256:b7d94d1067da4b1d7ec48bcfce531391bbe13202894692e6c37be358b95be843

Observation a20e40e5-ec2e-483d-b558-a50e0fc0c57f · outbound

This paper cites A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.927037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.927037Z digest=sha256:6baeeb5e294f20ea067bd43c35e800777f14acb7feae0f171dbbeff34762da35

Observation 8c018973-c0ee-4a62-9d21-bf2519e8e086 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.929776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.929776Z digest=sha256:a8503d04428467ef38349650bc1f3b34b4442aa33f2cf1558c0e1e27ba088847

Observation 0f16e29e-3e9a-41c1-9fcd-3cabd2f40520 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.932616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.932616Z digest=sha256:2b13cac7ee8624c64e2fb6b8b299c0cac7472b818e9430493197d56adea98ba9

Observation bdf3dd24-01e7-4af5-9c2b-c8b570da2546 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.935815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.935815Z digest=sha256:a18b1a5908276a7f349df9aa62874d5ecc3ef945a0bd1bdaca4a338e2cedb49d

Observation 39d689fd-4cd5-4e07-8e20-31bc538d89af · outbound

This paper cites Confident adaptive language modeling.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Confident adaptive language modeling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.513563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.938966Z digest=sha256:88e519c011b10b39d13ac84871bb467e4ddb1c0ab5d6c422afa765d8a4b3c84b

Observation 391d1b87-f156-45bc-b39e-55f78eb52bd6 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.941889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.941889Z digest=sha256:2a7d696a462dc6090a31d803bf9eb5548dbf4995b37ed826027040fb83458b5a

Observation 9e457abb-0556-4269-abc2-bfc9f1873c4e · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.944839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.944839Z digest=sha256:678caf39388c5ac02f9c2150ea9f99ffd290044e46def5c7bef91488d05179a6

Observation 54e6f198-ea37-4b4c-954b-1d271ad9dd34 · outbound

This paper cites A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.947280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.947280Z digest=sha256:a8661686b95092cf09bbfe64032d77539735245951944d58a8e14bcc74413efb

Observation 697491b3-5265-46f7-88e0-6c8627f8810b · outbound

This paper cites Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.505781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.949782Z digest=sha256:85b59c1a095b90ef92fa55bf6bc492976cf7c20680f39e4955156c55df363c69

Observation 73e114a5-6a81-4076-92a5-f43ba06e4e2d · outbound

This paper cites Self-distilled vision transformer for domain generalization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distilled vision transformer for domain generalization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.496992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.952400Z digest=sha256:5f20033fd1ecb2a5d97ded5e40c2442e2876bf2cdf985005314b7c62130c4cad

Observation d3b10aa2-14bb-43f3-b59c-0b57d06a790a · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.954669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.954669Z digest=sha256:257740af59cfa9efb682bf804840317189de79878f5789dd18dfd880bbd4769b

Observation 8312c520-0e96-41d9-b045-14fd40262c2b · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.957799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.957799Z digest=sha256:cf0cb015c46b4771b915701e18d62097f4c6ff375d502f1510aec95987512ac1

Observation b47718a4-6be5-4669-90a5-67c84787c257 · outbound

This paper cites Patch slimming for efficient vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patch slimming for efficient vision transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.960881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.960881Z digest=sha256:e1af1f7a52ee4c3b64e1061582cf9a86ee7e465ebe924aceaed27718ed92604f

Observation fc3cce1a-9e23-42a2-8fba-d94c5de0578b · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.483135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.963709Z digest=sha256:57295b9025957079af1c107401dbd4ebb1b86fb7ae2b2ea8f4cd3c8d0a06c202

Observation 8378a85f-e41e-4276-9c62-43ec2ac1af75 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training data-efficient image transformers & distillation through attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.966919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.966919Z digest=sha256:9fc8a9ce9da7b90ce2eabde1ffc00a9c4188463f2cac6779fa6af0cefa2553d6

Observation e3fa46d1-7012-4a75-910a-88ec06df7c40 · outbound

This paper cites Attention is all you need.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Attention is all you need

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.969832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.969832Z digest=sha256:d3d7ca1e038cde48ba2029cb7bffb0539fb8a6d70f49baac47961a93613e940a

Observation 03c70042-9d43-43c0-8c76-c5f78c24529f · outbound

This paper cites Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.464845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.972731Z digest=sha256:9a0dc85d50b8f1311e832a66555c9f14cb41de93568c81739d100c3bfaceb047

Observation 8adb33bb-3086-440e-b335-7cab20402372 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.975653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.975653Z digest=sha256:5628f177f5096a620321864498c7b9dc064ca2f57d0073412299374b1c740495

Observation b62e5c21-9f43-4b81-ab10-ab17702fb178 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.456906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.979027Z digest=sha256:9b5a828b95dc85ede99cc79f3f57c63da94b7161abda0bb3a8a013a821fb4759

Observation e24b8f23-3957-4d42-9397-6e0ec64b77d3 · outbound

This paper cites Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.448430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.982279Z digest=sha256:dddfc390e234e6152b79395b75643b7400a059a75934ff7388574f442e764428

Observation 1749f87a-2a79-4e35-ba69-ae658f8d1b8d · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tinyvit: Fast pretraining distillation for small vision transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.438892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.984954Z digest=sha256:43ec2d6c523ea981ff387dce7722bc3fed31dd8e73ff1313ef8b61610658379e

Observation 23cc25b4-72e7-4e58-bf94-be7ee936f505 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.987271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.987271Z digest=sha256:bdcc3adfde148a91a045e64123f6aaf0ddb353869f74d7bfe557c14f3b2e7d76

Observation 233e4718-6d55-4085-a644-6d508d92c1f0 · outbound

This paper cites Structured Pruning Learns Compact and Accurate Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Structured Pruning Learns Compact and Accurate Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.990020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.990020Z digest=sha256:00195b46da6e8f3a387f2d37838d442305b18ce849830c38b36e123acd3e2f31

Observation 33c2b9fc-d032-40a7-95a7-0384e79a0ba4 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.992893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.992893Z digest=sha256:9d98886a1d9a5360b348a8f37962b925ecf03905a9ca987e4d1d082331322699

Observation 33c4be17-a40d-4aa6-935b-01e311898b3d · outbound

This paper cites ThinK: Thinner Key Cache by Query-Driven Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.995190Z digest=sha256:c29065f72d37bd51854e84dfa934533a59e539cafa98f56858f10cd3fc5b8687

Observation 0273513d-35ee-4930-a305-18c2edda9540 · outbound

This paper cites X-pruner: explainable pruning for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation X-pruner: explainable pruning for vision transformers

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.424277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:51.998085Z digest=sha256:10f97d3e788bd0ecf7ba977adb067e0c6dc87dba5bff6358bb90544aef9f45ed

Observation 93c0355e-f848-45db-9b2f-12e6afd37318 · outbound

This paper cites Unified Visual Transformer Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unified Visual Transformer Compression

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.000864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.000864Z digest=sha256:49c38ce5ee4fddf5845426d7001af337107687066b6cfe90191fc9a59bdb675f

Observation 11a5dcf9-149d-4f7b-b713-3b8c2f8a1686 · outbound

This paper cites MoEfication: Transformer Feed-forward Layers are Mixtures of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEfication: Transformer Feed-forward Layers are Mixtures of Experts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.004555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.004555Z digest=sha256:00cccbc1649f5f0ff079198312e5783ab65c79df51295e008ef25651f3546295

Observation caea50ba-6c6e-424d-bb5c-0cc1ee2e5027 · outbound

This paper cites Knowledge distillation based on transformed teacher matching.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation based on transformed teacher matching

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.414808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:52.007595Z digest=sha256:4324a8145b41e3b58cdac2704508d38002d0f8acdbb05d0fabca01df121e2fd3

Observation 80958d0c-b3c7-4749-8de3-4432ec47e4ca · outbound

This paper cites The clock and the pizza: Two stories in mechanistic explanation of neural networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The clock and the pizza: Two stories in mechanistic explanation of neural networks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.010319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.010319Z digest=sha256:94ed62c846e317fecde6f11001f573812e30b2b85236fb9f93f5ac65af737909

Observation 19db123c-225d-4678-83db-042e11c80ad4 · outbound

This paper cites LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.013139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.013139Z digest=sha256:682bdf678d7ed1903ea8dacf41d4ed6edb0a4a2e47d65381f70fc2998f83b840

Observation e89b4d73-6cb9-4692-bfff-0be0ab6974aa · outbound

This paper cites MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.016407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.016407Z digest=sha256:32dcdda91445c2aeb87151ad2998d413e668d4ccab1e80b1e37abd2352bb2f80

Observation 2a226cfb-df33-4745-a251-066a3616abc2 · outbound

This paper cites Possible solutions could include:.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Possible solutions could include:

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.402343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:52.019705Z digest=sha256:d004407c3a81ec96e56906e9d4146b363483a1bf5f6a85c9805af6ede986426a

Observation da8cde43-a337-4086-a708-a2e5e522a007 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.394336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:52.022253Z digest=sha256:6afed9c1a6f959fce572c078258348639a939326abd5eb5daec8d177fc3d66c0

Observation 35c4f703-c447-4c33-ad9e-f9b61206eaa1 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.385463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T14:38:52.024717Z digest=sha256:0eb1eba45c9fcd0ba717d1c851cdd98f0bc453d3f3caad99690fe5df375e00e0

Pith citing papers

No inbound Pith citation observations are available.