Pith. sign in

Paper Citation Record · LEDGER

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation

As of 15 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15281 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:52.024717Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3d4b98e-6203-4ea3-84f8-481944c31916 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.774537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.774537Z digest=sha256:775499be3c7694a19eb346b9f217016a6752fd152bc60f81ddb19a0571f91354

Observation 6e161dd6-471d-400d-a50f-9a5dba66d90e · outbound

This paper cites Fluctuation-based adaptive structured pruning for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Fluctuation-based adaptive structured pruning for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.778052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.778052Z digest=sha256:30dde750ba36e8702e82e3d5919340cdec01b9a2666734d8ad82bb3f9dd5324f

Observation 02aab24a-4e03-4b51-b3af-43518215b7cd · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.780823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.780823Z digest=sha256:a88bfc12b5e73371d85b0ba0f6f9d4f1c44120a0fa6aa9241924e1223d759cb1

Observation b49bea46-440c-4585-bf77-5a4d907f760d · outbound

This paper cites A general language assistant as a laboratory for alignment, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A general language assistant as a laboratory for alignment, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.783833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.783833Z digest=sha256:d0ec44d7fe40ba82e595c923944831bb507f6e4e244c0aa57ba43db3a6012b42

Observation d4aa84aa-226c-415f-94ac-a90b796e9a9c · outbound

This paper cites Mitigat- ing open-vocabulary caption hallucinations, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mitigat- ing open-vocabulary caption hallucinations, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.787009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.787009Z digest=sha256:a1f46b7e74990462d5e5bbf1f92a6028333ca536c54466aa69dcc96fdbd3cf7b

Observation bac38b26-ef54-4f65-9184-f54dbebf4816 · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.790132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.790132Z digest=sha256:d7eb627fce46540e0fdaddd05c293a1b19eb6de07d0c24016a7e06145da1218e

Observation dfd730a8-8e37-4c5e-9891-b058b4c1b3cf · outbound

This paper cites Token Merging: Your ViT But Faster.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Token Merging: Your ViT But Faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.793216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.793216Z digest=sha256:33110527a478dc9136c5038e26adfff4535174e35dfb5140fb2d9e287a2b0764

Observation b91a4d79-f0b6-4bf1-a0bb-ffe534159e3a · outbound

This paper cites Emerging properties in self-supervised vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Emerging properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.796341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.796341Z digest=sha256:2077ace9ad893693600273246a63474f2c03f019ac9fb8d6d7c20a0389616f15

Observation 92479429-200b-4edf-9fc0-11b80eed0654 · outbound

This paper cites Vision transformer slimming: Multi-dimension searching in continuous optimization space.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Vision transformer slimming: Multi-dimension searching in continuous optimization space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.799238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.799238Z digest=sha256:d87658ac0d799c5a366e6c9b4772492c79d87cc051ebd6d031e6d3645dbc10c7

Observation b13b73c1-4e1b-42e8-b792-8ebfafee7847 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.802481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.802481Z digest=sha256:0ddade242854f23aefd74de27b2470673b517d90e148b4b4322636afdee9f840

Observation 01cc1d15-c44a-4cdf-a841-8966e90133a5 · outbound

This paper cites The lottery ticket hypothesis for pre-trained bert networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The lottery ticket hypothesis for pre-trained bert networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.698529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.805518Z digest=sha256:c704e3505e625d7974838832340f15f253d8578dc0120d948728123f7bab9046

Observation f86d788e-745b-4e07-9c79-257d048ac7ca · outbound

This paper cites The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.689473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.808327Z digest=sha256:e699ad8c335a0670b744594c5e9513edfc3335df113195a7b9fdccddbdf88ed6

Observation fee15ab5-e6a7-492a-9566-7a05fb5d132d · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A toy model of universality: Reverse engineering how networks learn group operations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.680942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.811213Z digest=sha256:1056522b268e1afbc451fd19ba98641db7679199333869d855cdfe2f1e582628

Observation dd75c708-4a6f-4750-a7f1-62719a01e8af · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.813816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.813816Z digest=sha256:dc5836d6b8bde9e91cf5a50f767cd91462e85e1926a041277b240ad8d49c1d52

Observation 3996ee05-cad3-4b1d-bff1-23af3959df71 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.816837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.816837Z digest=sha256:2ccaa3b4b01b3cae11d9c40f1178c227d6ce37a3b8b5b88b1240f4f8ded0fc61

Observation fbb301ab-1bab-4673-8ae0-54f5dd6d3740 · outbound

This paper cites Analyzing Redundancy in Pretrained Transformer Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Analyzing Redundancy in Pretrained Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.820203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.820203Z digest=sha256:6c21dc77b9ee7484b903d60b2e2540dfe555d92f0ea9081d360524aaa6d4d755

Observation 5538669b-0403-47af-876a-931014f27133 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Imagenet: A large- scale hierarchical image database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.822860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.822860Z digest=sha256:f5c603722112ac9b7a2f5e1ae6954bbab6b9f91f784d10943c4cad43a354cec2

Observation 650b0c88-dbfe-41b6-b3cb-405e0417c8d8 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Qlora: Efficient finetuning of quantized llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.825418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.825418Z digest=sha256:3e9b26e3ce69630dd14eac37f7086327b918287604d8db874a0b48b167f818d6

Observation f9609917-7f82-4be9-a255-a1b9eef0f188 · outbound

This paper cites Eventful transformers: leveraging temporal redundancy in vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Eventful transformers: leveraging temporal redundancy in vision transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.663161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.828153Z digest=sha256:8926a7b34c1327bcd1b9a6dae4bc92670f27cf527d97056ee4f3f4bbedf439ed

Observation a1f76e72-8720-422c-a01e-9072164333cf · outbound

This paper cites A mathematical framework for transformer circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A mathematical framework for transformer circuits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.830907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.830907Z digest=sha256:b5cb3b63321ad1315a92a0e346847f46cc7f49f5374f7a9fdc60ff7be305d2cc

Observation 199df78c-ee4c-4a08-bc03-183021e11d32 · outbound

This paper cites Depgraph: Towards any structural pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Depgraph: Towards any structural pruning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.833616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.833616Z digest=sha256:1c2726fdf6ea8543ab7b0f8eb84b2ac29a4c266d6c9db158449c5079b6e481cb

Observation 4907c096-c9f4-4529-82a1-5959b9676cdd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.836416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.836416Z digest=sha256:cc4bf6f0947301ac213c1a11ea935755d087f77f7a45b1f731de3adeee9532b0

Observation 7e1c5090-53ce-408f-834f-467a84021cfa · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Transformer Feed-Forward Layers Are Key-Value Memories

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.839216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.839216Z digest=sha256:b5b6283e3a7b3cb3fc9c99a2e9fad1612d22c5a6d3f20844db42b44342e3083a

Observation a0ba1185-ab48-45e3-9583-0794cbcd0c7e · outbound

This paper cites Successor Heads: Recurring, Interpretable Attention Heads In The Wild.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Successor Heads: Recurring, Interpretable Attention Heads In The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.842513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.842513Z digest=sha256:0549c51eed7665d1d11402025fccd3220ebe87e1ee8a464f8f0f35c0e2a87f50

Observation 26169097-f8ff-4db5-af4b-4befc0b14ab3 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MiniLLM: Knowledge distillation of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.845580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.845580Z digest=sha256:ce26e04dea2b50dcd8ace7259cd4a9ec1b2dfbe0800897bf9948c58bcb09b395

Observation 414e3cc2-5e12-4c2a-9203-cec92ff814e8 · outbound

This paper cites Learning efficient vision transformers via fine-grained manifold distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Learning efficient vision transformers via fine-grained manifold distillation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.634244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.848506Z digest=sha256:d610db08ec78bc11d5329f29c363a63514ba43cdd58d6ceadeaaadcce3ba6d7a

Observation 131feb79-a177-4bfb-a74d-01897e54cf75 · outbound

This paper cites Masked autoencoders are scalable vision learners, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Masked autoencoders are scalable vision learners, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.850924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.850924Z digest=sha256:c937115d88ef43fb98cd2d1b2e7e1e948a06b717d2e996be0e695cce799991f8

Observation 7485a9dc-058c-473a-8750-49079323c695 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation What Matters in Transformers? Not All Attention is Needed

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.853118Z digest=sha256:5870cb6ea00e1e0323eb05a31599b9973a30e89eed43260a3640944e42d198d9

Observation eb07f159-3abb-4775-ac42-e7edb1f28aaf · outbound

This paper cites Distilling the Knowledge in a Neural Network.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Distilling the Knowledge in a Neural Network

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.855575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.855575Z digest=sha256:b5d6034758b46a62c009c1be24f24de100b9f55d4c042b2d3de277c5a91d4f13

Observation 390edfd7-ede8-4410-9e30-cc7a3240712d · outbound

This paper cites Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:38:52.316704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.858131Z digest=sha256:3adf5344b0a2ee44ca62a2d6eb3c15102f7a0f47581b970103e294b3e4ae7427

Observation ad95b5a0-f5e3-493d-a44e-54b6f12f8fb1 · outbound

This paper cites Mixture of Nested Experts: Adaptive Processing of Visual Tokens.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.860966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.860966Z digest=sha256:68b63185de681b26eecb63b95847e0bd320270ef1993a58747709e48d20dcf9f

Observation a3be0e74-894d-484d-9fa7-405a2e6d81d2 · outbound

This paper cites Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.619430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.864184Z digest=sha256:d138d3d184f42ce6a0906b6ebf3e3e0649cc4ec9dff8f7acdc6ee032444b8b3d

Observation 55eab72c-802c-498f-82b7-411ff9c34e9a · outbound

This paper cites Expedited training of visual conditioned language generation via redundancy reduction.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Expedited training of visual conditioned language generation via redundancy reduction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.610412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.867019Z digest=sha256:938dea0fb593c083bb4e0e434832553d330a383ab51c9cc2a2efd0885fe82176

Observation 084e9234-9b0f-4e80-addb-0417bb2a341c · outbound

This paper cites Mixtral of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.869956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.869956Z digest=sha256:67b368971db677aafe6c7401c77aecfde3a9644379d38a6105f5a6f0cb223fd6

Observation 14383e2d-b807-48c9-b653-95ee6e53d585 · outbound

This paper cites Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit).

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.602038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.872978Z digest=sha256:30bc5f26158afd0cebafd26ccb4b89ccb3fe02a22a4662a8b14a924f0147e6ba

Observation 682ee29e-5d31-4b4a-9497-a16b0031564e · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation TinyBERT: Distilling BERT for Natural Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.875902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.875902Z digest=sha256:e1cbeec497ce3c63f3aeebf9f71b1582815eb28da48791f2e31bae5506a4bbb5

Observation 74b67fb5-c03e-424c-8183-147344f8fe73 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.879159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.879159Z digest=sha256:cf505b037a7bd26f47158e68760426c97ee46942d3e967bd04d4f680e9f60418

Observation e73ced7e-ee60-4cf1-998e-531e43d0c297 · outbound

This paper cites Self-Distillation for Further Pre-training of Transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-Distillation for Further Pre-training of Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.882149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.882149Z digest=sha256:a280dd052c41c5d0287fe3488938cb444956b397d3b23d72522df3b70a444db0

Observation 4113c3f4-3216-4e76-a6e1-48836cd0a659 · outbound

This paper cites Clustered imagenet labels for training production-friendly image classifier.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Clustered imagenet labels for training production-friendly image classifier

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.594211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.884764Z digest=sha256:4a018ab64ec9f2ca1733d4b91e71284fc414493da288d266e6e2ed0c10437351

Observation 5593dda4-a275-4e98-a5f3-4d9a0ef3351a · outbound

This paper cites Knowledge distillation via the target-aware transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation via the target-aware transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.585753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.887144Z digest=sha256:d662fe982417377869b0822ce7e6833d3be0f1a5e63a42b10c1e324449436599

Observation 8b84504f-67ac-4341-814b-58b247eef34d · outbound

This paper cites FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.889606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.889606Z digest=sha256:1df0829bebcefe96dece7d805b904767b05ee74ccf67be9f0bc600ccc5dc2801

Observation 50ac9757-c5a3-4dfb-8f42-4646693131ec · outbound

This paper cites Visual instruction tuning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.892356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.892356Z digest=sha256:e367eb71eef102cbf0f29757f4c1ead6a16cc9477f636707348314a3e0255ae4

Observation 3aa03e23-d79f-447a-8594-7510f5fce33e · outbound

This paper cites Oscillation-free quantization for low-bit vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Oscillation-free quantization for low-bit vision transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.572215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.894734Z digest=sha256:216598705f75749425ae5529a81512115939edce14e7ce2a4de5487049eda34b

Observation a0cf4fec-e8a2-451e-b519-75c6bbb1ef38 · outbound

This paper cites Post-training quantization for vision transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Post-training quantization for vision transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.896973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.896973Z digest=sha256:db4452aec360e3201e1a9c5e5b93df0dc6fad05ec5800044eb5c4b3f496af575

Observation 106afdf3-ea68-44a8-a341-4635adc29ec7 · outbound

This paper cites Anytime Dense Prediction with Confidence Adaptivity.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Anytime Dense Prediction with Confidence Adaptivity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.899751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.899751Z digest=sha256:6e55ef01c1e2277b54ff57a92e9a36777a01bd3973d9b143c37d1fb621e09862

Observation 7688bba1-7d21-40ad-8772-053af3629bdb · outbound

This paper cites A transformer-based model with self-distillation for multimodal emotion recognition in conversations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A transformer-based model with self-distillation for multimodal emotion recognition in conversations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.558394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.902822Z digest=sha256:717c941bc9af2c0b93b6c39fe1867331e8f651d577f261ea44751beb8f56e7e0

Observation 89d96134-36af-4bb5-a78c-36ab88e259ad · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Llm-pruner: On the structural pruning of large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.905570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.905570Z digest=sha256:afc8d12b10dfe0eb9484fa26479197d5f4ed21f9eb5a5cad181cfc3cdd4f10f4

Observation 8cc78494-02f8-4e58-ba88-8ed448c12c59 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Copy Suppression: Comprehensively Understanding an Attention Head

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.908556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.908556Z digest=sha256:27fd648bf05a539285ce82a9c5e0e6b7880f5de629b89e40e102ed0c58bfa7c6

Observation 0bd13a09-5670-464f-a3aa-340dc8418f03 · outbound

This paper cites Locating and editing factual associations in gpt.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Locating and editing factual associations in gpt

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.911656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.911656Z digest=sha256:b59f69b370e3dac1ecc6f4a556b8696fa7161444bfe04a89a54eaeedafd18c8c

Observation caa4206b-3f0a-453a-ae07-76d0bcfe2ca4 · outbound

This paper cites Circuit Component Reuse Across Tasks in Transformer Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Circuit Component Reuse Across Tasks in Transformer Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.914561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.914561Z digest=sha256:590aa6322a74774d284312a00155b92e871739ae8b1cee9645c32285eea76e1f

Observation 2baf1a4b-320a-45e0-82ef-de20738605bb · outbound

This paper cites Zoom in: An introduction to circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Zoom in: An introduction to circuits

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.917590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.917590Z digest=sha256:fe5c90a9e2cecca140d0577c1d9c0bb7e0b8b88a948d674ac416ecdce4729d58

Observation 0ad33b7e-34c7-416e-887b-3928d987a573 · outbound

This paper cites In-context Learning and Induction Heads.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation In-context Learning and Induction Heads

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.919899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.919899Z digest=sha256:c87c83bba5bc3af0a68bb71ee39753d74aea1b2f02043a5f233a1f9e589ad417

Observation f652ae6a-d2fb-4fc5-abb2-6ba6a042452c · outbound

This paper cites Ia-red 2: Interpretability-aware redundancy reduction for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Ia-red 2: Interpretability-aware redundancy reduction for vision transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.534655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.922482Z digest=sha256:c95b043c0c4d5116a877e63e4d510daf702deb492b1cda5992c5b165d2ad745f

Observation 5add9e11-2b4b-45b8-be8d-75586f889ba1 · outbound

This paper cites Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.526272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.924867Z digest=sha256:e3331d847203745169f32925f7c4bd4b4816e40e1a5355a95126fbc36c9591ca

Observation a20e40e5-ec2e-483d-b558-a50e0fc0c57f · outbound

This paper cites A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.927037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.927037Z digest=sha256:6baeeb5e294f20ea067bd43c35e800777f14acb7feae0f171dbbeff34762da35

Observation 8c018973-c0ee-4a62-9d21-bf2519e8e086 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.929776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.929776Z digest=sha256:a8503d04428467ef38349650bc1f3b34b4442aa33f2cf1558c0e1e27ba088847

Observation 0f16e29e-3e9a-41c1-9fcd-3cabd2f40520 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.932616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.932616Z digest=sha256:2b13cac7ee8624c64e2fb6b8b299c0cac7472b818e9430493197d56adea98ba9

Observation bdf3dd24-01e7-4af5-9c2b-c8b570da2546 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.935815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.935815Z digest=sha256:a18b1a5908276a7f349df9aa62874d5ecc3ef945a0bd1bdaca4a338e2cedb49d

Observation 39d689fd-4cd5-4e07-8e20-31bc538d89af · outbound

This paper cites Confident adaptive language modeling.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Confident adaptive language modeling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.513563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.938966Z digest=sha256:bd4f70ddea8f07715d65ffe1fd95b4790186c4e17ba2448affb8e2ec33048243

Observation 391d1b87-f156-45bc-b39e-55f78eb52bd6 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.941889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.941889Z digest=sha256:2a7d696a462dc6090a31d803bf9eb5548dbf4995b37ed826027040fb83458b5a

Observation 9e457abb-0556-4269-abc2-bfc9f1873c4e · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.944839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.944839Z digest=sha256:678caf39388c5ac02f9c2150ea9f99ffd290044e46def5c7bef91488d05179a6

Observation 54e6f198-ea37-4b4c-954b-1d271ad9dd34 · outbound

This paper cites A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.947280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.947280Z digest=sha256:a8661686b95092cf09bbfe64032d77539735245951944d58a8e14bcc74413efb

Observation 697491b3-5265-46f7-88e0-6c8627f8810b · outbound

This paper cites Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.505781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.949782Z digest=sha256:9dba41f942cb0f084fc26fd6bd4d508bbb95247e2ac8646dc1b7e10a77b00363

Observation 73e114a5-6a81-4076-92a5-f43ba06e4e2d · outbound

This paper cites Self-distilled vision transformer for domain generalization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distilled vision transformer for domain generalization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.496992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.952400Z digest=sha256:b02f0be33cc30404ad137089c3d705916fc03114424b5956774dc2b1489255d4

Observation d3b10aa2-14bb-43f3-b59c-0b57d06a790a · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.954669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.954669Z digest=sha256:257740af59cfa9efb682bf804840317189de79878f5789dd18dfd880bbd4769b

Observation 8312c520-0e96-41d9-b045-14fd40262c2b · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.957799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.957799Z digest=sha256:8aef01abeb9f0cb8c1cca9d64450f9d32d5d7be6712348bf1b4b4c34bef2c7b5

Observation b47718a4-6be5-4669-90a5-67c84787c257 · outbound

This paper cites Patch slimming for efficient vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patch slimming for efficient vision transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.960881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.960881Z digest=sha256:e1af1f7a52ee4c3b64e1061582cf9a86ee7e465ebe924aceaed27718ed92604f

Observation fc3cce1a-9e23-42a2-8fba-d94c5de0578b · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.483135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.963709Z digest=sha256:5dcd30fe50ad610fc361b0cf4586008852f3a5ad01653303f86c25e12dd91536

Observation 8378a85f-e41e-4276-9c62-43ec2ac1af75 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training data-efficient image transformers & distillation through attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.966919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.966919Z digest=sha256:9fc8a9ce9da7b90ce2eabde1ffc00a9c4188463f2cac6779fa6af0cefa2553d6

Observation e3fa46d1-7012-4a75-910a-88ec06df7c40 · outbound

This paper cites Attention is all you need.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Attention is all you need

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.969832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.969832Z digest=sha256:d3d7ca1e038cde48ba2029cb7bffb0539fb8a6d70f49baac47961a93613e940a

Observation 03c70042-9d43-43c0-8c76-c5f78c24529f · outbound

This paper cites Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.464845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.972731Z digest=sha256:ea0adc5a99daa63fc9f699f667189978d79bcae5b33e9ecebfd55150b1a1aecf

Observation 8adb33bb-3086-440e-b335-7cab20402372 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.975653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.975653Z digest=sha256:5628f177f5096a620321864498c7b9dc064ca2f57d0073412299374b1c740495

Observation b62e5c21-9f43-4b81-ab10-ab17702fb178 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.456906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.979027Z digest=sha256:99ae7f683a93a73c7e035a85b3119794b96709454613031b6ac07ea5c0d83cf2

Observation e24b8f23-3957-4d42-9397-6e0ec64b77d3 · outbound

This paper cites Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.448430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.982279Z digest=sha256:e6e444870ea9c2a50f1b64ca347df82ba782af35d208fff3420c9381b255bba6

Observation 1749f87a-2a79-4e35-ba69-ae658f8d1b8d · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tinyvit: Fast pretraining distillation for small vision transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.438892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.984954Z digest=sha256:eea785308e808d919c8dd991b93c49834e9061a6de926a678c359c47704fadd4

Observation 23cc25b4-72e7-4e58-bf94-be7ee936f505 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.987271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.987271Z digest=sha256:bdcc3adfde148a91a045e64123f6aaf0ddb353869f74d7bfe557c14f3b2e7d76

Observation 233e4718-6d55-4085-a644-6d508d92c1f0 · outbound

This paper cites Structured Pruning Learns Compact and Accurate Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Structured Pruning Learns Compact and Accurate Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.990020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.990020Z digest=sha256:00195b46da6e8f3a387f2d37838d442305b18ce849830c38b36e123acd3e2f31

Observation 33c2b9fc-d032-40a7-95a7-0384e79a0ba4 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.992893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.992893Z digest=sha256:9d98886a1d9a5360b348a8f37962b925ecf03905a9ca987e4d1d082331322699

Observation 33c4be17-a40d-4aa6-935b-01e311898b3d · outbound

This paper cites ThinK: Thinner Key Cache by Query-Driven Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.995190Z digest=sha256:c29065f72d37bd51854e84dfa934533a59e539cafa98f56858f10cd3fc5b8687

Observation 0273513d-35ee-4930-a305-18c2edda9540 · outbound

This paper cites X-pruner: explainable pruning for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation X-pruner: explainable pruning for vision transformers

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.424277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:51.998085Z digest=sha256:5ccba064397279d983021bae403f4c695590a771f707cd303a59141148ef3540

Observation 93c0355e-f848-45db-9b2f-12e6afd37318 · outbound

This paper cites Unified Visual Transformer Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unified Visual Transformer Compression

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.000864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.000864Z digest=sha256:49c38ce5ee4fddf5845426d7001af337107687066b6cfe90191fc9a59bdb675f

Observation 11a5dcf9-149d-4f7b-b713-3b8c2f8a1686 · outbound

This paper cites MoEfication: Transformer Feed-forward Layers are Mixtures of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEfication: Transformer Feed-forward Layers are Mixtures of Experts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.004555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.004555Z digest=sha256:00cccbc1649f5f0ff079198312e5783ab65c79df51295e008ef25651f3546295

Observation caea50ba-6c6e-424d-bb5c-0cc1ee2e5027 · outbound

This paper cites Knowledge distillation based on transformed teacher matching.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation based on transformed teacher matching

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.414808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:52.007595Z digest=sha256:b4b3fafd5c3c7ea39c1a450290f34f3c2e3d2356d214c8b871983c0c60c9f929

Observation 80958d0c-b3c7-4749-8de3-4432ec47e4ca · outbound

This paper cites The clock and the pizza: Two stories in mechanistic explanation of neural networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The clock and the pizza: Two stories in mechanistic explanation of neural networks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.010319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.010319Z digest=sha256:94ed62c846e317fecde6f11001f573812e30b2b85236fb9f93f5ac65af737909

Observation 19db123c-225d-4678-83db-042e11c80ad4 · outbound

This paper cites LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.013139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.013139Z digest=sha256:682bdf678d7ed1903ea8dacf41d4ed6edb0a4a2e47d65381f70fc2998f83b840

Observation e89b4d73-6cb9-4692-bfff-0be0ab6974aa · outbound

This paper cites MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.016407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.016407Z digest=sha256:32dcdda91445c2aeb87151ad2998d413e668d4ccab1e80b1e37abd2352bb2f80

Observation 2a226cfb-df33-4745-a251-066a3616abc2 · outbound

This paper cites Possible solutions could include:.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Possible solutions could include:

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.402343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:52.019705Z digest=sha256:7d06066943b02d5ab572be4f2b6f869727456a8717fd50d77d419c1a7bb5d99b

Observation da8cde43-a337-4086-a708-a2e5e522a007 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.394336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:52.022253Z digest=sha256:75c1072c062dd3a0a2668bd2456b97b655d625ed459d72d3514439dbce83dca3

Observation 35c4f703-c447-4c33-ad9e-f9b61206eaa1 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.385463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:38:52.024717Z digest=sha256:72096c70da6631cb119473a943a68f24f407d052746277b1f71681fcb1faedb5

Pith citing papers

No inbound Pith citation observations are available.