Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:01.600260Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2504.12471.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:01.600260Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d5bc583a-1d18-4d07-b544-3cfbc9c63eca · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04709667-ddbf-405f-a4bd-73688a44223d · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02205f2c-1793-4a17-99bf-9f8710723541 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6e31786-ba41-4a6b-a1f7-3ff7ca54effd · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Xlnet: Generalized autoregressive pretraining for language understanding,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e43c9f-4fc6-4ad5-ae42-b7508533dd6a · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1ef966-eeb9-4a2e-b753-49bee8ac847e · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d99227-abc4-49eb-a972-63038d69d19e · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Tokens-to-token vit: Training vision transformers from scratch on imagenet,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86087a9a-baec-4b2b-a6dc-ae24c91e7838 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Train big, then compress: Rethinking model size for efficient training and inference of transformers,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0e480de8-554a-4c8f-a501-7298829a108a · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Decoupled greedy learning of cnns,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8f5d6b-9443-41c8-bf3c-74974535856d · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Distributed learning of fully connected neural networks using indepen- dent subnet training,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c023a108-ae6c-4b7d-8f26-6d9c6cc591b6 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Decentralized training of foundation models in heterogeneous environments,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f041e34f-aae2-4916-9ead-f60fc8ea61d7 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6faf335d-8fd7-4e43-96fd-e2b04e8c3269 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models GSPMD: General and Scalable Parallelization for ML Computation Graphs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccaef75d-83c5-4b45-8cec-db44cf6b1fc8 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An efficient 2d method for training super-large deep learning models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4320ec30-7210-40a5-a002-fee5d1716e03 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Maximizing Parallelism in Distributed Training for Huge Neural Networks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b9a9a7-10df-4ba2-ae0c-5e966e4f8c9b · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Attention is all you need,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56549b7d-b47f-46af-a1bc-77c70bf23c34 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb481c18-cf18-4341-ad44-d177df2609b1 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc69d1d-4ee9-45f5-9ddb-999c73ec0a71 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Parameter-efficient transfer learning for nlp,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9f12e04-89a8-4707-a8c8-a7cea08d65a3 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Transformer in transformer,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeb14255-c24c-408c-8437-a2b1da53971e · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic Model Pruning with Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab2105b-d984-4bf5-a5f1-82535f07fb4e · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Single-shot pruning for pre-trained models: Rethinking the importance of magnitude pruning,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 175af5f5-08f8-440a-92db-bd0b027f9159 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Resource- efficient transformer pruning for finetuning of large models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a8d6db2d-79ac-45cb-ba58-e75a61e57f95 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Martello and P
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 10bfffd0-03f8-4671-b402-0ab5a23e6dd0 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models A class of generalized greedy algorithms for the multi-knapsack problem,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 33ba2ba8-7552-4e28-a629-1a4b909625ba · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic programming revisited: Improving knapsack algorithms,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56e54159-64e2-4182-ac57-1cf760711725 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An approximate dynamic programming ap- proach to multidimensional knapsack problems,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 894fd2b9-442b-47c7-a8bc-b9890dacd224 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Pytorch image models,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85a95a5f-ce58-4558-9e6a-e90d1ab14813 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Pytorch: An imperative style, high-performance deep learning library,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 71ca0b38-9469-4d83-9c36-cc4b77b0b685 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcab0ed5-41ea-4fbb-9e57-1557970f507c · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Where to pay attention in sparse training for feature selection?
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5ef3de64-789e-440e-8201-99704436b26f · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb564558-b981-491a-8465-66bf03934ff4 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bc1e12c-f978-4af8-a3a0-f45e3d1e71a4 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8373a6-2710-4d9c-ad0b-523a863bc711 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models M6-T: Exploring Sparse Expert Models and Beyond
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee6c30a-1aff-4e7b-9b8a-135de7eab909 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0050a150-c160-4e2e-b3ed-daaa81272531 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Hash layers for large sparse models,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2fdb58dd-8b11-4296-a9a1-d46c85de0cc2 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models SNIP: Single-shot Network Pruning based on Connection Sensitivity
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691ce3dc-2c8d-4d70-971d-cd7ec6a0d0c9 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Picking Winning Tickets Before Training by Preserving Gradient Flow
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a996e3-6500-4902-9882-304c8bec13aa · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Learning both weights and con- nections for efficient neural network,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f50ac5-1285-4a8d-b766-85a488798fb9 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic network surgery for efficient dnns,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1d97bb73-1bef-4b96-a880-e413ea0a24fc · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Compression-aware training of deep networks,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 758b4354-cabb-4fbd-b152-cb64229f978a · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models “learning-compression
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 49c00761-95bd-49c7-88a0-208a1f28b007 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b0f596-11a5-4078-b95f-c9a752d8a809 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Efficient lottery ticket finding: Less data is more,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2601157b-698e-4da9-97c2-fcb0decdc4d9 · outbound
You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models The lottery ticket hypothesis for object recognition,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.