Pith. sign in

Paper Citation Record · LEDGER

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models

As of 18 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2504.12471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12471 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:01.600260Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5bc583a-1d18-4d07-b544-3cfbc9c63eca · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.378085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.378085Z digest=sha256:30e8e7b59300e5d47cfe07a88d2dfb4f3608cb1e593eaf02eb3bc1dba61dde7d

Observation 04709667-ddbf-405f-a4bd-73688a44223d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.383768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.383768Z digest=sha256:b8fae66d674740ac4d2c8eb6f821bcfcc22843c5c972e04bd390176487e4a89c

Observation 02205f2c-1793-4a17-99bf-9f8710723541 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.389715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.389715Z digest=sha256:a4be4ed1a0d1ee6a6601677b4b9b39600837901dc9990947aef868c2e2aab53b

Observation d6e31786-ba41-4a6b-a1f7-3ff7ca54effd · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Xlnet: Generalized autoregressive pretraining for language understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.394718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.394718Z digest=sha256:e1c76ae01ee73e390b60f4bc19b6a56075c03fc0a2e8ee3adc86783b0ea4f871

Observation 82e43c9f-4fc6-4ad5-ae42-b7508533dd6a · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.399645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.399645Z digest=sha256:d994ed405294aa68d45bada912e66624605360d6ff51290e6729cafe037ab999

Observation 3a1ef966-eeb9-4a2e-b753-49bee8ac847e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.404894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.404894Z digest=sha256:8c0b9c7638430ca91171eb6f517283a81831dc532f802d423b29dce1decca913

Observation e3d99227-abc4-49eb-a972-63038d69d19e · outbound

This paper cites Tokens-to-token vit: Training vision transformers from scratch on imagenet,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Tokens-to-token vit: Training vision transformers from scratch on imagenet,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.410614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.410614Z digest=sha256:9e1ab4a501761486cc88d1c89215107740473113617bdb83c898c8d5012eeb23

Observation 86087a9a-baec-4b2b-a6dc-ae24c91e7838 · outbound

This paper cites Train big, then compress: Rethinking model size for efficient training and inference of transformers,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Train big, then compress: Rethinking model size for efficient training and inference of transformers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.272089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.415236Z digest=sha256:678f666f49d0367e4d619d54268f787d6623746473c637aade3f0db3c7bd7fd4

Observation 0e480de8-554a-4c8f-a501-7298829a108a · outbound

This paper cites Decoupled greedy learning of cnns,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Decoupled greedy learning of cnns,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.419869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.419869Z digest=sha256:6dab460a5435a4a847d34f0280613eafa234240ae70e6badedf371a4bf607113

Observation 4e8f5d6b-9443-41c8-bf3c-74974535856d · outbound

This paper cites Distributed learning of fully connected neural networks using indepen- dent subnet training,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Distributed learning of fully connected neural networks using indepen- dent subnet training,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.246274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.424539Z digest=sha256:c72ef06059ac64468ee93fd6a39a4ca33d253a21ceea6c36c5856b053aa25132

Observation c023a108-ae6c-4b7d-8f26-6d9c6cc591b6 · outbound

This paper cites Decentralized training of foundation models in heterogeneous environments,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Decentralized training of foundation models in heterogeneous environments,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.230088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.429256Z digest=sha256:d360e503690db8da82f038aac27b2520d6dc0495c9dec88dc4c1ab1e2e03453d

Observation f041e34f-aae2-4916-9ead-f60fc8ea61d7 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.434243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.434243Z digest=sha256:3f06e26c5296a046141a827f094f2f2783351d663fa7669cebe8b6326eeb939a

Observation 6faf335d-8fd7-4e43-96fd-e2b04e8c3269 · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.439349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.439349Z digest=sha256:fe75005cc4fffca855e179741d57c77b39ed7666154b80c02eaa7ad945baeb22

Observation ccaef75d-83c5-4b45-8cec-db44cf6b1fc8 · outbound

This paper cites An efficient 2d method for training super-large deep learning models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An efficient 2d method for training super-large deep learning models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.214409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.444343Z digest=sha256:3167cd2009c73a3b1f5411d2c1bb3a41bd6bfb7a45c9f2f11b5df71a92d295ee

Observation 4320ec30-7210-40a5-a002-fee5d1716e03 · outbound

This paper cites Maximizing Parallelism in Distributed Training for Huge Neural Networks.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Maximizing Parallelism in Distributed Training for Huge Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.449128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.449128Z digest=sha256:839e32857f69ba0d13fc10e07f5a89d040f971fbb7cde3f4f20e9d221c0fbba0

Observation 49b9a9a7-10df-4ba2-ae0c-5e966e4f8c9b · outbound

This paper cites Attention is all you need,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Attention is all you need,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.454113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.454113Z digest=sha256:1d8f2c55506dd088f014051f44ff69553f4627e7690fba8ec2327f7b1bc0cd26

Observation 56549b7d-b47f-46af-a1bc-77c70bf23c34 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.458824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.458824Z digest=sha256:01846911030dc22e936ae2d95f4976474d1581ce94d0a6f9629e9ffe7311f499

Observation eb481c18-cf18-4341-ad44-d177df2609b1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.463962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.463962Z digest=sha256:49295a656ec1881f0a0595b48e3b40d83456c379f1bcec90b94e395a1d96059a

Observation 4dc69d1d-4ee9-45f5-9ddb-999c73ec0a71 · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Parameter-efficient transfer learning for nlp,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.469195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.469195Z digest=sha256:979466ad90a3a2a2830821e721d4e3c0bd4127267eb8a3e24e95f88fbc42578a

Observation f9f12e04-89a8-4707-a8c8-a7cea08d65a3 · outbound

This paper cites Transformer in transformer,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Transformer in transformer,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.474722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.474722Z digest=sha256:0578d9de4382ed9447ba8521b8f9822fc36e26f0d977d96b280992f055d48373

Observation eeb14255-c24c-408c-8437-a2b1da53971e · outbound

This paper cites Dynamic Model Pruning with Feedback.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic Model Pruning with Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.479454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.479454Z digest=sha256:5851480a23c1e7791d9f6d743b4f135a22e088a13d16c041f68aeade864f8c8f

Observation 1ab2105b-d984-4bf5-a5f1-82535f07fb4e · outbound

This paper cites Single-shot pruning for pre-trained models: Rethinking the importance of magnitude pruning,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Single-shot pruning for pre-trained models: Rethinking the importance of magnitude pruning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.168168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.484437Z digest=sha256:79c17cb4b22bb94be2cbb493cbfd8ae88ae56fa4dc897f5ae24147bdd5f593a0

Observation 175af5f5-08f8-440a-92db-bd0b027f9159 · outbound

This paper cites Resource- efficient transformer pruning for finetuning of large models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Resource- efficient transformer pruning for finetuning of large models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.150004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.489089Z digest=sha256:97d1bc6262bdd4d72b549a16c654979705a3088ec975a4520b0fa29708d4c8cf

Observation a8d6db2d-79ac-45cb-ba58-e75a61e57f95 · outbound

This paper cites Martello and P.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Martello and P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.133939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.493907Z digest=sha256:2423709f59b5265ad57e7678d5d7c34e37dc78531db7594aeee239dbd2a9ce31

Observation 10bfffd0-03f8-4671-b402-0ab5a23e6dd0 · outbound

This paper cites A class of generalized greedy algorithms for the multi-knapsack problem,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models A class of generalized greedy algorithms for the multi-knapsack problem,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.118280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.498626Z digest=sha256:ba352235408eabb6716298e90ec35db0a641ea723f0e181d4027c3a50d431bc0

Observation 33ba2ba8-7552-4e28-a629-1a4b909625ba · outbound

This paper cites Dynamic programming revisited: Improving knapsack algorithms,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic programming revisited: Improving knapsack algorithms,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.102410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.503067Z digest=sha256:f4ea2a958ffecb41d9f94b5bf4abe9ebe605cebbeeab7d6df1b70cb54f8ade2d

Observation 56e54159-64e2-4182-ac57-1cf760711725 · outbound

This paper cites An approximate dynamic programming ap- proach to multidimensional knapsack problems,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An approximate dynamic programming ap- proach to multidimensional knapsack problems,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.087031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.507570Z digest=sha256:38c66591692bf38c98be43dd3733194b6d38ee40d87bca6f65944c52d5790cc4

Observation 894fd2b9-442b-47c7-a8bc-b9890dacd224 · outbound

This paper cites Pytorch image models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Pytorch image models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.512195Z digest=sha256:dc6ce7e0fa701b9945616e692ec83a68136115eff778569dd76190ad7cb277b1

Observation 85a95a5f-ce58-4558-9e6a-e90d1ab14813 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Pytorch: An imperative style, high-performance deep learning library,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.060041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.516879Z digest=sha256:d40ecf96e6068db2b25f64a7cfc72a48620c6a519cf55bb828277909abc121a4

Observation 71ca0b38-9469-4d83-9c36-cc4b77b0b685 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.521699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.521699Z digest=sha256:9c6c4c74c069145a448651b3fde14037b13c06904ec72200732d5482576006a4

Observation bcab0ed5-41ea-4fbb-9e57-1557970f507c · outbound

This paper cites Where to pay attention in sparse training for feature selection?.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Where to pay attention in sparse training for feature selection?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.044084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.526599Z digest=sha256:d6b57ced0fcb976e42b7bb3430b378b9dcb8d23949af08118e716e94d6a6c891

Observation 5ef3de64-789e-440e-8201-99704436b26f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.531151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.531151Z digest=sha256:591e20873d5382a7d803f79923b43867c11362ba244a4a7822bd712298b8d18d

Observation bb564558-b981-491a-8465-66bf03934ff4 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.536230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.536230Z digest=sha256:094c08e4ac7a9688aa2feb9a9581d1f4f28be6280bc7cd47063c38aec4378779

Observation 7bc1e12c-f978-4af8-a3a0-f45e3d1e71a4 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.540973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.540973Z digest=sha256:e68b0e8111d2edcd6d5be67dc681d6f17a4de850930a0b2ac91b11736fbbf40b

Observation 0c8373a6-2710-4d9c-ad0b-523a863bc711 · outbound

This paper cites M6-T: Exploring Sparse Expert Models and Beyond.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models M6-T: Exploring Sparse Expert Models and Beyond

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.546388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.546388Z digest=sha256:f902c91d3ac05e079da4639009f380f2f65cbf5057285453e66579b60547b991

Observation aee6c30a-1aff-4e7b-9b8a-135de7eab909 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Taming Sparsely Activated Transformer with Stochastic Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.551591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.551591Z digest=sha256:f94b91cb3798aae0999e2835e788cfe33bd88ea24b8be9daa2355da2db5f7074

Observation 0050a150-c160-4e2e-b3ed-daaa81272531 · outbound

This paper cites Hash layers for large sparse models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Hash layers for large sparse models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.006695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.556853Z digest=sha256:2a37be526b2bceb622c83ef659ff43ba1053a7a10f14403014271f0445acf6a6

Observation 2fdb58dd-8b11-4296-a9a1-d46c85de0cc2 · outbound

This paper cites SNIP: Single-shot Network Pruning based on Connection Sensitivity.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models SNIP: Single-shot Network Pruning based on Connection Sensitivity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.561592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.561592Z digest=sha256:35567ae67047f372c0d63f2cc09250940e121cca057bca9faef7202d6e9c4748

Observation 691ce3dc-2c8d-4d70-971d-cd7ec6a0d0c9 · outbound

This paper cites Picking Winning Tickets Before Training by Preserving Gradient Flow.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Picking Winning Tickets Before Training by Preserving Gradient Flow

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.566591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.566591Z digest=sha256:d092f77a8c91b85e96568d06e0af0d0298c16bd827f25c9a6ee5e1ba27cd3450

Observation 87a996e3-6500-4902-9882-304c8bec13aa · outbound

This paper cites Learning both weights and con- nections for efficient neural network,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Learning both weights and con- nections for efficient neural network,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.571670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.571670Z digest=sha256:bb6a1ca9d57ac1a88ea0a69760f6cefd7020acbdded4126cc1c32431930744e8

Observation a2f50ac5-1285-4a8d-b766-85a488798fb9 · outbound

This paper cites Dynamic network surgery for efficient dnns,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic network surgery for efficient dnns,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.978419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.576409Z digest=sha256:cd03ff17cf2ee4ce76acae5288cdb9b7373173ce9b419f465af521cb96e6e1ae

Observation 1d97bb73-1bef-4b96-a880-e413ea0a24fc · outbound

This paper cites Compression-aware training of deep networks,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Compression-aware training of deep networks,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.962721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.581071Z digest=sha256:cc1d189bcf0c4f93ca40cf9d82f877d64d9a6deee94a07cc3777b676ebf6ec7e

Observation 758b4354-cabb-4fbd-b152-cb64229f978a · outbound

This paper cites “learning-compression.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models “learning-compression

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.947063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.585807Z digest=sha256:9f866dc9665720e99231d78ed9882d7e2013dc24f0510ade2f5460de70e483b1

Observation 49c00761-95bd-49c7-88a0-208a1f28b007 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.590570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.590570Z digest=sha256:320645817b96fd0a907d04fa5e9a58f6092c3d62a2a6b78647dea834114c5a3f

Observation 02b0f596-11a5-4078-b95f-c9a752d8a809 · outbound

This paper cites Efficient lottery ticket finding: Less data is more,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Efficient lottery ticket finding: Less data is more,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.930852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.595807Z digest=sha256:39d1a7a2fae30909fe0179fe5e1f14a01efe875ab46082e3659af7e6a6c37ba1

Observation 2601157b-698e-4da9-97c2-fcb0decdc4d9 · outbound

This paper cites The lottery ticket hypothesis for object recognition,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models The lottery ticket hypothesis for object recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.914443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:40:01.600260Z digest=sha256:e1c429daa47d8e1630ba51fe8347802c1817785c444cdf5f1b5caf14f1d4fce7

Pith citing papers

No inbound Pith citation observations are available.