Pith. sign in

Paper Citation Record · LEDGER

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models

As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2504.12471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12471 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:01.600260Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5bc583a-1d18-4d07-b544-3cfbc9c63eca · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.378085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.378085Z digest=sha256:d7a5335e543b2a2822f06d3359b833958e1c7f0b72971e23b7f2d19faed97298

Observation 04709667-ddbf-405f-a4bd-73688a44223d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.383768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.383768Z digest=sha256:ef92bea96f86f060466fed0dd8bfaa970e3ccdbc138642f29390e2c0bae6060c

Observation 02205f2c-1793-4a17-99bf-9f8710723541 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.389715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.389715Z digest=sha256:37f252b824b13d0fe3bf316f5e5cf8f51eb82fcfd88bf6302c0e62289286aef9

Observation d6e31786-ba41-4a6b-a1f7-3ff7ca54effd · outbound

This paper cites Xlnet: Generalized autoregressive pretraining for language understanding,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Xlnet: Generalized autoregressive pretraining for language understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.394718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.394718Z digest=sha256:1957d8f0f40ff92e93892e95cb6ed21a9cfa1691614ffda7a112bcc0ccb42a9d

Observation 82e43c9f-4fc6-4ad5-ae42-b7508533dd6a · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.399645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.399645Z digest=sha256:973813f93b72700a9344698360f1d22642a0f263d16dcc28b7522b66c68b9e81

Observation 3a1ef966-eeb9-4a2e-b753-49bee8ac847e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.404894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.404894Z digest=sha256:fd4db8de964fe79b0e92d434ceeefb4476a2753b3d3cf12261240b2ae88615c0

Observation e3d99227-abc4-49eb-a972-63038d69d19e · outbound

This paper cites Tokens-to-token vit: Training vision transformers from scratch on imagenet,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Tokens-to-token vit: Training vision transformers from scratch on imagenet,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.410614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.410614Z digest=sha256:620d95186b0559fd859ca94cb3e1b605e86c7c7377b0bdf6b40774329d281943

Observation 86087a9a-baec-4b2b-a6dc-ae24c91e7838 · outbound

This paper cites Train big, then compress: Rethinking model size for efficient training and inference of transformers,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Train big, then compress: Rethinking model size for efficient training and inference of transformers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.272089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.415236Z digest=sha256:4fe197282c4cc27b35c97ad67388681e9263e77c34d53fe70e55a63f2ced37a3

Observation 0e480de8-554a-4c8f-a501-7298829a108a · outbound

This paper cites Decoupled greedy learning of cnns,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Decoupled greedy learning of cnns,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.419869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.419869Z digest=sha256:f869111820f62f2eb2e8a7f97624f613ef2ed1cb745b56a58a34bca895968878

Observation 4e8f5d6b-9443-41c8-bf3c-74974535856d · outbound

This paper cites Distributed learning of fully connected neural networks using indepen- dent subnet training,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Distributed learning of fully connected neural networks using indepen- dent subnet training,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.246274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.424539Z digest=sha256:7830781d5a14d4051a64fafdd96b875e9392e603f8867683b16b31912f80495c

Observation c023a108-ae6c-4b7d-8f26-6d9c6cc591b6 · outbound

This paper cites Decentralized training of foundation models in heterogeneous environments,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Decentralized training of foundation models in heterogeneous environments,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.230088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.429256Z digest=sha256:2e287a1c4cd09e743a04971d7e97c53348ef91d291857750452c321dc0c90845

Observation f041e34f-aae2-4916-9ead-f60fc8ea61d7 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.434243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.434243Z digest=sha256:69e658008fdec13a7721aa94aa61a7f6c2a4bf04916dffa8a43143fdfb658d1b

Observation 6faf335d-8fd7-4e43-96fd-e2b04e8c3269 · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.439349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.439349Z digest=sha256:a9d370b3a58309528031ca1452b87e021b4f25cc7647485ac3e5c0c9b0de5268

Observation ccaef75d-83c5-4b45-8cec-db44cf6b1fc8 · outbound

This paper cites An efficient 2d method for training super-large deep learning models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An efficient 2d method for training super-large deep learning models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.214409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.444343Z digest=sha256:23196835c1ca6168e9938a018a3c8a9db9218dad9a76c7bcd58b38893454c455

Observation 4320ec30-7210-40a5-a002-fee5d1716e03 · outbound

This paper cites Maximizing Parallelism in Distributed Training for Huge Neural Networks.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Maximizing Parallelism in Distributed Training for Huge Neural Networks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.449128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.449128Z digest=sha256:11cf75df1d3a5a2154cffa98093949832c6a689139a1f6dc2487909f0f0b5bdf

Observation 49b9a9a7-10df-4ba2-ae0c-5e966e4f8c9b · outbound

This paper cites Attention is all you need,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Attention is all you need,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.454113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.454113Z digest=sha256:3a233e5792c1dab9a913d67549e113a8e78d23831f1a761c509b1501b2c05320

Observation 56549b7d-b47f-46af-a1bc-77c70bf23c34 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.458824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.458824Z digest=sha256:f2aeb353030f60cf5d17ac1ca7c51685882649a9b238bbfc7aa0bb4d9af184bb

Observation eb481c18-cf18-4341-ad44-d177df2609b1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.463962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.463962Z digest=sha256:c83d5e49c3e816732f275949dda346382a2f152f2b3821894e12865ffd06f728

Observation 4dc69d1d-4ee9-45f5-9ddb-999c73ec0a71 · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Parameter-efficient transfer learning for nlp,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.469195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.469195Z digest=sha256:85e824875b43ac3f790377d6c0f179dcd97df7ca66b1638d6e4dd11e8f6d65f6

Observation f9f12e04-89a8-4707-a8c8-a7cea08d65a3 · outbound

This paper cites Transformer in transformer,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Transformer in transformer,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.474722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.474722Z digest=sha256:5bc6f462575a9b3ad6c58fb63ae7be44f3ef10232eb49566c63a461d1ffd0d4f

Observation eeb14255-c24c-408c-8437-a2b1da53971e · outbound

This paper cites Dynamic Model Pruning with Feedback.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic Model Pruning with Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.479454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.479454Z digest=sha256:1753d4a2ce6ac62f16c463aa94e261093d65e01f78038e50ad018235b58104d5

Observation 1ab2105b-d984-4bf5-a5f1-82535f07fb4e · outbound

This paper cites Single-shot pruning for pre-trained models: Rethinking the importance of magnitude pruning,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Single-shot pruning for pre-trained models: Rethinking the importance of magnitude pruning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.168168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.484437Z digest=sha256:1b645428954ac6fcf773425bedb021201f63094802470e2badf34746b370b377

Observation 175af5f5-08f8-440a-92db-bd0b027f9159 · outbound

This paper cites Resource- efficient transformer pruning for finetuning of large models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Resource- efficient transformer pruning for finetuning of large models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.150004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.489089Z digest=sha256:02d2895fdeb82d05847349a8f745e7634bd87143d8aff203f7ba467034390fba

Observation a8d6db2d-79ac-45cb-ba58-e75a61e57f95 · outbound

This paper cites Martello and P.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Martello and P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.133939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.493907Z digest=sha256:ec9a443583fb29e0403f29bda91a5bafbf38990091a6b119105a271d8f49f302

Observation 10bfffd0-03f8-4671-b402-0ab5a23e6dd0 · outbound

This paper cites A class of generalized greedy algorithms for the multi-knapsack problem,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models A class of generalized greedy algorithms for the multi-knapsack problem,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.118280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.498626Z digest=sha256:57495ae30be6c9b92f2ba10e61c4ed89b57017277c8f3c80cc7753284133f5f0

Observation 33ba2ba8-7552-4e28-a629-1a4b909625ba · outbound

This paper cites Dynamic programming revisited: Improving knapsack algorithms,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic programming revisited: Improving knapsack algorithms,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.102410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.503067Z digest=sha256:91d2e3fa4e2f5f2ca244d397b10a7940e59f6fc2ae11fddd3661c96bf61c18b6

Observation 56e54159-64e2-4182-ac57-1cf760711725 · outbound

This paper cites An approximate dynamic programming ap- proach to multidimensional knapsack problems,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models An approximate dynamic programming ap- proach to multidimensional knapsack problems,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.087031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.507570Z digest=sha256:5fdd785625f216750aa72147f33f4f3826b583c2a2c82c17d070bfc192f20f97

Observation 894fd2b9-442b-47c7-a8bc-b9890dacd224 · outbound

This paper cites Pytorch image models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Pytorch image models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.512195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.512195Z digest=sha256:fe2895e684c9d37b95825f3c65694fe55237fdf0b83bda7b2f7ab275b9bd2a33

Observation 85a95a5f-ce58-4558-9e6a-e90d1ab14813 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Pytorch: An imperative style, high-performance deep learning library,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.060041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.516879Z digest=sha256:e21f28479add04f090805dd6609adec659a7daa6bd3975f4af6ae96bb6ef853f

Observation 71ca0b38-9469-4d83-9c36-cc4b77b0b685 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.521699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.521699Z digest=sha256:1f777dc79d7e1c725b3aceb30f94da39e8c6177d004c87ea993a66737bd9edde

Observation bcab0ed5-41ea-4fbb-9e57-1557970f507c · outbound

This paper cites Where to pay attention in sparse training for feature selection?.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Where to pay attention in sparse training for feature selection?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.044084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.526599Z digest=sha256:b31e1bfbe831dcf122b8f5b705cc7ffd293003799704ed393e6e37f290206df2

Observation 5ef3de64-789e-440e-8201-99704436b26f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.531151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.531151Z digest=sha256:595ed107121a3b832c7b8d2ec2ebc3636c584775e1e91e139cb5a9668453b77b

Observation bb564558-b981-491a-8465-66bf03934ff4 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.536230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.536230Z digest=sha256:e4471d4ebf3f82dc24c8c532987e0492ec4231c7eba969edf118cd69dc948ba7

Observation 7bc1e12c-f978-4af8-a3a0-f45e3d1e71a4 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.540973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.540973Z digest=sha256:e83aabf626fc89b73f6f217f806fe301614b0e1185e9e02256fb6e2224ebf399

Observation 0c8373a6-2710-4d9c-ad0b-523a863bc711 · outbound

This paper cites M6-T: Exploring Sparse Expert Models and Beyond.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models M6-T: Exploring Sparse Expert Models and Beyond

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.546388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.546388Z digest=sha256:0d6536a208d3dbbbd7e8ebd8e925311cdbfc7ccd0c17a03b56bf495645d7446c

Observation aee6c30a-1aff-4e7b-9b8a-135de7eab909 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Taming Sparsely Activated Transformer with Stochastic Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.551591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.551591Z digest=sha256:0c569321d7b0ddecdd608da6b4121027e0bb3b15c6a2f0b729ae6a2956c6e8ff

Observation 0050a150-c160-4e2e-b3ed-daaa81272531 · outbound

This paper cites Hash layers for large sparse models,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Hash layers for large sparse models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:02.006695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.556853Z digest=sha256:454b592e5739a36d6306d005403fe04445cb9e30eb3c57b5f60acf34731a466e

Observation 2fdb58dd-8b11-4296-a9a1-d46c85de0cc2 · outbound

This paper cites SNIP: Single-shot Network Pruning based on Connection Sensitivity.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models SNIP: Single-shot Network Pruning based on Connection Sensitivity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.561592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.561592Z digest=sha256:628120b638dc504c7049cc5dff9f4931737dd22169a2cc27475968ebc8fd77e6

Observation 691ce3dc-2c8d-4d70-971d-cd7ec6a0d0c9 · outbound

This paper cites Picking Winning Tickets Before Training by Preserving Gradient Flow.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Picking Winning Tickets Before Training by Preserving Gradient Flow

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.566591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.566591Z digest=sha256:12d1775d94a9324ef83ec19ed603c3ccf28bc6ad6811d92ff64d2718fcc97cf0

Observation 87a996e3-6500-4902-9882-304c8bec13aa · outbound

This paper cites Learning both weights and con- nections for efficient neural network,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Learning both weights and con- nections for efficient neural network,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.571670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.571670Z digest=sha256:71b43494f84e54dfcdbe633e430e7c2ce65ce1fb2b7c63b6b57a3e3fa235e4ae

Observation a2f50ac5-1285-4a8d-b766-85a488798fb9 · outbound

This paper cites Dynamic network surgery for efficient dnns,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Dynamic network surgery for efficient dnns,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.978419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.576409Z digest=sha256:b6467c3301f849417db02605fda168cf4c02282c981bbff44139c0f8862b0094

Observation 1d97bb73-1bef-4b96-a880-e413ea0a24fc · outbound

This paper cites Compression-aware training of deep networks,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Compression-aware training of deep networks,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.962721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.581071Z digest=sha256:46ac613e8175163507fec8cb86bee4cafd2d62a0954d89d655e644725a22d658

Observation 758b4354-cabb-4fbd-b152-cb64229f978a · outbound

This paper cites “learning-compression.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models “learning-compression

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.947063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.585807Z digest=sha256:f79b47ce228cbd5f564646fc636771f2c223daa9e38a15c33f3c67f3c058fc7c

Observation 49c00761-95bd-49c7-88a0-208a1f28b007 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:01.590570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:01.590570Z digest=sha256:407671635d7f9c0c0fbdad75255816deb443dbfaf7c4550d73a6a592c8f07b31

Observation 02b0f596-11a5-4078-b95f-c9a752d8a809 · outbound

This paper cites Efficient lottery ticket finding: Less data is more,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models Efficient lottery ticket finding: Less data is more,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.930852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.595807Z digest=sha256:5cc6b9f33fb845faa3891c9f9bdc0db99324925034e957197220b22dee9e8a2d

Observation 2601157b-698e-4da9-97c2-fcb0decdc4d9 · outbound

This paper cites The lottery ticket hypothesis for object recognition,.

You Don't Need All Attentions: Distributed Dynamic Fine-Tuning for Foundation Models The lottery ticket hypothesis for object recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:40:01.914443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T12:40:01.600260Z digest=sha256:493fde2de094de1cefcc7e74c09aa58b2ec8d8896a4513f13cc1d3627524f761

Pith citing papers

No inbound Pith citation observations are available.