Pith. sign in

Paper Citation Record · LEDGER

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts

As of 20 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2509.10530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10530 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:32:03.727373Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b21f112d-6e84-4952-a97d-2b5314cff2ee · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.102282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.102282Z digest=sha256:f99d1ee9755088d70c3f7ad638dc4b922dc35c30840acc24917c518ba9634527

Observation 418fec02-e729-4cc7-b677-502c07e9dccf · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.114581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.114581Z digest=sha256:312cf9d52620bf8515d1df6f3aa62331900a4a0760c5a8200b9fa2fae443dfac

Observation a2cfd6db-3aca-4ed1-a360-c97afcc8c9aa · outbound

This paper cites Scaling Laws for Neural Language Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Scaling Laws for Neural Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.122984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.122984Z digest=sha256:61362ab2ae9261b818daccee92e0760e5c8c831c8c8390638979753bc706e3e1

Observation 370dd306-3b37-448e-8b88-6fbb2d425ed9 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.132827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.132827Z digest=sha256:f6840e842f99557ceb4da5a5f63dcbe5652e7a77470134c85e2010eb8b2fb542

Observation 7f772533-9eaf-4953-b708-019590d4c3b6 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.141013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.141013Z digest=sha256:78aa0a6ca9b85c8aa125dbe92263cdfc1dffc1668b008d85b4c4f255c585724a

Observation c026a01c-5db7-416b-af46-ef1937360e92 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.152543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.152543Z digest=sha256:de2b8b63f02b10747201e5b6558bd0121ad8974631811997fec55fc620df5757

Observation 07fe1b71-d834-4af7-9164-0fc1304e57e5 · outbound

This paper cites Adaptive mixtures of local experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Adaptive mixtures of local experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.161637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.161637Z digest=sha256:87c49d16ff91bea980f4558bfcb606802d4b0a1e42ebf4273397a5ef14d1db09

Observation 0b9c3a84-db0d-4c4a-a20f-374f2e2e96e0 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts What Does BERT Look At? An Analysis of BERT's Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.168149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.168149Z digest=sha256:b1960fcad8c33fcd88abe343912e10c7c83786759f7cc341ddc665f7ee0aaac7

Observation e2464f07-ac6f-458e-b872-bbbfc014d0b7 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.174237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.174237Z digest=sha256:a47638c64d0ef757244dd5276e5ca1ad8e1a248475e02e6d7399a49cf8f89511

Observation 845cbdcc-ed51-4c1f-bd48-6711d3a0c56c · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.178360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.181778Z digest=sha256:8d5415323ad6322d74104693896e976383fcd5764ce00789ebed9c667287e53f

Observation 380e7d19-998e-4c41-86d2-dc1cbe62254f · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, Nov 2022.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Chatgpt: Optimizing language models for dialogue, Nov 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.160707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.188074Z digest=sha256:e97d15bd5efc875f0d93cd7277f1495b2126897e0d41247a0cad551d9b63b4c5

Observation bc83fb04-e1a6-4ac1-905c-270fe3c952d3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts LLaMA: Open and Efficient Foundation Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.196590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.196590Z digest=sha256:95319a206c0263ed3eef75aa33f522a61b0f3a4dcef9b428dbf24d9113aa66f3

Observation e0d0b965-b4b4-4de5-b6ca-918a8a3aaf06 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Glam: Efficient scaling of language models with mixture-of-experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.205921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.205921Z digest=sha256:c885e23112897169980aa0c4cda5e2eb3690e542a1510b4aa1635ea09099bc71

Observation 81c2cb9a-e169-4e83-92dc-b58f18f78bbf · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Glam: Efficient scaling of language models with mixture-of-experts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.124029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.215508Z digest=sha256:4b4aba0c88d1a170bdf4a699892b38bc83e5674dc5e04c1167b0a50362ece2f0

Observation 722a6d10-20c9-480f-9b15-5bc66e284559 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.226383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.226383Z digest=sha256:b16e39c61484fda880c50e3c292e30fc4338200e240134ad2b1969517db5abb3

Observation bcdbf71e-4747-4e51-ab4e-fb3743e3a3ea · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Mixture-of-Experts with Expert Choice Routing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.233617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.233617Z digest=sha256:60be68e023f0d2cb1c527b14ce37bb886de94fa02eb957029cd9c1510cf9f5e6

Observation 86591840-0333-4746-b4ed-fd9cdb30bf34 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Palm: Scaling language modeling with pathways

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.249006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.249006Z digest=sha256:4b852a6c036835984b3031a16d449a789440ae98f91002a64a7fa29f43c44748

Observation 72e07bab-da7d-412b-82f1-329f27358bbf · outbound

This paper cites The Efficiency Misnomer.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Efficiency Misnomer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.259828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.259828Z digest=sha256:e87de6f7e5f4c6a21b10be8a031815822d9cf10db9eab6fa9fbfa63068d52fa1

Observation b0f5f81b-586a-43af-821e-6494e872ac9f · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.268556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.268556Z digest=sha256:49f1d9c69a8024f741cd3cc76e49a9e8881c85f842e26bd175d787267fa6dcba

Observation 382d5947-2e0c-49bb-bc48-93a6cff44cf3 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297, 2020.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297, 2020

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.278164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.278164Z digest=sha256:7ffd549173b28d80027d2e381f159167858fead4af7a85b66f962c27d27e26f9

Observation 990c5cbe-0781-49b6-a3b7-c4fe2765a371 · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Long Range Arena: A Benchmark for Efficient Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.286629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.286629Z digest=sha256:d4c0a683aee4ae03773b36a1435eee582e99a42bfbcdb7f45d8e23e345045037

Observation 7036773e-43d0-4468-a9c7-c562bbda26f5 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Base layers: Simplifying training of large, sparse models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.070594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.293221Z digest=sha256:16cf954837259c7201e4b05256e7690c8a5aea84333777cc433ab315e3f80f5f

Observation 9b4d9c3d-c9c6-486e-8b0b-6a74e6c8c037 · outbound

This paper cites Sparse is enough in scaling transformers.Advances in Neural Information Processing Systems, 34:9895–9907, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Sparse is enough in scaling transformers.Advances in Neural Information Processing Systems, 34:9895–9907, 2021

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.041059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.300884Z digest=sha256:a3a162fef2b24be7cd013df02c90f76d4ea77fcd369b806c1ff229c945de29dc

Observation a66adfdd-9e1a-4cc7-ae92-830454ea166e · outbound

This paper cites Hash layers for large sparse models.advances in neural information processing systems, 34:17555–17566, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Hash layers for large sparse models.advances in neural information processing systems, 34:17555–17566, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.313419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.313419Z digest=sha256:611bfcfd8edcb880e479f5ba4981b83a382c1b5b0d4ed6db4513ca9ac53aec8c

Observation 62556ec8-19e2-4526-b831-7f93dfe37da5 · outbound

This paper cites Synthesizer: Rethinking self-attention for transformer models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Synthesizer: Rethinking self-attention for transformer models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.994516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.323430Z digest=sha256:0bd249056fb3416e96cb023ec6747a566bfffe2a9a7ea80a909b14434cd35652

Observation 0554f798-9c13-4e51-881c-71471344b083 · outbound

This paper cites Random Feature Attention.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Random Feature Attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.332917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.332917Z digest=sha256:15d44706128ed87ccf27eb40c78fd78204a84c9ba5ca2669f1cf6056c6a0e3c4

Observation da608b2d-cd50-4437-8ab0-3fccb3b16c01 · outbound

This paper cites Skyformer: Remodel self-attention with gaussian kernel and nystr\" om method.Advances in Neural Information Processing Systems, 34:2122–2135, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Skyformer: Remodel self-attention with gaussian kernel and nystr\" om method.Advances in Neural Information Processing Systems, 34:2122–2135, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.340929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.340929Z digest=sha256:b51775c6bdde92b5d4576c0103be4f1cf3d5c9476177935f84e242231a41662e

Observation 58f8d87c-1034-4894-a51a-7c7ba397f406 · outbound

This paper cites Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.354871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.354871Z digest=sha256:69eb9b4fb9505cf98a99bf15d5af30b34e2a0b2c7cbb151f39cd71a2eb48e540

Observation d3dd0fbf-0933-4c95-937e-b3e43f08fe74 · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.365892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.365892Z digest=sha256:563034c6bdbce95fe3227d6880e7623a161b97e6c8f9c9c0ad94b01a3b213669

Observation e0d273da-fc0a-4fcb-811e-48df01efc227 · outbound

This paper cites Flex- moe: Scaling large-scale sparse pre-trained model training via dynamic device placement.Proceedings of the ACM on Management of Data, 1(1):1–19, 2023.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Flex- moe: Scaling large-scale sparse pre-trained model training via dynamic device placement.Proceedings of the ACM on Management of Data, 1(1):1–19, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.947004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.381550Z digest=sha256:ae6430166b7d57f291d123f9d4942cfbf19b796c3912e3ee4b2c2c9156894c9e

Observation 89e83ad8-8caf-4f34-9518-5d356499d39d · outbound

This paper cites Go wider instead of deeper.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Go wider instead of deeper

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.923186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.388315Z digest=sha256:aac4d9e54be15ac24501e725f7bfd49bfe1402b2c854a4ce0a5ffd31909fee15

Observation 14c12484-f515-47ff-be29-bd1b956b6f9f · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.398538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.398538Z digest=sha256:0531f5d8142c2f9f42089cf6d70184f4b2bc6d04f56a1904a032a1b0a9c412b9

Observation 9bdc8378-f738-437a-8556-561e68ffe5d0 · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.406558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.406558Z digest=sha256:7593803bc71ef63b714ebdd19c79d2902dbb2998a3a4d2f887c04959f9588e47

Observation 62fac901-0581-4402-8615-6888bd3ccf09 · outbound

This paper cites MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.413681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.413681Z digest=sha256:53e0bdf5030b0f8232148b29336e340cea2ea7bee4b132174d018d81194f58ea

Observation 22b8ed44-bf7b-4baa-891a-8711216a334a · outbound

This paper cites Dolan and Chris Brockett.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Dolan and Chris Brockett

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.422607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.422607Z digest=sha256:a674579b351ce7f56fa5f13ef191785dca2d114b655e1382147ce845c39a20ee

Observation 21cc405f-e474-4077-9ae3-3fe9aaad2f20 · outbound

This paper cites Efficient Large Scale Language Modeling with Mixtures of Experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.432673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.432673Z digest=sha256:a646c8baa86e948853553d91f5707f813653072c128966ea9a7f0333a9a7535b

Observation d8b27b0d-4fc7-4af9-a25c-c2708bd6bcc1 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.440111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.440111Z digest=sha256:78bb6557a3bd30791511a79a3e5b5de2b9d013d53a9eda6fa7c23f4ddb4c83ea

Observation 6f302862-a9a5-4880-bb6a-98b23e3ba248 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linformer: Self-Attention with Linear Complexity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.462100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.462100Z digest=sha256:076717668c46037d7a713e4a50d3e3ea5b96cbd7964bb7bfc0951ffcbbd61b29

Observation 94e832ac-c52b-4e3c-be74-d4fbef186d24 · outbound

This paper cites Reformer: The Efficient Transformer.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Reformer: The Efficient Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.473597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.473597Z digest=sha256:9bebfa48cbd585f2ff8148c592be2b8a21db3949451a001688f91cc40b922127

Observation 923ac743-556d-46f9-a7ea-c4b6c7c65b90 · outbound

This paper cites Rethinking Attention with Performers.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Rethinking Attention with Performers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.482557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.482557Z digest=sha256:11caa1ac8a211199550467d806450d1834600b2e922ed73d000539c4053bd837

Observation 9c1f67ac-16a8-476c-82fa-13aa06552c53 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.505998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.505998Z digest=sha256:616cdaf9bb525d25372646e6ce511a7c74e5ae00716da16851d546676273ce97

Observation 39b74526-d5c2-4629-bf7c-b2179ccb8030 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.516804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.516804Z digest=sha256:46d8d85fcb788d41f53ebc6b2f13403520ac18dcb20942f9952d32ffe560aa8d

Observation fea4b75b-7025-4fd8-b160-97d4ce5e80af · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.530148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.530148Z digest=sha256:2fc2dd819826063f2466c842874d66f78f598972c7547f16bbd3d1a0ec65ad5a

Observation 67704a54-058a-4ec7-8875-311ae2f4964d · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.Advances in neural information processing systems, 32, 2019.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Superglue: A stickier benchmark for general-purpose language understanding systems.Advances in neural information processing systems, 32, 2019

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.542392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.542392Z digest=sha256:2849e5440452325ae9aac68f15d5c9a9db508bf3cde6f82c03dbe6aed47cb11b

Observation eb2e7070-68de-4329-a7f9-3679667f5e83 · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations, 2021

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.842222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.549030Z digest=sha256:b67e2e32e204289b1529c9d06013e229abe7bda9865f7b393255073f97a73b89

Observation 7767dbbf-f17c-4873-bfe0-d3410182ba68 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts CMMLU: Measuring massive multitask language understanding in Chinese

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.557780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.557780Z digest=sha256:b279bb17ed0c2b63561995bceb67c109212aacd1a1e36afda8e21076a7136cca

Observation 4ccf242e-1ea1-4eeb-9979-d6badaa23c22 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.566309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.566309Z digest=sha256:09b5d2ab46a48cf751d5306b3bf24ede2ed65db37c06a840425482e80778a3a9

Observation 20cff156-4cbb-4f87-9db4-a7e43b3891fa · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.573573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.573573Z digest=sha256:c692e4b538a5558920a2ed93c9b49f525c472573b3d66b11ac307f10e31fd380

Observation ef53fdfa-93b6-4e3a-9271-161471d16a54 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.580103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.580103Z digest=sha256:6d928516ad529a5e9e3f2b5e85af5a1add2e4810faead7d7f363e5b93b63ff98

Observation ebc47581-586c-4aff-8f88-cc0b5bc31742 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Measuring Mathematical Problem Solving With the MATH Dataset

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.587345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.587345Z digest=sha256:73a5686313c5422efaa86fbf8b9db26c4d389dc9bcfa6f9c9337a620bb986904

Observation 81b3adbe-ecf5-4878-91fb-9c0b0dc9c85a · outbound

This paper cites Program Synthesis with Large Language Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Program Synthesis with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.598148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.598148Z digest=sha256:ef269330dc268e49dbbe089f8e55716f69fe8ab53692e43b0868865e53f9bce1

Observation f5338c63-003a-49a5-83da-7e989e0e3965 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Evaluating Large Language Models Trained on Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.605123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.605123Z digest=sha256:a558d9be30c85113940806c7eae7a3b47a08650f14666895e9e31b53cb148488

Observation d36177ae-036b-41b6-8caa-36cf8829469b · outbound

This paper cites Qwen3 technical report, 2025.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Qwen3 technical report, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.613500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.613500Z digest=sha256:6dc6c17bec5668d73e1da400c992bc06d0c78d3fac02b0f2a3ff928878cb77fc

Observation 23e9b094-a44d-403a-a3ec-2ea2b7caa112 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.619572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.619572Z digest=sha256:00a0e3f0071d7f1a9d8488adceba0ce4a0344d4ef4e7bcdf948413525adb385d

Observation 4aec8be8-f622-4277-958d-21d921fe120e · outbound

This paper cites Gemma 3 technical report, 2025.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Gemma 3 technical report, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.798914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.625468Z digest=sha256:fe3862583830c17411fa22dbf39ca012361b177e9bb147470cb48e5264f9b2ed

Observation 5bcd0223-bb1a-419a-9456-10ab6af55da1 · outbound

This paper cites The Llama 3 Herd of Models, 2024.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Llama 3 Herd of Models, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.777934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.631773Z digest=sha256:7aa8f06a55983508f321b88ed4022fc1cc3c08c938f9db3d7ef2eff95f075625

Observation 37a12c06-9543-4e0c-bb9f-ddcf54f92780 · outbound

This paper cites Hewett, Mojan Javaheripi, Piero Kauffmann, James R.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Hewett, Mojan Javaheripi, Piero Kauffmann, James R

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.754191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.640184Z digest=sha256:afa0028214bb0539c8c6bca8855bf5341d608e55d49b3c40131069c6b83e2138

Observation 8f61c776-484d-4836-ab5f-6d27171eb856 · outbound

This paper cites Multi-Scale Dense Networks for Resource Efficient Image Classification.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Multi-Scale Dense Networks for Resource Efficient Image Classification

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.650932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.650932Z digest=sha256:6b49c6b679e6b7229e8264c9ee2cf638b1d3c4da64cefe78238adb5a5ec47bab

Observation 368fc9a0-32e7-48e7-81f4-69853d82ec85 · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Manning, Andrew Ng, and Christopher Potts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.664409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.664409Z digest=sha256:bb742fbfd4f49ee3e1bea2a4e20c1e49a5ba2d71a41cb379b7e7d23ac90cdd98

Observation 6b4450c4-38db-4f31-b009-ecea4e70e44d · outbound

This paper cites an unresolved cited work.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:32:04.720488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.675879Z digest=sha256:5665e9a83638b7c4f46626e43fbb32cceaa52f805e601f24bd1f0c8b40fca5bb

Observation 6e7cf706-c3cc-450b-9e5a-57ffaaadb444 · outbound

This paper cites an unresolved cited work.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:32:04.700027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.681799Z digest=sha256:6d5c95270a4361ae57756ab06e5d3e180268a4d0b1c1b9a231b379bb4a30d22a

Observation a72df6b2-87ca-4b4b-a0f8-eb6ad8848330 · outbound

This paper cites Squad: 100,000+ questions for machine com- prehension of text.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Squad: 100,000+ questions for machine com- prehension of text

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.681840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.687831Z digest=sha256:effd92c4374530eb7ce859dbd2488a6179e04b6c473600342413a3b6bcc69a0b

Observation f607e9a1-5cd4-4e2f-84d4-3f398bb577c9 · outbound

This paper cites The pascal recognising textual entailment challenge.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The pascal recognising textual entailment challenge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.662745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T16:32:03.693730Z digest=sha256:07319f3590a06f7dcad47fdfae73c002e0e8369d9a7492f18c7b23c4ff27cee4

Observation 8f91b630-358f-4061-8801-cb20e554f439 · outbound

This paper cites PaLM 2 Technical Report.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts PaLM 2 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.698662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.698662Z digest=sha256:8ec2e21c98818929bb4535aa34a56a5a365a7e6db3045700f07bb770c5de9a6f

Observation 8a9d747b-9a42-4583-9d1b-02b07b4e6363 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.704458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.704458Z digest=sha256:2ef031d6296791f405af89086c2561fb1ad8c550afdd4f3c48ef40ee153936cd

Observation fc30f7b7-3e5a-4e4a-80ed-637a98deb910 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.714918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.714918Z digest=sha256:616754200c38f109e12bb6434002b8bd03e3804553e936eb20d85fbd1fcd166d

Observation 440c65a9-df59-4eaa-99f9-ad67e76f2045 · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts FNet: Mixing Tokens with Fourier Transforms

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.720600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.720600Z digest=sha256:c7ecd6ea83e616f06ad6e08ab435f65af87fc859ed75dc65bf6e1e5c64a48d6b

Observation 598c47b0-3365-4852-a613-acafa17a55f4 · outbound

This paper cites Densely Connected Convolutional Networks.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Densely Connected Convolutional Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.727373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.727373Z digest=sha256:2903f730469e406f745e65f296004cf3e132b0c05025bda3858625a7e207929e

Pith citing papers

No inbound Pith citation observations are available.