Pith. sign in

Paper Citation Record · LEDGER

Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2402.14800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14800 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:53:54.748614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T19:08:50.012079Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be36b2e7-608d-4942-b512-7d9e33db634f · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.169470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:e11912e3012518ca2f9e6f22ecba9a04f476ff062aed0a89c755c509d61725b8

Observation 3cce7e1d-3d45-4a35-92ce-7874c9a6a402 · inbound

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection cites this paper.

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T17:05:42.981077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T17:04:13.905401Z digest=sha256:42ffe0b4cf273464202033fcfd5e7487df1da40f321a0e23df8c46fef9edb884

Observation 04cc05ad-d4f1-40d2-a50f-9be1250037d4 · inbound

Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning cites this paper.

Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:13:14.061601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-23T17:10:06.053994Z digest=sha256:e36c6b4707408671d06f4021a34127474c7f3198fcb73e9b5c45c5b035ce7c42

Observation be861a79-27c5-4833-b505-0ca987992fb0 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.784417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.784417Z digest=sha256:22a86d3931dc22fdc4802a7f38a36aad82af5b0877521994fa31e30ccd570b21

Observation 18b9bd51-1a3c-4b17-8def-39ac652d051d · inbound

ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing cites this paper.

ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T12:03:42.933715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:03:42.933715Z digest=sha256:bcd44664ba6acbdffa60de9a25bad89dcb9621bfa186b27fd33d92711b103e89

Observation 7c5fdd40-eae3-4dce-9bf1-da10091dabad · inbound

MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training cites this paper.

MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:19.102680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:49:19.102680Z digest=sha256:bb4f7d951953ca25ab2ef443d4364ba3b30016737551871694c1e2631a61baa5

Observation b3dba00c-5bcb-4b78-bc43-ea49ad7fb54d · inbound

DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference cites this paper.

DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:53:54.748614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:53:54.748614Z digest=sha256:2b7f93566fb07024ea659f7cbdd33b270469dd765a1d83f593fa9f1af1a2cc5f

Observation 1e936820-78cb-4947-8138-fb9c67f9fa18 · inbound

Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis cites this paper.

Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:12:31.024128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T04:08:29.089438Z digest=sha256:c2328880c0f3e9fc6afc1e922b1219194dcb76c9f850e788262fb70a4c60d766

Observation 67490233-5793-44a9-b3c2-f78e74f865f3 · inbound

Mixture-of-Experts for Personalized and Semantic-Aware Next Location Prediction cites this paper.

Mixture-of-Experts for Personalized and Semantic-Aware Next Location Prediction Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:24:26.135538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:24:26.135538Z digest=sha256:7021d55edf9158be3bf1fb2919ad19019d3f8df7c904eceadbaf5ae0e9f6dc14

Observation c4a8d045-f5ad-4705-bb1d-3a339f719c1c · inbound

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference cites this paper.

LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:13.500526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:13.500526Z digest=sha256:a4cd8325208e6536f1f02c4ff00acd7d0ce8569bbc5207d8c8a65239a8153d53

Observation 822ff542-9253-4c11-9d77-6273aba3864c · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:05.721834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:05.721834Z digest=sha256:3750a92cf64c146727859762b1d318d8602ba34b4b643170e73b4747ab0777f2

Observation e503e305-eb19-4b77-b620-c66d5a4e5d81 · inbound

Unified Start, Personalized End: Progressive Pruning for Efficient 3D Medical Image Segmentation cites this paper.

Unified Start, Personalized End: Progressive Pruning for Efficient 3D Medical Image Segmentation Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:27:19.418335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:27:19.418335Z digest=sha256:4114ba95dbd54e8c83f7479eb9e3b9167a9a1fe4330d321a298f4dc70253ce44

Observation 0283ff76-3549-4e48-8be1-a5c9d2fb79ea · inbound

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs cites this paper.

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:25.160350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:25.160350Z digest=sha256:88c6afe7963fa0ae48d93ff082bb980e163d2032d0035662e0f5c20ba7f2e0da

Observation d35c89df-12f2-4f4e-85cb-b42da094b9be · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.803615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:cb02adea7288fecd948179504f03e4f9043587bda0402daa550298d82db385f3

Observation 1cd25960-a6da-476c-8de5-adbda4992407 · inbound

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference cites this paper.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.898841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.898841Z digest=sha256:804006c763ec856479b4fc7c2c504326dfaa169408f6eff2f02ed9bdf98db7a7

Observation 601b09f5-644d-4f07-a968-e1c972f721ed · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 199

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.829536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:4605f06dd695ec632200e206063da3f9b44ece421abf576aae5af059eaeb08b3

Observation a8dc26dd-950c-417a-9d1e-cd7f1aa050c4 · inbound

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE cites this paper.

EvoESAP: Non-Uniform Expert Pruning for Sparse MoE Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:35:55.505976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T14:34:48.524592Z digest=sha256:6ae11f208400aa24bf1f069665128c497042c5be67fa66ee44cc09faa49c2e32

Observation b3989a0d-4e18-4db4-8d27-4aa2352b11a4 · inbound

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving cites this paper.

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:18:13.308753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T20:16:16.466375Z digest=sha256:1cb9015343ce211bcaad4624b270a0b59a2a04db981581c976be36d5b6fb59b2

Observation 46f94e26-bb0d-4f72-a72a-8bc6e83b628f · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.482695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T03:29:16.555166Z digest=sha256:607fa951e8b5402303ee865c5183c5b8feb9219ade41f9c77d7d7e4c6d75ec6a

Observation cac8bb70-77c3-4759-b6f5-0b739b619197 · inbound

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts cites this paper.

Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:15.362277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T02:03:02.654035Z digest=sha256:87ec1fe3bf044ed22c42dcde5e79b1c309802fb7a28d36a4523144092fcfa9b5

Observation 94f2ab8f-a75a-49cb-acad-49e2270e6646 · inbound

Temporally Extended Mixture-of-Experts Models cites this paper.

Temporally Extended Mixture-of-Experts Models Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:39:48.314205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:39:39.492135Z digest=sha256:e216b5b6bb69aa59bc3950d7b4fb8b33988bb381e7671475e8b6b63ec8b97676

Observation c9372f65-d6c0-4c07-addc-dcaba4a9384f · inbound

Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning cites this paper.

Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.646812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-08T12:03:58.279499Z digest=sha256:a4c914e6d05bf793c653fa9669db7c7550fcb1fdeadfbc6181be5af312978c9f

Observation 042d1a9f-7455-48ff-93a5-397f2be0765d · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.405151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T06:47:25.155437Z digest=sha256:f41948ef79eb35d895787744538eed8c69e49705e625fa79b28b71a827a84e1d

Observation a72b52a1-6a79-4870-9cc3-a46ce8901e92 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.053900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:d2ecd3469b322074e577cf9a82691c02601857ab2ddd603ac7a5ece333e16f96

Observation 8b1db52f-d37f-4c5f-951d-fda88f01870b · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.485473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T03:34:10.370956Z digest=sha256:4d21bf1f460b2d3813b60ab82d8c5d5fdaacd5f6b170413f87229fd7a6d9f4ca

Observation 8b1d5189-8889-42d4-8edc-44ec47436a21 · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:23:51.219912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T23:22:51.808346Z digest=sha256:e9ef8f5712573730049dde6598d4d0600f3db6e6928c6c350bf6e17dd75096c9

Observation a8f8732b-b5f8-480f-a672-2627725cc172 · inbound

When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing cites this paper.

When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:22:39.608550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T16:21:02.198882Z digest=sha256:0e5bd3a11eb945b066bf516c13485833a6c244095ff43c74ea47b8813abb695f

Observation 8fe14920-7977-42ae-b05a-7b40f3205cab · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.868047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:ea43061f5ac8a7f67d87b32007cd2d2b314e07c9bd32e63f827c10aa847c1686

Observation 4122790f-c394-4632-9378-fc589a2f737f · inbound

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization cites this paper.

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.441458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T15:38:18.616792Z digest=sha256:6b9f351dfbe6e6af4b58ab0d3113608bb4e0384565d9fd240686d603e331d94c

Observation cd5e468b-b1a5-4354-b53c-5e4c6d067459 · inbound

Expert-Aware Refusal Steering cites this paper.

Expert-Aware Refusal Steering Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.122065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T10:04:13.562338Z digest=sha256:04517bd6a3851a596ae62e3bbec32b449b0d20943ce4a31f5bc4d79e2edca9d4

Observation 4959bbd2-0486-4e39-af8f-e7e1a649a509 · inbound

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression cites this paper.

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 13

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T19:08:50.013861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T02:07:44.237002Z digest=sha256:a6e904c56c9a6abde54df94abbabab080f15a05a3cb34426b07c0d45008e9855

Observation 39e3a7b9-a421-459d-a781-95c484fb4e55 · inbound

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference cites this paper.

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:14:57.435789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T04:27:02.915854Z digest=sha256:1767d4b7cfcf2420fc92775da0f7ae51bd37d9a260b6d78d9dd11a21df7347e2

Observation c6e63b0b-f3f9-4024-b036-bbd514680b1d · inbound

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference cites this paper.

Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T08:35:22.347459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T08:35:22.347459Z digest=sha256:309e5b75f16caddf7f2350059c34f3d29cf809fac3daea3f1a54304c13b43422