Pith. sign in

Paper Citation Record · LEDGER

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2511.04805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.04805 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:41:52.952675Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:34.294269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:43:25.643835Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01d8d632-0f50-416c-a227-1bb7d1a31f30 · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.834106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.834106Z digest=sha256:44c85914288d308fdcb3d5ff1adf279131065d214cc69f830196597b923eddde

Observation 08b778f4-0e17-4462-a0bc-d74535b28087 · outbound

This paper cites Retraining-free merging of sparse moe via hierarchical clustering, 2025.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Retraining-free merging of sparse moe via hierarchical clustering, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.837795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.837795Z digest=sha256:15e361dd0986c01163f4f83ea6fca6f895da086aa7010c0f3f5015f157fc9d74

Observation 991a0df8-9dd3-4937-bcc5-32c20e3974ef · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.840823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.840823Z digest=sha256:f945bf6d932d25009530bbda71ae0f136057988be7298c34575c5bccd6709a02

Observation 045acc1c-6900-406b-a9f8-8b148b8ae78c · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.843975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.843975Z digest=sha256:2273a3c6dad254bc663854938c3f456cb3b8c82365efdc1db50924c0d266f3f4

Observation 90d95ac3-731f-4fa0-908f-8216daf7813a · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.847092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.847092Z digest=sha256:ab30b34b5106e88c1a889e66382bbf1f9ec9b65dfd59cb18660b1a7f8175d9ab

Observation a1e71420-08dd-4415-bec4-d6d31c269da7 · outbound

This paper cites MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.850528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.850528Z digest=sha256:bf20e0258e7586513da24dfdb890735562a00af914a20140bf26aca98498452c

Observation d306b013-b76d-4eef-89b4-564c7cad30cc · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.853842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.853842Z digest=sha256:0f568ef00dde4d30d9edd52223538a110f21a3b042738bb0e241bdba75e155e5

Observation a02cae48-d646-42a7-943f-7854944a495e · outbound

This paper cites Delta Decompression for MoE-based LLMs Compression.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Delta Decompression for MoE-based LLMs Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.856931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.856931Z digest=sha256:e58fac1558420381e6faa6babaa16ffae70c7843a8e8ce6b306da0323e935b4c

Observation 00087d0a-f498-4b2b-8434-ae3b5b9bc33d · outbound

This paper cites Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.860011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.860011Z digest=sha256:e8285e194959e191521460f4c790ec0ede02d7c11dae8a932c3e2e6da9797204

Observation 7f8a18a1-2ca8-4822-a646-55e5456bf810 · outbound

This paper cites Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.862989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.862989Z digest=sha256:d3daab506dd844c014210fb419bf52e976d7e97f6af3ab1f6ab29e938e1e90f9

Observation 6576e63a-5b29-4b7a-b325-5a756d272590 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.865923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.865923Z digest=sha256:05ee279914e3da7e904186859818e110f29a14bee5a2080c453069703d576e8d

Observation 8a063dfd-ee8a-45a6-869c-13db3188df59 · outbound

This paper cites MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.868779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.868779Z digest=sha256:cf59e66da1a0a94554f07cbcf5e1b94f2ec99ded6fc5a2a409bb891ba1577114

Observation ee6e218f-ac0d-4a91-ad32-5c40fc7a4fff · outbound

This paper cites MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.871652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.871652Z digest=sha256:293b83b6c4840115710528719f4476f718464b7a1e02bac7452bb2edfa1648c6

Observation 9ca93892-cceb-4536-a2cb-23743e9981af · outbound

This paper cites Mixtral of Experts.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Mixtral of Experts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.874759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.874759Z digest=sha256:59fd0d5fa62b64b81e19e581b6412e27a141746104a8b7ad45cd066bb0719ac0

Observation 5345532e-07ac-4710-acc3-c0145d49c1b6 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.877669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.877669Z digest=sha256:bf8e84c1fef1a6635e8e01e329b365ef140c908862208f2be23c9bc6f128e1b4

Observation 1f19fcee-eeb6-46a2-9a2f-59a5a0f17840 · outbound

This paper cites Compressed sparse tiles for memory-efficient unstructured and semi-structured sparsity.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Compressed sparse tiles for memory-efficient unstructured and semi-structured sparsity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.880459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.880459Z digest=sha256:0c6058b279afb3d65138a525e73336399f1f99e0a5e5ae8ab363dff3b0e83069

Observation d8d09d90-56db-404f-952d-062cf6df0437 · outbound

This paper cites an unresolved cited work.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.883263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.883263Z digest=sha256:c3a9541ffb559dc77bb830aa6e6ed7d305bd998098466b92da0ebf9f5fa1e427

Observation e5a99e4a-4446-4a66-978b-cd849d8a7891 · outbound

This paper cites STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.885793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.885793Z digest=sha256:393a93ffaa53257ad4441ca36c54e95ba9da72267ccdb910e96b1992254d202f

Observation aa53dd31-ba35-43bd-a312-d9718e0365cf · outbound

This paper cites Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.889236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.889236Z digest=sha256:23ffba494a422a8938a4c8438e493bc5f1851c5ad29b8c26c798138667565e35

Observation d48c4271-7f56-4a8e-b8f5-47e105aec39a · outbound

This paper cites Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.892552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.892552Z digest=sha256:1e9fd1a9be7b546da2c54056427ebee8506222d304803ef4bd6a1cd16d45ad51

Observation b5df031a-ec70-4ec2-b6c7-fb95b160fc42 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.895649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.895649Z digest=sha256:fb1c16f1f17d9d6d28f3c774016ec371efd6c11adbe59e5ddbefd69e70cec368

Observation 1cd25960-a6da-476c-8de5-adbda4992407 · outbound

This paper cites Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.898841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.898841Z digest=sha256:8a30f4cd3694d7233bda3400eb4278fd19670005b5c5761d7ab6692268380559

Observation da4f4a16-d4d7-4b78-845a-cd58e8b87437 · outbound

This paper cites Pointer sentinel mixture models, 2016.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Pointer sentinel mixture models, 2016

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.902228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.902228Z digest=sha256:06360b5de2a470138a798cf6e41c128058f2f91ff7a732638844f7143f4c3e0b

Observation 025e935a-b34d-42c3-8a2b-1a75f9a5c104 · outbound

This paper cites Ronald Miller.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Ronald Miller

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.905552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.905552Z digest=sha256:441afeaca7d649cc12dffa6761fab9b603a15dffe257a4b7c697082a3cff394f

Observation 3fe2b623-33f9-4e78-96a5-feb718932f01 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.908779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.908779Z digest=sha256:63a449d24c3b8503bda9e627d6169ba099448e3e65087c22d5724037fba36ec6

Observation acdb41fb-8c1e-448f-beb2-a547391a88cb · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.911960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.911960Z digest=sha256:4eab4b84aecc06d40b35107e8e9cfeea1c95d028becd465baee13ea4ccc27d8f

Observation 8a676dbe-75d5-4327-b584-e9c86751f416 · outbound

This paper cites Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.915255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.915255Z digest=sha256:00e22f0d60d82c7f86f848cd7d49d8de2b48c1d65f9f1bde3d5fe2beceb42444

Observation 8b98c728-1004-4379-a87a-1534c1aad3cb · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference A Simple and Effective Pruning Approach for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.918257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.918257Z digest=sha256:9c0e8030225d2873a481d55de964c7da87a7a1ebf90321454a13f17c2c7a2c6c

Observation 0b123788-4dea-4741-9699-a2a1fcd7b1fb · outbound

This paper cites Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.921149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.921149Z digest=sha256:0fac79b16c7afad00b506a87914ff394134a1ff77e5d79d7439c0ca226744f22

Observation 011be3d4-145b-4233-99e7-dfedbfc066f5 · outbound

This paper cites Qwen3 Technical Report.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.924296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.924296Z digest=sha256:b8c8dfb80cb5bc92cd128c7ef9356d6de6e9b6908d1ddec506044fa418500404

Observation c986254d-4e84-45a7-81e5-e58639686ea9 · outbound

This paper cites MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.927468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.927468Z digest=sha256:8801d95d4de4209c30045f304c11ef49b24e731e40f285a7d034dec8701fda51

Observation 926af6f5-b454-492f-bdc7-4c4613c6064b · outbound

This paper cites Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.930376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.930376Z digest=sha256:e0e5ef7c97106e420e8ce83713a7f4105783db0996c5a90565acfff1f55a853e

Observation 021d6ffe-eb1a-4971-a46c-8aa30a6331d9 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.933402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.933402Z digest=sha256:3f2e7ae782faa93542a98f7e08c5f47b0b3fcfcbbbf35d501a9773ec4521cea9

Observation c332b9a1-ca43-443a-bc6a-f37421567945 · outbound

This paper cites 70 URL https://arxiv.org/abs/2504.11651.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference 70 URL https://arxiv.org/abs/2504.11651

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.936512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.936512Z digest=sha256:93aa68b2705aba0a1bd01348e4f8848a5e546de82ba75be61150e1778706ffe7

Observation 35a97476-723d-43b8-bf54-07c85b7fceb1 · outbound

This paper cites Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.939720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.939720Z digest=sha256:040a1a63ed71e830d4a8457b41e51c8c78e9bc4a23c348c37c13c0129a2f4978

Observation 3de22f6b-4fb0-432a-9842-f726d01d94dd · outbound

This paper cites write newline.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.942847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.942847Z digest=sha256:06eeb5d6c94270722645b5a952ccb04af14ee54f0e28946b255dc4a4e37c97ba

Observation d3cbb099-45b1-45f8-a50e-aff714f8f1d8 · outbound

This paper cites @esa (Ref.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference @esa (Ref

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.946582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.946582Z digest=sha256:b2e55d97497356b94d3899452233a353953a5728c3df72bd546a16f45770173d

Observation 0ece3df0-42ae-408b-bd84-f23684b18c69 · outbound

This paper cites an unresolved cited work.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.949651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.949651Z digest=sha256:0cb4d265c9e58d0613f1531d069956c1493c81a5ffd01ad212312ea91d1927e6

Observation 4c352232-09da-40ee-a84f-2162e5063b0b · outbound

This paper cites However, their widespread deployment remains limited due to the high memory overhead associated with storing all expert parameters, particularly as the number of experts increases.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference However, their widespread deployment remains limited due to the high memory overhead associated with storing all expert parameters, particularly as the number of experts increases

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.952675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.952675Z digest=sha256:6b0779544e56758f5d854a9e7e018e65feafa5ab19fea42f602b33f59ae6ae9d

Pith citing papers

Observation 35e6eba3-7f8f-4c47-9480-71d5faa9e20b · inbound

ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits cites this paper.

ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:32:55.640537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:32:55.640537Z digest=sha256:bc6efb4cd28b323b01a6b8412202222126cd7f790e87101afecd85946b44ad81

Observation 046135be-7fb7-4b0c-bec4-6263e30c6054 · inbound

Pruning and Distilling Mixture-of-Experts into Dense Language Models cites this paper.

Pruning and Distilling Mixture-of-Experts into Dense Language Models PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-07T04:17:56.460609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T12:39:25.535897Z digest=sha256:2482d21442d4c25cb05b5b8b47ebf7d99927fe7ca06c0a03868742a61c8e0807

Observation 98ebed38-9329-46d0-b19d-3d444bf03a0f · inbound

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models cites this paper.

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:34.294269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:51:34.294269Z digest=sha256:fae5be4f520de549ce85748958c10a0e84bbf84ab9fe16c67b3b20acf43b1e19