Pith. sign in

Paper Citation Record · LEDGER

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

As of 16 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 8 inbound Pith citation observations for arXiv:2505.05799.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05799 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:01:20.181224Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:23.206783Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:44.791368Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fae7f16a-4076-4d7c-b3ac-4397f61496f8 · outbound

This paper cites QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.057459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.057459Z digest=sha256:e0094c19b2707febb0e01e68a131b57cd38350fdfb447a00629683a54bebcae3

Observation 442e46a2-6bc2-411c-84c1-41317f8a6c3e · outbound

This paper cites Low-bit quantization of neural networks for efficient inference.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Low-bit quantization of neural networks for efficient inference

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.523687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.062112Z digest=sha256:861a107077e3981567043c942b1d241b04b56031f2bcd26d10a53f2afcaec467

Observation 2bbf663d-b615-4203-9c62-2ac844c7cb79 · outbound

This paper cites an unresolved cited work.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.065538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.065538Z digest=sha256:56caecb3bd5619c73186893011eb54b10c8f1c5b844ed74d9b21c40b62320d81

Observation 13d58736-2609-45fd-a29c-c156249b2ea0 · outbound

This paper cites W., and Keutzer, K.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design W., and Keutzer, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.506413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.069871Z digest=sha256:bbb623e196ec9ca1a6e08e0c94ee4a97377c611563ea803c004c53de499a67e9

Observation 3401228c-f079-463d-ada3-8a6ca46e1288 · outbound

This paper cites SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.073536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.073536Z digest=sha256:01077d27f19a8f71e6d529f1cc5af84d729fde785d58c60c355b07ce2fa653e1

Observation 6c448d4d-c0f8-492a-94d2-40e883942945 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.077587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.077587Z digest=sha256:b6c49900c595211f76d78fd5a055ccf893cdf0e88f4766cd341e6948ab5d0fcc

Observation 3d60be69-3fb9-49a9-b18a-05b46e5b8c58 · outbound

This paper cites MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.081983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.081983Z digest=sha256:e09a7234847f9f13cb50fe78d3160da290c71730da4afa249b39eb5901bcc4d4

Observation 953946d5-49a6-46f0-b501-a78f5152ee61 · outbound

This paper cites Bounds on multiprocessing timing anomalies.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Bounds on multiprocessing timing anomalies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.494395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.085706Z digest=sha256:555996d1502f2ad5be1042b9828bd7ad050deec76a870995cb64b6805331e3b7

Observation e648adde-0231-4344-914a-9972409ac979 · outbound

This paper cites KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.089202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.089202Z digest=sha256:d5b65a8099eb6607c066e5bf2caab614a83003fd01bb88563ea9c489ecc1f690

Observation 029dab77-d234-4476-b5fd-912743227866 · outbound

This paper cites Mixture Compressor for Mixture-of-Experts LLMs Gains More.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Mixture Compressor for Mixture-of-Experts LLMs Gains More

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.093251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.093251Z digest=sha256:57902d78117b8edc701e86f99a22a494fefbebdd7028c781783a7140b7633ad9

Observation 6b93a2e4-05c1-46cd-afb1-76b41fa05113 · outbound

This paper cites Mixtral of Experts.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Mixtral of Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.097145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.097145Z digest=sha256:5b0cf1297beec7c3a56a2c487eb63443ae7b5a0923971b701805d465d491d7a9

Observation 41c61031-866a-4907-bf8c-50104bd6c46d · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SqueezeLLM: Dense-and-Sparse Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.100645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.100645Z digest=sha256:b3ecafcc21616d850554cde26d172aad954820169ce4ccb9e74f5b5b1d27d9be

Observation 7d6270ae-746d-425f-beab-4b82702672fd · outbound

This paper cites Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.108195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.108195Z digest=sha256:20e66f107be3b601dc567f66933d635331fc82015d2418162b5ce1348b93ae17

Observation fcc5b887-e0e2-4a17-912a-fc68b7a8fbed · outbound

This paper cites H., Gonzalez, J.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design H., Gonzalez, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.111677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.111677Z digest=sha256:eb53c1185a692eaa09a6f103b92ae6f20a6e92c1aa55e6f6364582b0fdc70429

Observation 9aa0f757-8182-4331-bd24-960600a3dd8e · outbound

This paper cites QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.115100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.115100Z digest=sha256:30f99206b49d29eb25e37bb5072ffa0e0ef82499f2ec6f7324a86dbc2e46173e

Observation 9b0cd6e3-4513-46c0-9e10-5dd0c9b3bcb1 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.476199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.118660Z digest=sha256:99822630c3d168d6c5ef8f6f03d9219588beceebeaa9403dbab7bdf68f3b9e11

Observation 82e02a0e-00ff-4469-a085-791b4132f00a · outbound

This paper cites QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.121741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.121741Z digest=sha256:0e6515566a68b06f88afcb675438ad3241c020312cdfa932e0426c8974d73c0d

Observation 8045fb33-47be-48ae-be19-9c8c422b18f5 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.125386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.125386Z digest=sha256:f24578172ec8375dfe1067103d9ab75f8b78a017f70e351097ceba676fcbeb0e

Observation c11c59ea-be45-44f3-ba3b-8b0c62cdea3c · outbound

This paper cites DeepSeek-V3 Technical Report.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.128688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.128688Z digest=sha256:b43f0c2592428f92dfebb9c99edc3470c9155d1ac66ec85a0c893bdc048382cb

Observation 9c216f39-8978-4db8-b4d6-5add35b02b26 · outbound

This paper cites KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.131856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.131856Z digest=sha256:21ce1e48d393413a8888075678b425d3666aa9dc6efad51f3c635f94afd5c5df

Observation 5546954f-c393-4fed-baf4-296e1ba9c294 · outbound

This paper cites Pointer sentinel mixture models, 2016.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Pointer sentinel mixture models, 2016

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.135454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.135454Z digest=sha256:6edaf1bdabb1cb6a7fa26a72b06c8b0d9e3d6cea594043c82d312e2049650eb9

Observation d406c120-e30f-4666-a268-cd6cdd663e32 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design OLMoE: Open Mixture-of-Experts Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.138698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.138698Z digest=sha256:42efda57be5ac2808cd0b53bd7ea151f46595bd5dd3dc07a47787abc6e466c8f

Observation bc90ad85-5426-4828-a92f-99e0c69fa85e · outbound

This paper cites Massive Activations in Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Massive Activations in Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.142490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.142490Z digest=sha256:6d1eb5e4410a668b81455a4d944662705abbc96a5e3fe327313d7a5268027ff3

Observation 40010b15-ec52-4f49-b234-e096bfb73003 · outbound

This paper cites HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.145976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.145976Z digest=sha256:9276e20de03c49b2df255611e70e7202dbe20d01f4152fbe229a558b1c6e3266

Observation 94be0a7a-2c94-49e8-bf56-a42d2a7e4936 · outbound

This paper cites CUTLASS , January 2023.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design CUTLASS , January 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.149211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.149211Z digest=sha256:3937481a8df3a439bd8f95a57cdb18ba406443dd105cc851654e1d599cc52e6e

Observation 8077026f-2022-4705-8605-6602c0fb99bd · outbound

This paper cites Introducing DBRX: A New State-of-the-Art Open LLM , 2024.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Introducing DBRX: A New State-of-the-Art Open LLM , 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.453101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.152779Z digest=sha256:e1a05eff2a68a682dafd214b079dbfa8ac8d7555f3baaa88748a7554c4cb3cd9

Observation 766fb9b8-4356-41e6-a632-aadb94f155d4 · outbound

This paper cites Haq: Hardware-aware automated quantization with mixed precision.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Haq: Hardware-aware automated quantization with mixed precision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.442414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.156164Z digest=sha256:25f8f0826410f7cacca10a4e1345e01c5f8551ecbae004102e48d996a1adf87d

Observation 7a091d6f-6b83-4fd1-8f30-a549748a27f1 · outbound

This paper cites Roofline: an insightful visual performance model for multicore architectures.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Roofline: an insightful visual performance model for multicore architectures

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:01:20.431000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T23:01:20.159576Z digest=sha256:591ad8e0681fe13dc3fecd508140d3c81699861c85cba6736e3fd997bfc8c471

Observation d898436a-7226-44cc-ae98-1d0efa43961d · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.163033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.163033Z digest=sha256:6b441ef76cfa4957ee66be0f7e9e61aa01f27cd8b241522bb909eb13d6341b31

Observation eddbd31d-f171-4b83-b690-17228831ae6f · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.166912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.166912Z digest=sha256:1234141dc03cd693b0e4096e56d1c64fe0e4ebc5e49634238fe25dbd23c6abcf

Observation bd7beb66-6f9b-40f3-a631-1496e62c40c7 · outbound

This paper cites Qwen2.5 Technical Report.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.170464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.170464Z digest=sha256:cc9b06872ddb4fdcd2b94a87d52ddb40acba45a1bc5399a7f82868e7bfd71a87

Observation 96252e40-ad26-42ad-b7a9-d50829ebd01a · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.173955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.173955Z digest=sha256:2b12640f4e248d63abfbe43d47e69fd9f1bf8c1866bda920884690817d5f070f

Observation 1d7a95aa-ce8c-4afb-b10d-3c0656c2ece8 · outbound

This paper cites Atom: Low-bit quantization for efficient and accurate llm serving.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Atom: Low-bit quantization for efficient and accurate llm serving

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.177528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.177528Z digest=sha256:905c0d1edcdc9e6f6b7344c17e40d3e342509888a8c8df29825f6f0403617f1d

Observation 2d12f9fc-d808-40e2-a996-746172485f14 · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T23:01:20.181224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:01:20.181224Z digest=sha256:5abdba041c7bf073acda7bfbc126a39cfed8218d9fa246ddc3a98f87e56e8b53

Pith citing papers

Observation aac9bb44-f36d-488a-8da0-8cff6d5dd7b3 · inbound

Get Experience from Practice: LLM Agents with Record & Replay cites this paper.

Get Experience from Practice: LLM Agents with Record & Replay MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:23.206783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:23.206783Z digest=sha256:f9b4c544ace44e739d6258281df3c67f185025fe668e2a37f8fcf7da4e4741b7

Observation 4e76c524-1d36-4d4b-b0e7-1723665d2ccc · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:03.992193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:03.992193Z digest=sha256:13ddf4de319cad2ed122de4685f881c59983332b8ed78d45f27c1bb9422f9254

Observation 3f50dad6-e6cb-4bef-a7a2-38beaa5ceeba · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.798871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:784bfacfea960523612acf3469f61df535634518d2be20060df0f693f905d1d0

Observation a1e71420-08dd-4415-bec4-d6d31c269da7 · inbound

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference cites this paper.

PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T23:41:52.850528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:41:52.850528Z digest=sha256:4166b7a8f6fbfdef83c1501279bb62742c4c2fbcad6612a77b1704ac31f64e14

Observation 1861e9f3-1d25-473e-916f-a8c08c0f001a · inbound

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering cites this paper.

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.205305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T07:09:48.239662Z digest=sha256:ca3258d875b8b1ee9045f9581da8896882f87a6a9ca9339407fbe7245b169994

Observation abbf53c6-7b83-45a6-bc60-0514dfa14d96 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.704030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:d27343fcee34d589db967cbb61d71b9b9826d74569406b10188a3ca2cd662b2c

Observation 4315bd31-77f7-43ef-971f-904869774e49 · inbound

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs cites this paper.

GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.963748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-25T05:33:06.719954Z digest=sha256:21e9d22cf2779425c22101202db2ae9ea71111b2c62ca093e870149ab7eb83da

Observation f994a790-dfd9-4824-a686-9f7f09ec4ad9 · inbound

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization cites this paper.

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.793200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T07:06:32.220604Z digest=sha256:a01afd9b5f22ef169f38d5285a42b36261b4734936dbd7794b39b6ed88a73880