Pith. sign in

Paper Citation Record · LEDGER

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2602.19938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.19938 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:32:47.917300Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c5649613-9af0-490d-aee7-85a81c5d09e0 · outbound

This paper cites Deep Rewiring: Training very sparse deep networks.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Deep Rewiring: Training very sparse deep networks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:44.948013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:44.948013Z digest=sha256:9c1077aa613e2507dbedf0cc1c913ad285623802bfadd19539620ebd68be2664

Observation 26f304c9-6c42-4ac7-8b60-0fc3bfc08edf · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.301159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.301159Z digest=sha256:60a2d95f4b6cfb4f87db1996b8bea1d9d4bc6282e488b44aa3c7f35f636c48dc

Observation 3fba3b89-ab0e-4b0e-a054-683de3c8271a · outbound

This paper cites Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.535224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.535224Z digest=sha256:b7b7efcc3758dbdd7ac6ff24faffecd63e236d45ee01c9791fbfa9eb04ce8136

Observation be9d72b0-5d09-4624-a639-712d70c7d671 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs MathPrompter: Mathematical Reasoning using Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.586588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.586588Z digest=sha256:fb1cab2a7efcbc47dbb2510601d31adcb9a8045d85a5d61e404e18310fe22dcd

Observation a16666b4-85b6-4f95-8b37-244edddc5327 · outbound

This paper cites Scaling Laws for Neural Language Models.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Scaling Laws for Neural Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.815295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.815295Z digest=sha256:49bbc81bcf9200e4ca99081a84b60e54889ba7c3880cb03d528385bd81dd784b

Observation bf7cb693-3216-476d-aae0-c18d55c558fc · outbound

This paper cites SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.891720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.891720Z digest=sha256:3f472d41188c30f210524f03e591bb4c2b6ad04d47d4b434be799edd08f5593e

Observation 17699da9-e3dd-44f2-b44e-2c75b053d9d0 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.963798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.963798Z digest=sha256:a0d2afd8f4a9b4a8e3890c59e6a4b9db2dae35bd18d70cdc056724a5250852a6

Observation ecd4fd35-43a1-4ae2-ad23-d9c4901028dd · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.028389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.028389Z digest=sha256:5a7294513a46da27efd1cdd23b83631146c00af5dabb037c65b6cfaf860331df

Observation 1e87c76d-2f96-4b48-a16d-4167011f5afb · outbound

This paper cites AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.101296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.101296Z digest=sha256:acd1d41a9ae033e944b824e9d8b250a2308838e0ae9a682d937300f3bf5594d1

Observation 9e41dd5c-4aa4-40b2-b200-12d63ee3e27c · outbound

This paper cites LLM-Pruner: On the Structural Pruning of Large Language Models.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs LLM-Pruner: On the Structural Pruning of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.231928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.231928Z digest=sha256:2f836a609197b51bffaa926e4dbd72c024e3c52c2f1c0ac696639c36cf57cf5b

Observation 7a8a08df-5f80-47b7-9ffc-862559cd3f3f · outbound

This paper cites Pruning Convolutional Neural Networks for Resource Efficient Inference.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Pruning Convolutional Neural Networks for Resource Efficient Inference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.323377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.323377Z digest=sha256:327eb1980836a8937407063720b975a61724fbf392b37c5728add361e40eca7d

Observation f3ec516e-e2c7-4ccf-849a-da6909f3db15 · outbound

This paper cites https://aclanthology.org/Q19-1016/.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs https://aclanthology.org/Q19-1016/

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.565024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.565024Z digest=sha256:ca914f2e270b5638e3b057bd926abedce71b253f61602035121818436c8f9451

Observation 0240ea56-1bcd-4876-9813-92b052dcdf30 · outbound

This paper cites Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Systematic Characterization of LLM Quantization: A Performance, Energy, and Quality Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.741030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.741030Z digest=sha256:abce4f45b21293b29ec2eba592681900b453ce2f5286de06784658b3d44d3220

Observation 640d47b3-c827-4530-8609-4c1171b90e33 · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.952975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.952975Z digest=sha256:93e3672040b96205dd56b79a9e1e0764f40ff16e8eb17cddcc2989d16fdc28cc

Observation 5747fa34-0ae2-40c9-a46c-b9e4e25a1ec5 · outbound

This paper cites Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:47.069093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.069093Z digest=sha256:b7e88ab9431f3a3d6994888bbfc7e9d17c332c37522006da4640b9e2caf072fd

Observation dd601836-1c38-4c44-8c73-b5ac115df73c · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-02T21:32:47.190057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.190057Z digest=sha256:88675f686359dcf3d0b09934f10437faaedd7ece9100d5b1734bd199d10cfab5

Observation 1dab9dbf-2ff2-4733-a74f-14cdf17ce36f · outbound

This paper cites Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:47.700690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.700690Z digest=sha256:937107afea2060e22bfe7bf73d0f78f1993680a636ee112c0643aeb83fb70438

Observation e0bd5800-dffa-442c-b380-5b2bef93cdc7 · outbound

This paper cites Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:47.822382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.822382Z digest=sha256:4e4000abba025518e8acc8fb34438b6c701c9ce606873c25f3127430c0093313

Observation 72bf5afb-77d3-4ae0-af99-d779c209ab0a · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:47.917300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.917300Z digest=sha256:d74ad2978f83118b6f619fb58b57be5441c068513f51d6d386889d1a7c1d7c9c

Observation a40d2930-fdbf-445b-9d2c-2df26913bf5a · outbound

This paper cites Mixtral of Experts.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Mixtral of Experts

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.765434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.765434Z digest=sha256:62dff0ee9f4afb045d9128f967745142a113b51e398aa0113ce36e3e8904114a

Observation 6ba0ceff-4272-4239-832b-8fa7681de94f · outbound

This paper cites GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.476307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.476307Z digest=sha256:ccfacb3225714589c53d1ad0f387fbd98d54bdf7b263e7ce312713d4dc9f5fdb

Observation 2c82cef0-921c-4e14-b83f-4b266164cd0b · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:47.576363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.576363Z digest=sha256:288f474f5a20ee1476139b7e3329715320487fdf200c89e6e8d250f6d675154c

Observation 12f5d8b2-ac3f-44c6-8603-a738968efac1 · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.022568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.022568Z digest=sha256:c129b69fe432c4fa4fd703653f143e3d8f9ab6aae5428a89c4f13b00ed39d84f

Observation 30dd69dd-16ed-488e-9087-118f33bdf256 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.369082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.369082Z digest=sha256:be30fbe1d4fc649ff122b435050ca1d4ee99445aac4f9908e0922d227e81b3c1

Observation a13b915e-d7a6-4cbf-8d5e-9333396d095c · outbound

This paper cites Load balancing mixture of experts with similarity preserving routers, 2025.https://arxiv.org/abs/2506.14038.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Load balancing mixture of experts with similarity preserving routers, 2025.https://arxiv.org/abs/2506.14038

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.495634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.495634Z digest=sha256:66244fb014df2f09461437baf7bd360fe90f4541b28166ec94f51b2b29a7d363

Observation 38e2c223-ba81-467f-ae84-d88585e7db20 · outbound

This paper cites Sergey Zagoruyko and Nikos Komodakis.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Sergey Zagoruyko and Nikos Komodakis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:47.342765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:47.342765Z digest=sha256:4b1219a8a380f7bb94e29b628ea009c0dea76bde0fc0f6e21ce76b658b665b9e

Observation 48e7364a-526e-43d0-b9ca-bdd52b718f6d · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:46.661601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:46.661601Z digest=sha256:46dac481eeda7f1dcb3a287994143197fb6a66a17e1698f40b98f61a1f94dde3

Observation 91ba37c3-0a0d-4940-9a6c-5cc4c030616b · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs PaLM: Scaling Language Modeling with Pathways

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.108043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.108043Z digest=sha256:ac20ed978b8cf61732a320d512c9506948cce449cf96716558886d43fe74492b

Observation 60a34feb-a324-4079-a8c2-30d24acc0b46 · outbound

This paper cites Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.681616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.681616Z digest=sha256:436c411656d9be66c9f5d4a82387c0459a2535d77cc85a4f5687de91f7851e1c

Observation aa8540e2-59da-45f2-89cc-8733b036db1e · outbound

This paper cites LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.167870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.167870Z digest=sha256:c9f8e1bb0fea5a7ff05f14222e1899ffc10e650f8dc7bcc5325e4153d65229f6

Observation f5098165-2ac2-4d9e-aede-d79affeade7e · outbound

This paper cites Zachary Doucet, Rishi Sharma, Martijn de Vos, Rafael Pires, Anne-Marie Kermarrec, and Oana Balmau.

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs Zachary Doucet, Rishi Sharma, Martijn de Vos, Rafael Pires, Anne-Marie Kermarrec, and Oana Balmau

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T21:32:45.229607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:32:45.229607Z digest=sha256:89d89054b67d31af558e2585bd47db1c238d8df3013e913e15500c438aaa2e4f

Pith citing papers

No inbound Pith citation observations are available.