Pith. sign in

Paper Citation Record · LEDGER

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2505.06481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06481 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:45:12.253030Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3cb88a68-5a9c-499f-a41d-47a13801a8cd · outbound

This paper cites GPT-4 Technical Report.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.108866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.108866Z digest=sha256:6abf62df5c90c5c3fb3e228314322555bec902b9d31550adc1bc0f67de424616

Observation 2977fc75-2541-4c08-9382-1b5a6ebbfae7 · outbound

This paper cites Instruction Tuning for Secure Code Generation.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Instruction Tuning for Secure Code Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.140853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.140853Z digest=sha256:17be58cc6932a6b3411d6c350ec904e10151768815862a70ab2a4bfaca990917

Observation 926e7fdb-af0b-4f6d-a7fd-264c558c3acf · outbound

This paper cites Mixture of Experts with Mixture of Precisions for Tuning Quality of Service.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Mixture of Experts with Mixture of Precisions for Tuning Quality of Service

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.154979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.154979Z digest=sha256:f0829b384cf5d8ea6f9b8aa04962ebffc4736e4466ed72508d59902bbc2aca6b

Observation 0ad6bc08-7f39-478d-90c6-43f97f14df6f · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Averaging Weights Leads to Wider Optima and Better Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.159574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.159574Z digest=sha256:3fb317435eedd0173f1bcef6f751c3d2dfbe03c21c38e782a57d391f8584ab9c

Observation 67993dd1-3c23-42bc-81de-66d53d5ebf02 · outbound

This paper cites Dataless Knowledge Fusion by Merging Weights of Language Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Dataless Knowledge Fusion by Merging Weights of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.168712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.168712Z digest=sha256:f5d25c6f7f211474a74d7a9640c9bb361c9d9cc55ae25e26ce451c0e82179583

Observation 180a2d0f-3b78-4c18-a98c-20115a7492bd · outbound

This paper cites Scalable and Efficient MoE Training for Multitask Multilingual Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Scalable and Efficient MoE Training for Multitask Multilingual Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.177984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.177984Z digest=sha256:afa885c7f94c8de3d35ba74a8be46359874c8e2ae0d5f3af25ad9e2ee3343c16

Observation bd693f6d-3a48-4236-8e8e-3dc39cbb3c5f · outbound

This paper cites K., El-Araby, E., and El-Ghazawi, T.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration K., El-Araby, E., and El-Ghazawi, T

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:12.977588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:12.182655Z digest=sha256:582b2c54583a36a92745be5dd4608497c9febe78f80bc9e37072a234366073a8

Observation 1fc3f0ee-fe10-4aa9-a96a-fa192e03b7b9 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.191723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.191723Z digest=sha256:a734c2e245e86afcefc9bec67cab3a5a7a2b2fa89c09106ff1ac937ec489f0ca

Observation e50477da-2adb-4129-8311-01b9b4a8f7cb · outbound

This paper cites A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:12.961418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:12.196486Z digest=sha256:a361bfb6cee7f4c6d183d771665f003ca0549b9e9ba832b6fd3d7336217e3a64

Observation e75e35e0-ff73-4501-ab77-ae34bc89b0ca · outbound

This paper cites Pointer Sentinel Mixture Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Pointer Sentinel Mixture Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.201264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.201264Z digest=sha256:0d39e88f50da313706fe7554e4ae0f342fe51c13ffcbf2690763b43564ac2a7a

Observation f706776d-9051-4c0e-9b59-a2ef268d3098 · outbound

This paper cites Learning More Generalized Experts by Merging Experts in Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Learning More Generalized Experts by Merging Experts in Mixture-of-Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.210312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.210312Z digest=sha256:f3933a50abe63d1c769e4e30be5ddd1c15df0b8859891a5265fda4a95c2dd2ae

Observation 28b23937-7fa0-437e-a296-b90d895a38f0 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.214771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.214771Z digest=sha256:2008bd503bfef3a84149a6d5bf3e1d1add227af10af3622aaa05fec70409df62

Observation 675643da-3237-4297-8020-47c5b86e76c0 · outbound

This paper cites Wortsman, M., Ilharco, G., Gadre, S.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Wortsman, M., Ilharco, G., Gadre, S

Reference 26

Resolution
verified exact
raw_fallback, observed 2026-08-15T22:45:12.562004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:12.228731Z digest=sha256:8fbd13679cb10bb265110905c9ea5ba05a539bd8c45a834351f59db610de89e4

Observation 97984be8-d51f-433c-b35f-b5b268f12e12 · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.233874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.233874Z digest=sha256:37208f760f3c2616dc0d770e8a9edb893c44c0edfea341ae25227ce99dee687a

Observation 90bcabaa-a40b-4ee2-80a9-1e0eaeb73d1f · outbound

This paper cites TIES-Merging: Resolving Interference When Merging Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration TIES-Merging: Resolving Interference When Merging Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.238873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.238873Z digest=sha256:2369203585fec2a0f31746812329b29d1ffad0ef475fc64383f85d299a5b4fb3

Observation 4635aa8b-5752-4e4d-be90-438bdaba0c1d · outbound

This paper cites MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.243625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.243625Z digest=sha256:d5d167db60d09f6092e1781975191b5a051e77c07c73e85ac48e01519f4a2161

Observation 51732976-2fb8-4f82-a106-ec094726bdf9 · outbound

This paper cites SurgeryV2: Bridging the Gap Between Model Merging and Multi-Task Learning with Deep Representation Surgery.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration SurgeryV2: Bridging the Gap Between Model Merging and Multi-Task Learning with Deep Representation Surgery

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.248175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.248175Z digest=sha256:2fba6423ec9b583c925d2e7626d61ec21717df81aa16fcdae67c4d199d06e82f

Observation 4cbcd22f-5233-4ee4-ac32-627f4265ded2 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Taming Sparsely Activated Transformer with Stochastic Experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.253030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.253030Z digest=sha256:e4f0beb7842dad39c7ef56a264e0a0096fbf5ae6656d6cfded7a2575d7565132

Observation d90afda7-6451-4fa6-a7cc-66e96165ae23 · outbound

This paper cites Mixtral of Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Mixtral of Experts

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.164099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.164099Z digest=sha256:0c0eb6557aaaebb8beab05963f6ae8209e83b082e2e7378671fe26571b985f13

Observation ebd18d54-e81c-405d-a247-9128159767ba · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.173343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.173343Z digest=sha256:cb13b80e66c8af5e3fd38f5e6c061d7188a30cdea17852a6da57fa2fac18dcd2

Observation 1e9d518d-105e-415c-bf48-de4c9ac13204 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.135322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.135322Z digest=sha256:566f794bdb4e1294049d0647fba73e31f9a0d4d0fe4fcdb34c7e330f53680ff1

Observation e560976b-51aa-4a73-8b09-283b64fb4c1c · outbound

This paper cites BiLLM: Pushing the Limit of Post-Training Quantization for LLMs.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.150327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.150327Z digest=sha256:f970448da33d0327b9651b3a13d8b598679ae2f654a73a4e36c84eccea768951

Observation 1c64c3ed-d511-460e-a6db-c20dfd1dabe7 · outbound

This paper cites M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.187026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.187026Z digest=sha256:d59283288fa38b61f29d4de05b89b21e649408c819432bd9edb7d49aff995ddd

Observation 58335d6e-ac25-4847-819c-062afed78d62 · outbound

This paper cites SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.205697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.205697Z digest=sha256:e129831ddfc90b88b26dc08affe10f2c315f6679dda163e23b44d23549324ee1

Observation 2f83e62e-312f-41df-b7c3-fb75f1f40703 · outbound

This paper cites Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.219620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.219620Z digest=sha256:cbe3f4e13960bbb2ebeafc5d4a4b0f9776a3a6e4df1df71ae0036ef888839c65

Observation 7000e4ab-0d46-4197-a195-5babe867c87d · outbound

This paper cites Exploring in-memory accelerators and fp- gas for latency-sensitive dnn inference on edge servers.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Exploring in-memory accelerators and fp- gas for latency-sensitive dnn inference on edge servers

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:45:12.946045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T22:45:12.224258Z digest=sha256:5770a4d6e96921369a4708fba6171a7068ea93fec77f455d75070675c9b47f63

Observation a47a333e-07c6-4081-a2f7-b1d752c1d7cd · outbound

This paper cites A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.119688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.119688Z digest=sha256:0482b1caa8309aec85b2c1ebfecec56aa677d517b63795b72b96da92aad8fc8e

Observation fcd730d3-e048-4e14-97a4-e59d216fa2a9 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.125138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.125138Z digest=sha256:58702aad7b9411931b2e5c2f542b203dcebe007ec4f3c3f363853f2ff880b845

Observation b72950c3-4cfe-439e-af6d-2931c5b715e3 · outbound

This paper cites Parameter Competition Balancing for Model Merging.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Parameter Competition Balancing for Model Merging

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.130419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.130419Z digest=sha256:957a7da1e90742271712329de45451b74c4cda9860308a51f663ac78c1eff874

Observation 5ac90b88-ef90-4698-8c4e-5138cb000122 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Measuring Massive Multitask Language Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.145498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.145498Z digest=sha256:9da519d4deef77faa52e6daa2a1c3827d375bad2ce4eeb91f212df659dc26ed8

Observation 3b4c9814-32b5-4088-869a-dafd627ad67b · outbound

This paper cites Task-Specific Expert Pruning for Sparse Mixture-of-Experts.

QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration Task-Specific Expert Pruning for Sparse Mixture-of-Experts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T22:45:12.114554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:45:12.114554Z digest=sha256:54c0b36675bf7000146ceb896f3e77db7a5ee4ae38eacc2972617bc7e3224a07

Pith citing papers

No inbound Pith citation observations are available.