Pith. sign in

Paper Citation Record · LEDGER

Fast Inference of Mixture-of-Experts Language Models with Offloading

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2312.17238.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17238 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:48:38.249590Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:39:56.479057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b7b9dc76-8433-4408-b2b3-edbc5d6a2998 · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T04:13:53.768565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:547282b2d135565d88599d0d246b013861059ab1e4c1cba10eb9a0672fc44359

Observation 57cbe660-d453-4c4b-b85b-fd89038d3988 · inbound

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection cites this paper.

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T17:05:42.962447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-23T17:04:13.905401Z digest=sha256:c1948f6e8622f9d571321234ba4f5405e88364121f78becddd221af33a608aba

Observation 05763e97-7237-497d-8947-ffbd0783a2d1 · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.244176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.244176Z digest=sha256:cf945a72d0ec39744095f98bcb985b94faefdde78e83d869e460c29d54f02abc

Observation d6be3993-e01f-4bed-890f-9fe9708749d7 · inbound

A Survey on Inference Optimization Techniques for Mixture of Experts Models cites this paper.

A Survey on Inference Optimization Techniques for Mixture of Experts Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:35.308441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:35.308441Z digest=sha256:7f6cefd4d98ee2dbe1aadb4f17fab68369a954b93e51dd47dcf63e9795bc06ed

Observation 0aa618c3-99d0-451c-93a7-67e6393df335 · inbound

DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference cites this paper.

DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:53:54.783360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:53:54.783360Z digest=sha256:e5e14199b5863f36d2b8025b51e44c62719e77ebab4d12cf5ba5378eb2184958

Observation f5067132-b344-48c8-9b72-3580ed10c4d3 · inbound

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline cites this paper.

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:55:54.011111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:55:54.011111Z digest=sha256:37aca688431c8a67804bd445f1f633162de31313425bfeb5e90e244ac192fa8e

Observation 55ab2dfb-5b33-4648-897b-1a4f4fa7ad5f · inbound

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention cites this paper.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:49.567474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:49.567474Z digest=sha256:496ffc196e36af079681bfc9cefe8cf12fd304fac708e30cecbe385af14acef3

Observation 6a3c1a15-ecce-4ede-abb9-dc516e42cffe · inbound

Cache Management for Mixture-of-Experts LLMs -- extended version cites this paper.

Cache Management for Mixture-of-Experts LLMs -- extended version Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:48:30.265238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:48:30.265238Z digest=sha256:9795f72f1ef0b67cecfc5b691600c944d5caee97156854b76035b550e5f0d780

Observation 50f6c9cc-1a75-4bd0-85ca-829cd62620d5 · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:04.085632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:04.085632Z digest=sha256:f3a16a012ea255a60edd0f2380649346e4e349fb2a7c4edaa8ff977795abce78

Observation 875b6dce-b865-416e-833d-d33894ade016 · inbound

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism cites this paper.

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:55:18.824902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:55:18.824902Z digest=sha256:ea076d81df35bce9a75c71bf359b5a6564700558c4bd83a85416b417e766dbc9

Observation 97cdaeee-e4d1-4928-81bc-99f65c599ee5 · inbound

Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference cites this paper.

Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T09:56:13.364339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T09:52:48.027958Z digest=sha256:57de0028681826c85f8f9aa8c593bbbbacd4f0cc18e6ead7ea1d24e329554795

Observation 55c79c1e-5af5-4994-ab70-2257e512a681 · inbound

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems cites this paper.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.908768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.908768Z digest=sha256:682cd33a0eb23d4d258ad509fb2e03e9bc2afef56859d41b5541fa6d194571dc

Observation f5ea1806-5533-4d32-b8f8-4fa5b08bc178 · inbound

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling cites this paper.

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.258625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T07:43:24.937953Z digest=sha256:2d139698dba7922cf671de896296f5d560dd85d506ebbca4bb8ca9869912a0ce

Observation 279b913d-45ce-4afd-96e1-dee673e1c68f · inbound

Temporally Extended Mixture-of-Experts Models cites this paper.

Temporally Extended Mixture-of-Experts Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:39:48.246838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T00:39:39.492135Z digest=sha256:177471b73d83d3464754e06681e536d2228a9e90248006f21b9d19a5b2639d0e

Observation f0d89220-01c7-4c6a-8d56-df888f7fefab · inbound

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns cites this paper.

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.827410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T08:29:17.149710Z digest=sha256:d6e0078d4418274eafae0d00d4870459c4369e2253f6f7cc198db8ca2d01e7ff

Observation d8d31969-13bf-4095-be85-c883a2df96c4 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.453365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:0c819e35ef966a7c9fa81cd1c63692815d63676f18164c59ef49dabffb2e0642

Observation 27eaf72f-a32c-4b38-86b6-2340c878028f · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.710574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:3d6a5a70f6c323c9969fe7c686a9dae50ce7002ea12eb16979af25e0331720b1

Observation a20efcfd-4283-4bd5-b036-b37f7283b9fc · inbound

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models cites this paper.

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:09:33.685626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T04:06:32.227076Z digest=sha256:b6ae8e59bd6e0980f28ca191144a37d5381f21da298588af9ef6220b2b428d8c

Observation dd0db0fd-e662-4314-b0aa-de69969497cd · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.570551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:1a34bf337d7f2544ff1d63ec0c9759cbb1c3c70e5d99bfd6f46438839133e11c

Observation 77fe61f6-e4f1-485d-a0f7-e05a193591eb · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:48f68c9c7925924fa3252c1253d56d3233de0b7a074eae56917c070e449cf3e3

Observation 0ad6a6fe-0125-4b4e-8410-e7ef99d3df00 · inbound

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware cites this paper.

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:39.572738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T12:40:53.103233Z digest=sha256:c75b944c1831173673db8854c0ab998f8bd931466269f1b08752bf07d21ad86f

Observation b47a5de4-1eeb-449d-bbfc-1e1b32b55484 · inbound

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models cites this paper.

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.480301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T01:32:40.435742Z digest=sha256:ca2247559605a38472f5a58742a30bd69b942fd7da2aeaf43896389cf4bb513a

Observation 4b72e20b-0132-473b-bccb-7cc1e8163784 · inbound

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference cites this paper.

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:14:57.447292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T04:27:02.915854Z digest=sha256:eda18b4fd0d08c13c6fefbee05dd7c6d23c1ff7158ccff2a5ebccfd51b4ee589

Observation 778f6169-e786-4b4a-afd2-887bcfc3461e · inbound

Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM Inference cites this paper.

Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:01:16.896518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:01:16.896518Z digest=sha256:abc644748325d981fe75bbde21f2b778b1e9e7b7ea73229a1a97288b9871ae7b

Observation 49f8e92b-3cea-42da-90fc-e13ea8c702a2 · inbound

Shape Mutating Expert Compression:LorExperts and BTExperts cites this paper.

Shape Mutating Expert Compression:LorExperts and BTExperts Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:18:21.193154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:18:21.193154Z digest=sha256:f9cef2b33a5dac6786d6e65fc44114ae06d4ed9e08a98807ff299e4e5c9792b2

Observation 7bd38b41-23a1-4121-bcc9-af8ab39fe66f · inbound

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes cites this paper.

When Does Trace-Driven Evaluation Mislead MoE Expert Caching? Replay Semantics, Workload Contamination, and Operating Regimes Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:48:38.249590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:48:38.249590Z digest=sha256:e9b41f0119ac095e2845b929d5ba80cb695ef1741cfd9a52a0e85845c7f69b77