Pith. sign in

Paper Citation Record · LEDGER

Fast Inference of Mixture-of-Experts Language Models with Offloading

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.17238.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17238 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:58:49.567474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:39:56.479057Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b7b9dc76-8433-4408-b2b3-edbc5d6a2998 · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T04:13:53.768565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:3713cb17c11a718d837853856761347334e58fdea0d14f1e8ae15216bf8690b3

Observation 57cbe660-d453-4c4b-b85b-fd89038d3988 · inbound

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection cites this paper.

Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T17:05:42.962447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T17:04:13.905401Z digest=sha256:396220c9332d690d91b88095ca9ebc3be91830f68ae800535c89ce92f047249c

Observation 55ab2dfb-5b33-4648-897b-1a4f4fa7ad5f · inbound

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention cites this paper.

On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:49.567474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:58:49.567474Z digest=sha256:9ce1b9a448d059742a5c4f02cde53cf729c0bf8ac41c665d3d6967efb92ccc2b

Observation 6a3c1a15-ecce-4ede-abb9-dc516e42cffe · inbound

Cache Management for Mixture-of-Experts LLMs -- extended version cites this paper.

Cache Management for Mixture-of-Experts LLMs -- extended version Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:48:30.265238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:48:30.265238Z digest=sha256:8253553a965e70f9dcd0cefe326c781bebb3a1732fdfe4eab9b2e83ed41ab558

Observation 50f6c9cc-1a75-4bd0-85ca-829cd62620d5 · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:04.085632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:04.085632Z digest=sha256:2042352bd7b6846a7c3331dc4aa0787c042bade9879964da266176ee17933991

Observation 875b6dce-b865-416e-833d-d33894ade016 · inbound

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism cites this paper.

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:55:18.824902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:55:18.824902Z digest=sha256:6a6e90f29a43e20ac606450bb1ad91b842bda27092ad77b1a11c5cc720ecee25

Observation 97cdaeee-e4d1-4928-81bc-99f65c599ee5 · inbound

Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference cites this paper.

Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T09:56:13.364339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T09:52:48.027958Z digest=sha256:5a060f7896fd2869020eb414c80d73cf7da20fa89deb609889ded5ec84853fff

Observation 55c79c1e-5af5-4994-ab70-2257e512a681 · inbound

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems cites this paper.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.908768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.908768Z digest=sha256:aa9f09eda55a19529418e1037e556d86a55ca3584bc860cae98c6885469bb0b2

Observation f5ea1806-5533-4d32-b8f8-4fa5b08bc178 · inbound

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling cites this paper.

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.258625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T07:43:24.937953Z digest=sha256:88267eae177493fbc6f9b45b7a4bdee247d493b11bf29b21b042955d932e33da

Observation 279b913d-45ce-4afd-96e1-dee673e1c68f · inbound

Temporally Extended Mixture-of-Experts Models cites this paper.

Temporally Extended Mixture-of-Experts Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:39:48.246838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:39:39.492135Z digest=sha256:f3aed25cf68408eaad88abb2946086b8b25825e502cf64d7c21c3750897f5374

Observation f0d89220-01c7-4c6a-8d56-df888f7fefab · inbound

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns cites this paper.

Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:10.827410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:29:17.149710Z digest=sha256:136341548110b2c89dc0644ccfed79a6754e9aa137ef234d9c09eb73f9ac4e78

Observation d8d31969-13bf-4095-be85-c883a2df96c4 · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.453365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:1fa6154e54a1385e5ee21f4c628b30e466a088b77d6099c912e2f484f6b82992

Observation 27eaf72f-a32c-4b38-86b6-2340c878028f · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.710574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:7c60d67b998fc23483ab158c4441facb7e2a3716f370215fde1d1c595715aede

Observation a20efcfd-4283-4bd5-b036-b37f7283b9fc · inbound

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models cites this paper.

PALS: Power-Aware LLM Serving for Mixture-of-Experts Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:09:33.685626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T04:06:32.227076Z digest=sha256:c7060289fda07beb7b56fbb2b924a84914d893b6fb8e3090b07aef3f9b63e782

Observation dd0db0fd-e662-4314-b0aa-de69969497cd · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.570551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:85cf5fd5e2fd26b79152847c4c0d3fc1abf7a07eed18e508917721ffb51e62b4

Observation 77fe61f6-e4f1-485d-a0f7-e05a193591eb · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:539196684a1be923cfd7b25727217289d7a84e04e014984ee9fdacfaed69a348

Observation 0ad6a6fe-0125-4b4e-8410-e7ef99d3df00 · inbound

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware cites this paper.

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:39.572738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T12:40:53.103233Z digest=sha256:f0666c014f053ee7734e834974c56c03be76f26ca701e667c8aadbd4d57894d5

Observation b47a5de4-1eeb-449d-bbfc-1e1b32b55484 · inbound

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models cites this paper.

GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T15:39:56.480301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:32:40.435742Z digest=sha256:7520cddfcb282b7102568682db76515546a04bc1c040bddfc8e76807dc758bf5

Observation 4b72e20b-0132-473b-bccb-7cc1e8163784 · inbound

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference cites this paper.

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:14:57.447292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T04:27:02.915854Z digest=sha256:a0a824f28687811484c2e72a28e8190f2f400e7b275495470ec283a7d146d52b

Observation 778f6169-e786-4b4a-afd2-887bcfc3461e · inbound

Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM Inference cites this paper.

Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM Inference Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:01:16.896518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:01:16.896518Z digest=sha256:b6475f81cb7bcc14f38d9c679a1d3f2b737589e5c5219c11b861facefefa4597