Pith. sign in

Paper Citation Record · LEDGER

MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2401.14361.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.14361 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:55:54.162523Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.550340Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6e448380-db6b-4462-b9ba-53343266b24c · inbound

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline cites this paper.

Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T17:55:54.162523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:55:54.162523Z digest=sha256:cbb52b2414359137ac7d50294fd7ec7ba41831ef9caf4e9774166027cb2db890

Observation 07f98ea7-7877-4bc3-bf7e-72b0de36a718 · inbound

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement cites this paper.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:52:51.896416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:401f5238eeb7f1d0c712562c20cf0a1bc23c4fdd73415836881f7229d4c6df08

Observation 94c46255-a666-451d-b290-ce5b4aa13816 · inbound

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices cites this paper.

MoE-Beyond: Learning-Based Expert Activation Prediction on Edge Devices MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:03:42.658889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:03:42.658889Z digest=sha256:4de4b800d48d75ef6d939c39493195602855e4edc026291aa72ae604032e6d2b

Observation ae163c1c-a6a1-43c3-b44a-7a64dd5ea6a2 · inbound

DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance cites this paper.

DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:36:44.231748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T18:33:55.795906Z digest=sha256:ea273505e3855cdbe7ce5596a0add481158964f1f0fbef35d7026c5362d2980c

Observation 7a2237d3-2723-437f-bcdc-a3fb5ea9870c · inbound

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? cites this paper.

MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model? MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T21:51:06.401918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:51:06.401918Z digest=sha256:b13ad84ef125a345e9a0f8eda189ff2668cd43303eacea5ace178b678289f797

Observation 16a1edf2-a8f0-4291-aec4-96a3df5d55e6 · inbound

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism cites this paper.

Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:55:19.584295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:55:19.584295Z digest=sha256:5249d95ae2c97aa64e51836b84dbfa1114dc38debb3fe790c24ad8c803695d73

Observation 3696400c-f729-49d2-83b5-1c5d6cbeb036 · inbound

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers cites this paper.

LayerScope: Predictive Cross-Layer Scheduling for Efficient Multi-Batch MoE Inference on Legacy Servers MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:41:22.790153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T12:38:31.783807Z digest=sha256:9585614bc55df99802140612c91217205de78dbc95bf6a9f5a54246345ee2609

Observation 050f97cd-542a-4526-96d4-9f81eac3ebd9 · inbound

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling cites this paper.

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.207159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T07:43:24.937953Z digest=sha256:dffabb554c51949ba0b01ce1bb7864cdbbe143e6ecb918aeb1530045fc224b1b

Observation da4e28d0-5ecc-43b7-a955-3f4bbfc5c878 · inbound

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving cites this paper.

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:18:13.312025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:16:16.466375Z digest=sha256:c7ccba4481b1b900decaf00d2a493210cf973d76aa47ebabb64024737968d426

Observation 0f51cc86-66b4-4384-adfd-2d1f08f02b43 · inbound

ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving cites this paper.

ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.627584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:28:26.169161Z digest=sha256:f572836398e1566fffcc6df087d5d3a2d482ad02cb68f0112d4c2c7255cd2b39

Observation 4c43c914-c0c0-44b0-afd7-2ed143816c41 · inbound

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy cites this paper.

Layer-wise MoE Routing Locality under Shared-Prefix Code Generation: Token-Identity Decomposition and Compile-Equivalent Fork Redundancy MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.447419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T06:46:52.811371Z digest=sha256:4f459a4178a358d0a20a1177fb4e0a2f60e3b079c5b2a45a2928dcf7828754fd

Observation 5fbd1d3d-e3b1-4435-8059-2abfea64ee3b · inbound

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs cites this paper.

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.174701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:31:25.826205Z digest=sha256:e8a1096ad4bca12b818fd658db82c802be71e486e35b89cf6a8724ba856d5940

Observation 794d1f12-71e5-4b83-9186-1f73e578f0ab · inbound

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading cites this paper.

VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.468196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T16:10:22.588945Z digest=sha256:7e75354e535833d28aed8335d0c2a35ae3b78d98b665639d0d46a0b42f5cdec9

Observation 7a8d0425-079a-4154-afa6-63511dc03319 · inbound

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution cites this paper.

CoX-MoE: Coalesced Expert Execution for High-Throughput MoE Inference with AMX-Enabled CPU-GPU Co-Execution MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.396632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:07:53.045082Z digest=sha256:aca3a138d5d805adebde3de0ac3f90f1f005440e87548977ecaf04718ce1c9c7

Observation 81743ddd-ab30-4031-9717-3c676997fcbc · inbound

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG cites this paper.

C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T02:12:58.326430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T02:10:57.582345Z digest=sha256:b9c9f5fcd203385b125b477335b5f6eba16086e0f3f8f698299d7d9f0d558902

Observation b4175d1d-845f-4f1c-b391-4bfa1a17e30f · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.693206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:6efcc1e0e710d9ed6a9a12417931a2063e73c3273a01284255f8c6bf3cd565b4

Observation 7b1639c6-c879-4556-8598-1c975e26c877 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.552150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:9c5e07fcfcde84a8a57ef750bcde8f9148fd54715986a53aab95f39b9a5d3b34

Observation 98e4b80d-de56-429b-b932-23f96c90d71d · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:c6341abff3aaf69017ef7ba163cf886f5bf4a3eafa6b8a569451dabb67b0968c

Observation 7feb9fdd-86cd-4459-b7ab-e676637d5b6d · inbound

BatchGen: An Architecture for Scalable and Efficient Batch Inference cites this paper.

BatchGen: An Architecture for Scalable and Efficient Batch Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:39.555548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:58:25.319255Z digest=sha256:f20ecb0228bc0c8307658e1ecdb5cb91f334e09fc9b1d4068bc9ce45993caaaa

Observation 0586afed-f7b3-4eb8-9d54-1716a183b7d3 · inbound

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware cites this paper.

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:39.570088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:40:53.103233Z digest=sha256:5894cd9b22ed1de130944167882f5b24b5323d2ab3ed9c53ac06ba93363c9bd6

Observation 6863ad4a-8a91-4c67-b247-f838f9f292f3 · inbound

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices cites this paper.

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T13:40:50.742149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:40:50.742149Z digest=sha256:a146072b12096f200e0ee24b3d8d498082a717d4a225fb6f0392390a66dae48c

Observation 43ad2203-067a-49c0-a69e-efe1f2f3fcfc · inbound

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference cites this paper.

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:37.574777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:37.574777Z digest=sha256:999676721b0dd6f0d20af3e8cbc531abc94fd27126ce71de2599d0a936a1587b

Observation 7dbd1280-86be-4709-a542-8d32f1fd8085 · inbound

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference cites this paper.

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T00:37:29.052904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:37:29.052904Z digest=sha256:cf0ea34d8ca11638ffdf114fe48ae10c749b73f3d890b8653f41f9b194f12ae3

Observation b5f181eb-5125-4db1-8aeb-b2973b4e5ba1 · inbound

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models cites this paper.

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:37.745029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:51:37.745029Z digest=sha256:ffe6c8dfd4fb84b7ee13df21fce62389b8710d93f1cbffafb59c83196d985a53