Pith. sign in

Paper Citation Record · LEDGER

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

As of 21 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 4 inbound Pith citation observations for arXiv:2508.12851.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12851 v4

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T22:52:39.416575Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:37:29.070663Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:07:44.131237Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact13
  • verified fuzzy21
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4fb10cc8-a015-4a03-974c-0d137f463bd2 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.049412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ce7d312d70b6a67b4ea11c110b4d74fe535bd1575890c611ab1b755a74574b2b

Observation a87e8e74-f757-4436-a62d-6ad928348af8 · outbound

This paper cites Mixtral of Experts.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Mixtral of Experts

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.955108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:2371e83feafa196699bfd562eb8553bdc69479eccbf3fb0e90379705c92247db

Observation 2d376eaf-8462-4e61-b09d-5c46021642ce · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement DeepSeek-V3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.936457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:51c322b25d21524f5cf425ae3b664fee2d4f7b92bad2641fe11b5254b5a7ef1d

Observation 23bff876-6a79-4ee9-bd75-5de700e664aa · outbound

This paper cites Geforce rtx 40 series.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Geforce rtx 40 series

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.985010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:239b9d0d7d1c6cb2fa12f86509c12c6c9adb3e4cf9cb0778069b917e7f8817c3

Observation 19c09be9-0738-434a-94fc-4d3c69197c24 · outbound

This paper cites Gpunion: Autonomous gpu sharing on campus.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Gpunion: Autonomous gpu sharing on campus

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.950085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:aeaa092038b72dde36ed4d5a461e0a5f7dec8b0ac3496b89c6f6d66931c7077a

Observation 79250885-432f-42c5-807f-b29e73cd6a49 · outbound

This paper cites Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.874632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:622a380c7d0bd64521b34803683640b813f438d5a2ec0cd5031c3992d8d678a6

Observation 5a23cd60-650e-4d0c-a848-4dd8e1d088cf · outbound

This paper cites MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.924894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:1f2f3cb43405328d07a28b4fccacd45712dd9c96ec95dcb5d28de733dc1bfe0c

Observation 28f25519-5370-4395-94e5-818027fc61df · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.942278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:f157aa0f62c6528e3cddfe9c693d7d9f28c1318bc537314ea8ce218e5b0ff4b6

Observation 352ef6f0-0f0e-41ad-aa33-87cb1bd15739 · outbound

This paper cites Expert Parallelism Load Balancer (EPLB).

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Expert Parallelism Load Balancer (EPLB)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.017792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:9e237d20b6e97514d19821a81344bc639baac3dda1da08cbcde434f642ac4c94

Observation 70e646ec-bad1-4f76-ad27-0e0e41e1e4ce · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.995958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ea5e1366d1357bed63107098ff6e79d1edb251293a005d7b9c54cc8d1a1eabc3

Observation 42834221-0771-4a8d-a968-82bfcca28b9d · outbound

This paper cites Moe-infinity: Efficient moe inference on personal machines with sparsity-aware expert cache.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Moe-infinity: Efficient moe inference on personal machines with sparsity-aware expert cache

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.028077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:89362ce489ab96b927a91539f2c589be627843bb213a3f25c3b196d81d76e173

Observation 88c9b676-31a2-46a5-a55c-19bee0e0cc4c · outbound

This paper cites Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Mooncake: Trading more storage for less computation — a KVCache-centric architecture for serving LLM chatbot

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.981756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:027de0495a3933f80bf994bcc0b977301e2df5d35a4b7e4526899e049a1c1104

Observation b5426018-7783-4318-9668-586020acbabd · outbound

This paper cites an unresolved cited work.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-18T22:52:51.968571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:a7ecb5bfb700ca8cd60389200a8843a0bcc34a8c50ec92cf7457a73551def0a9

Observation f80b0c76-ef3c-4d15-9b6a-9029202c8e7d · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.919074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:285f53087fb328330a89e4f260ede5e064f821f687a2f4ed0f7920cc0a7aae4d

Observation 33b53e28-96ef-4d18-8419-7549eb4b4098 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.881624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:9d77c014b7ce389d52bb51bbc6058393d709af3d2f02115fc25507318279a025

Observation 1653027f-19ef-47dd-a00b-ef4e2d8659da · outbound

This paper cites Pointer sentinel mixture models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pointer sentinel mixture models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.062287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:7f4d2db56e420f0beb629f70cd699fc24080ec0fb0ad6cc7333a397af9ae69e9

Observation 162d2908-54e7-4457-b0be-4ee977d78c01 · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement TACO: Topics in Algorithmic COde generation dataset

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.908258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:ae323337f49519f3f17dcb257237bee7458137873a8af3c063b7ecae83236407

Observation 7ff6cb8c-caa1-41f1-add2-96631a92d043 · outbound

This paper cites {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement {SmartMoE}: Efficiently training {Sparsely-Activated} models through combining offline and online parallelization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.992346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:acc16c6daf3639137148de2b8176a7e0f44eb05e5c62ff984e68fbcdfd59a50b

Observation 27729f97-78dd-44f1-a1cf-f9b3cba67c3c · outbound

This paper cites Joint application placement and request routing optimization for dynamic edge computing service management.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Joint application placement and request routing optimization for dynamic edge computing service management

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.022246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:c276d62d9ff1bfbe58266be90a78153fdedd3c71e0dd79366f41c6314a89c8cf

Observation 46ec477a-dcc3-4e53-aa45-6acc8f0b077d · outbound

This paper cites Task placement and resource allocation for edge machine learning: A gnn- based multi-agent reinforcement learning paradigm.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Task placement and resource allocation for edge machine learning: A gnn- based multi-agent reinforcement learning paradigm

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.973069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:f4658d14ee77cdc0569dc7b9d4732c96d9469b079c707e32f117089c2e71c361

Observation 94421f3a-aee4-4696-bc92-8a53e4996ba2 · outbound

This paper cites Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Tapfinger: Task place- ment and fine-grained resource allocation for edge machine learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.044212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:dacf6cb2647920e76bbd36b011953fb1c962a6a259c246d6c13ed6a39d0074bd

Observation a4d1369b-13fd-4c1f-bf09-04ef30249892 · outbound

This paper cites Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Faster- moe: modeling and optimizing training of large-scale dynamic pre- trained models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.057367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:2ba0de2ffcfb81c41b76262025204d88bdd5e26ddd943db950a98ef3a394b21e

Observation 295ed844-d8fd-44c3-abae-acb58ffc8c93 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.068135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:a5de0bc6f64d20c0cdd3bb97cbf64e2e46366d79a88c2cd935e2936c9b41cf37

Observation a59f0663-8ab0-4e3e-9bba-eacf0386e657 · outbound

This paper cites Prophet: Fine-grained load balancing for parallel training of large- scale moe models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Prophet: Fine-grained load balancing for parallel training of large- scale moe models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.038428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:028ae7d8e0234c7d4491a5711e2ee754655c73fc914cc4ba0b92486da4c17cf7

Observation 2bbbc822-bdac-496b-b4ae-c8a673107c8d · outbound

This paper cites Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.888643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:b37856bcacb61eed62a4475e5536f2698ae9be88f598a79bf1be4b21857a59c6

Observation 8223bee1-211b-4122-962f-abcb7fa044a7 · outbound

This paper cites Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.001256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:a35e217a252bfe47731503d8a38159f1b7cafea7bbe29b3ddc5a5b96fe2719d3

Observation 189a01be-0957-4ce8-acbf-ca8f9b7ad920 · outbound

This paper cites Accelerating distributed {MoE} training and inference with lina.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Accelerating distributed {MoE} training and inference with lina

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.009910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:7da56668e0c36ae23bf962bb4749bca318730af7e217a32ba82a82c7de406701

Observation 07f98ea7-7877-4bc3-bf7e-72b0de36a718 · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:52:51.896416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:1fe3f02c32e5fcd678a3d87a743275047664323faa69a3ec227511c95287f551

Observation c8181ca5-3127-45e8-8eeb-d2c31f7b72f3 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.931759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:642dac6ecc36b82a99531b8c71b695bec5d780d7faa4e3974f9b2c25988db415

Observation 91c32161-0484-4cd7-88be-11e5b78f1e3f · outbound

This paper cites AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.914141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:db709c8dedaead997a64e9951d5d7f9ad5fc9131a237f6eaa41972230da2732e

Observation 63ab3982-2d7b-4a8d-aac7-51adcfe8dbe5 · outbound

This paper cites Swapmoe: Serving off-the-shelf moe-based large language models with tunable memory budget.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Swapmoe: Serving off-the-shelf moe-based large language models with tunable memory budget

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.988367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:b4d01f7fb1daefbfde255a98e19b289b284006da0ec203bd1ce432b540b2822a

Observation 0a824cf5-fbcf-4467-af21-b6c80fe85591 · outbound

This paper cites Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Sida: Sparsity-inspired data-aware serving for efficient and scalable large mixture-of-experts models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:51.977772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:69e7d619be65f1064b94ef84a2213a8f71657b44d982fd95a4ae72bd0864a177

Observation 8926653f-8995-4258-9dae-2495de2d0545 · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:52:51.902779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:c84800fbcc8e4518bd8c1971580644dfc85fab19482b79a7218fb48f47305b17

Observation df5f0a61-7f5b-4617-b094-43cf6ee7c39e · outbound

This paper cites Pipemoe: Accelerating mixture- of-experts through adaptive pipelining.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Pipemoe: Accelerating mixture- of-experts through adaptive pipelining

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.033806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:990177b89324811b5d9b8a2fa19891c9a2a31f791444a3d2be2de981416a5600

Observation 6b75f761-5b45-4cd4-bc95-07f70bce0efb · outbound

This paper cites Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.013559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:623a5f6e93bc276e3f99186b32adc4d2d26d5ccd15e8f8c4dfc31dd4e92b246e

Observation 5a664ba4-c476-4d49-b517-72779e4d8c19 · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement Tutel: Adaptive mixture-of-experts at scale

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T22:52:52.005928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:9efcc355bf8116c9c3fef2b95de72380087d8f9cf449d0a1e9f3b249debb0054

Pith citing papers

Observation b7337072-9053-4129-8709-020fe34ac8a7 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-08T13:44:54.930922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T13:43:17.950000Z digest=sha256:ba16f7266a0f6c24c9a602d69172af9829d5dc7903f5c57821843f08c48d704a

Observation 67f9a1aa-d74f-4562-b519-63ae177f5500 · inbound

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods cites this paper.

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:07:44.161601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-11T01:06:46.426582Z digest=sha256:f8d9024dd4b0d6b856828f33cd1b6d0b57588c15c680e76365ee35f58fdf5acb

Observation f0ba5b1d-6ddb-4152-bc87-f0d5573af3fa · inbound

OrderMoE: An expert similarity driven distributed edge MoE inference cites this paper.

OrderMoE: An expert similarity driven distributed edge MoE inference Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:19.766584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:19.766584Z digest=sha256:f0a3739f9d3e3f7053b2f6082f45ae46f48b58f65a6ae138a8a300413a6d11b3

Observation a8f20f64-188c-45fc-926e-a12c82a7d1d4 · inbound

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference cites this paper.

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T00:37:29.070663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:37:29.070663Z digest=sha256:7f5c28dc8dca918d32f2444c0258bf24681a12007c0d79ce5e285c0426e54ddd