Pith. sign in

Paper Citation Record · LEDGER

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2508.19373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19373 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:53:31.121991Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:16:13.946445Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:19:54.305666Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78aa2a76-a5c6-4126-907b-2a602040a306 · outbound

This paper cites Efficient Large Scale Language Modeling with Mixtures of Experts.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.504969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.504969Z digest=sha256:5b251795e8d8ea46493d28ec2a622ff62732ad52e522054c7e9410033e81206c

Observation 8df2bef5-66bd-425c-944c-7d07216a5636 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.577800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.577800Z digest=sha256:42217b5a3eeead7389446801c56173b679bf33203c6288f88b46fe65c546fea4

Observation a84a559f-a641-4df7-897c-a4e92ed5c209 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.650393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.650393Z digest=sha256:92afd3fcda9d12921be951561842df8ff7f15a9d313f6dc1820db8f7f24c0069

Observation 2840d992-add2-4378-bd70-3d352bdd3536 · outbound

This paper cites DeepSeek-V3 Technical Report.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:28.788647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:28.788647Z digest=sha256:8901fef4f72671f657ca63497f38dc29d897db1e8f4b65f08881bf8d87cfe083

Observation de6209a7-9317-4477-ae13-1488bb143802 · outbound

This paper cites Scaling vision with sparse mixture of experts,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Scaling vision with sparse mixture of experts,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:35.471837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:28.889836Z digest=sha256:f38f27de813a8521b57c05d2e508fa071ccbf034e5058312b821d3adac450616

Observation 1fe8b34c-17fa-42fa-bd32-2838f16b6e75 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Efficient memory management for large language model serving with pagedattention,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:35.213876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:28.944744Z digest=sha256:ab5d0350e8a95736de68f4a0f2d2ae83e478c17b0c936057ec587985beed62c2

Observation 804a2c0c-1dc3-47dc-a9f9-82497e7d3e21 · outbound

This paper cites DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.009532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.009532Z digest=sha256:b5bcbdfd2f7a27056273a4c085ab8272e3789c0597318235b7d5edb5d8eb4cd9

Observation 2267c1e9-1983-472e-a626-811df387e571 · outbound

This paper cites Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:34.893015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:29.122972Z digest=sha256:f3d50b38d3dd60c5f13734aae449979a25ed1cc19a81a76b9ee9a462778a3734

Observation d9a93d69-9e96-4857-b601-668103d04040 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.246910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.246910Z digest=sha256:07b393f0b1f0af869ed14e3e16eaafc773acebe2f52e696eaa0b242c5358d5ba

Observation 0891a351-fee2-4949-8b0e-eadd58e3fc64 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.303551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.303551Z digest=sha256:90754897cac73dc00a1401837ce34eae286a5de549110660c1d177fee8754522

Observation 32bea133-c8ac-4fac-88ad-5fcaa90b9656 · outbound

This paper cites Mixtral of Experts.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Mixtral of Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.456526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.456526Z digest=sha256:4a4f6b17bb1cba72aa5c3d562fe0372840be819415b7c9825f998f6f39e0e17a

Observation 67d1fffb-a61a-40c8-a54d-fda1aebfd43b · outbound

This paper cites Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:34.649050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:29.581788Z digest=sha256:034eaeee8fe66bb68a6d981800270e0bdd5b58bab4540e449bf8ebf4ca3fa9cc

Observation b4145053-23c4-4042-9b99-8e4fbb40d428 · outbound

This paper cites Qwen2 technical report,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Qwen2 technical report,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:34.383939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:29.683861Z digest=sha256:42de691dba1507bccddc4c700321735295125165fb037ce4bf1297d87e5041a3

Observation bebc2fe6-f70e-47f9-b054-42ebf961b062 · outbound

This paper cites A tensorrt toolbox for optimized large language model inference.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference A tensorrt toolbox for optimized large language model inference

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:34.150158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:29.789935Z digest=sha256:314d4fdc30622594466d47cafe821ff20377c53d584d282bbc9466eb6ce57178

Observation 7defcc0e-ab5d-4898-bf95-3bc49fd3b840 · outbound

This paper cites Sglang: Efficient execution of structured language model programs,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Sglang: Efficient execution of structured language model programs,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:33.863906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:29.885147Z digest=sha256:77dda6fc92457b6cd9e6e22131fc4d61ce0e9cba7b62346de2d472e4e78886a1

Observation 8a2df785-e639-439d-b2dd-6cafd657326f · outbound

This paper cites Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:29.967927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:29.967927Z digest=sha256:df00444b32e09ceaccb2edc9fb0da716bcb5b88c60353049fe800110614331aa

Observation bf46c29f-504f-411e-800f-c48a25a3b836 · outbound

This paper cites Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:30.143109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:30.143109Z digest=sha256:a83d67af3c7e9c27fc4a603e06b441f71165dce95fd391bd007122de325138b5

Observation 10fe3b85-88c6-4681-9a38-de8394dae73e · outbound

This paper cites Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Deepspeed-moe: Advancing mixture- of-experts inference and training to power next-generation ai scale,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:33.554501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.284288Z digest=sha256:efb70c0e6d0bbc10175f4e6b9ddf45a2e4a7a20606923abdb3c8a5b9f42882d5

Observation 80bbc26a-33c2-4fec-9cf4-ddeb9c3bd5f6 · outbound

This paper cites A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference A hybrid tensor-expert-data parallelism approach to optimize mixture- of-experts training,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:33.247641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.429204Z digest=sha256:574ff221d0801c9a0ce890c03d0c11756f8337f5d7199ad8a9d55b7eead1b1c6

Observation d4006313-5169-4be0-93d9-b7a563c8e2de · outbound

This paper cites Deepep: an efficient expert-parallel communication library.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Deepep: an efficient expert-parallel communication library

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:32.922786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.553332Z digest=sha256:02a26adaa3457eb01152c86a4a8c6ca837d600c5f37e2c92ac14011c9f7db369

Observation abc71dfc-ed1b-48f9-8d2e-e53006875e2d · outbound

This paper cites Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Alpa: Automating inter-and {Intra- Operator} parallelism for distributed deep learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:32.709061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.716371Z digest=sha256:7560fff9dd4ae853c001eb85e8f17adecf65e26dd407981990146ba94736be10

Observation c9593ac8-d628-4788-b72f-54a186ca65ff · outbound

This paper cites Tutel: Adaptive mixture-of-experts at scale,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Tutel: Adaptive mixture-of-experts at scale,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:32.422585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.782234Z digest=sha256:aa97bf69386a804535af4f53326c31bc38ec7faf9d39e533bc66e21e44f3bb48

Observation 3184221e-1e66-418f-8983-9e307435cec1 · outbound

This paper cites Pcie gen-5 design challenges of high-speed servers,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Pcie gen-5 design challenges of high-speed servers,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:32.187333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.865671Z digest=sha256:2af1665d56338e5b9d1548e3a5ca7b53fccaf2704126a696cc973ad76e6fccdd

Observation 1bc0417a-b19b-4dc5-a8fc-f6eb5be375ca · outbound

This paper cites Nvidia a100 gpu: Performance & innova- tion for gpu computing,.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Nvidia a100 gpu: Performance & innova- tion for gpu computing,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:31.855990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:30.959463Z digest=sha256:328e64ea6accb41203805f1fb07c9357270a0db10c08120b20b4674b41f3b96e

Observation e15fa19d-f1e5-42fe-8275-28e067506fe9 · outbound

This paper cites Bitsandbytes: a lightweight python wrapper around cuda custom func- tions.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference Bitsandbytes: a lightweight python wrapper around cuda custom func- tions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:53:31.652787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:53:31.050897Z digest=sha256:fab10116a24ba943ce299914d614ee68f15d925a8cfbe74fe22320a4a76e719a

Observation ddd5ed51-d969-4652-bf2e-66513c54a65a · outbound

This paper cites A White Paper on Neural Network Quantization.

HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference A White Paper on Neural Network Quantization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:31.121991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:31.121991Z digest=sha256:923b9aa8e3f9f75d5b0013dfe1491b8527c0c9cd374b3ca7d28ce36f872ac517

Pith citing papers

Observation 605ca0d3-c892-4868-8dad-b1eab5e2fa33 · inbound

Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch cites this paper.

Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:19:54.307203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T04:06:13.426379Z digest=sha256:d5a4c5a9b0553bd60c96df90c05c15cb8a0d7b8118459eeaee67461241c2ecd1

Observation 658b5d45-7a36-4a0d-8404-f8307c332c67 · inbound

Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference cites this paper.

Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T19:16:13.946445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:16:13.946445Z digest=sha256:589a29d15bef5e67b4f1052982c50c28cb0e32f801135f0b0ada2443a9757f26