Pith. sign in

Paper Citation Record · LEDGER

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend

As of 10 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.08215.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08215 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T11:11:55.188720Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact15
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 897c8722-3ab0-4963-9a71-eba72526e343 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T11:17:03.651965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:80e9da36bd450b0b0fe68c6dfa2530106b00ab752926e30dc92782c8307785e3

Observation e66b49e2-54d3-4c08-875f-b7759226bc93 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.647232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:c59fb6d26054a5644c84509ef543029c83ad814fcfc2d6ff0f8b8d2efcd8bd49

Observation 80283ae9-d170-43b4-b891-24b98ccdd472 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.951895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:51ae8cb18c71562b20cf3b14263c90aa1ad15cdd61469bdf4208def7e52ff7e1

Observation fbbf6edf-a691-4e0f-a2dc-2d08c5b65790 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.654359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:e57aafa8492ddef4309689106d4403da398158a1a503b2684d9398eeab31abad

Observation ac465b48-fb90-404e-895c-b07fd8522086 · outbound

This paper cites DeepSeek-V3 Technical Report.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DeepSeek-V3 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.656488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:dcd6d07b318bbf3006d1a9ffd54ca10b12833eeac2d08a6da8feee78ca6d6125

Observation 075d6415-14e4-4c52-9554-8eacd184263e · outbound

This paper cites Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.955259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:15780dc847f910c0fc059d9f1ac362b349162b15ce062fd9c7f4a453545f54d9

Observation f4b94852-e38b-4d51-bc2a-87c2046b499b · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.658979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:2237e4f5549e2c8a4f1b975dcc2dd76ba242c6bdaea3c3ec29ce2ccf81b183cc

Observation b9970f5c-570b-4a96-a4b1-97c1a9088a8a · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Le, Yonghui Wu, and Zhifeng Chen

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.958961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:1856eb0174cdf66edfbdc7e51f95b2fce447adaf0a979e8472650079e5c41992

Observation 0babd0a7-68ce-4a4c-a87a-51bc8ee76706 · outbound

This paper cites CANN: Compute architecture for neural networks — documentation.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend CANN: Compute architecture for neural networks — documentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.962697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:73d8d41db096c3c1def02e035173fd5f8a8d4d1f1716e7d2e1ff20c23eb6aaa3

Observation 52f93f76-9e2c-49a1-a96f-f8c95a9f0ff4 · outbound

This paper cites Mixtral of Experts.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Mixtral of Experts

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.639556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:0323921c1f2d34693f704765a1e1acad06a2e82f601f46be263d5ce3efcfe8dd

Observation cbb4e218-af68-4820-aced-9213ad750dcb · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T11:17:03.600722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:3ea06c3a1210463a73083a37c7ae2244cbe9b008b78ba3507d614714e77f2e48

Observation ff08e0d9-58a2-4217-8ea1-a6155940bf6c · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.627839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:b06aa42e480db2af6a5f63dfac4a3af7e458c07144d54537da73c1d124f0118e

Observation 2383e7e1-9dca-4091-bd4d-ec137da3ee53 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Fast Inference from Transformers via Speculative Decoding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.661344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:9728481284c5de42c42495ae8f30c269830bd9047a7433b6909be4fbe10a1aa7

Observation af5ff6bf-dfc2-4cae-8b5e-f6d934e557ed · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.632648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:572257de2464f46a55775946cb9bc267faff1bf168c13ab6d697efcc52e9c7ac

Observation 474e0dc6-8c1c-429f-b6ea-8ab81a78b075 · outbound

This paper cites DaVinci: A scalable architecture for neural network computing.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DaVinci: A scalable architecture for neural network computing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.957134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:f3572fac0e07dc229d3517fa1c27d88646d8bfa1e19a1eaa9739e3a59fbd0ccb

Observation bc4c22cc-3733-4152-bbcb-16b3bd90348b · outbound

This paper cites Ascend: A scalable and unified architecture for ubiquitous deep neural network computing — industry track paper.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Ascend: A scalable and unified architecture for ubiquitous deep neural network computing — industry track paper

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.960802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d96a56946d5f697da734cb504af8e83e42a2eac41141cf375c949f35276caa22

Observation ef1077cc-dc58-4c02-98c2-a8748651db36 · outbound

This paper cites Villa et al.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Villa et al

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T11:17:03.598028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d724c66a212bd54b4072f5016378440577b63cf1c8b49ad44de215439b439683

Observation 0c805e96-a852-43d1-a064-730d043576f1 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.642054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d7093c7ad0a5bfcdbe32627c375207051a55053a23c247a2480c23615b63d13c

Observation 0e2ac552-fbc7-4114-909c-a3122c01d833 · outbound

This paper cites Visual Instruction Tuning.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Visual Instruction Tuning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.637295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:0ce8ac511c4e898d1bb5449ab885cc01f19c2b6daa07515104af6621ec207f85

Observation 520683a6-f800-4144-b9bb-cdfd752c4fa7 · outbound

This paper cites PyTorch: An imperative style, high- performance deep learning library.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend PyTorch: An imperative style, high- performance deep learning library

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.950078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:4a248cd191048fbcff8212422a8765e44114fb642edd3dd8a4c5d6f3a863d52e

Observation 7d6012de-f9a1-497b-81d4-145711defedc · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.635016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:1b6d221e01f7c98e4b4489472e16b2ff9dc0861c80cc4ebd54c17e8e107d3a84

Observation be493855-4b23-41f1-a0f3-057dc266a059 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.644722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:de988d334795baebec9cb1c2c6f84fbddab898e582a89cbce7d72e045cd85b46

Observation 8914cb5e-0ea8-437e-822a-44a85d158b13 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.649613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:4f9b8d3cdfb495b5e89fa120dc99351cce9631e87c4c2db3400f94ea84331f5d

Observation 30170d5b-f11a-479d-a145-a5c690ad7408 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.953518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:e937dde3c21b3579345426892c31f487e03aed665fbb8631dc0e1819da9fc552

Observation e20f6114-626b-4f78-a0c9-7d43a4ab4a73 · outbound

This paper cites vLLM Ascend plugin (vllm-ascend) documentation.https://vllm-ascend.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend vLLM Ascend plugin (vllm-ascend) documentation.https://vllm-ascend

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.948244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:a80717736055f5b996925c23a73d38145657fe5c6654d384ea10095505915175

Observation d2aee196-2cde-4ef4-85af-ce384a7e2106 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.630296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:9b68e2f73ab639e23b2385c24be7f34c11eb1a37076ffed43148373e9851df4d

Observation 619f3c47-2f2b-44a3-8616-3e195866da7f · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.625272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:e7d2a9dca30aa417dada0e48d9a9478fe78d603154b03ebb23c69d38128a67d0

Observation b52fa409-2221-4a6d-81cc-6993095aa59c · outbound

This paper cites Orca: A distributed serving system for Transformer-based generative models.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Orca: A distributed serving system for Transformer-based generative models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.946484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:e32c6cdca036c7718fb671b4af715311a356623031ebc5badbda0ec9d096219b

Pith citing papers

No inbound Pith citation observations are available.