Pith. sign in

Paper Citation Record · LEDGER

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2607.08215.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08215 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T11:11:55.188720Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact15
  • verified fuzzy10
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 897c8722-3ab0-4963-9a71-eba72526e343 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T11:17:03.651965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:cca7122c58cb975faa757255239324c9fc5d6b81cf365a320d30778cf938a3ad

Observation e66b49e2-54d3-4c08-875f-b7759226bc93 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.647232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:3720c0044e675a597c3210811e5116c8c465907dab209223cfeb51bad4a7fa9e

Observation 80283ae9-d170-43b4-b891-24b98ccdd472 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.951895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:adcae3a06f46e783c78a0f16759fc926a99e163f28c6ed481efa80df32829bd3

Observation fbbf6edf-a691-4e0f-a2dc-2d08c5b65790 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.654359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:99d47897eec58598f22d48c11042aab50592c397c912107ee65d8b4834cacf7c

Observation ac465b48-fb90-404e-895c-b07fd8522086 · outbound

This paper cites DeepSeek-V3 Technical Report.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DeepSeek-V3 Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.656488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:1999dda7f57398335ec68c7b60cafd72bfdd0ba139c3437fbafe73d9eb69e157

Observation 075d6415-14e4-4c52-9554-8eacd184263e · outbound

This paper cites Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.955259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:16e55e5a0f88c084aeb470d5a431a981efa9b2495a33af025480c4241e339968

Observation f4b94852-e38b-4d51-bc2a-87c2046b499b · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.658979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:18d25a53634cb938af00b81713976acdb26fc7d42b1f034f30b5c4124c031733

Observation b9970f5c-570b-4a96-a4b1-97c1a9088a8a · outbound

This paper cites Le, Yonghui Wu, and Zhifeng Chen.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Le, Yonghui Wu, and Zhifeng Chen

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.958961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:b1fe1e737fed2e982245831ddd74a756856b1e362e5e42c72feba954345fe98c

Observation 0babd0a7-68ce-4a4c-a87a-51bc8ee76706 · outbound

This paper cites CANN: Compute architecture for neural networks — documentation.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend CANN: Compute architecture for neural networks — documentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.962697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d6fc66c96e5624d47d3d091eb042fba19b454c40ea306c0399c7c159a69b8d56

Observation 52f93f76-9e2c-49a1-a96f-f8c95a9f0ff4 · outbound

This paper cites Mixtral of Experts.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Mixtral of Experts

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.639556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d514352638f1fe2c61cf0a592f3d701896c4849e90971dba02fb880981d5805a

Observation cbb4e218-af68-4820-aced-9213ad750dcb · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T11:17:03.600722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:7e74fd4647b5d6975f7ca2513c5349ad5f4bd3395e6f49a7398b6051bb6539b3

Observation ff08e0d9-58a2-4217-8ea1-a6155940bf6c · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.627839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d482b795efc1c46419109316ca9d607d4cd580499d68bea87abd1e72c8820d10

Observation 2383e7e1-9dca-4091-bd4d-ec137da3ee53 · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Fast Inference from Transformers via Speculative Decoding

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.661344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:c09f31f36d16a06905a555709993a35678681b2a8b27e0efc094aab5a43f80be

Observation af5ff6bf-dfc2-4cae-8b5e-f6d934e557ed · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.632648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:cb41029c8f6b367a234b2007f5a423b850047607c38ce24adf14c079ebb72a30

Observation 474e0dc6-8c1c-429f-b6ea-8ab81a78b075 · outbound

This paper cites DaVinci: A scalable architecture for neural network computing.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend DaVinci: A scalable architecture for neural network computing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.957134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:b67f4c80e1c83b45c47c7b3f272fe12e14ecf744f6ce23d82cdc1c77755ea87d

Observation bc4c22cc-3733-4152-bbcb-16b3bd90348b · outbound

This paper cites Ascend: A scalable and unified architecture for ubiquitous deep neural network computing — industry track paper.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Ascend: A scalable and unified architecture for ubiquitous deep neural network computing — industry track paper

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.960802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:6cda582552eb66356ccc5fa2de842a9ac6715732790e341ddb056aca2dce55f9

Observation ef1077cc-dc58-4c02-98c2-a8748651db36 · outbound

This paper cites Villa et al.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Villa et al

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T11:17:03.598028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:ee6271b223be07d87f19a8f7cd6c93550dbf5e6e985cde1815a415eb82de4066

Observation 0c805e96-a852-43d1-a064-730d043576f1 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.642054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:d63706cdfc9a0ab802df55e8b007046ba520c68af9940b85e68eeb01bec36fdc

Observation 0e2ac552-fbc7-4114-909c-a3122c01d833 · outbound

This paper cites Visual Instruction Tuning.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Visual Instruction Tuning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.637295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:1821698869b7a7b0961e9a5488672b7195a7083b328677a94371ced1e30f3df1

Observation 520683a6-f800-4144-b9bb-cdfd752c4fa7 · outbound

This paper cites PyTorch: An imperative style, high- performance deep learning library.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend PyTorch: An imperative style, high- performance deep learning library

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.950078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:bb3ebd0018741f606281e513ed6ecde31a25a61b3892abcce5e6bdc24cd6623b

Observation 7d6012de-f9a1-497b-81d4-145711defedc · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.635016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:ffacf84abcf273cdb366a97c1f77ec47fc13aaeec63f00561b931f51eeb47ee0

Observation be493855-4b23-41f1-a0f3-057dc266a059 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.644722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:5285a04e6d9501a07753398626dc51f6d4f33b0bc757ac3add59e78592b4a0b0

Observation 8914cb5e-0ea8-437e-822a-44a85d158b13 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.649613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:7f8bbfa9342deba607892200226c675720f8bf470a3c88a458d2b157414838ba

Observation 30170d5b-f11a-479d-a145-a5c690ad7408 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.953518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:ecd48a8aa93ad3e5a0722ae65ef1b7ceead06199538765387b970ca469b773a6

Observation e20f6114-626b-4f78-a0c9-7d43a4ab4a73 · outbound

This paper cites vLLM Ascend plugin (vllm-ascend) documentation.https://vllm-ascend.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend vLLM Ascend plugin (vllm-ascend) documentation.https://vllm-ascend

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.948244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:62505b569a4f7bc63ec5e9486cd0d7811ff01d9c1eb3b8f8209fb46c7beb79e9

Observation d2aee196-2cde-4ef4-85af-ce384a7e2106 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.630296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:fc9590c965e07764ba3a5aa0e49c250fb8cd1eccc760d9096b160141d95a78ef

Observation 619f3c47-2f2b-44a3-8616-3e195866da7f · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:17:03.625272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:5e04dd4edf6c2c8d66c6ada6bf1ea4e13fb87f75acb4e56a611a9170ca4f36f4

Observation b52fa409-2221-4a6d-81cc-6993095aa59c · outbound

This paper cites Orca: A distributed serving system for Transformer-based generative models.

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend Orca: A distributed serving system for Transformer-based generative models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T11:17:03.946484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T11:11:55.188720Z digest=sha256:b6ea1514aa321bd34459a65b28b71986dedf65aab1d2c1aa3d107a995c3a8775

Pith citing papers

No inbound Pith citation observations are available.