Pith. sign in

Paper Citation Record · LEDGER

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2604.10152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10152 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:38:30.250542Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T15:02:37.509164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy59
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b566f263-596f-4861-b051-ab2de4d0446a · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.475973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:49dddbabd9c2ef7ee710a27e0b0556c6d429fec89177762024ed55726431e11e

Observation 37c68beb-f224-4c9a-ab1f-2d5145135533 · outbound

This paper cites GQA: Training Generalized Multi-query Transformer Models from Multi-Head Checkpoints.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding GQA: Training Generalized Multi-query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.463271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:86a2e3b8629f75d538688309630b1b34c1baa2f6e204230da37f21f17d7cbcaf

Observation 0355760d-2113-457e-ab29-17d67028646e · outbound

This paper cites DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Un- precedented Scale.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Un- precedented Scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.422803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:cbeebb64b6ecbde0f7dc1e447ef2b96381086c31d54c781b1a7296fdb1ad3858

Observation 405fb333-0b98-4080-bbdf-ff93d89518a9 · outbound

This paper cites Findings of the 2014 Workshop on Statistical Ma- chine Translation.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Findings of the 2014 Workshop on Statistical Ma- chine Translation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.426003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:b3be25429c598662e0feb89454739983d12ba1607e1fea0e42b27d1fecffedd1

Observation 5ef2a298-8e66-4a77-aff4-bf70cb0eac91 · outbound

This paper cites Language Models are Few-shot Learners.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Language Models are Few-shot Learners

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.432864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:343c979148137971baecec0b88790b5f4108f9978511d85d46d09943b1ac5336

Observation 691ceb76-dfb3-45f2-899e-dd57412a02f0 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.439607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:c9f025881e26a5adaa46930dfd406b2b655e2de7ab5fd89633997910c714673e

Observation a457f1cb-78c2-43cb-8439-7b8517015f3b · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sam- pling.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sam- pling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.416059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:458ab167b09001b7764bb63cf1dcd8b7c1c2c225a90936375090b53856658fb6

Observation 9d54a26b-0d96-48f0-ad37-9915d17a105f · outbound

This paper cites Punica: Multi-tenant LoRA Serving.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Punica: Multi-tenant LoRA Serving

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.412494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:1001449d66478076194e1c2e85136a680d0202b6d3aa0b66537ee9f4e355fc8b

Observation e90832ac-51a3-404d-9e9c-7b7fb628efa9 · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.429488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:7ce9ca9325deeba6c52560a45de53c5f3e761fcf59005611e047db57aff0b1cf

Observation 63ef1615-8a70-43f7-8ea9-5e3381c3c2b8 · outbound

This paper cites DeepSeek- R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding DeepSeek- R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.397886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:bc1851afea89d014cc51484c692c993721585e7717dd770d29269afe8276eec5

Observation 4e2c2884-dad3-4c97-a583-5128aa759f66 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-speculative Decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-speculative Decoding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.401644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:1e076946ecd9ea7e2d37d06154cedfe0a7d3016a05e074a61b49783a42c6ce43

Observation 316409e3-8e20-4ea0-b8a0-bf173dfa743f · outbound

This paper cites Fast Inference of Mixture-of-Experts Lan- guage Models with Offloading.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Fast Inference of Mixture-of-Experts Lan- guage Models with Offloading

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.405055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:94c47ff1ec86f3819aab5abd2917843474f9750514a7323533119276f9a0cb21

Observation 9b469f28-fafa-4ff3-a3c1-53de9ea7df4f · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.376410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:29d0aa7615d0cfaf95023fe57ed66975f8965858f5cb36200c91b5babe363e9d

Observation 00b850ca-407a-49c9-9aa6-e4ec8a787b66 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.379747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:f4c05e13b55ae9d98deb86fb27d3ae1a7738d17f5dcb0d2d26258af59bb9f069

Observation fc612821-20af-46cb-aad8-e3437c17930c · outbound

This paper cites Gemini 1.5: Unlocking Multimodal Understanding Across Mil- lions of Tokens of Context.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Gemini 1.5: Unlocking Multimodal Understanding Across Mil- lions of Tokens of Context

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.383398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:89dae5b961054fbcbb6118e8a1627695b411aae3709e192d3242a42d30b48e43

Observation 1b2bbd69-c14a-4a2a-856a-d4f77c5cc20b · outbound

This paper cites Teaching Machines to Read and Comprehend.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Teaching Machines to Read and Comprehend

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.369300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:ed948a1c3f9f9aa3aa7a7e35e4299dc5c5ff1307000a6bc63740e2f2b008fd1a

Observation c20af1fb-17fe-4ce9-8b5f-51c9fe90cec8 · outbound

This paper cites Training Compute-Optimal Large Language Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Training Compute-Optimal Large Language Models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.372262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:b6de6ee6c676f7ad109fa553e2c1495766d1559adb54a0374edf7f8e48320270

Observation bc0d8b24-2a96-449d-81d9-fdbec2afffc9 · outbound

This paper cites Speed: Speculative Pipelined Execution for Efficient Decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Speed: Speculative Pipelined Execution for Efficient Decoding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.386939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:9f8df89ad9eac3025d3a7ee9592f9d0c9b0209b2e1a73c84c2ba7f81c4fdb241

Observation 3273372c-dfa3-447e-ae51-585c7ab9126c · outbound

This paper cites Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.390734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:51da55044545087d30d71af46e3c5d19dd6369eccbf1d55301c2d28d73ab5a97

Observation 0e7d6da9-72ab-479a-a2cc-53779fcf006c · outbound

This paper cites Hugging Face Accelerate.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Hugging Face Accelerate

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.435783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:2a93cad8e078e374a4abf4484653037c153a21380af822e4a160fdae8e7a7a53

Observation ca1d037e-de87-406d-b988-b0b447cb7f0b · outbound

This paper cites Tutel: Adaptive Mixture-of-Experts at Scale.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Tutel: Adaptive Mixture-of-Experts at Scale

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.482714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:7bea4bcc47f6461eb879d52d390c20db62d1c8982b63770b324900568a19e2d6

Observation 380ee061-47b4-441b-b21d-bb33afbbaddf · outbound

This paper cites Pre-gated MoE: An Algorithm-system Co-design for Fast and Scal- able Mixture-of-Expert Inference.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Pre-gated MoE: An Algorithm-system Co-design for Fast and Scal- able Mixture-of-Expert Inference

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.689471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:e3f16aa8506d28cc70a793e0561954906cadeef3e40b1f41ab291ed66ca8da29

Observation 9f212b4f-49d4-4af5-872d-fe58e011a35e · outbound

This paper cites Mixtral of Experts.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Mixtral of Experts

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.679247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:bd967f920750a10e3a4a398d327496626e8b722f824fcd2fcb25be947e2e2e6c

Observation 3207587f-8095-457b-84b0-abf69c6fac66 · outbound

This paper cites Scaling Laws for Neural Language Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling Laws for Neural Language Models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.672578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:7596f9fd560591a9dfd9da9198f37091011b9a6c3bc493132528d11637446ab0

Observation 45d4180b-0a65-4969-9a44-dd2e13215463 · outbound

This paper cites Scaling Laws for Neural Language Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling Laws for Neural Language Models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.491827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:84d11b63240ca322f3096fbe05f65fdb91ec9896010bf9efc9ea41a6f8b4272c

Observation 44fd884e-aff6-4fab-ba46-55f5d7f7203e · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.496131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:8958b43a23555d2e82b0d80e7907533ab794ddd8d2352412fa1f52847df69f26

Observation 57bc9b2f-70d8-4dc6-9d71-2ee6c47007c9 · outbound

This paper cites GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.500340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:f8f792356689edb28feeb0288a861040aca80552d79eab623e188bf1dfbf8275

Observation e373bc3f-c164-4c10-9f82-151aeb10e8e2 · outbound

This paper cites Fast Inference from Trans- formers via Speculative Decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Fast Inference from Trans- formers via Speculative Decoding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.504183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:2b8f0d9ff2a41b76d3f65d55efc9cfa988568754a29a6e5878f409c952ce919c

Observation e0320d12-49a5-4b9f-9b93-ae742bba1fcc · outbound

This paper cites EAGLE: Speculative Sam- pling Requires Rethinking Feature Uncertainty.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding EAGLE: Speculative Sam- pling Requires Rethinking Feature Uncertainty

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.508176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:60e07884f645e3cbb139f2a03caf7933b2ba9b0bf081301908782cfcf1f73fa6

Observation 3f0aa89f-3f17-47d7-a6cd-da36cc2aedd4 · outbound

This paper cites Llama 4 and Multimodal AI: Expanding Intelligence Across Modalities.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Llama 4 and Multimodal AI: Expanding Intelligence Across Modalities

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.511888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:976d8ef1c273ddb265324488695135bf1e68775bef9e4db10ea8d5d12a0e5556

Observation cc43ec45-b1d8-4d1f-936b-c447b47639dc · outbound

This paper cites Specinfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Specinfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.515980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:90235d97c7a0dd7db47f50e52e963f429a29f148d144828dddd3e60330f65d92

Observation c8b8c081-b2bd-41c2-9e38-d86730793e35 · outbound

This paper cites GPT-4 Technical Report.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding GPT-4 Technical Report

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.692508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:94d2e26b1259976e0ed770c2b1b7dbb282acb70323ba077785b144fffba26bfa

Observation a1297ea4-aa9c-4bfb-ba42-6aba885d4899 · outbound

This paper cites Characterizing Power Management Opportunities for LLMs in the Cloud.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Characterizing Power Management Opportunities for LLMs in the Cloud

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.456058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:d1a608f597a9ec835d898d5fc494717782e0a82f1a5eebc51b8319324b21d87a

Observation 77769e0d-fda0-465a-ae90-0b268f419368 · outbound

This paper cites Splitwise: Efficient Generative LLM Inference using Phase Splitting.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Splitwise: Efficient Generative LLM Inference using Phase Splitting

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.459939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:14ebce6be679e51431c60754453ffb3ede30455620f0ee5ce2a8aa2bae764cc2

Observation 64822a6a-ed74-421d-88a9-4a9cbe0ab2a0 · outbound

This paper cites Scaling Language Mod- els: Methods, Analysis & Insights from Training Gopher.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling Language Mod- els: Methods, Analysis & Insights from Training Gopher

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.452825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:5fa80e50ba32fc49de666ed0304467806510c0f77d0b903f988716acd1f160f1

Observation b56a02c1-932a-4554-9718-f0f37cbe7fba · outbound

This paper cites Deepspeed-MoE: Advancing Mixture-of- Experts Inference and Training to Power Next-generation AI Scale.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Deepspeed-MoE: Advancing Mixture-of- Experts Inference and Training to Power Next-generation AI Scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.469923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:3ca192ea6d1dc1aebf4d8b0bfd62fae4f4f223f998d0672895ddbe98c259c4e4

Observation 86722ba3-7bc6-4f07-ab14-6c430ea7080f · outbound

This paper cites Zero- Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Zero- Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.472899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:d9ab666187cd5afd96f35784f7920bd48a54db57a17ca29ea185269b2689d104

Observation 1fff10f7-fcb5-465e-be75-3a766bab5524 · outbound

This paper cites ZeRO-Offload: Democratizing Billion-scale Model Training.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding ZeRO-Offload: Democratizing Billion-scale Model Training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.486277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:50761a806eb63e562e2848d8943c6c8b0fb0969f96ece0abbb7c7692837c9f42

Observation 7cff7c62-b0a1-45bf-ad7d-9d0bedf159fb · outbound

This paper cites Speculative decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Speculative decoding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.419368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:87abf1131814a067f47105aadf159a4c14c7532808c05db9edd25179fe231f04

Observation 4b96bcb2-3a70-4688-9f1e-fbac3f6bf4f7 · outbound

This paper cites Get to The Point: Summarization with Pointer-generator Networks.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Get to The Point: Summarization with Pointer-generator Networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.442967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:7ea06eabdd3675c2a2897a116ba62d82852af5ca01003e706542d010da1f8120

Observation 20547573-269e-4efc-a69e-3eb014b69054 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-gated Mixture-of-Experts Layer.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Outrageously Large Neural Networks: The Sparsely-gated Mixture-of-Experts Layer

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.408980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:43a7e4060ac85ff627debcb0742b1dece18cad76e6d01a748a7547112d23567f

Observation 06d6bb69-ca1b-428f-ab40-45e8f0c88523 · outbound

This paper cites S-LoRA: Serving Thousands of Concurrent Lora Adapters.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding S-LoRA: Serving Thousands of Concurrent Lora Adapters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.393962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:67958f8636f849b07c174aca90b41966204aa637fc1fe68c1a1cdaf280fcb02e

Observation 60790074-9311-4c6c-9e43-64ef6c24e27e · outbound

This paper cites FlexGen: High-throughput Generative Inference of Large Language Models with a Single GPU.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding FlexGen: High-throughput Generative Inference of Large Language Models with a Single GPU

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.446198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:0518af31d131aafe761cd0b1f278699532d1cfa1a3b72c94e9f193be254d3e67

Observation 6229c04a-c89f-4a4b-bf40-d45ff7b91e61 · outbound

This paper cites Scaling LLM Test-time Compute Optimally can be More Effective Than Scaling Model Parameters.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling LLM Test-time Compute Optimally can be More Effective Than Scaling Model Parameters

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.479285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:97dd52f56537c425ea4a178f71eda7cffd70971756504c100b3425b498f986b9

Observation d8f3d2e9-fe74-4de4-9918-f22dc97f60b7 · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.647013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:81dde57026505fba6357014ec584b86343616f5b1e2d433131cbf770a2a45df0

Observation b92f4e30-bb73-49cd-b78a-f03a85dc7ef8 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Accelerating LLM Inference with Staged Speculative Decoding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.669024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:23fa4389d613c60e61a847d4df9cfbc0d0a42277efc3a3caadae8f6049f20f27

Observation abdc8994-7c0a-4b0c-bae2-fb8455d9c68e · outbound

This paper cites Blockwise Parallel Decoding for Deep Autoregressive Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Blockwise Parallel Decoding for Deep Autoregressive Models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.656890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:71cc0fc1ca07b205ade83d0d4ce38fdadbd9b831c12d7dc6563bf8274d4524a6

Observation ff9b6ab1-76cb-44d5-b0db-34c072dd1000 · outbound

This paper cites SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer Devices.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer Devices

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.652480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:d20a5de03906b965b7c2fae792b31f73cadd9625ac59b4c7183f36535d96aae8

Observation 226645b6-bd55-4f95-a397-6a318d1a2b47 · outbound

This paper cites No Language Left Behind: Scaling Human-centered Machine Translation.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding No Language Left Behind: Scaling Human-centered Machine Translation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.682631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:4077aa1c4b7b2a4b7d0525853d10b3fdd917359c1c8b851ed0e2af052c21bfa4

Observation e954ce5e-ffd4-4a18-967d-19b206295255 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.686123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:05f7a23e4f7a09f8b637447e79d1345ea96906643128448afb18975b22b43160

Observation 4d9233e4-5e1d-4f94-aea8-8666986847d9 · outbound

This paper cites APTMoE: Affinity-aware Pipeline Tuning for MoE Models on Bandwidth-constrained GPU Nodes.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding APTMoE: Affinity-aware Pipeline Tuning for MoE Models on Bandwidth-constrained GPU Nodes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.466586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:01334b1b0d740c1dd09261dd8a6ea297b4e9f5dfd72814716e7a182d51c0191b

Observation 94b21548-da04-47bf-acc5-9c6098344bc2 · outbound

This paper cites HuggingFace’s Transformers: State-of-the-art Natural Language Processing.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding HuggingFace’s Transformers: State-of-the-art Natural Language Processing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.449500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:6dfa9a3d16c17fe5c4ff015d6fa16221b3f8b30fd8ce07507791922ce5c848cb

Observation 3538104c-2d4b-4b64-8f90-746e9fde1d4a · outbound

This paper cites {dLoRA}: Dynamically Orchestrating Requests and Adapters for{LoRA}{LLM} Serving.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding {dLoRA}: Dynamically Orchestrating Requests and Adapters for{LoRA}{LLM} Serving

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.695468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:04d5ef9b0697697e204ae374edd68fd2764aa3fb0385194e1b5857ba70848607

Observation 14eab501-0dd4-4503-ac73-39a6295c2de9 · outbound

This paper cites EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.664824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:0c8b3336c61c3395061f77272f0d9a3aa140ad1ad8ba0d288dbb380ba0b62c16

Observation 9cb31b67-b91f-4c60-81d2-22040506c662 · outbound

This paper cites EdgeMoE: Fast On-device Inference of MoE-based Large Language Models.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding EdgeMoE: Fast On-device Inference of MoE-based Large Language Models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.661084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:b6248973b6679e74920f3d8d05c6e11a10c478aafb914a2049c64f3ff4e782ea

Observation 317ec3a9-c46a-4606-b4f5-0a5642e00339 · outbound

This paper cites MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.527756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:4ed045d4bbbab155be645783657c8a2d91f9b7661056df03fbc7261d85df77e6

Observation 591e4733-b9f1-4fbe-859e-8518bda0c982 · outbound

This paper cites Orca: A Distributed Serving System for{Transformer-based}Generative Mod- els.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Orca: A Distributed Serving System for{Transformer-based}Generative Mod- els

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.519934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:432b661460f653c41f414f737357128d0d80d84baea0252b495843a69b32b757

Observation c2a44426-12d3-4750-90a2-5704b2b52f10 · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:19:46.524136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:a47ac99c3b9fdf0a6a3050645ee855626ee251b0cd66199ede809c7c8c15349f

Observation 1817edb4-684e-4db7-b8d0-c049b052165a · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self- speculative decoding.

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self- speculative decoding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T14:21:51.675824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:30.250542Z digest=sha256:e8b3b38f0cd6d0db54edcf3959cbcc2914b22bf3ae10e7835fcd90a3f075aa00

Pith citing papers

Observation 49d73e67-0429-40a3-8a4e-9807b34d1f08 · inbound

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference cites this paper.

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:37.509164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:37.509164Z digest=sha256:5777ec57011a3b4047cc7ac7b6e5eac64a11c7cd153393a4c000f85545e983ed