Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:38:30.250542Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 1 inbound Pith citation observation for arXiv:2604.10152.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:38:30.250542Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T15:02:37.509164Z
A source-named dated measurement, never combined with another source.
Source: cited_works
59 of 59 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b566f263-596f-4861-b051-ab2de4d0446a · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37c68beb-f224-4c9a-ab1f-2d5145135533 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding GQA: Training Generalized Multi-query Transformer Models from Multi-Head Checkpoints
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0355760d-2113-457e-ab29-17d67028646e · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Un- precedented Scale
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 405fb333-0b98-4080-bbdf-ff93d89518a9 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Findings of the 2014 Workshop on Statistical Ma- chine Translation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ef2a298-8e66-4a77-aff4-bf70cb0eac91 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Language Models are Few-shot Learners
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 691ceb76-dfb3-45f2-899e-dd57412a02f0 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a457f1cb-78c2-43cb-8439-7b8517015f3b · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sam- pling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9d54a26b-0d96-48f0-ad37-9915d17a105f · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Punica: Multi-tenant LoRA Serving
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e90832ac-51a3-404d-9e9c-7b7fb628efa9 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 63ef1615-8a70-43f7-8ea9-5e3381c3c2b8 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding DeepSeek- R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e2c2884-dad3-4c97-a583-5128aa759f66 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-speculative Decoding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 316409e3-8e20-4ea0-b8a0-bf173dfa743f · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Fast Inference of Mixture-of-Experts Lan- guage Models with Offloading
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b469f28-fafa-4ff3-a3c1-53de9ea7df4f · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 00b850ca-407a-49c9-9aa6-e4ec8a787b66 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Gemini: A Family of Highly Capable Multimodal Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fc612821-20af-46cb-aad8-e3437c17930c · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Gemini 1.5: Unlocking Multimodal Understanding Across Mil- lions of Tokens of Context
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b2bbd69-c14a-4a2a-856a-d4f77c5cc20b · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Teaching Machines to Read and Comprehend
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c20af1fb-17fe-4ce9-8b5f-51c9fe90cec8 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Training Compute-Optimal Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc0d8b24-2a96-449d-81d9-fdbec2afffc9 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Speed: Speculative Pipelined Execution for Efficient Decoding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3273372c-dfa3-447e-ae51-585c7ab9126c · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0e7d6da9-72ab-479a-a2cc-53779fcf006c · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Hugging Face Accelerate
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca1d037e-de87-406d-b988-b0b447cb7f0b · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Tutel: Adaptive Mixture-of-Experts at Scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 380ee061-47b4-441b-b21d-bb33afbbaddf · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Pre-gated MoE: An Algorithm-system Co-design for Fast and Scal- able Mixture-of-Expert Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9f212b4f-49d4-4af5-872d-fe58e011a35e · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Mixtral of Experts
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3207587f-8095-457b-84b0-abf69c6fac66 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling Laws for Neural Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45d4180b-0a65-4969-9a44-dd2e13215463 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling Laws for Neural Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44fd884e-aff6-4fab-ba46-55f5d7f7203e · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 57bc9b2f-70d8-4dc6-9d71-2ee6c47007c9 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e373bc3f-c164-4c10-9f82-151aeb10e8e2 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Fast Inference from Trans- formers via Speculative Decoding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e0320d12-49a5-4b9f-9b93-ae742bba1fcc · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding EAGLE: Speculative Sam- pling Requires Rethinking Feature Uncertainty
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f0aa89f-3f17-47d7-a6cd-da36cc2aedd4 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Llama 4 and Multimodal AI: Expanding Intelligence Across Modalities
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc43ec45-b1d8-4d1f-936b-c447b47639dc · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Specinfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c8b8c081-b2bd-41c2-9e38-d86730793e35 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding GPT-4 Technical Report
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a1297ea4-aa9c-4bfb-ba42-6aba885d4899 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Characterizing Power Management Opportunities for LLMs in the Cloud
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77769e0d-fda0-465a-ae90-0b268f419368 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Splitwise: Efficient Generative LLM Inference using Phase Splitting
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64822a6a-ed74-421d-88a9-4a9cbe0ab2a0 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling Language Mod- els: Methods, Analysis & Insights from Training Gopher
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b56a02c1-932a-4554-9718-f0f37cbe7fba · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Deepspeed-MoE: Advancing Mixture-of- Experts Inference and Training to Power Next-generation AI Scale
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 86722ba3-7bc6-4f07-ab14-6c430ea7080f · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Zero- Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1fff10f7-fcb5-465e-be75-3a766bab5524 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding ZeRO-Offload: Democratizing Billion-scale Model Training
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7cff7c62-b0a1-45bf-ad7d-9d0bedf159fb · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Speculative decoding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b96bcb2-3a70-4688-9f1e-fbac3f6bf4f7 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Get to The Point: Summarization with Pointer-generator Networks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 20547573-269e-4efc-a69e-3eb014b69054 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Outrageously Large Neural Networks: The Sparsely-gated Mixture-of-Experts Layer
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06d6bb69-ca1b-428f-ab40-45e8f0c88523 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding S-LoRA: Serving Thousands of Concurrent Lora Adapters
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 60790074-9311-4c6c-9e43-64ef6c24e27e · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding FlexGen: High-throughput Generative Inference of Large Language Models with a Single GPU
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6229c04a-c89f-4a4b-bf40-d45ff7b91e61 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Scaling LLM Test-time Compute Optimally can be More Effective Than Scaling Model Parameters
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8f3d2e9-fe74-4de4-9918-f22dc97f60b7 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b92f4e30-bb73-49cd-b78a-f03a85dc7ef8 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Accelerating LLM Inference with Staged Speculative Decoding
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation abdc8994-7c0a-4b0c-bae2-fb8455d9c68e · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Blockwise Parallel Decoding for Deep Autoregressive Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ff9b6ab1-76cb-44d5-b0db-34c072dd1000 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding SpecExec: Massively Parallel Speculative Decoding For Interactive LLM Inference on Consumer Devices
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 226645b6-bd55-4f95-a397-6a318d1a2b47 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding No Language Left Behind: Scaling Human-centered Machine Translation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e954ce5e-ffd4-4a18-967d-19b206295255 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding LLaMA: Open and Efficient Foundation Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4d9233e4-5e1d-4f94-aea8-8666986847d9 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding APTMoE: Affinity-aware Pipeline Tuning for MoE Models on Bandwidth-constrained GPU Nodes
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 94b21548-da04-47bf-acc5-9c6098344bc2 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3538104c-2d4b-4b64-8f90-746e9fde1d4a · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding {dLoRA}: Dynamically Orchestrating Requests and Adapters for{LoRA}{LLM} Serving
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 14eab501-0dd4-4503-ac73-39a6295c2de9 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9cb31b67-b91f-4c60-81d2-22040506c662 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding EdgeMoE: Fast On-device Inference of MoE-based Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 317ec3a9-c46a-4606-b4f5-0a5642e00339 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 591e4733-b9f1-4fbe-859e-8518bda0c982 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Orca: A Distributed Serving System for{Transformer-based}Generative Mod- els
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c2a44426-12d3-4750-90a2-5704b2b52f10 · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1817edb4-684e-4db7-b8d0-c049b052165a · outbound
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self- speculative decoding
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49d73e67-0429-40a3-8a4e-9807b34d1f08 · inbound
DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.