Pith. sign in

Paper Citation Record · LEDGER

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

As of 22 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2411.11217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11217 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:53:58.740609Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:17.763102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T23:00:19.435484Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8da072c8-5949-4691-b6a2-fef7e5d32b8c · outbound

This paper cites Flashinfer: Kernel library for llm serving.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flashinfer: Kernel library for llm serving

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.210066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.569013Z digest=sha256:8607d23582cc956d8c3c4a33dd54e29c3599e954412436356dcaefd87390d7cf

Observation fa5341d7-fbc9-4a3d-bbef-2b1e985a54e5 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.572830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.572830Z digest=sha256:c1aa5b1e210bd83f281b177138dc75e8cb8e48a84eee9a3fa5e79f80d8fbc848

Observation 009a2d51-8c25-4fdc-a176-d01a4b12d137 · outbound

This paper cites Llm in a flash: Efficient large language model inference with limited memory, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm in a flash: Efficient large language model inference with limited memory, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.201165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.576262Z digest=sha256:e5908f394bef457e847beebf589e5960e87c2ccda7837f2a7288d9434b9f7b3c

Observation 9967325f-6af9-423a-b845-3c356f178664 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.579622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.579622Z digest=sha256:2712481d6d1dbbca38f069f6b491694787a8082b5a885e3146ac488e585a5a9a

Observation b0c8aec7-1d4f-4c0a-95cd-375f17aa571f · outbound

This paper cites Accelerating large language model decoding with speculative sampling, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating large language model decoding with speculative sampling, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.582688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.582688Z digest=sha256:7d498316bbd1b3d255dc342b9f8dc3ce269c8764e064c1b802e3ee831181492f

Observation 6f2e38fb-5ad5-4bbe-ab60-b9f11d7fcc7f · outbound

This paper cites Lifelong language pretraining with distribution-specialized experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Lifelong language pretraining with distribution-specialized experts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.181927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.585827Z digest=sha256:3d0f597066e112f2bcc685748413420b541240f8c9b0b38767ff8e71b3b86e20

Observation 1a75a12b-1a28-4e22-ac01-7b08d93f3396 · outbound

This paper cites Spreadsheetcoder: Formula prediction from semi-structured context.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Spreadsheetcoder: Formula prediction from semi-structured context

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.172382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.589148Z digest=sha256:837067b1682c2419e70bdc29d5d0488aec528c03a2b58e40daaf3722fbadac9c

Observation cb96245e-f471-4cfe-8707-587d6323d85d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.592078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.592078Z digest=sha256:cc693cbe4d6a2f0439eddd89ced5a41b3deb2cef1dd3b8f09437512d15eb4d20

Observation 8ce5be6e-b065-41b1-88db-eda92ce41cc2 · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Generating long sequences with sparse transformers, 2019

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.595355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.595355Z digest=sha256:2ee77d11276ca18b9cba258b4d0156e8ee92791a3e64948250360955d1af88f4

Observation beb32e99-611c-4b2a-b42a-16685f72a4aa · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.598074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.598074Z digest=sha256:38963f54b826594262e1acbbb7e63a9f0561c66a067609cbf8d4e122c0c1b677

Observation 7a11b14a-0f39-440e-86f8-c9814b8fa096 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.158305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.601425Z digest=sha256:049c94622a87e4b18c9d1f102f327e7c88758225e6d234441c917d27b6fa8e7d

Observation 68aa2d8d-7e73-42d5-8009-6502458b876e · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.149069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.604536Z digest=sha256:d5e1d687d0bc468175a9df1f2470158bc94dfa365c98950d109f6dbfa27e4878

Observation 7cb30671-87f3-47e3-a01a-61078df5b0cb · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Glam: Efficient scaling of language models with mixture-of-experts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.140456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.607616Z digest=sha256:9df2021c2aed5c7c5e11df93c2f75f4cf2253da2df96ac529ce6c9fe69e56b47

Observation 41d24f27-b2cf-4e63-88ba-f6b921fbf4ff · outbound

This paper cites The Llama 3 Herd of Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.610574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.610574Z digest=sha256:b95eae0d4d3a03e6def2ec04f89e81e304d349e873a2b240c6e43a91f4e4f805

Observation 5eac09d1-0975-4c86-9970-0e15d2df344a · outbound

This paper cites Fast inference of mixture-of-experts language models with offloading, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference of mixture-of-experts language models with offloading, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.131718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.613730Z digest=sha256:121004928cc19f8d658fa25909fba6f597d6bfb67890f3d2bf7456b808e821bc

Observation 8fe4204f-144d-4044-8fdc-8b1f16d15c1d · outbound

This paper cites Dap- ple: A pipelined data parallel approach for training large models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Dap- ple: A pipelined data parallel approach for training large models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.122947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.616760Z digest=sha256:e4ed8c1b909503ca4bbdd1c017b6e4be75188c4d3f761bc69f779e6bf68e5dfe

Observation 43ead05a-ff96-4587-b4b1-ce2d915667f8 · outbound

This paper cites Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.113719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.619628Z digest=sha256:27a9cf4bc8801106fab6362ce30ea10dacbec84ebaa326bd6ac1a77f855f2139

Observation 4e0e2679-1602-4ac0-a4db-611b73ba8ab5 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.622530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.622530Z digest=sha256:b2a865d0bc8900b6c87326024782694972113a5fc8a202252cb625369316a7f7

Observation eb5a6805-341d-47e0-ac34-f47d35eca29d · outbound

This paper cites Hugging face accelerate.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hugging face accelerate

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.098032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.625275Z digest=sha256:be89929028c8280b0f828e0141bbbf974e17b5a3619e50531082e757134d2e43

Observation 42c246a0-b8cd-48f7-b8f5-c3d9d776cfe8 · outbound

This paper cites Intel(r) oneapi math kernel library (onemkl).

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Intel(r) oneapi math kernel library (onemkl)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.088910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.628410Z digest=sha256:880032c7df537ef862eaca7873b7df6de19406ddda728010db515b6ecf4a39eb

Observation ed49b8cd-f4f6-43d6-9ad9-1a9a7b7870d4 · outbound

This paper cites Jacobs, Michael I.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Jacobs, Michael I

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.079406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.631285Z digest=sha256:776305ddfcd461dd1a60e0d4dbf41eeb2d7f9bfa54afbf5eedc5921a2dfbf858

Observation 2cddb5ee-09a3-4e40-a6ba-974e76d19948 · outbound

This paper cites Mixtral of Experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.634169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.634169Z digest=sha256:bd2a129c1830e070a382137f61d2914eacc841bd944dee96696da3c92162ea57

Observation d529eedc-00ec-4315-ba86-75a4ad3b9ed9 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hierarchical mixtures of experts and the em algorithm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.637627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.637627Z digest=sha256:88960f8c32a41cefbdaf5278d893740f55c4d91ed3a4fa5159789d1fbf3bf1db

Observation d2e2f8a2-b873-41f7-ae23-52fc6bd3699d · outbound

This paper cites Fu, Christo- pher Ré, and Azalia Mirhoseini.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Christo- pher Ré, and Azalia Mirhoseini

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.064941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.640464Z digest=sha256:942956b9aed65a32b78acfdc534a9005e0d7589340da2010108e86bfa8e992f2

Observation 27c18c25-64a7-4d95-96f2-a8546200e836 · outbound

This paper cites Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.055522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.643198Z digest=sha256:fc4b34a955f07a24a9b5585eb3fa26b183b1c44d0eec5b8c741009c0d647dcfe

Observation bf8ed4c5-ecf2-4a2a-a662-4c00a31e70b5 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient memory management for large language model serving with pagedattention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.646145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.646145Z digest=sha256:a1482a2c799973be0ae158a4d1c0a8f8444e9b76df5194e7d4559b794c1b3291

Observation b60c6726-b44a-453b-8d85-dd8ad830cae8 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.649016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.649016Z digest=sha256:8915d81a174260bfd2b25d9aa162e81e8d92fb906096b0cf342ae080dcd96a2b

Observation 50cc8e15-4e9d-43c8-bd38-bf08e7a406a5 · outbound

This paper cites Fast inference from transformers via speculative decoding, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference from transformers via speculative decoding, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.652418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.652418Z digest=sha256:c0fb5c9a20d12c1b85425d76ccd4d1b05ac2dfe2efa6186729624c98fee8d5a2

Observation 854a1e9f-aa4b-4011-9d9e-4a05206611dc · outbound

This paper cites Holistic Evaluation of Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Holistic Evaluation of Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.655265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.655265Z digest=sha256:02104baa96d8c690777089b3d234737b0e3352bf203f81f3854ceb9e8db5dae4

Observation 0c52fa16-3192-49a1-959a-6a51a5ec991c · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.036481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.659136Z digest=sha256:82066413bf8f8cd28c35f6cec953eba9c180dc0d71dd3b67f5cdc913732afb0d

Observation ab0a877a-ac83-429c-9448-a8b5ceeeed51 · outbound

This paper cites Gonzalez, Ion Stoica, and Matei Zaharia.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Ion Stoica, and Matei Zaharia

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.661945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.661945Z digest=sha256:a95fc2fc8472584344c676919162707aa5ed617f8326d98614be5446894582c4

Observation 0ddcab18-85f7-4c14-b8f9-f58b2fed74fa · outbound

This paper cites https://mistral.ai/news/mixtral-8x22b/, April 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs https://mistral.ai/news/mixtral-8x22b/, April 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.022040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.664995Z digest=sha256:2760d1b2086b5e008c57790f888a9393f3d7a760542e75835ca3f09cb1e6009f

Observation 98ae3737-33c3-4c1c-acb9-a09cb2ce5d7a · outbound

This paper cites Can Foundation Models Wrangle Your Data?.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Can Foundation Models Wrangle Your Data?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.667936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.667936Z digest=sha256:e246e29abe0b943635253036f0176c84f8bd0348138cc231cd4f453fec997512

Observation caa36d69-df5c-4da8-ba84-fd06f4860082 · outbound

This paper cites Pipedream: generalized pipeline parallelism for dnn train- ing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pipedream: generalized pipeline parallelism for dnn train- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.013260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.671082Z digest=sha256:acc142629c5b5858d51aa280a9f01b4dda7d20c5907cab43b8f3afe1c55c45f4

Observation 0f061996-066e-48ae-b67f-4d82d6e4132a · outbound

This paper cites Efficient large- scale language model training on gpu clusters using megatron-lm.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient large- scale language model training on gpu clusters using megatron-lm

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.673854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.673854Z digest=sha256:28e0036798fc2d6f4a84962f3d7dd3977fcb7f0b6556778e48033cb61a980ee0

Observation bbcac705-de64-4bce-89a2-325d5c149bab · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pytorch: An imperative style, high- performance deep learning library

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.676884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.676884Z digest=sha256:1136c5fdb739d2d4099ac6ff07f98716713ad55a00d259f920039f122ae4c2b1

Observation a48a4139-6f90-4c9a-b9df-fa69800813b2 · outbound

This paper cites Efficiently scaling transformer inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficiently scaling transformer inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.994340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.679584Z digest=sha256:a79550ba02d3fbae4ce1bbfd24de982d2c805ed7750838aa9596a719421a178d

Observation 3c40886d-2f0c-49a8-ba95-f17e086b9b8d · outbound

This paper cites Int4 decoding gqa cuda optimizations for llm inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Int4 decoding gqa cuda optimizations for llm inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.985437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.682282Z digest=sha256:cf4ae336fb8c2994c89489fec3657716499fc2fae61d2d02eb8799285314f8e8

Observation 27ca07b3-ed1d-4218-815e-0adcd9ae4157 · outbound

This paper cites Accelerating transformer inference for translation via parallel decod- ing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating transformer inference for translation via parallel decod- ing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.976464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.685155Z digest=sha256:589b4215a3951bd270e1b8a6f76847077d8c35ce74dc718d62d83adb4d6307c2

Observation 529a1180-301d-4b3f-b41a-d892698429ce · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.688100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.688100Z digest=sha256:0643a71d26206b75a99ea73a646664af780d1e47d630769008f611fc305f706a

Observation c9d347d1-5518-461c-afb6-97737c26e6c1 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.691288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.691288Z digest=sha256:30f2b9cbb92d3cb00d165944542263e2cedcea1c1a02f7346aaaa699b19b9b77

Observation d3d95062-058a-4ec2-9761-03a9e6ee0b21 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.694423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.694423Z digest=sha256:1756e07c43de4578785fea5c7c6dc794b83be99c8ad680dd86f9e9774094b9e4

Observation fe41c939-776a-4a03-ab5d-d5bde01a1396 · outbound

This paper cites Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.961970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.697301Z digest=sha256:d8286b298c9e2645942143464e829be7f85701bb14e59a96136ae933ddf36ff6

Observation 01f3f65f-3a72-4fab-81e1-462b883fb1c3 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models, 2018.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Blockwise parallel decoding for deep autoregressive models, 2018

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.951762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.700423Z digest=sha256:902356721ac6dfb79448f2718af8206c2ced3d4292e43407e9f010474f85625c

Observation 65f40479-2123-4938-a987-30448f8f6236 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.703312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.703312Z digest=sha256:df92f933b4716c89bd71e3b0e82b779b02f0b6aaa1860c74ca56e46a8110bd41

Observation bc4e239d-4295-49b9-a2d4-0f7d79d70e05 · outbound

This paper cites Introducing dbrx: A new state-of-the-art open llm, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Introducing dbrx: A new state-of-the-art open llm, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.942656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.706692Z digest=sha256:1f6c665d8b5a4a5d46b687e81adb33c7818daf8a4994467ad37085a94aa157f8

Observation 8512b705-473b-4585-8a91-61a75c353df6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.709940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.709940Z digest=sha256:b6034cbf2b87184205582a20e8dd0da2f124f92a2f146faf22da9bc7d5f834e4

Observation 82f6bb67-35da-44ff-9b32-512a4dd9306a · outbound

This paper cites Patterson.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Patterson

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.933257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.712972Z digest=sha256:38f37ab4e354c031937ee27430353ddd3bca498f18eb4b6f6529d863a2dbc216

Observation 426a9ebc-8270-4434-89e8-8c451dcce4e3 · outbound

This paper cites Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.924325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.715786Z digest=sha256:e7438492990c687e23bd7f80a10c744730b727eef5d93ba09743edd85bbfcaab

Observation 4d0a4246-cd10-47af-a22e-b2ef80cf4cbb · outbound

This paper cites Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.914493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.719047Z digest=sha256:4f8f2d1e2e670a1a7309929b9bd15cdff28880cc9c5307a8486009fbe1b6fd61

Observation 92fdc5f8-79f9-45cc-8fce-0ead5940fce9 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Orca: A distributed serving system for {Transformer-Based} generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.722241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.722241Z digest=sha256:c5c53637cfe41220208be57e88b73bb4fab95c450674cd7dd51089e5fc9d1207

Observation 091b585a-d928-4c63-9bef-ad4582ae5f35 · outbound

This paper cites Llm inference unveiled: Survey and roofline model insights, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm inference unveiled: Survey and roofline model insights, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.725433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.725433Z digest=sha256:ef55c88d89ca91fc972598881f486c36bc143393944a3888e7b877a10bbe43d4

Observation 96114ba4-97b0-4a2e-82ab-d502a5f31661 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs OPT: Open Pre-trained Transformer Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.728292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.728292Z digest=sha256:d4073013330a41a696df81625bf76ee5ec37d1a445eabe5b130621fc9afca4bc

Observation edf02297-787e-4ff1-b456-6570380d7f81 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative in- ference of large language models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs H2o: Heavy-hitter oracle for efficient generative in- ference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.894605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.731529Z digest=sha256:756161b76d203218c23173954acef06ab3732b7254c8bc16a22261db78301e91

Observation 91bdd8b7-1edd-4155-b51d-1767b4b4eae9 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.885475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.734590Z digest=sha256:4e983b4bd2d5dc722749f3364a080b0f2a5dd4ce84da6f7f350030a802095448

Observation 07fb519d-789a-43b0-a989-ce9514a7e432 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Clark Barrett, and Ying Sheng

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.737467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.737467Z digest=sha256:baf1169a64a1e54affa5f62970875a5f636baf36e03e72d47c0b1743985777ce

Observation 859ce6be-be1c-4784-b280-8697531af128 · outbound

This paper cites Mixture-of- experts with expert choice routing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixture-of- experts with expert choice routing

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.870916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T18:53:58.740609Z digest=sha256:fdafd9425efc39ca99b974a37194d169f945d5d0cf5ae28caa883c80e5d12e8c

Pith citing papers

Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · inbound

FloE: On-the-Fly MoE Inference on Memory-constrained GPU cites this paper.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:00:19.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.763102Z digest=sha256:7fec6a41394372e68431b027ea8b82d9289353627f756ebd5deb0780758bb1f1