Pith. sign in

Paper Citation Record · LEDGER

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

As of 22 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2411.11217.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11217 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:53:58.740609Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:17.763102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T23:00:19.435484Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8da072c8-5949-4691-b6a2-fef7e5d32b8c · outbound

This paper cites Flashinfer: Kernel library for llm serving.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flashinfer: Kernel library for llm serving

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.210066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.569013Z digest=sha256:d9d7101bbdd956612779082339cbeb689fdf7950ffb4e4b563a83be38d28df21

Observation fa5341d7-fbc9-4a3d-bbef-2b1e985a54e5 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.572830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.572830Z digest=sha256:c1aa5b1e210bd83f281b177138dc75e8cb8e48a84eee9a3fa5e79f80d8fbc848

Observation 009a2d51-8c25-4fdc-a176-d01a4b12d137 · outbound

This paper cites Llm in a flash: Efficient large language model inference with limited memory, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm in a flash: Efficient large language model inference with limited memory, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.201165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.576262Z digest=sha256:80c7fb1eafe6c99df4209beb751bd377ee3a1d68caffd6a792a0eb6dc2894cbd

Observation 9967325f-6af9-423a-b845-3c356f178664 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.579622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.579622Z digest=sha256:2712481d6d1dbbca38f069f6b491694787a8082b5a885e3146ac488e585a5a9a

Observation b0c8aec7-1d4f-4c0a-95cd-375f17aa571f · outbound

This paper cites Accelerating large language model decoding with speculative sampling, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating large language model decoding with speculative sampling, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.582688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.582688Z digest=sha256:7d498316bbd1b3d255dc342b9f8dc3ce269c8764e064c1b802e3ee831181492f

Observation 6f2e38fb-5ad5-4bbe-ab60-b9f11d7fcc7f · outbound

This paper cites Lifelong language pretraining with distribution-specialized experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Lifelong language pretraining with distribution-specialized experts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.181927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.585827Z digest=sha256:8496fcfa854cf187ff34b3df1ff9b4c0098ed72c8ecf175a24bc6332249a308d

Observation 1a75a12b-1a28-4e22-ac01-7b08d93f3396 · outbound

This paper cites Spreadsheetcoder: Formula prediction from semi-structured context.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Spreadsheetcoder: Formula prediction from semi-structured context

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.172382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.589148Z digest=sha256:f52cafdecfd701fd171b90ea7cfe524d0ae658661712308f054b780a1739ac36

Observation cb96245e-f471-4cfe-8707-587d6323d85d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.592078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.592078Z digest=sha256:cc693cbe4d6a2f0439eddd89ced5a41b3deb2cef1dd3b8f09437512d15eb4d20

Observation 8ce5be6e-b065-41b1-88db-eda92ce41cc2 · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Generating long sequences with sparse transformers, 2019

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.595355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.595355Z digest=sha256:2ee77d11276ca18b9cba258b4d0156e8ee92791a3e64948250360955d1af88f4

Observation beb32e99-611c-4b2a-b42a-16685f72a4aa · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.598074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.598074Z digest=sha256:38963f54b826594262e1acbbb7e63a9f0561c66a067609cbf8d4e122c0c1b677

Observation 7a11b14a-0f39-440e-86f8-c9814b8fa096 · outbound

This paper cites FlashAttention-2: Faster attention with better parallelism and work partitioning.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs FlashAttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.158305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.601425Z digest=sha256:d0da3800b6efe4147e161f02cf02c015250a6f651b9f01ed95568327674571e9

Observation 68aa2d8d-7e73-42d5-8009-6502458b876e · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.149069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.604536Z digest=sha256:c93c9dc68101e4c507f7bec8e27c0acdd3d5ad57411dcb024831c6183c77c0ff

Observation 7cb30671-87f3-47e3-a01a-61078df5b0cb · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Glam: Efficient scaling of language models with mixture-of-experts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.140456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.607616Z digest=sha256:531b1278b8531e2f75ea3a6c0b39a6d76abb6aaf755a40025bd6165ac98e19ba

Observation 41d24f27-b2cf-4e63-88ba-f6b921fbf4ff · outbound

This paper cites The Llama 3 Herd of Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.610574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.610574Z digest=sha256:b95eae0d4d3a03e6def2ec04f89e81e304d349e873a2b240c6e43a91f4e4f805

Observation 5eac09d1-0975-4c86-9970-0e15d2df344a · outbound

This paper cites Fast inference of mixture-of-experts language models with offloading, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference of mixture-of-experts language models with offloading, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.131718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.613730Z digest=sha256:cbedab45769fc554ba238a6cfe9e972351f75fafccbbbf08c9f0d4ddbdf7af06

Observation 8fe4204f-144d-4044-8fdc-8b1f16d15c1d · outbound

This paper cites Dap- ple: A pipelined data parallel approach for training large models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Dap- ple: A pipelined data parallel approach for training large models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.122947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.616760Z digest=sha256:43a5261b62132e9641431c33abb4807ffcd98432f3ae3b3bdd4181bbd2feb8d4

Observation 43ead05a-ff96-4587-b4b1-ce2d915667f8 · outbound

This paper cites Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fastdecode: High-throughput gpu-efficient llm serving using heterogeneous pipelines, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.113719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.619628Z digest=sha256:7e87fd9ff087e6ae886b994d18f4de1d932105d8f15a5159e01ea47425dbb3d5

Observation 4e0e2679-1602-4ac0-a4db-611b73ba8ab5 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.622530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.622530Z digest=sha256:b2a865d0bc8900b6c87326024782694972113a5fc8a202252cb625369316a7f7

Observation eb5a6805-341d-47e0-ac34-f47d35eca29d · outbound

This paper cites Hugging face accelerate.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hugging face accelerate

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.098032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.625275Z digest=sha256:770a379d9f05acd1a69a70c50afd0d4e0d1c590a55376a4e44df3670313f3718

Observation 42c246a0-b8cd-48f7-b8f5-c3d9d776cfe8 · outbound

This paper cites Intel(r) oneapi math kernel library (onemkl).

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Intel(r) oneapi math kernel library (onemkl)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.088910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.628410Z digest=sha256:7203cd0e48d1c0d7eda6311f4c0e6614bb711ff47c1df5bdf9b4b721e2b570e4

Observation ed49b8cd-f4f6-43d6-9ad9-1a9a7b7870d4 · outbound

This paper cites Jacobs, Michael I.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Jacobs, Michael I

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.079406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.631285Z digest=sha256:ac3dabb10d2dc65027bbd3230f6a4ae64123d4d7168c49164f39898f0348184e

Observation 2cddb5ee-09a3-4e40-a6ba-974e76d19948 · outbound

This paper cites Mixtral of Experts.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.634169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.634169Z digest=sha256:bd2a129c1830e070a382137f61d2914eacc841bd944dee96696da3c92162ea57

Observation d529eedc-00ec-4315-ba86-75a4ad3b9ed9 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hierarchical mixtures of experts and the em algorithm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.637627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.637627Z digest=sha256:88960f8c32a41cefbdaf5278d893740f55c4d91ed3a4fa5159789d1fbf3bf1db

Observation d2e2f8a2-b873-41f7-ae23-52fc6bd3699d · outbound

This paper cites Fu, Christo- pher Ré, and Azalia Mirhoseini.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fu, Christo- pher Ré, and Azalia Mirhoseini

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.064941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.640464Z digest=sha256:37989767f4c06e45221606521da627989663e4e5942adc353219f7e3ab77cae6

Observation 27c18c25-64a7-4d95-96f2-a8546200e836 · outbound

This paper cites Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fiddler: Cpu- gpu orchestration for fast inference of mixture-of-experts models, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.055522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.643198Z digest=sha256:7b9191aa5958d9efd99bf0697b8de350f8432b676d77e4a7294959ebd771ef1f

Observation bf8ed4c5-ecf2-4a2a-a662-4c00a31e70b5 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient memory management for large language model serving with pagedattention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.646145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.646145Z digest=sha256:a1482a2c799973be0ae158a4d1c0a8f8444e9b76df5194e7d4559b794c1b3291

Observation b60c6726-b44a-453b-8d85-dd8ad830cae8 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.649016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.649016Z digest=sha256:8915d81a174260bfd2b25d9aa162e81e8d92fb906096b0cf342ae080dcd96a2b

Observation 50cc8e15-4e9d-43c8-bd38-bf08e7a406a5 · outbound

This paper cites Fast inference from transformers via speculative decoding, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Fast inference from transformers via speculative decoding, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.652418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.652418Z digest=sha256:c0fb5c9a20d12c1b85425d76ccd4d1b05ac2dfe2efa6186729624c98fee8d5a2

Observation 854a1e9f-aa4b-4011-9d9e-4a05206611dc · outbound

This paper cites Holistic Evaluation of Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Holistic Evaluation of Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.655265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.655265Z digest=sha256:02104baa96d8c690777089b3d234737b0e3352bf203f81f3854ceb9e8db5dae4

Observation 0c52fa16-3192-49a1-959a-6a51a5ec991c · outbound

This paper cites Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.036481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.659136Z digest=sha256:4a26679a7599502b01b45cad22f9bf20c920fa66d95d3d8f84bcc6c512bdff7f

Observation ab0a877a-ac83-429c-9448-a8b5ceeeed51 · outbound

This paper cites Gonzalez, Ion Stoica, and Matei Zaharia.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Ion Stoica, and Matei Zaharia

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.661945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.661945Z digest=sha256:a95fc2fc8472584344c676919162707aa5ed617f8326d98614be5446894582c4

Observation 0ddcab18-85f7-4c14-b8f9-f58b2fed74fa · outbound

This paper cites https://mistral.ai/news/mixtral-8x22b/, April 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs https://mistral.ai/news/mixtral-8x22b/, April 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.022040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.664995Z digest=sha256:f695e8f3ea9c985f5eec5c6930de5e2e6b7f57cb24838c10de2766e8b5f4009f

Observation 98ae3737-33c3-4c1c-acb9-a09cb2ce5d7a · outbound

This paper cites Can Foundation Models Wrangle Your Data?.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Can Foundation Models Wrangle Your Data?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.667936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.667936Z digest=sha256:e246e29abe0b943635253036f0176c84f8bd0348138cc231cd4f453fec997512

Observation caa36d69-df5c-4da8-ba84-fd06f4860082 · outbound

This paper cites Pipedream: generalized pipeline parallelism for dnn train- ing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pipedream: generalized pipeline parallelism for dnn train- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:59.013260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.671082Z digest=sha256:55ba2948d8f243b1d88a34cb6cfd67e02b6367e073fdcf6203d205dd2738a8e1

Observation 0f061996-066e-48ae-b67f-4d82d6e4132a · outbound

This paper cites Efficient large- scale language model training on gpu clusters using megatron-lm.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficient large- scale language model training on gpu clusters using megatron-lm

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.673854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.673854Z digest=sha256:28e0036798fc2d6f4a84962f3d7dd3977fcb7f0b6556778e48033cb61a980ee0

Observation bbcac705-de64-4bce-89a2-325d5c149bab · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Pytorch: An imperative style, high- performance deep learning library

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.676884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.676884Z digest=sha256:1136c5fdb739d2d4099ac6ff07f98716713ad55a00d259f920039f122ae4c2b1

Observation a48a4139-6f90-4c9a-b9df-fa69800813b2 · outbound

This paper cites Efficiently scaling transformer inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Efficiently scaling transformer inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.994340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.679584Z digest=sha256:582580de1444cd6402dbf30e53e69e802b179cc099112bab7b21997c86910f68

Observation 3c40886d-2f0c-49a8-ba95-f17e086b9b8d · outbound

This paper cites Int4 decoding gqa cuda optimizations for llm inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Int4 decoding gqa cuda optimizations for llm inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.985437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.682282Z digest=sha256:9e5da6fc3843226039a5cbd1cb104a7fb2eebc889ba44ac42deb5964e344ea6c

Observation 27ca07b3-ed1d-4218-815e-0adcd9ae4157 · outbound

This paper cites Accelerating transformer inference for translation via parallel decod- ing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Accelerating transformer inference for translation via parallel decod- ing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.976464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.685155Z digest=sha256:ca1cf32687e1a7ac10cbdd41674e4829b307454e55cc83fe43547ae472cc91af

Observation 529a1180-301d-4b3f-b41a-d892698429ce · outbound

This paper cites BLOOM: A 176B-Parameter Open-Access Multilingual Language Model.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.688100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.688100Z digest=sha256:0643a71d26206b75a99ea73a646664af780d1e47d630769008f611fc305f706a

Observation c9d347d1-5518-461c-afb6-97737c26e6c1 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.691288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.691288Z digest=sha256:30f2b9cbb92d3cb00d165944542263e2cedcea1c1a02f7346aaaa699b19b9b77

Observation d3d95062-058a-4ec2-9761-03a9e6ee0b21 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.694423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.694423Z digest=sha256:1756e07c43de4578785fea5c7c6dc794b83be99c8ad680dd86f9e9774094b9e4

Observation fe41c939-776a-4a03-ab5d-d5bde01a1396 · outbound

This paper cites Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.961970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.697301Z digest=sha256:7c3f5a269257bee1911494048e84ceba79f1eea3529d8fc7a245cc766b7c9152

Observation 01f3f65f-3a72-4fab-81e1-462b883fb1c3 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models, 2018.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Blockwise parallel decoding for deep autoregressive models, 2018

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.951762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.700423Z digest=sha256:68145a3a01d0cb24bec3a6c8d775e2d3500f5088da61b25973c405b31bfd2df0

Observation 65f40479-2123-4938-a987-30448f8f6236 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.703312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.703312Z digest=sha256:df92f933b4716c89bd71e3b0e82b779b02f0b6aaa1860c74ca56e46a8110bd41

Observation bc4e239d-4295-49b9-a2d4-0f7d79d70e05 · outbound

This paper cites Introducing dbrx: A new state-of-the-art open llm, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Introducing dbrx: A new state-of-the-art open llm, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.942656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.706692Z digest=sha256:2aadbede024d53f1f84a7853923ebc36c287984bd3384b6f0e4e8f0463d2cc3c

Observation 8512b705-473b-4585-8a91-61a75c353df6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs LLaMA: Open and Efficient Foundation Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.709940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.709940Z digest=sha256:b6034cbf2b87184205582a20e8dd0da2f124f92a2f146faf22da9bc7d5f834e4

Observation 82f6bb67-35da-44ff-9b32-512a4dd9306a · outbound

This paper cites Patterson.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Patterson

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.933257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.712972Z digest=sha256:304dba0e898bcaabc8f2075084573997d5f12dc27e26714e11ba12d194be88c6

Observation 426a9ebc-8270-4434-89e8-8c451dcce4e3 · outbound

This paper cites Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Hetegen: Efficient heterogeneous parallel inference for large language models on resource-constrained devices

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.924325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.715786Z digest=sha256:926d95727d483b0254a552022f87fcfb725f567f987d99aa578ea94fc571fa6b

Observation 4d0a4246-cd10-47af-a22e-b2ef80cf4cbb · outbound

This paper cites Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Moe- infinity: Activation-aware expert offloading for efficient moe serving, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.914493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.719047Z digest=sha256:100d45fbf345c2e1d47ac6b5a2c11876958518a489f9b5ca3672271f58cab641

Observation 92fdc5f8-79f9-45cc-8fce-0ead5940fce9 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Orca: A distributed serving system for {Transformer-Based} generative models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.722241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.722241Z digest=sha256:c5c53637cfe41220208be57e88b73bb4fab95c450674cd7dd51089e5fc9d1207

Observation 091b585a-d928-4c63-9bef-ad4582ae5f35 · outbound

This paper cites Llm inference unveiled: Survey and roofline model insights, 2024.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Llm inference unveiled: Survey and roofline model insights, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.725433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.725433Z digest=sha256:ef55c88d89ca91fc972598881f486c36bc143393944a3888e7b877a10bbe43d4

Observation 96114ba4-97b0-4a2e-82ab-d502a5f31661 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs OPT: Open Pre-trained Transformer Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.728292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.728292Z digest=sha256:d4073013330a41a696df81625bf76ee5ec37d1a445eabe5b130621fc9afca4bc

Observation edf02297-787e-4ff1-b456-6570380d7f81 · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative in- ference of large language models.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs H2o: Heavy-hitter oracle for efficient generative in- ference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.894605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.731529Z digest=sha256:38166c2968a7c97feefb9007401e70ccd6d0cfb65370958692884b2c44c27a45

Observation 91bdd8b7-1edd-4155-b51d-1767b4b4eae9 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.885475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.734590Z digest=sha256:dd018cafe7992a83e0a49399b5af33c900247041612a964e5366c459f3b8bc6e

Observation 07fb519d-789a-43b0-a989-ce9514a7e432 · outbound

This paper cites Gonzalez, Clark Barrett, and Ying Sheng.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Gonzalez, Clark Barrett, and Ying Sheng

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:53:58.737467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:53:58.737467Z digest=sha256:baf1169a64a1e54affa5f62970875a5f636baf36e03e72d47c0b1743985777ce

Observation 859ce6be-be1c-4784-b280-8697531af128 · outbound

This paper cites Mixture-of- experts with expert choice routing.

MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Mixture-of- experts with expert choice routing

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:53:58.870916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T18:53:58.740609Z digest=sha256:c69857a15e6b793f9641bc455ead3f7092d6155ae59eedba09c4999c570dab17

Pith citing papers

Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · inbound

FloE: On-the-Fly MoE Inference on Memory-constrained GPU cites this paper.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:00:19.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.763102Z digest=sha256:d2c44f9a0c53f53b1598a4294b7207964b9370952be07f388a815537c0c7d22e