Pith. sign in

Paper Citation Record · LEDGER

FloE: On-the-Fly MoE Inference on Memory-constrained GPU

As of 22 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2505.05950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05950 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:00:18.175589Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-07T10:32:50.809236Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:36:25.890370Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 677af2a7-ab45-4cc1-a6cf-26fd2a0eb536 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.731611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.731611Z digest=sha256:ffa27e5ef7054000996f4873093ef772738238ff18bd0043d0b2964469129133

Observation bab1eeb2-937e-47af-a381-8421766f1a6b · outbound

This paper cites Phi-4 Technical Report.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Phi-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.739273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.739273Z digest=sha256:acb733bbe95a2240aabe67534bfd56006d934c24b4f404b75aaee351931a1f62

Observation 3a684dc7-b459-4971-9310-dbc1ee9f3a79 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.745925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.745925Z digest=sha256:15c9dd30b47083b5d5af69524fa36c2722a89c26c5ee268bff2cb94ef42a97cb

Observation e76686a2-ee19-4dfa-b3fa-d00fbaa16392 · outbound

This paper cites Y., Rajbhandari, S., Awan, A.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Y., Rajbhandari, S., Awan, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.751415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.751415Z digest=sha256:54f52461f7437d8900fa49cfc8470907587bfed90db090c3c169a81924c55eb2

Observation d76b3fe7-ac24-4bae-bad1-10d4219a1f68 · outbound

This paper cites and Shaji, A.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Shaji, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.224588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.757214Z digest=sha256:d7ec8f45c7890b1b08908374ba5c2fd22b3256b0f00e25fe536f6568fa8e76d0

Observation 3ce825a4-85b8-4bf8-b9a9-d13331b3daf9 · outbound

This paper cites MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T23:00:19.443779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.763102Z digest=sha256:d2c44f9a0c53f53b1598a4294b7207964b9370952be07f388a815537c0c7d22e

Observation c5e29aa8-6157-4dfb-9fbe-580f2760c38a · outbound

This paper cites Active multi-task representation learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Active multi-task representation learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.199982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.769821Z digest=sha256:816fb795775e7c2e0a8c9f6507c28036fa2f543843837e24fb06fc363f146c18

Observation bfe644cd-5cd9-4aa9-abd9-01ff6d4eabea · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.775307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.775307Z digest=sha256:881101ef32ddf3ff72c295513d57b93d8340eb248642a864a933445dc0359ae5

Observation 52f07ce0-f174-4036-9d54-5bba9f6fc698 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.782233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.782233Z digest=sha256:02294e4088bee5c4ed99a994bd0fae517724293ccc1711bec41aff2e9adea7c1

Observation 327c293c-9f4a-49ae-95b0-424966f745af · outbound

This paper cites D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU D eep S eek M o E : Towards ultimate expert specialization in mixture-of-experts language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.791097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.791097Z digest=sha256:9f4e0bb2d410256fe4f48502a0c24990484ae82fb2d14ddac34946a05dfbf101

Observation 727dac6a-b239-4e5c-b196-f08320653ecf · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.799618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.799618Z digest=sha256:91e553566e4dc0a6da866c682e53ef0acf95a3e0bd3a3f3bf8a19c01c385ba0a

Observation 47029fd0-e2e1-477d-88f6-0b58449e6d6d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.806520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.806520Z digest=sha256:88f6a2c85d785d40c40bd97d7f550f5aa2108f9ec5f27e20612727305af4a7bd

Observation 7a525660-27d3-4001-9f46-2fccff125050 · outbound

This paper cites Few-Shot Learning via Learning the Representation, Provably.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Few-Shot Learning via Learning the Representation, Provably

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.816788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.816788Z digest=sha256:4aa74e432e8598a55e4fc27d7b0a217fd90c85d12c412a3676751de1749c20c0

Observation cbdf6ffa-d052-4180-8de0-69afaa65e337 · outbound

This paper cites The Llama 3 Herd of Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.823318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.823318Z digest=sha256:d23180df4246ca49f79aae3f833086dc4d1a29f03ddb88d93a9dabe25d44ab59

Observation 725c3151-30e9-4283-8e9a-ce509f742b17 · outbound

This paper cites mixtral-offloading.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU mixtral-offloading

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.148601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.828971Z digest=sha256:0625c7060094cbe9729b0541b34460b5aa0d9862610f1b619bda3c30e84b52de

Observation 24a75ecc-bda2-479a-a402-2e51d3870d32 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.834440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.834440Z digest=sha256:cfc1d20caaabfef4a1639548473ff2b96e900c86fd557fe7f9e14e27fbed866a

Observation 84ee202f-5eb4-4359-804f-d3be1ea19b8d · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.840069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.840069Z digest=sha256:8476a289ca02542f50c8621c631d15c2ba7e4b3504be38235fdd9a14df498f31

Observation c143d9d5-eff0-4231-8ad8-99bb8ed62d2c · outbound

This paper cites and Alistarh, D.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Alistarh, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.108972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.846718Z digest=sha256:ca691232a3d4610232e6b5836d8f6230102f4f4ac938c7cde9e7754dee473708

Observation bab7de65-c6fd-478b-9bbd-083636a3e1e6 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU A framework for few-shot language model evaluation, 07 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.852078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.852078Z digest=sha256:c0616b6b75a87d8a3f58d1765fa95087568ca95cf7a594e955c6af4a6e11a2b6

Observation 90db61db-82a7-4a3e-83e0-43e8603c7159 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Transformer feed-forward layers are key-value memories

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.857019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.857019Z digest=sha256:5e906a836b4ac5a6cbfc8f2c9669be3735fa54d53fdeca755ce764e17faa32fd

Observation 28e977d2-584b-4820-88f5-c5ecb758d738 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.863141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.863141Z digest=sha256:9ada038c80e75d13a996eb210c5f55b79d718904190a80df0ac71fbcc4fdfc34

Observation bb5fd3d0-45d8-4e20-b9d1-6287cebae18b · outbound

This paper cites Accelerate: Training and inference at scale made simple, efficient and adaptable.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Accelerate: Training and inference at scale made simple, efficient and adaptable

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.041897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.869011Z digest=sha256:54827d2c703ff909cde55d8d56da698dff15a6379d5bd33c27286d75affbbd98

Observation 41ee817b-2f93-4290-b901-4f84ed2fac7c · outbound

This paper cites CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:00:19.156537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.875990Z digest=sha256:45e680eda572517c016aa16a0812e1c3b8ccc48337d2949587eb85fc5280b979

Observation 3b58a5f5-e855-4a40-b0c3-80347f028c66 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Measuring Massive Multitask Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.881727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.881727Z digest=sha256:343bd13196cbad5a99a0a90b93fb3773c0bbb1f5ff9c9c9bc12ac6eadbd08e5b

Observation 0480f9b6-fb16-4211-b02d-6eb2c6daaf25 · outbound

This paper cites Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pre-gated moe: An algorithm-system co-design for fast and scalable mixture-of-expert inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:20.020055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.887270Z digest=sha256:91c712ccd77b2767d99315d0a386b955953d58d079955c8674170425912a2395

Observation 86d25de5-e46b-41be-8515-7729a727dab2 · outbound

This paper cites Mixtral of Experts.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.896175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.896175Z digest=sha256:1a7c47cae97ff739b6c2bf0ca507a3e3519aef56fdfdcb894d5a9e2c4ee2f216

Observation c6483c91-a401-4d11-a347-f66fe7148c14 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.902621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.902621Z digest=sha256:14d98374c6c05b4a15ae94b432fe5329448a47aebc0c99279355af71f9f05b80

Observation 1ede6a3e-f848-411e-952d-fd84810af759 · outbound

This paper cites Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.909669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.909669Z digest=sha256:7f8bbcab683243cef5570528655c7f4ecb09dd362f41a5efdd99817bf6ad2a30

Observation b9869fc6-dd19-4af2-89c0-e12f69f572a6 · outbound

This paper cites S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU S wap M o E : Serving off-the-shelf M o E -based large language models with tunable memory budget

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.964733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.918029Z digest=sha256:633c9a6d4b0b812c7448aae22f5b65d071cb6fa726262d6e17c825d988a5673f

Observation 2f39fe42-3c72-43c4-b54d-70ead58ce9c4 · outbound

This paper cites CATS : Context-aware thresholding for sparsity in large language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU CATS : Context-aware thresholding for sparsity in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.937766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.923383Z digest=sha256:b3a8a25438ff93e7643a5e9addc7032c09ba36d43cffa02fd1e3c7fa7787cff6

Observation ccd9857e-4fa1-42d1-835d-18a12ca7d368 · outbound

This paper cites InfiniGen : Efficient generative inference of large language models with dynamic KV cache management.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU InfiniGen : Efficient generative inference of large language models with dynamic KV cache management

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.910507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.930729Z digest=sha256:ab2a57bfc60f064b4da9d65c5553c105740705519df2130ca48071e7000ed6c9

Observation 721b679f-5e19-4d5b-8847-d32c9d840572 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Training-Free Activation Sparsity in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.937817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.937817Z digest=sha256:bbb77a9b64a884ca4af418845c7ba0c2db1511fe227193f746de4aea979ca0d0

Observation 90acd159-9eae-44a2-be46-2df4ef83b4a5 · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deja vu: Contextual sparsity for efficient llms at inference time

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.883330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.944379Z digest=sha256:90bef984d2fd36649d67831e507429e3a6b5de0d0cf9a5b1b29f533a3f350625

Observation 12c6a7b1-854a-4a6f-9f98-cd25bd05f0a2 · outbound

This paper cites llama.cpp.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU llama.cpp

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.857973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.952242Z digest=sha256:e3f7b0acd3417e65b65e78442eabffdad811c388cb40f956e381f36241232c6b

Observation 2e1be1af-7e6a-4334-80b9-d51fcc0d4362 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Llm-pruner: On the structural pruning of large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.829297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.960800Z digest=sha256:a33f8042fadd9b074f4004328114be8d1bd7b3b89cc3f9ac8134e902f2ff3220

Observation fa720d6d-75bb-44eb-b3f7-6cdfe7b87cfb · outbound

This paper cites Pointer sentinel mixture models, 2016.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pointer sentinel mixture models, 2016

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.966839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.966839Z digest=sha256:8db5aa8ac3d7ba8d368a3bfa83683b1bf572d2f0a96d99ac4732dcf526147c94

Observation 8298ebcf-7fd1-4210-a33b-d5de6894be62 · outbound

This paper cites Deepspeed-mii.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Deepspeed-mii

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.787328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.973800Z digest=sha256:6ebd01eac380bacb6948f42936b30a0f796adaa365c4ce39ccc73995e4e31777

Observation 4996bea5-75f8-4329-9f3b-c6a878ea5b9c · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.983480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.983480Z digest=sha256:9506c7a7b9f78daa23ee58ea27fa943cfe4083259eb78bc291af474e53d05a8e

Observation cdc1ba83-bfda-4c95-86b3-d3ca038e1f71 · outbound

This paper cites GPT-4 Technical Report.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:17.989790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:17.989790Z digest=sha256:ef481ef94d9df867ea0eab758d85b87af3b40068d93e06de28bfbddc81e68500

Observation d91d0e5e-58af-4a06-8580-15f3e665cb1e · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Pytorch: An imperative style, high-performance deep learning library

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.759512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:17.995847Z digest=sha256:e7076536ce94e51e5d71e0668e73570957e46a9eac4b8a981c5a09d48a09d83c

Observation 02c19407-6e5e-4461-b6bc-a9ab78ee0377 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.005803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.005803Z digest=sha256:43cd15801d3ff70c0ab55055e10527ef0b4698df31396ba5b1377b2c825453f9

Observation 9a578f77-24fc-484b-916a-488691c6535a · outbound

This paper cites Zero-infinity: breaking the gpu memory wall for extreme scale deep learning.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Zero-infinity: breaking the gpu memory wall for extreme scale deep learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.015066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.015066Z digest=sha256:b045b78ef879cb36128854a5a00821368e5d10659c9d2600c5a854e2d18960c0

Observation 247c01b2-1b5e-4d2e-a3be-0c67cf207d01 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.021867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.021867Z digest=sha256:903fb6eb1e02a91ca1a612f017aedae668340dd20f52d157eafc8809eff56812

Observation 379bb1b7-2012-488b-85e8-dd56140309c9 · outbound

This paper cites Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Edge-moe: Memory-efficient multi-task vision transformer architecture with task-level sparsity via mixture-of-experts

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.716571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.027790Z digest=sha256:92c1861efb130be2e8029ce49a3ad499bda11656035e8a5f294ec6a46823bcca

Observation 743ac543-05d0-4955-94e1-eb0d5c5b87b2 · outbound

This paper cites Sharegpt.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sharegpt

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.694207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.033475Z digest=sha256:9a6cb409633a9f4588401fec51ce2b3f0a36e42eab232578d11879630df38412

Observation 3229b3bf-af95-4555-b7c4-e6dd604ab33e · outbound

This paper cites GLU Variants Improve Transformer.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU GLU Variants Improve Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.040770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.040770Z digest=sha256:250b3eae509ac743a2ff9e1502c6c897b548e7c1b165f3fa15052d005f3e6050

Observation 6c7eeb3d-bca8-4529-b9cc-75833d975c4d · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Flexgen: High-throughput generative inference of large language models with a single gpu

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.670024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.046116Z digest=sha256:819cf62ab5cd27eba90a9db758ebe2dcda484f32210afebc27d0f0f4a455345c

Observation 79b653c8-7925-4814-85c9-3af318e79a75 · outbound

This paper cites SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.053935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.053935Z digest=sha256:7207f2a4d979efa0b293d1cfc69000a856990995cbc586ca3ee0c23be779307a

Observation 020fa894-8bee-4aa0-81da-03f28f859508 · outbound

This paper cites ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.060634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.060634Z digest=sha256:04d195c2adb28e5f0b27e36748af0c36d22a5ad6d2733a0ee7792fec1459386d

Observation c8f6bcbe-b074-4bbf-b084-00f63654d330 · outbound

This paper cites ProMoE: Fast MoE-based LLM Serving using Proactive Caching.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ProMoE: Fast MoE-based LLM Serving using Proactive Caching

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.068102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.068102Z digest=sha256:9e05084f329a1925798c88a88f1657064bee510ba02cbf8c88d398856568a4a4

Observation 84f9b36d-d5ba-4aa4-8804-3ff60ef17578 · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.080046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.080046Z digest=sha256:6a605d93cf38c1a3e39ae3736d267d47fb264566497ad52d0261a9cb8c2e361d

Observation fc400941-896c-477b-a247-7d5c13026a17 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU A Simple and Effective Pruning Approach for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.088388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.088388Z digest=sha256:ba6ec1d38a79efd8f6dd22bfcc9542137d3b52bbf54e0276ffffe9b796552f11

Observation 020ce44c-6145-4969-b276-bd807a80f7a2 · outbound

This paper cites HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.099183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.099183Z digest=sha256:82d5be63c88e30074a5f1a7d4524e337891b986767f4f373f00e4499e2ea68bb

Observation c5e2268d-4170-42b0-8a56-6860916b885c · outbound

This paper cites Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters", February 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.105373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.105373Z digest=sha256:a337fbf593b31c62e6041fe1a8071153ab41d019eb5e1503bd626e5e47eba493

Observation 34b0b7a3-6ade-4a8f-bae6-82b815c30366 · outbound

This paper cites Sample Efficient Linear Meta-Learning by Alternating Minimization.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Sample Efficient Linear Meta-Learning by Alternating Minimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.110331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.110331Z digest=sha256:db65f2d1ce60a98da2b62ac5ad1de3778a564efb3a073e4f6d51f514c3906a63

Observation 9e445290-c782-4165-9ed4-658d4a891137 · outbound

This paper cites T., and Cox, D.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU T., and Cox, D

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.117468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.117468Z digest=sha256:68b2e12d1f1a15d8b8d950d95bad6c8f8c6bd06ad9192ba26081775059525d3d

Observation e16e6808-ffae-4924-abbe-99bd90483fdc · outbound

This paper cites On the theory of transfer learning: The importance of task diversity.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU On the theory of transfer learning: The importance of task diversity

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.626650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.125925Z digest=sha256:e5a8008385921c6ab9f1fed5bd9227587b241c3dc9f5a4d090959b9bff91c681

Observation ab80d049-fdd9-4250-8266-c4ec26090e30 · outbound

This paper cites Provable meta-learning of linear representations.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Provable meta-learning of linear representations

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:00:19.600418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.133442Z digest=sha256:dd93ad86df93d04e6ced64c7471f55999c0aa28eeed5366e8b71fbb2fcf3baa1

Observation 2300d282-1841-401c-aa6a-b3c072c8df32 · outbound

This paper cites an unresolved cited work.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:00:19.577922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T23:00:18.140080Z digest=sha256:4553aa2e7456e2fbac3b2aae1f001ce3b9f8be8858a40055c1877acd8a8c85c1

Observation adff6728-d4ce-4fc3-aca1-7179da0c7abd · outbound

This paper cites MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.147537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.147537Z digest=sha256:750f0652030c247b2665eec419d6a5ab069f55c56e197f858afe1b9ee004fcc8

Observation fc9a2071-2c3c-4899-ad54-d6cbf007e740 · outbound

This paper cites PowerInfer-2: Fast Large Language Model Inference on a Smartphone.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU PowerInfer-2: Fast Large Language Model Inference on a Smartphone

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.153988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.153988Z digest=sha256:26dd275d93578db223b05c6c80ef153715285396d3e3d8ede6aece0bacb2cf54

Observation 9f415b45-5a5a-40e4-9337-3a562c093362 · outbound

This paper cites and Ananiadou, S.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU and Ananiadou, S

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.161527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.161527Z digest=sha256:a0830a41dc898154998ccbd1abd846b4e9ab1c2cf0322ab51667d60ff1c814f3

Observation cc91911b-4e65-4197-9ca3-932ce986b3eb · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.167395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.167395Z digest=sha256:8c171f590d1e8ef1549cd5d29197d11e6fabf9742623192e13bcb8412d33dec2

Observation 9ccd1920-6cc0-40bd-85c3-af1cdb5f9d27 · outbound

This paper cites write newline.

FloE: On-the-Fly MoE Inference on Memory-constrained GPU write newline

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T23:00:18.175589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:00:18.175589Z digest=sha256:491d42480408bb4c7b8e267ddee69dc7c4ca39b25bf15bfaabd3180ff30e25d5

Pith citing papers

Observation 513369c9-9c91-4e0b-ad07-1138bdc3b9cd · inbound

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving cites this paper.

FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving FloE: On-the-Fly MoE Inference on Memory-constrained GPU

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:25.893028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T10:32:50.809236Z digest=sha256:3a92319f87495ba4f1da34693f735e27b8b34dbc584e17fa1d1df0466bd95b9c