Pith. sign in

Paper Citation Record · LEDGER

HarMoEny: Efficient Multi-GPU Inference of MoE Models

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.12417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12417 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:32.963734Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T04:31:11.271729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T04:33:03.804063Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved11
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b83d114-1784-493f-9690-744ac60e2df6 · outbound

This paper cites GPT-4 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.813852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.813852Z digest=sha256:7ff1d3a22af50222a748b35598f10e8e312301d3c4fb92d7c3b2e9256b5ab8ef

Observation f80582c9-8953-4a5a-bc28-185de4c14044 · outbound

This paper cites Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.504866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.818850Z digest=sha256:0c9da7ab68b9e06b822aece7d711313efb359e356043ff6875bc20ae67273b35

Observation 59dddad4-1d4c-4ed6-b0c5-62a42d571780 · outbound

This paper cites A neural probabilistic language model.

HarMoEny: Efficient Multi-GPU Inference of MoE Models A neural probabilistic language model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.491982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.822866Z digest=sha256:fca0a47cef8ca9f6ae989ae6edbe5227e614f9a36ae720611ca6a6dd2286e59e

Observation 71b4bcfb-9f00-4e92-afa8-ad95ab622190 · outbound

This paper cites Datacenter power and energy management: past, present, and future.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Datacenter power and energy management: past, present, and future

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.479177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.827146Z digest=sha256:e023d09fb68135a19f48d3d2845ab13d5fa82ebe5e1ac3b8862d46162843ca77

Observation d8ead566-47c9-44d8-9279-225b98678ce9 · outbound

This paper cites Language models are few-shot learners.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.831078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.831078Z digest=sha256:907b6f6a3c1ef5ab3e03b85abd89ee271f62d95fafdb4ae098fc89178146eb3a

Observation 2fe6c22e-b416-4b89-bbbe-cbcd728c75d7 · outbound

This paper cites Large scale distributed deep networks.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Large scale distributed deep networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.457504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.834897Z digest=sha256:4e2e694a081902d42bf93280c911b6cf058a86959d3969272f87a441d2b05dc4

Observation bf21b206-e2b7-45a7-a376-5c26f5da9c52 · outbound

This paper cites Quasar: Resource-efficient and qos-aware cluster management.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Quasar: Resource-efficient and qos-aware cluster management

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.444890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.839116Z digest=sha256:e25e469bc5a031257d5a553ad077e193c9884764f60ba4b982c56e6478fc04e6

Observation bad8eb6c-46f3-481c-80fb-f768b78f85bb · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.432026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.842730Z digest=sha256:bd9b15878d85d156012910b8b168d87bb833d37ca3f89cab56b6e1f601aabecb

Observation c1678eaa-a4de-4c35-b151-fd9c69584ff5 · outbound

This paper cites Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.417764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.846300Z digest=sha256:4fb0dad0b6d394b7785762ddef4c49d7a48e2eff99859ead317b5fc0a851b694

Observation 46e54108-df6f-4716-a6b2-03918b500a69 · outbound

This paper cites Megablocks: Efficient sparse training with mixture-of-experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Megablocks: Efficient sparse training with mixture-of-experts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.404964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.850070Z digest=sha256:7ebc46ffe9094d4f87d8caf33973bf827b550f9d8adedd424c6d679f2045f8fd

Observation 482cdcd6-4f65-4613-aab7-5327551c3ef7 · outbound

This paper cites Character-based NMT with Transformer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Character-based NMT with Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:33.086400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.853759Z digest=sha256:d999f2466f48da9a1b8f8e1faa91d8bc8eb95a7b0a3907de6591514a3e6160d3

Observation d4340967-aa22-4a69-aad8-fd6b8cf1e4c5 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

HarMoEny: Efficient Multi-GPU Inference of MoE Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.858076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.858076Z digest=sha256:ff04f25537173f7155eef31517ac1a0803abd5a143d604102c0748c46472cee4

Observation 97e82041-bfbd-4914-b069-7b631afd185a · outbound

This paper cites Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.392572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.862163Z digest=sha256:e43ab6fb0bcf2b763dff784ad382abec0f6356be4e6f1ae0d99ccecf7f70e696

Observation 5affdc3e-01fd-4096-a213-3589403847e8 · outbound

This paper cites Tutel: Adaptive mixture-of- experts at scale.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Tutel: Adaptive mixture-of- experts at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.380243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.865806Z digest=sha256:041ee0d89dad726f1fc6b6b6e2235d75877eba87ed13c4293549838be054ade2

Observation eb3db8f3-6b32-46db-b5a8-175c2bbc2252 · outbound

This paper cites Adaptive mixtures of local experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Adaptive mixtures of local experts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.367056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.869364Z digest=sha256:302fc0fe4ff917fa5e5c43c51e27c56fb5e4d384c03b8de6b37bb5ffde624ce7

Observation 880a320b-7746-4f15-a62a-0d8f087a5101 · outbound

This paper cites Mixtral of Experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Mixtral of Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.872804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.872804Z digest=sha256:8bed0022974172df46d58d5335b6cffa2ba5ea870d87b578ebeec67cf8a25247

Observation 60a95c6e-66a4-40b7-ba63-0318e79fa7ca · outbound

This paper cites Scaling Laws for Neural Language Models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.876451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.876451Z digest=sha256:21c47ba79ba1fd5a2d66910f3a7b752c2a1de5cdb8de6a119a3cbf9332fd62b4

Observation 4915bd77-dafe-4b2c-8794-96d864e8cee6 · outbound

This paper cites Learning multiple layers of features from tiny images.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Learning multiple layers of features from tiny images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.880325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.880325Z digest=sha256:93222d03fb038f5b8d5ce1feee5d82c5803ef5524b96887a7c8b801f8a0c77ad

Observation 899c9652-4ce1-4421-91cf-c81dd3cd9781 · outbound

This paper cites Subword regularization: Improving neural network translation models with multiple subword candidates.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Subword regularization: Improving neural network translation models with multiple subword candidates

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.344951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.883909Z digest=sha256:8b712d51fe69c9a87739c23d6e01944b70aa9c832f06f17bbf5130226832af00

Observation 9903387d-0a07-4e20-8ae6-13c96a2f4168 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.332584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.887516Z digest=sha256:361dc0d09b33697ac9dfb41da09a00c0b78a5b7abf72143e8ad2c9dd1a7960af

Observation a53ad049-1517-4573-a92b-deb66a80e2bb · outbound

This paper cites AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019.

HarMoEny: Efficient Multi-GPU Inference of MoE Models AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.320398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.891089Z digest=sha256:817445db651cd82fd468bda84009f283f4cc77d83ab3ad0657b2af389c72d7bb

Observation a1b134cd-08ca-4b04-9b4a-9d45491791e4 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.307403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.894473Z digest=sha256:98a493348d30a836337195034194ec80e95451074eabebbe9867571574ac2ad8

Observation b49519fd-b6e4-4747-858f-eba378d20bbd · outbound

This paper cites Accelerating distributed MoE training and inference with lina.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Accelerating distributed MoE training and inference with lina

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.295114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.898045Z digest=sha256:56c6dd3791627916b48a120eec2f1261d2fe70e027d7cf9379a39962607ddf67

Observation d509d5f2-c70d-49d2-b1aa-6b4a4438f67b · outbound

This paper cites DeepSeek-V3 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.902330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.902330Z digest=sha256:8f2e2bc611db96d09d9e93290127a7441758b96637f23c3e3ae7c90ce82ac73d

Observation 434cd7b8-ca7e-4e57-98d7-209d76973cb4 · outbound

This paper cites Pointer sentinel mixture models, 2016.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Pointer sentinel mixture models, 2016

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.905968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.905968Z digest=sha256:8f8353e0779e42e117f1761a85309234f0ece2aca9f3fb07ca56fa437fa840b9

Observation ee18e1d9-862c-4881-b67c-c18d9894008d · outbound

This paper cites Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.273719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.909635Z digest=sha256:f489c6b00edd67a3affb26e8323f2b435410a6cf72ba50e5283c346e4baa396f

Observation 95b2e0a3-21d7-4511-9f68-f2f545fc8063 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Pytorch: An imperative style, high-performance deep learning library

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.917296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.917296Z digest=sha256:a17947003fcf3c3b77268663d96d48dbf3c3358866887189e1f383ae2f5c098a

Observation 3195126a-8790-45de-982a-aefa27955565 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.239815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.920856Z digest=sha256:d2ae2bbf956499d170f486c122bd96fee87c308e90bdcea9dda0323da7db1970

Observation 797a5ad8-591e-4927-a66e-8fffdc192611 · outbound

This paper cites DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale.

HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.227715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.924343Z digest=sha256:ff49688971acb929b771f8ce739f140b17200beb0cef9f9d47ef313662b277d4

Observation aef745bd-b596-4ae7-a9ee-7af36f6fdcb7 · outbound

This paper cites Outrageously large neural networks: The sparsely- gated mixture-of-experts layer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Outrageously large neural networks: The sparsely- gated mixture-of-experts layer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.215236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.927724Z digest=sha256:2ef6538a8dc2a50d4fae188e72dd87da9dc5ed9936e68414e1a4a5824814c779

Observation 461ce789-0910-464d-b777-7519b71a159c · outbound

This paper cites Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.202658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.931223Z digest=sha256:c174851ebf7985543f51d4a333ac529254616aa6263cf9b5729b1a8bfa930d20

Observation a1711014-cffb-4582-8751-962448885974 · outbound

This paper cites Borg: the next generation.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Borg: the next generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.190565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.934674Z digest=sha256:793e490440111854f1fef328fa35083f44706bb461d1e0b0e28e1f3c3879f178

Observation e3d7be92-1c27-4073-bfd5-2ddf0befb776 · outbound

This paper cites Attention is all you need.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.177817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.938064Z digest=sha256:c3f5a584279856f1556ca22f2a0813dcd92d4b9a71ad529af63ef8040b84f8d6

Observation b515975f-470e-4213-941d-a8e93a5a0525 · outbound

This paper cites Prophet: Fine-grained load balancing for parallel training of large-scale moe models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Prophet: Fine-grained load balancing for parallel training of large-scale moe models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.164132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.941528Z digest=sha256:0d360003fd2199710d6529b33042c2ee199bd80647e8fc494804527f6d713dd3

Observation 09193130-6c13-4302-a06f-cbb3dbe6ed3c · outbound

This paper cites Resource-efficient algorithms and systems of foundation models: A survey.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Resource-efficient algorithms and systems of foundation models: A survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.151534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.945320Z digest=sha256:260ec914a1dcd0872695230a1ec116462317513cb275966c63d5b17f63bd11d8

Observation 70c8a7b8-7217-4b16-a8ca-4bee9cf1fc76 · outbound

This paper cites Qwen2.5 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Qwen2.5 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.948934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.948934Z digest=sha256:b81b2b708ef8d159f8eb41ed3db5f7b1f9b6c5c64c2ff3a2e1fabc14043c991c

Observation f13c2bcf-08ac-4428-969e-df9c5f0bc56f · outbound

This paper cites Harnessing the power of LLMs in practice: A survey on chatgpt and beyond.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Harnessing the power of LLMs in practice: A survey on chatgpt and beyond

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.139056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.952538Z digest=sha256:7d0fa9cda052f2b8efcbafd662481b110f686ca0e0d9ccefeed8dea9cf4547a3

Observation 33d07119-e7e6-4e3f-a417-77c4e826e6de · outbound

This paper cites Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.126477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.956355Z digest=sha256:fd68355797a3abbe6fc42d6c3c54228b1ac1a6f9675ab8b3e7995f40cfbed921

Observation 25f8865e-cbd2-42ad-909a-9a5eeac9efb0 · outbound

This paper cites SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization.

HarMoEny: Efficient Multi-GPU Inference of MoE Models SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.113747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.959823Z digest=sha256:980a48cdbe2d19cbc6c70960b02fa776e33934f7c6ab3dbf0ea86cb9db2ee67a

Observation c367ef27-b1e2-4b77-9f1f-ef08e899122a · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.963734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.963734Z digest=sha256:4a992bcc59823fc5f464583efc2a6975d2639a83393997ebe1169f82d4ed39ab

Observation 32e082bc-2146-42df-8ca6-c3ee6fb2de9b · outbound

This paper cites an unresolved cited work.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T00:57:33.261045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:57:32.913583Z digest=sha256:8f990fb103bd14348208d7c62c142f5b509d2ee10ad568785d56db71dbdeb7f0

Pith citing papers

Observation 0685dec3-b41b-42bf-afb6-58de81036b73 · inbound

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems cites this paper.

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems HarMoEny: Efficient Multi-GPU Inference of MoE Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:33:03.808270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T04:31:11.271729Z digest=sha256:2d55e69b23e17a5c4e71a6c070032d203873d99599c87dd40ed2a5de6459ba1a