Pith. sign in

Paper Citation Record · LEDGER

HarMoEny: Efficient Multi-GPU Inference of MoE Models

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2506.12417.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12417 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:32.963734Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:43:48.882267Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T04:33:03.804063Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved11
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b83d114-1784-493f-9690-744ac60e2df6 · outbound

This paper cites GPT-4 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.813852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.813852Z digest=sha256:e1a1c859b8bae50c746e47c481870f25d538f10c1b93f1b2116757376013e8a0

Observation f80582c9-8953-4a5a-bc28-185de4c14044 · outbound

This paper cites Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Amazon EC2 update – inf1 instances with AWS inferentia chips for high performance cost-effective inferencing, 2019

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.504866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.818850Z digest=sha256:80ddd870bb48ba037a1bf9511a6fa40a7d7ed772017de1eda5749b91135527b4

Observation 59dddad4-1d4c-4ed6-b0c5-62a42d571780 · outbound

This paper cites A neural probabilistic language model.

HarMoEny: Efficient Multi-GPU Inference of MoE Models A neural probabilistic language model

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.491982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.822866Z digest=sha256:76ea234d4a0cdeac8e35b99c0accba98a95cdd54896e0113ef5ae56193830bc5

Observation 71b4bcfb-9f00-4e92-afa8-ad95ab622190 · outbound

This paper cites Datacenter power and energy management: past, present, and future.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Datacenter power and energy management: past, present, and future

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.479177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.827146Z digest=sha256:c1ef6751295fdfded170dcf3d0e201c9d5005b46f4f702671f713d4c4e7044ee

Observation d8ead566-47c9-44d8-9279-225b98678ce9 · outbound

This paper cites Language models are few-shot learners.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.831078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.831078Z digest=sha256:cf67a89edd364bb977487d17cd2eee937c04c328648d101a0b89a2d540e1747e

Observation 2fe6c22e-b416-4b89-bbbe-cbcd728c75d7 · outbound

This paper cites Large scale distributed deep networks.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Large scale distributed deep networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.457504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.834897Z digest=sha256:24f296f8c3ce1ca4cab1bd42f23895f898a813644b5e7e7f97c07ad131a39c11

Observation bf21b206-e2b7-45a7-a376-5c26f5da9c52 · outbound

This paper cites Quasar: Resource-efficient and qos-aware cluster management.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Quasar: Resource-efficient and qos-aware cluster management

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.444890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.839116Z digest=sha256:d31a74a456752762adae12d6288db2ee5bbe5352c7cccaa1a676a1b951e46a6c

Observation bad8eb6c-46f3-481c-80fb-f768b78f85bb · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.432026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.842730Z digest=sha256:8b6947e1b2a0d531dc9a34b1cc63a31614cb66840ad8acb3ee05ff499856fc0f

Observation c1678eaa-a4de-4c35-b151-fd9c69584ff5 · outbound

This paper cites Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Acl 2019 fourth conference on machine translation (wmt19), shared task: Machine translation of news

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.417764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.846300Z digest=sha256:86592aa8f4bbdc757f44536075b516b8f0c236158a1b7efff2b9f7a3c965f7c1

Observation 46e54108-df6f-4716-a6b2-03918b500a69 · outbound

This paper cites Megablocks: Efficient sparse training with mixture-of-experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Megablocks: Efficient sparse training with mixture-of-experts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.404964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.850070Z digest=sha256:4ca3acd0f431833abcd12a2bc78180b4f692185124a474ebb14d73bb14af3d7a

Observation 482cdcd6-4f65-4613-aab7-5327551c3ef7 · outbound

This paper cites Character-based NMT with Transformer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Character-based NMT with Transformer

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:33.086400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.853759Z digest=sha256:4ae4f044979d78cab2c431e27d3866c00bb7ef1a095b9547e85556947a8e14a9

Observation d4340967-aa22-4a69-aad8-fd6b8cf1e4c5 · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

HarMoEny: Efficient Multi-GPU Inference of MoE Models FastMoE: A Fast Mixture-of-Expert Training System

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.858076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.858076Z digest=sha256:069cd861d96a2e22e4ed5ade3391238253d0c76a24b10defcee8fd7655dbc269

Observation 97e82041-bfbd-4914-b069-7b631afd185a · outbound

This paper cites Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.392572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.862163Z digest=sha256:be719b7aaa2602dbfc06bdd775c9a5bc27dc9131c4baa44dee5cba9f0ca89c56

Observation 5affdc3e-01fd-4096-a213-3589403847e8 · outbound

This paper cites Tutel: Adaptive mixture-of- experts at scale.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Tutel: Adaptive mixture-of- experts at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.380243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.865806Z digest=sha256:67c22a8976535afa8378484d76eee78b99d0f8e9baa8e06905b99d844da45fc3

Observation eb3db8f3-6b32-46db-b5a8-175c2bbc2252 · outbound

This paper cites Adaptive mixtures of local experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Adaptive mixtures of local experts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.367056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.869364Z digest=sha256:0660c5c9f7d4a2f01506469456be7411a71bbe3a55a7ff227d740747d8f579ac

Observation 880a320b-7746-4f15-a62a-0d8f087a5101 · outbound

This paper cites Mixtral of Experts.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Mixtral of Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.872804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.872804Z digest=sha256:2b6c5e7cc5c27488bb79e17c765a3993d995adeab0b09ee43654d1e398d08be4

Observation 60a95c6e-66a4-40b7-ba63-0318e79fa7ca · outbound

This paper cites Scaling Laws for Neural Language Models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.876451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.876451Z digest=sha256:0e61874dbb084191f101ce12f27536fc47be42a8f7480669c134af749105b70d

Observation 4915bd77-dafe-4b2c-8794-96d864e8cee6 · outbound

This paper cites Learning multiple layers of features from tiny images.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Learning multiple layers of features from tiny images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.880325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.880325Z digest=sha256:6cad8421db185fdcca961e9804b89f8ef2a26b8834b7c2d7226da21e57fc61b9

Observation 899c9652-4ce1-4421-91cf-c81dd3cd9781 · outbound

This paper cites Subword regularization: Improving neural network translation models with multiple subword candidates.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Subword regularization: Improving neural network translation models with multiple subword candidates

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.344951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.883909Z digest=sha256:9916db3c083504f1a607c7f25132d96db9026ef1c43bb170058c3b5dc9f71fd1

Observation 9903387d-0a07-4e20-8ae6-13c96a2f4168 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.332584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.887516Z digest=sha256:22d6996be9a6cd5604b7525986e34c6f58a7ad3a136f1951b716994cba2aef04

Observation a53ad049-1517-4573-a92b-deb66a80e2bb · outbound

This paper cites AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019.

HarMoEny: Efficient Multi-GPU Inference of MoE Models AWS to offer nvidia’s t4 GPUs for AI inferencing, 2019

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.320398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.891089Z digest=sha256:01b706da26cec02a4de0297d9ad143099c25034cdad21c006b8c8729095e1939

Observation a1b134cd-08ca-4b04-9b4a-9d45491791e4 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Gshard: Scaling giant models with conditional computation and automatic sharding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.307403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.894473Z digest=sha256:09309a6f5a539c115eec11780723d735531bedd1d5e7f8beab1f28f48bf86a53

Observation b49519fd-b6e4-4747-858f-eba378d20bbd · outbound

This paper cites Accelerating distributed MoE training and inference with lina.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Accelerating distributed MoE training and inference with lina

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.295114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.898045Z digest=sha256:a14bb416cf94fa60d3ed9f35d027e0c4b314b4191c1ad7af6e290b4450edbcad

Observation d509d5f2-c70d-49d2-b1aa-6b4a4438f67b · outbound

This paper cites DeepSeek-V3 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSeek-V3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.902330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.902330Z digest=sha256:db95cf4004acbb0fd7657a3934801cbfdb1d172a916f47d8bef22c76355b8859

Observation 434cd7b8-ca7e-4e57-98d7-209d76973cb4 · outbound

This paper cites Pointer sentinel mixture models, 2016.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Pointer sentinel mixture models, 2016

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.905968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.905968Z digest=sha256:a8355833e1549718e7b02f87c8f22ca82a7f34c6e5af336f399eeb174eb6394f

Observation ee18e1d9-862c-4881-b67c-c18d9894008d · outbound

This paper cites Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Deepspeed-mii: Mii makes low-latency and high-throughput inference possible, powered by deepspeed

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.273719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.909635Z digest=sha256:0d8115e49e560eace3a9866bd91d3489927fbd693222fb6829d7371ccb957891

Observation 95b2e0a3-21d7-4511-9f68-f2f545fc8063 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Pytorch: An imperative style, high-performance deep learning library

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.917296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.917296Z digest=sha256:4a84bf0c5de9064e4cb2b7db1b8adc4d595323c133a617de930965d9ac2c755b

Observation 3195126a-8790-45de-982a-aefa27955565 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.239815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.920856Z digest=sha256:0198bbb4b429cd2b8a39147334c7a6b8799d801a74c2527e40d480890c45d8cc

Observation 797a5ad8-591e-4927-a66e-8fffdc192611 · outbound

This paper cites DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale.

HarMoEny: Efficient Multi-GPU Inference of MoE Models DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.227715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.924343Z digest=sha256:be7cec75bb976b000716718b60fb2ba6c3022fef40e162eef82b16d9b3ec2194

Observation aef745bd-b596-4ae7-a9ee-7af36f6fdcb7 · outbound

This paper cites Outrageously large neural networks: The sparsely- gated mixture-of-experts layer.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Outrageously large neural networks: The sparsely- gated mixture-of-experts layer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.215236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.927724Z digest=sha256:416f5dd4806470fe774c309735b84f4467bd003e143f41fd430215237bec1106

Observation 461ce789-0910-464d-b777-7519b71a159c · outbound

This paper cites Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.202658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.931223Z digest=sha256:bd9365b3ed3f8cf67fb04c8aa6beccf5d58a634d92d4450ec6f45f2027c2e763

Observation a1711014-cffb-4582-8751-962448885974 · outbound

This paper cites Borg: the next generation.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Borg: the next generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.190565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.934674Z digest=sha256:f66daf71487a221ff00b0eb1db1263fb768e3e48f614b9a20910875be4f44e0b

Observation e3d7be92-1c27-4073-bfd5-2ddf0befb776 · outbound

This paper cites Attention is all you need.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.177817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.938064Z digest=sha256:db010066f1ea394f2a7351f208054a52032822f8120a529a37abdb5a3a5ee6fa

Observation b515975f-470e-4213-941d-a8e93a5a0525 · outbound

This paper cites Prophet: Fine-grained load balancing for parallel training of large-scale moe models.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Prophet: Fine-grained load balancing for parallel training of large-scale moe models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.164132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.941528Z digest=sha256:e02c9f778639f92e017b42dfc43e0ec3cdc1ec90b086e66f931ce46c7ebd707f

Observation 09193130-6c13-4302-a06f-cbb3dbe6ed3c · outbound

This paper cites Resource-efficient algorithms and systems of foundation models: A survey.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Resource-efficient algorithms and systems of foundation models: A survey

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.151534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.945320Z digest=sha256:65717cec500fbd17dff0b25eb37496592471fd6c8606b5c77bb4fe005bfa744f

Observation 70c8a7b8-7217-4b16-a8ca-4bee9cf1fc76 · outbound

This paper cites Qwen2.5 Technical Report.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Qwen2.5 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.948934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.948934Z digest=sha256:aebbc391e06698d5c57ba6acbea6aa62586b373cec71f5bbdab5d17b10167e25

Observation f13c2bcf-08ac-4428-969e-df9c5f0bc56f · outbound

This paper cites Harnessing the power of LLMs in practice: A survey on chatgpt and beyond.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Harnessing the power of LLMs in practice: A survey on chatgpt and beyond

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.139056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.952538Z digest=sha256:fc1fd2215c25c7a51b3c996f3343f49be59e6038d84e3455c6dc869834a38f58

Observation 33d07119-e7e6-4e3f-a417-77c4e826e6de · outbound

This paper cites Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Exploiting inter-layer expert affinity for accelerating mixture-of- experts model inference

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.126477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.956355Z digest=sha256:3b71ee4490b8045644d8ff87a9622a2617a34616694be316d7b4964f01bb657d

Observation 25f8865e-cbd2-42ad-909a-9a5eeac9efb0 · outbound

This paper cites SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization.

HarMoEny: Efficient Multi-GPU Inference of MoE Models SmartMoE: Efficiently training Sparsely-Activated models through combining offline and online parallelization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:33.113747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.959823Z digest=sha256:cba044f5d5bfcd748a799b1c8e4c4cd8e11be2603b83fd8bc7442ffdd44a0452

Observation c367ef27-b1e2-4b77-9f1f-ef08e899122a · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:32.963734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:57:32.963734Z digest=sha256:bbfc6a19076b8c75988fc8ac6742132e08a4f747421a1d2dc7a0456cf0a12337

Observation 32e082bc-2146-42df-8ca6-c3ee6fb2de9b · outbound

This paper cites an unresolved cited work.

HarMoEny: Efficient Multi-GPU Inference of MoE Models Unresolved cited work

Reference 2022

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T00:57:33.261045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:57:32.913583Z digest=sha256:af50b77a765bb589d0960345db74a8066e648e53ff0bc0f654a592e96244adba

Pith citing papers

Observation 0685dec3-b41b-42bf-afb6-58de81036b73 · inbound

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems cites this paper.

GEM: GPU-Variability-Aware Expert to GPU Mapping for MoE Systems HarMoEny: Efficient Multi-GPU Inference of MoE Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T04:33:03.808270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T04:31:11.271729Z digest=sha256:cd69fffbfe3386c91a2cf0d15bcb94702a48d7f926eef933814cc4b89db59381

Observation dd622958-19a7-4268-afc3-70aea1dd789a · inbound

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference cites this paper.

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference HarMoEny: Efficient Multi-GPU Inference of MoE Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:43:48.882267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:43:48.882267Z digest=sha256:fc666c272eba1d7af0a76917442779ce16bf2081dcbe20c349396e414ce748fb