Pith. sign in

Paper Citation Record · LEDGER

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems

As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2601.03992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.03992 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:12:49.662460Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb908f54-7c62-445c-904c-a84efd3f76aa · outbound

This paper cites Attention is all you need,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.246249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.246249Z digest=sha256:bed1bb01984cedd5783b4d7be2d5f4ed6b8c0ee6484c8ede841bc80c96b3ebd9

Observation 9c069319-6e8e-4f6b-b780-899828ee8ad3 · outbound

This paper cites Adaptive mixtures of local experts,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Adaptive mixtures of local experts,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.330770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.330770Z digest=sha256:410e3fcac03f2989b6c1d78ef44e80befac30d0c08ff8d0b6aab31a4125e1358

Observation 040eb670-725b-49f8-aeab-6dc394e1cbd8 · outbound

This paper cites Outrageously large neural networks: The sparsely- gated mixture-of-experts layer,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Outrageously large neural networks: The sparsely- gated mixture-of-experts layer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.481094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.481094Z digest=sha256:9ff2828afe0a90317c4f6fecb20d5f580709325cc090312a51a440a524acd500

Observation c8ba9011-0790-4f7f-900c-a4856e6ce203 · outbound

This paper cites GPT-4 Technical Report.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.584804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.584804Z digest=sha256:7cf621f874bb1267d443de2e4f50ac5583ca3515d3551dfbd1b1403900c78ff3

Observation 798e843b-507d-46db-9c8a-5795a5da64a2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.682060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.682060Z digest=sha256:366d2d31708e0df9cbd95c49d6bf35c60781a0dc65da1c84fc7664b8eaf7794f

Observation b9f9bbad-12eb-4684-9c98-b3d3884a54c3 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems OPT: Open Pre-trained Transformer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.857396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.857396Z digest=sha256:1655312f55ae08ef246d6c70a04bc08fa0857765bf11b01f200293a3fb6377b6

Observation 77d96968-f5a7-424e-a377-1b2cd3680072 · outbound

This paper cites Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.032169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.032169Z digest=sha256:75e565d5f1a517476eb4183526b23dbdcfd23311eb9676cad6dc01b50e31a60c

Observation c2580717-12ba-485c-a366-3a679e19a3af · outbound

This paper cites A survey on mixture of experts in large language models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems A survey on mixture of experts in large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.199042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.199042Z digest=sha256:7a5a608b4d33a05aeb143187663645262e0ae6c1fd64b424eaa458bc74f395ad

Observation dd75ece3-f377-482d-9eb4-1c96fe40361f · outbound

This paper cites (2025) GeForce RTX 5080.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems (2025) GeForce RTX 5080

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.372783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.372783Z digest=sha256:7cf9be84465e97724c188bb7552f3635127ca61b19f941638281904db042974f

Observation 6f3bd0c9-0b32-4f35-bf22-b8e86318e18c · outbound

This paper cites Qwen3 Technical Report.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Qwen3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.501611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.501611Z digest=sha256:12b45c0b354e95dc861d836c2afe9ea7d568f9945baadbedb5977d0ecac8c501

Observation 4f595b27-c512-46d0-8bb1-d8aa7b4ae2b0 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.637460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.637460Z digest=sha256:4c335d8c87704f5c5d02ae6265506eb24f708ed5781eed7e1e6b99175b61b499

Observation 7d60b81f-dad6-4b85-90da-0421a5b5a144 · outbound

This paper cites Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.731223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.731223Z digest=sha256:d740626dad118f425f810164a7eb17271d1a84e8e65b14716eac90479a6f20f2

Observation 55c79c1e-5af5-4994-ab70-2257e512a681 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.908768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.908768Z digest=sha256:aa9f09eda55a19529418e1037e556d86a55ca3584bc860cae98c6885469bb0b2

Observation b6305259-e3db-461a-8787-04ef004c4b67 · outbound

This paper cites DAOP: Data-aware offloading and predictive pre- calculation for efficient moe inference,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems DAOP: Data-aware offloading and predictive pre- calculation for efficient moe inference,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.073480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.073480Z digest=sha256:92e04aa114a9479b479ac2381aa7c449f7324f3b458da9722d6aea4b3dbe512e

Observation c78c45cc-9a64-4796-bc27-61a70a4db24a · outbound

This paper cites Moe-lightning: High-throughput moe inference on memory-constrained gpus,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Moe-lightning: High-throughput moe inference on memory-constrained gpus,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.134788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.134788Z digest=sha256:e3367026befefcce81fbe4f5c406e8889f39c15bacfc2100eeb3406fd36301c6

Observation dfb7a272-16ae-4c75-8266-941ea208f832 · outbound

This paper cites Fiddler: CPU-GPU orchestration for fast inference of mixture-of-experts models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fiddler: CPU-GPU orchestration for fast inference of mixture-of-experts models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.176571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.176571Z digest=sha256:b031b2cdea7e8e3afb3d2b161015ef29aa9639290b947d312c7fbb798bfb62fd

Observation d018a23c-8b61-4c64-8ffa-a5e838a69b61 · outbound

This paper cites Monde: Mixture of near-data experts for large-scale sparse models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Monde: Mixture of near-data experts for large-scale sparse models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.263000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.263000Z digest=sha256:4380797a12d8c26b32bb91aeaacfe81125d2c1ee484449ef1c476c523754603f

Observation 5095e32c-e81e-4822-bcdc-f2f23f1d6f58 · outbound

This paper cites Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.351344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.351344Z digest=sha256:74a86e9ae6e7ae2624b631fd97495a751f0a7ca64650ccf083d917bac3a0ce98

Observation b453858d-4086-434b-a15a-744210aede6d · outbound

This paper cites Co-designing binarized transformer and hardware accel- erator for efficient end-to-end edge deployment,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Co-designing binarized transformer and hardware accel- erator for efficient end-to-end edge deployment,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.419157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.419157Z digest=sha256:74df7585ae621859412624a13de0932e85c9e2ed9b75ae12b1eccf93e9a6d4e4

Observation 9156e491-9016-40bb-9e00-5ef4c67a2c51 · outbound

This paper cites Language models at the edge: A survey on techniques, challenges, and applications,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Language models at the edge: A survey on techniques, challenges, and applications,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.512321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.512321Z digest=sha256:3347880912b331567f8363323ae3cee739c056002ddb186065f7a08d9d72bed7

Observation 1fdc3f5b-7058-474d-8dd4-8571e2cec33d · outbound

This paper cites A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.604670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.604670Z digest=sha256:ebf82f5cb666edb34ad589187a60bb795c8ef17b2e63d9f6825f170cc2371358

Observation f9aab894-9168-4329-a9c2-3d5d1ab96678 · outbound

This paper cites Spark: Scalable and precision-aware acceleration of neural networks via efficient encoding,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Spark: Scalable and precision-aware acceleration of neural networks via efficient encoding,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.700109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.700109Z digest=sha256:db4546717b326fd6c53677b29df84f45ace2c54bfdf950d104b9d89db7fb52ef

Observation 72c14f80-ecbd-435b-aa5f-51eebba247bd · outbound

This paper cites Medusa: Simple llm inference acceleration framework with multiple decoding heads,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Medusa: Simple llm inference acceleration framework with multiple decoding heads,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.790247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.790247Z digest=sha256:619cab4305edcdf4a451f606a6ccf7653086b4f46c57ba5103bcf633edd9d0b6

Observation 2cc085ba-503a-4bcf-bbbe-134d1b3aef5e · outbound

This paper cites Recnmp: Accelerating personalized recommendation with near-memory processing,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Recnmp: Accelerating personalized recommendation with near-memory processing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.878841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.878841Z digest=sha256:d7673e5a7df6f7ad541ae7fc516aa8ccec7f867baa8dbf68da6242104e3f161d

Observation 6b0bf0c4-6d8e-4874-af4e-e49eca8e287a · outbound

This paper cites Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.035954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.035954Z digest=sha256:d92215e15915e800f62eff1f55bcd8544daedea997b5b06bc1a78a3b7dbb1311

Observation 23489549-cbba-4acd-bd44-1773d0a8c3ab · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer-based generative model inference,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Attacc! unleashing the power of pim for batched transformer-based generative model inference,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.119842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.119842Z digest=sha256:f01ee1646655d482a4fa8213209806991991c40d0a3ea4719b9bdbf8de5152e6

Observation 8baebdbd-3461-4ad5-af9d-2e813169c9ca · outbound

This paper cites PIMoE: Towards efficient moe transformer deployment on npu-pim system through throttle-aware task offloading,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems PIMoE: Towards efficient moe transformer deployment on npu-pim system through throttle-aware task offloading,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.223436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.223436Z digest=sha256:635c3b94c82e86137ef0d04627d9938d0a05f83e7de394275933b692188b2788

Observation 87a6ccdb-83b2-44d8-8705-3b301705dfe3 · outbound

This paper cites Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.299533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.299533Z digest=sha256:835395fbe42d1de812d1e0a964ee625b0b20a3d34b8fbbfc3cf1190b81cb16dc

Observation d36e300a-70f9-4da5-b29a-b73369e0cc3b · outbound

This paper cites The true processing in memory accelerator,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems The true processing in memory accelerator,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.429188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.429188Z digest=sha256:f878523e6c5ba307de99658ad718bf0133f0d48471b19d0871413842f91ac68f

Observation 4b8310dd-19bb-4e6b-ae5c-e2295098443e · outbound

This paper cites LP-Spec: Leveraging lpddr pim for efficient llm mobile speculative inference with architecture-dataflow co-optimization,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems LP-Spec: Leveraging lpddr pim for efficient llm mobile speculative inference with architecture-dataflow co-optimization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.502980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.502980Z digest=sha256:0aec1c502dfd2e8c44605f422c181843543395f283d5b21e6eb9788dde793d8a

Observation 79d59443-36f5-41f6-9dac-53b5672c11c8 · outbound

This paper cites Ndpage: Efficient address translation for near-data pro- cessing architectures via tailored page table,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Ndpage: Efficient address translation for near-data pro- cessing architectures via tailored page table,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.608565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.608565Z digest=sha256:6e18ff6642e45dc26b1b1b0390acdf4d2cbde4834b0fc936f01939520dddb37c

Observation 43dcc1e0-9a50-4d7d-9a2b-346fdf6e1b1f · outbound

This paper cites Hyqa: Hybrid near-data processing platform for embed- ding based question answering system,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Hyqa: Hybrid near-data processing platform for embed- ding based question answering system,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.740231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.740231Z digest=sha256:ed27c06951d3995afd4f8656bef26a5ab694b930e9267a7b80301d3e315789a0

Observation 2d22f975-30f8-4062-8ad2-fe8b650b36ae · outbound

This paper cites Near-memory parallel indexing and coalescing: Enabling highly efficient indirect access for spmv,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Near-memory parallel indexing and coalescing: Enabling highly efficient indirect access for spmv,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.840345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.840345Z digest=sha256:1a9fab60ad4c8587b05e58b3f1e0a3ed68033b5593cbfe1cbe3d675715a3b1f6

Observation 8cc330cc-7484-4e7e-bbb6-1184a963bc1a · outbound

This paper cites Um-pim: Dram-based pim with uniform & shared memory space,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Um-pim: Dram-based pim with uniform & shared memory space,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.904066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.904066Z digest=sha256:7bfedc1e86a0ec7bce750a65fc1effcc0f02aef061dd19961d514afab9ef529f

Observation 9877c3c1-1c39-4302-b36c-c18530c54ea0 · outbound

This paper cites Bramac: Compute-in-bram architectures for multiply- accumulate on fpgas,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Bramac: Compute-in-bram architectures for multiply- accumulate on fpgas,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.006935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.006935Z digest=sha256:dc65b7803277ac8daea06caef4e48356806dd0c4a4147adcec1790dcb660d8df

Observation 320f2894-342f-4de9-af40-4b84f23cdf2e · outbound

This paper cites An overview of processing-in-memory circuits for artificial intelligence and machine learning,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems An overview of processing-in-memory circuits for artificial intelligence and machine learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.108408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.108408Z digest=sha256:567573369b1557a0d792265816e42ccc2b93550b5b6d8b1b4cd2fd76eea89860

Observation 68b97508-1b2b-421f-b195-90957737fd48 · outbound

This paper cites Mixtral of Experts.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Mixtral of Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.179344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.179344Z digest=sha256:f65ad97bbb1a3306be03e5203fa885269e70c5e202d26d2be4eac6fa6fcec5f6

Observation d58f14cc-db82-4710-a1c3-ea80e7dcac6d · outbound

This paper cites Aim: accelerating computational genomics through scalable and noninvasive accelerator-interposed memory,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Aim: accelerating computational genomics through scalable and noninvasive accelerator-interposed memory,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.293176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.293176Z digest=sha256:e311f98fa775186050dda9c2e10e0950dd77dd2bf81eededb6955b457f6361bd

Observation bc86f0e8-9fb7-4aff-9fbc-1eb2e95a0d7e · outbound

This paper cites (2025) intel-core-i7-14700.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems (2025) intel-core-i7-14700

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.350877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.350877Z digest=sha256:145640188797d34cf6418c7deef0ea61a3f4edd3379c1bf224ab064bd03d2996

Observation 5ce93888-8b09-4c08-beed-03005f3e05ea · outbound

This paper cites Ramulator 2.0: A modern, modular, and extensible dram simulator,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Ramulator 2.0: A modern, modular, and extensible dram simulator,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.452877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.452877Z digest=sha256:48cad18de1280153aa06f4085304d07fce3723fb672283476a771db9a3122ea2

Observation 3da4ecce-5822-4ad8-9ef6-7ccbcdd31636 · outbound

This paper cites DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.560097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.560097Z digest=sha256:4c1df172c1c997d199df57ef97fe7745581865b1ee6d4a0082eb6cc04368a0bc

Observation 07c7e0cb-b925-4b9e-8dd4-6bfa3fa1c13d · outbound

This paper cites Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.662460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.662460Z digest=sha256:739ef4026e7dfc81820871e44bd6d891a01bf442b6fc6a459fb29118a16f3a0e

Pith citing papers

No inbound Pith citation observations are available.