Pith. sign in

Paper Citation Record · LEDGER

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems

As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2601.03992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.03992 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:12:49.662460Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fb908f54-7c62-445c-904c-a84efd3f76aa · outbound

This paper cites Attention is all you need,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Attention is all you need,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.246249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.246249Z digest=sha256:30a8e299286ff64966062731c6839cd92e3d5e8644371b429ee17bbb4b5fa9a5

Observation 9c069319-6e8e-4f6b-b780-899828ee8ad3 · outbound

This paper cites Adaptive mixtures of local experts,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Adaptive mixtures of local experts,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.330770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.330770Z digest=sha256:e03c09bd8adbb3b35d64b1ee4e86549913d633f92fb278d910bad8ce9e1303a5

Observation 040eb670-725b-49f8-aeab-6dc394e1cbd8 · outbound

This paper cites Outrageously large neural networks: The sparsely- gated mixture-of-experts layer,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Outrageously large neural networks: The sparsely- gated mixture-of-experts layer,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.481094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.481094Z digest=sha256:84ac04dbfe7d8fc76e89da0929461718f8cded62721a3be7e7dd8e2ca9e4e6ae

Observation c8ba9011-0790-4f7f-900c-a4856e6ce203 · outbound

This paper cites GPT-4 Technical Report.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.584804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.584804Z digest=sha256:605a16cc0b79ee89abdec38f794ed1667962f5b45500c6cc845f1fa3098c67b1

Observation 798e843b-507d-46db-9c8a-5795a5da64a2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems LLaMA: Open and Efficient Foundation Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.682060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.682060Z digest=sha256:b0f81cee113c4debf6f640889b336ce6dc752c36b33e91393e1ee024ab26c15f

Observation b9f9bbad-12eb-4684-9c98-b3d3884a54c3 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems OPT: Open Pre-trained Transformer Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:45.857396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:45.857396Z digest=sha256:3117e1281222e260d2468863996e684a3d4a6b88bcb406360b56dcc6c036c989

Observation 77d96968-f5a7-424e-a377-1b2cd3680072 · outbound

This paper cites Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Anda: Unlocking efficient llm inference with a variable- length grouped activation data format,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.032169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.032169Z digest=sha256:f8bee8a457a946156a97a958397336d375cd3c356b4fc617e08e84cc026a33ee

Observation c2580717-12ba-485c-a366-3a679e19a3af · outbound

This paper cites A survey on mixture of experts in large language models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems A survey on mixture of experts in large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.199042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.199042Z digest=sha256:383168a404d89adfd5b9a04f49e74f2aaa5295d09f3c4189e8e764a3a2c93c8c

Observation dd75ece3-f377-482d-9eb4-1c96fe40361f · outbound

This paper cites (2025) GeForce RTX 5080.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems (2025) GeForce RTX 5080

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.372783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.372783Z digest=sha256:8dc1b831f9bfb115ba961f7fd64fe4beac7135d3c0abf4c6695ca97b199908eb

Observation 6f3bd0c9-0b32-4f35-bf22-b8e86318e18c · outbound

This paper cites Qwen3 Technical Report.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Qwen3 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.501611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.501611Z digest=sha256:d483fc49ba19fd73b82af6d73171d23c2f5e0fa6faf2c22fcd338d22b268af40

Observation 4f595b27-c512-46d0-8bb1-d8aa7b4ae2b0 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.637460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.637460Z digest=sha256:3c976f60ac03acf492f6cc341f028ac58c4d36280f067a510b06bd483c931a17

Observation 7d60b81f-dad6-4b85-90da-0421a5b5a144 · outbound

This paper cites Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Klotski: Efficient mixture-of-expert inference via expert- aware multi-batch pipeline,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.731223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.731223Z digest=sha256:159a3bdcf3f2af79d005dfa6c39804f20f7ded23eb4013aa8dbb40db4d7c3893

Observation 55c79c1e-5af5-4994-ab70-2257e512a681 · outbound

This paper cites Fast Inference of Mixture-of-Experts Language Models with Offloading.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fast Inference of Mixture-of-Experts Language Models with Offloading

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:46.908768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:46.908768Z digest=sha256:682cd33a0eb23d4d258ad509fb2e03e9bc2afef56859d41b5541fa6d194571dc

Observation b6305259-e3db-461a-8787-04ef004c4b67 · outbound

This paper cites DAOP: Data-aware offloading and predictive pre- calculation for efficient moe inference,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems DAOP: Data-aware offloading and predictive pre- calculation for efficient moe inference,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.073480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.073480Z digest=sha256:dc9179f0e63187606e9760f3988208f3ada34c7172e44b79b5d1c3d11eb175e3

Observation c78c45cc-9a64-4796-bc27-61a70a4db24a · outbound

This paper cites Moe-lightning: High-throughput moe inference on memory-constrained gpus,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Moe-lightning: High-throughput moe inference on memory-constrained gpus,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.134788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.134788Z digest=sha256:af7e8b24cd39ab7eb5f39e587a19b5142a61ffecc58adb2e4ff5ea56319489b4

Observation dfb7a272-16ae-4c75-8266-941ea208f832 · outbound

This paper cites Fiddler: CPU-GPU orchestration for fast inference of mixture-of-experts models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Fiddler: CPU-GPU orchestration for fast inference of mixture-of-experts models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.176571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.176571Z digest=sha256:d4f1641639848a4abe20b4e266abd5baa4239332fa682d783b4e8284dc228b71

Observation d018a23c-8b61-4c64-8ffa-a5e838a69b61 · outbound

This paper cites Monde: Mixture of near-data experts for large-scale sparse models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Monde: Mixture of near-data experts for large-scale sparse models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.263000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.263000Z digest=sha256:0754dd45998c67515fa730824ae9b819328687143005de6c55a1a5d711a1f571

Observation 5095e32c-e81e-4822-bcdc-f2f23f1d6f58 · outbound

This paper cites Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.351344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.351344Z digest=sha256:6301c2bbebc570790ec1c93596281ba8c21b76bb119d813c19f9f1ba5f615814

Observation b453858d-4086-434b-a15a-744210aede6d · outbound

This paper cites Co-designing binarized transformer and hardware accel- erator for efficient end-to-end edge deployment,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Co-designing binarized transformer and hardware accel- erator for efficient end-to-end edge deployment,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.419157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.419157Z digest=sha256:ed06a81f6e4a3e56614e85bbc0355b85c3efbea699cd19045d1f7b2d067f9306

Observation 9156e491-9016-40bb-9e00-5ef4c67a2c51 · outbound

This paper cites Language models at the edge: A survey on techniques, challenges, and applications,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Language models at the edge: A survey on techniques, challenges, and applications,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.512321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.512321Z digest=sha256:877d5906bf560a73476b6c5a0ff958a4c52b670482e642a8d894f381586797ae

Observation 1fdc3f5b-7058-474d-8dd4-8571e2cec33d · outbound

This paper cites A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems A precision-scalable risc-v dnn processor with on-device learning capability at the extreme edge,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.604670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.604670Z digest=sha256:85f169f9b4cdb070755d1d5999ceef28c3e85a9e4fec115706a49892c020d4e7

Observation f9aab894-9168-4329-a9c2-3d5d1ab96678 · outbound

This paper cites Spark: Scalable and precision-aware acceleration of neural networks via efficient encoding,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Spark: Scalable and precision-aware acceleration of neural networks via efficient encoding,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.700109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.700109Z digest=sha256:e4923eacc95d6e7a2859923905a270fbf29c415660392fa9d9c17ec1ab4a01a5

Observation 72c14f80-ecbd-435b-aa5f-51eebba247bd · outbound

This paper cites Medusa: Simple llm inference acceleration framework with multiple decoding heads,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Medusa: Simple llm inference acceleration framework with multiple decoding heads,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.790247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.790247Z digest=sha256:db7cdd24d59853a0c81fefae406b5bf00472940c324ffca2e5d817270999a71b

Observation 2cc085ba-503a-4bcf-bbbe-134d1b3aef5e · outbound

This paper cites Recnmp: Accelerating personalized recommendation with near-memory processing,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Recnmp: Accelerating personalized recommendation with near-memory processing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:47.878841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:47.878841Z digest=sha256:0bbf929d48dad455b02dfbf2a0895f3976f896b86c2ab96d1d068c48438455d4

Observation 6b0bf0c4-6d8e-4874-af4e-e49eca8e287a · outbound

This paper cites Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.035954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.035954Z digest=sha256:55c192cf87537ff84c90264c3eb79261d229557f96fe59f844f9450f054bb591

Observation 23489549-cbba-4acd-bd44-1773d0a8c3ab · outbound

This paper cites Attacc! unleashing the power of pim for batched transformer-based generative model inference,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Attacc! unleashing the power of pim for batched transformer-based generative model inference,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.119842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.119842Z digest=sha256:5be9ab5758d83951c1189eccb574b8532f2d85a13c825583ed4038ee0aad3c7c

Observation 8baebdbd-3461-4ad5-af9d-2e813169c9ca · outbound

This paper cites PIMoE: Towards efficient moe transformer deployment on npu-pim system through throttle-aware task offloading,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems PIMoE: Towards efficient moe transformer deployment on npu-pim system through throttle-aware task offloading,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.223436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.223436Z digest=sha256:501021e8419f510bfde0655ae5831ffcb5c8f71dc44607ca0cd7bce47baad4a0

Observation 87a6ccdb-83b2-44d8-8705-3b301705dfe3 · outbound

This paper cites Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Make llm inference affordable to everyone: Augmenting gpu memory with ndp-dimm,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.299533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.299533Z digest=sha256:4346cd7d01816df80ff8263753c30322ea10dd13b8f9328803916c0a694fa4bf

Observation d36e300a-70f9-4da5-b29a-b73369e0cc3b · outbound

This paper cites The true processing in memory accelerator,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems The true processing in memory accelerator,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.429188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.429188Z digest=sha256:884b11aafaadb38efbe32642c13b8a20cef59a7274a3224c5c2641109c9d79a4

Observation 4b8310dd-19bb-4e6b-ae5c-e2295098443e · outbound

This paper cites LP-Spec: Leveraging lpddr pim for efficient llm mobile speculative inference with architecture-dataflow co-optimization,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems LP-Spec: Leveraging lpddr pim for efficient llm mobile speculative inference with architecture-dataflow co-optimization,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.502980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.502980Z digest=sha256:c20a62b9248a25b834ec5e48646f9eb436c29ad6aa921428ad289b4cd953139d

Observation 79d59443-36f5-41f6-9dac-53b5672c11c8 · outbound

This paper cites Ndpage: Efficient address translation for near-data pro- cessing architectures via tailored page table,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Ndpage: Efficient address translation for near-data pro- cessing architectures via tailored page table,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.608565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.608565Z digest=sha256:d91b0e3e336944351afc60331db1bf3a790269bd6c8f65b7c46efc5973564bab

Observation 43dcc1e0-9a50-4d7d-9a2b-346fdf6e1b1f · outbound

This paper cites Hyqa: Hybrid near-data processing platform for embed- ding based question answering system,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Hyqa: Hybrid near-data processing platform for embed- ding based question answering system,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.740231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.740231Z digest=sha256:26f09b43cc71a2b8e6314c9a509ef57245f3a76928af5562ba9cdc701bdeb5c5

Observation 2d22f975-30f8-4062-8ad2-fe8b650b36ae · outbound

This paper cites Near-memory parallel indexing and coalescing: Enabling highly efficient indirect access for spmv,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Near-memory parallel indexing and coalescing: Enabling highly efficient indirect access for spmv,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.840345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.840345Z digest=sha256:06d235fb008450dbc11dd9e2d7c86c9c672b99664b7aee150e15138df1ddc598

Observation 8cc330cc-7484-4e7e-bbb6-1184a963bc1a · outbound

This paper cites Um-pim: Dram-based pim with uniform & shared memory space,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Um-pim: Dram-based pim with uniform & shared memory space,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:48.904066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:48.904066Z digest=sha256:454e2dbbeb88826656d73e4995dc4f4a6712efbfa7125c379679f45e1eb3cdec

Observation 9877c3c1-1c39-4302-b36c-c18530c54ea0 · outbound

This paper cites Bramac: Compute-in-bram architectures for multiply- accumulate on fpgas,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Bramac: Compute-in-bram architectures for multiply- accumulate on fpgas,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.006935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.006935Z digest=sha256:ecd3755d309e82f5aa5aee976659c0946ecf4599b5115d6532e768875bb375ce

Observation 320f2894-342f-4de9-af40-4b84f23cdf2e · outbound

This paper cites An overview of processing-in-memory circuits for artificial intelligence and machine learning,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems An overview of processing-in-memory circuits for artificial intelligence and machine learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.108408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.108408Z digest=sha256:dde9bf5a09656506857a4a798571c68ae63e015e34c497dd6c00a0c297712787

Observation 68b97508-1b2b-421f-b195-90957737fd48 · outbound

This paper cites Mixtral of Experts.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Mixtral of Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.179344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.179344Z digest=sha256:77a802ea523e3440ce4652009258d9fb5bf47e33bdb26eb3701e48a3313dd0db

Observation d58f14cc-db82-4710-a1c3-ea80e7dcac6d · outbound

This paper cites Aim: accelerating computational genomics through scalable and noninvasive accelerator-interposed memory,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Aim: accelerating computational genomics through scalable and noninvasive accelerator-interposed memory,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.293176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.293176Z digest=sha256:f4739d134c730aa2049e8900de7d4da96d152b1a8cd0c7491e5af0ec0ab9408b

Observation bc86f0e8-9fb7-4aff-9fbc-1eb2e95a0d7e · outbound

This paper cites (2025) intel-core-i7-14700.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems (2025) intel-core-i7-14700

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.350877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.350877Z digest=sha256:98b9ffa4f14006e3c7122ccc8a89552178a58d5b37134ef6f9313dd941de1703

Observation 5ce93888-8b09-4c08-beed-03005f3e05ea · outbound

This paper cites Ramulator 2.0: A modern, modular, and extensible dram simulator,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Ramulator 2.0: A modern, modular, and extensible dram simulator,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.452877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.452877Z digest=sha256:66928c1deacc749d4fc8c51cc0cc1a7f0fe1af0fd55115e607dbc4907f15e97f

Observation 3da4ecce-5822-4ad8-9ef6-7ccbcdd31636 · outbound

This paper cites DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.560097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.560097Z digest=sha256:cc124749ca0b45020755643e9b852293a87284f96d3660f34c020fedee239c01

Observation 07c7e0cb-b925-4b9e-8dd4-6bfa3fa1c13d · outbound

This paper cites Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle.

A Scheduling Framework for Efficient MoE Inference on Edge GPU-NDP Systems Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T12:12:49.662460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:12:49.662460Z digest=sha256:83d39e20090e6259690ea0a51b23cadd495f65b432e1584b7999d24d74b27b95

Pith citing papers

No inbound Pith citation observations are available.