Pith. sign in

Paper Citation Record · LEDGER

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

As of 17 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2608.05303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05303 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:48:05.592441Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact5
  • verified fuzzy25
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed695317-8d7c-41bc-837b-f2d21b5eca00 · outbound

This paper cites MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.370776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.370776Z digest=sha256:47a21708af96b454c4a8cf5815bb843b8b68da26113cb25395b30e542ffd1b9c

Observation f3f5d51a-8acc-4b44-ae49-dbfe0854c943 · outbound

This paper cites EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.701101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.374909Z digest=sha256:ae3edfa500fb4bb45c703b0fed2a92f890b8ad9ec5f2bbdea9d606d507a07cfb

Observation ae263b66-9294-4466-8a4f-00936893b512 · outbound

This paper cites SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SMoLPU: 122.1µJ/Token Sparse MoE- Based Speculative Decoding Language Processing Unit with Adaptive- Offload NPU-CIM Core,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.693024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.378530Z digest=sha256:237471159af89598e9811a80ff3bee7c67dd6f6a44f4edf3e07c564584650b3c

Observation 5b2ecdfb-c0dc-4ed0-a45a-9bcbf43469f4 · outbound

This paper cites GPT-4 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.381623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.381623Z digest=sha256:a0a476aefba3e3ef3a61512c042d60a1fc8cd89101aaf41740b01fdbce3257fa

Observation db22c0f0-b175-463e-9034-0cf8ee242061 · outbound

This paper cites The Llama 3 Herd of Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.385130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.385130Z digest=sha256:8ee9781bd40ce838e92030d5ec1ac1f2a7e8783626138dab63c372ecfbcef1d2

Observation 051c5fe7-3177-483f-a8ef-098e51f1b1e2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.388512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.388512Z digest=sha256:b1cc71eccfbce3ab9f19894d6febef31ae8348ed10b55a3cdc4dee339892f40a

Observation 4e9b8a5e-85bd-4193-bf90-dce9372c11d4 · outbound

This paper cites Squeezed atten- tion: Accelerating long context length llm inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Squeezed atten- tion: Accelerating long context length llm inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.392025Z digest=sha256:4eed1db6677a6c82a3c471c7c1a247aa585881f41d29519d02d2f45325ff8649

Observation 809ce551-d722-49fb-b103-277daee70d60 · outbound

This paper cites ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.685342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.394901Z digest=sha256:5adedb8aa3217376470f73bcc6c8090d3faf0c28300bc75dfcde5e3715c4c1e7

Observation ac65b07a-5a20-485d-aca6-501ffe8b44d5 · outbound

This paper cites 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 23.7 BROCA: A 52.4-to-559.2mW Mobile Social Agent System-on-Chip with Adaptive Bit-Truncate Unit and Acoustic-Cluster Bit Grouping,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.676879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.398249Z digest=sha256:db669ecc74be6d5e639152a1b31a3ba5732a08cec07bb3335f1edbf294c757b4

Observation ba28f471-836c-4fd3-a67c-ef42ce85d1c0 · outbound

This paper cites Fast on-device LLM inference with npus,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast on-device LLM inference with npus,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.667360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.401099Z digest=sha256:5c5a8afb7b999cdaeaff21e6e430e813edd884c8fb9c53bfbaf66c0d29ce1473

Observation 511be75b-39ab-4849-b3e2-28be5a08d4d8 · outbound

This paper cites C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding C-Transformer: An Energy-Efficient Homogeneous DNN- Transformer/SNN-Transformer Processor for Large Language Models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.658221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.404815Z digest=sha256:62a92dd5ce3d5dd2965b6fc47a9386db8f60241e917c08ad1ac306890a699dc6

Observation c4fb3c66-d9e3-42e2-8ef6-7be519996abf · outbound

This paper cites MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix Partition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.648874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.407945Z digest=sha256:36af487927dc6abd07168cd67a774e2d8c258c02cd180a9f75786be4b00dcbfa

Observation 1838530b-5a93-4da0-8371-ce4af794968f · outbound

This paper cites Granite 3.0 Language Models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Granite 3.0 Language Models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.639501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.411220Z digest=sha256:76f4de625e7795407254e84385f27f1c92b6914710e26a9c1aa370540c4d3d05

Observation 4abd0f5e-0d08-47be-a10b-d0e55a78211a · outbound

This paper cites Qwen3 Technical Report.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.417780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.417780Z digest=sha256:ac3ffb17cb7a1d8950c0dd8cc78b4af35742dc9efc628803453464d8829fd2ff

Observation a5d2494a-4a8b-45d7-94f5-a145dd962520 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.420966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.420966Z digest=sha256:ff42caaa6e9bd6ed9b95ce9ad6886d717e8d17b6c0e58f8762591134047cd05a

Observation b6b9e935-9268-43a5-bac9-02781e31ef48 · outbound

This paper cites GlaM: Efficient scaling of language models with mixture-of-experts,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding GlaM: Efficient scaling of language models with mixture-of-experts,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.621676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.424336Z digest=sha256:b41b6b7fb89435ed149913729afac931caf758fc28bad93d69e3e21bf5911526

Observation 664ae81a-5b4f-4a52-919c-51ae0d8ffa87 · outbound

This paper cites Mixtral of Experts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixtral of Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.427000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.427000Z digest=sha256:62fa55f20db75068a09382f23609405da6c1d8fcf079b025a4584e0740c3b79a

Observation 7a846a52-6e18-4057-9cca-fb3633c892d5 · outbound

This paper cites LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LLaMA-MoE: Building mixture-of-experts from llama with continual pre-training,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.612630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.430825Z digest=sha256:e0a929fb6548c10a8df347982a0074ffe3273c100cae42c28ec1644ce955afb5

Observation 88b87ab0-4e6a-4802-9d8f-310c81853211 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.433834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.433834Z digest=sha256:7391316d06701aeb183c812db2e27cf88aca6f20bdfea7250c09f740c55ecbb5

Observation 1823bac8-8314-4eac-823e-f9ef4c1401c8 · outbound

This paper cites Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.437417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.437417Z digest=sha256:16fbebe7ad2e26c2f90a9c5f500727653ec1556723927379085de3b96a7d97f8

Observation 8c9a80ed-3034-4919-b947-ed19cdf9a461 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast inference from transform- ers via speculative decoding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.440658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.440658Z digest=sha256:a58da0e3f7cb63702dc9f9239410f94832afe74d819eda12c8c2ddb7566dc045

Observation 961ebce8-44ad-4258-a45a-2eba2a942fc6 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.444304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.444304Z digest=sha256:c28abb3344d6f63218f079f525e8780aed0aa65afe36d6ca151730455dd51f0a

Observation e6f76a28-735c-4b5f-9efa-b3e29193f342 · outbound

This paper cites LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.447443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.447443Z digest=sha256:4a18738e7d9972be39d3e2a42432be5d5e39edb4618f115c4fd07c91af9a9465

Observation 7f16ba3b-926f-441d-91b0-9aa9211f3dee · outbound

This paper cites Speculative decoding with big little decoder,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Speculative decoding with big little decoder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.597995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.450810Z digest=sha256:0bb55416a857ab157c4f16c49833a8dceacd25f65d466546d3bdfdf723f16118

Observation eed58b92-de0e-40ab-9dd9-621f04b8eca5 · outbound

This paper cites EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.453468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.453468Z digest=sha256:1e20337ccdc01d62f62b2857f40091ac5111aed9f1b600c35c63a1d93313b5c4

Observation 85835bdf-8a21-4e13-bce0-e53038b268d5 · outbound

This paper cites ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.456995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.456995Z digest=sha256:1afed9ea5f6b8ab0bab54013c3f6ec242017d9cd0fd40588cef9655b1acd22f6

Observation fb2305c1-23da-42da-8dc2-8b7077efd693 · outbound

This paper cites EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EdgeLLM: Fast On-Device LLM Inference With Speculative Decoding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.588032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.460225Z digest=sha256:dde9bc99b05e50ce40366c67d07a0cb754238cff0c0f607e208b2de915bf8ca0

Observation cc9a7068-a057-40fd-b7fe-df79c43812e4 · outbound

This paper cites SpecMemo: Speculative Decoding is in Your Pocket.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecMemo: Speculative Decoding is in Your Pocket

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:48:07.063112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.463273Z digest=sha256:833e80a4154d4c744c19d4500accfc4c1182bbe7fd16d03b0d2485786bcdc051

Observation 82ee0902-2a62-4b5e-90af-0586845a430d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.466359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.466359Z digest=sha256:19386d0a24501d0d679764bd47b4523abb6228da746c6438a73d76cec98f18e5

Observation aa1bf1a1-79ba-4f9a-860e-05bd8e02feb6 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.469674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.469674Z digest=sha256:0f239589fa13af80bfcddba64050da5a87f7f3486c2d99334aa3d3cdfb582005

Observation abf13411-7af7-4b04-82e5-104d3b16802a · outbound

This paper cites MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoESD: Unveil Speculative Decoding’s Potential for Accelerating Sparse MoE,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.472875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.472875Z digest=sha256:299fc97d10129739e27bc320d6986d4fcdc9e56c28661bafb68c99cb74688748

Observation dfe7d41f-1e8f-4dcc-9ad0-2d43a8c8b853 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.475775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.475775Z digest=sha256:7f76affe0abc7eb97df9eefc2f6fc88889671b1e849ba9e90aa420ad47491644

Observation fed56ccb-64d4-4da0-84f1-215bb62753bd · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding OLMoE: Open Mixture-of-Experts Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.478932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.478932Z digest=sha256:25f461a4e0bc746877d85847f0e99c7aa1766cce1e3ad4950b24ff6bdf53ac59

Observation a17e14cd-43c7-43ad-a5bd-340afe275ac9 · outbound

This paper cites Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.482175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.482175Z digest=sha256:a2c85df883c2d00dcef3a4e88cac13feea623956343ecd9fa36afc5dbfc3ebbb

Observation ea6b6aec-b0aa-4b41-9583-7c5cc053537a · outbound

This paper cites Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.578776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.485453Z digest=sha256:12d353821d9376fa3133c43ef4f26ef4a2b47798544fbaa4dd5bd15c90aa9ebf

Observation 0251c8c5-b9f9-48cc-ab79-2af095ca6c76 · outbound

This paper cites Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.568814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.488860Z digest=sha256:3e4b95f8f2f5d7c7f617b27d189e52460f5d043139df836c19d6494c2f09cdc1

Observation e67b1ddc-3bec-4b49-9202-303e48de200a · outbound

This paper cites 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 20.8 Space-Mate: A 303.5mW Real-Time Sparse Mixture-of-Experts- Based NeRF-SLAM Processor for Mobile Spatial Computing,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.559672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.491756Z digest=sha256:50eed271c0afd6399cc733a80f37f7ba79f0caa43eb6be4b265303a376ca5d1a

Observation 8c15e677-9f27-46b1-bb99-04049c5179a9 · outbound

This paper cites SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification,

Reference 38

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.818755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.495329Z digest=sha256:2b03fee0a3439ed3cf99d308867465a1c6ced076dafae54fc7d3547ef63e3a75

Observation 1de07ba8-0bac-4dc7-a0db-c8acf3181940 · outbound

This paper cites SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.498093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.498093Z digest=sha256:d52e21879e337eb28a176e4b5762bc97f717e8ebf911d9eb8eefca921a31aa22

Observation d7160d2e-1d4a-4abe-a1ef-4aeb4170b491 · outbound

This paper cites Fast best-of-n decoding via speculative rejection,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Fast best-of-n decoding via speculative rejection,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.550479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.501944Z digest=sha256:9eb373d451e691bad3468bc483b1d3f4084fadbbf7f23e84ce578939f5166b70

Observation 2055f79f-efb3-4781-b2ee-6dea2c53cc7a · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.504790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.504790Z digest=sha256:e4e3191eb74f1f46c8555da8bfa359ea2a9131c2271c2869a34cd299ee9d22f8

Observation 5934ef46-3f3a-4b75-92bd-25a4b590bc57 · outbound

This paper cites Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.541433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.508383Z digest=sha256:1ed93374e2605ed2797aede433375622ffafff1fc70ebed6d101076362857ad8

Observation 8385853b-6233-4470-88db-f6eccc8cde97 · outbound

This paper cites Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Moe-i2: Compressing mixture of experts models through inter-expert pruning and intra-expert low-rank decom- position,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.532354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.511406Z digest=sha256:6e9e3a0315cc0410461a3ca4932d00901c72892ce703d8df08226faf62db4e11

Observation 5ed985b9-0324-446d-8bd4-5eb1f49936ce · outbound

This paper cites Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.515013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.515013Z digest=sha256:3c15026370b11ff8903431ff0d5ebd1e63f1b992435be586f14707a00c6f9821

Observation c2592ca9-03d6-4912-86e9-b7e0c254c2f5 · outbound

This paper cites Self-speculative decoding for on-device moe acceleration,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Self-speculative decoding for on-device moe acceleration,

Reference 45

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.627756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.518084Z digest=sha256:bd6b92487a2e6636299d59525b06fa3ccb6d3e2433d988adb50f6e61d828fdbc

Observation 5b962831-68d3-4223-a6e7-fc29abcb80b6 · outbound

This paper cites MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MoE-Spec: Ex- pert Budgeting for Efficient Speculative Decoding,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.521241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.521241Z digest=sha256:ddf7c702245badbec36eeba745a3047b111c099584d34851cefc6c1455985c41

Observation 671a1651-b0a5-4294-a807-7dd54cf9ccbf · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.524086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.524086Z digest=sha256:1ca8b44ffb6d3070372948e225ce14806ddb72471ab7aa44a24b64c3a7f2d3a0

Observation f549ddb7-a3b5-4147-a11c-f38f85d2d55e · outbound

This paper cites Emerging properties in self-supervised vision transformers,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Emerging properties in self-supervised vision transformers,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.527731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.527731Z digest=sha256:fa1e82ae5940a74246865a66e47f8bfd5dc97f4abc12316891858c2d924a92bb

Observation a764ec8c-961b-4968-96a1-09507f4cd167 · outbound

This paper cites Dense passage retrieval for open-domain question answering,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Dense passage retrieval for open-domain question answering,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.530625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.530625Z digest=sha256:a3bc4a865db4d10cac03e75a9f92348917c7a478fb3ab976dfe78f5ad9a001b8

Observation 00f90a41-8bbb-41cb-a3f2-5cb76184f4b1 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.534196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.534196Z digest=sha256:1b347f69db0c097151211863b49577c9510ab5ac4cf47ad8e3afcf0293576557

Observation 825ed448-8683-47ed-81eb-7260bb47b5e6 · outbound

This paper cites HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Ab- stract),

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.537368Z digest=sha256:5abf36810ffcc2a006e46162f1f19904cf5872c52e829e40b8c05c2d02809d2f

Observation 50900402-f295-4453-b4d4-e4d01cb9b06d · outbound

This paper cites Atp-llava: Adaptive token pruning for large vision language models,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Atp-llava: Adaptive token pruning for large vision language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.498714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.540750Z digest=sha256:913e8c7d36e6e53e06fda01271edea8b994053cd69e7a57eaedb7663811066b1

Observation 6bc4df81-91b6-4a07-8c93-e95bf1956905 · outbound

This paper cites all-minilm-l6-v2: Sentence transformers model,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding all-minilm-l6-v2: Sentence transformers model,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.489593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.543761Z digest=sha256:b9f8a79dff1e7141c5772d84918e43e2be9d60c5cc00d921489ceb302c5e573f

Observation 29504228-7919-4bce-8800-7b54eac0093d · outbound

This paper cites Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference,

Reference 54

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:06.180319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.547508Z digest=sha256:4bd1c64d96497ead675c0a29e38c319de97725dcd880c7c8834e37864d6cd0a0

Observation c2805c7d-735b-49ee-baec-2298e994a166 · outbound

This paper cites Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Sigma: A sparse and irregular gemm ac- celerator with flexible interconnects for dnn training,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.550292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.550292Z digest=sha256:8f4d7e119f4bb6c436b2073eff518c4077c55307f71a4c908b53676521268bd6

Observation 81e16d87-8484-431b-a7dc-94a1f2e2fd3d · outbound

This paper cites 2022, rev.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding 2022, rev

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.475928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.554023Z digest=sha256:2ba2a3cf5d3c88f8d10891aef97940209fa27a559307d81059aea351d478fb93

Observation d188d3c3-0f45-45c3-9001-e3110bebfdb0 · outbound

This paper cites SCNN: An accelerator for compressed-sparse convolutional neural networks,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding SCNN: An accelerator for compressed-sparse convolutional neural networks,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.467223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.556825Z digest=sha256:9f9b1a33c64f32b043c17dee16a9168e0dae7c5a6f15161845851968c7edbe56

Observation d61f043c-d1fe-4200-a1e9-596323a909b5 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.560519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.560519Z digest=sha256:d0215afe3f5f018579e75183c41fb854af1c6955676914dbf2d27d95f4173419

Observation 49e0ac6f-c82f-421f-870a-ebb14c9e39f8 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.563489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.563489Z digest=sha256:c0b8c44eedab4d8525232cbbad961b567b1d7757db14e4b3b1dfc5aeb77210a0

Observation ff4a0b63-6cce-403f-be61-5e4fa02f69d8 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Measuring Massive Multitask Language Understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.566935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.566935Z digest=sha256:28ca96fcd51adc653406f0a3908a5900aa44e5a3a152aa24714fe64017c536b1

Observation 78b6eddc-74c7-4cde-9c81-9e55b7cab3fd · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.569995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.569995Z digest=sha256:527103e486a8cf3544af8224d289c9f1039e23106bd9edcbbdc605d96e1ed47f

Observation 3a9f32e8-afbd-456d-870b-e8a4b73faf5d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.573207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.573207Z digest=sha256:0639650a6f8c5e062062aca5b110297702529ccea8f82a2635c3e1beffd8e150

Observation 69a296cd-70d0-443a-9e69-0552bbec13ea · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Winogrande: An adversarial winograd schema challenge at scale,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.576315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.576315Z digest=sha256:267ac26e923dc9ea20ec7abb8cc1cf1d52dc351ff884dc9d048fc140905577d0

Observation 88be2782-5a1d-43a9-b183-231a003128c7 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Piqa: Reasoning about physical commonsense in natural language,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.579324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.579324Z digest=sha256:c9491dc9ee615a3d7fac7a4b5101ac9c2fab569538ff651436be575a3e740ec5

Observation 5a9cb099-3d82-4296-8501-7d41cfc99c69 · outbound

This paper cites Pointer Sentinel Mixture Models.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Pointer Sentinel Mixture Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.582063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.582063Z digest=sha256:87721ee00b1250773305b01d1734a0be571f49cd8c395e5045c014caacfee076

Observation 3c6fcc8c-25af-4813-b23b-866bc30e9d7f · outbound

This paper cites MLPerf Inference interactive benchmark,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding MLPerf Inference interactive benchmark,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.443859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.585984Z digest=sha256:843b0e1d579b05e5b068e62c2594a36fe66b1568fde06de0250b8b01da418680

Observation dc73aff7-80fb-4e2a-8b99-4cc40d8d1e95 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Spatten: Efficient sparse attention architecture with cascade token and head pruning,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:48:05.589074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:48:05.589074Z digest=sha256:2f559554627fe8408600d911436aa2032ad5aeb6de1c7c5a42871ccf58285df6

Observation 0dbda71c-3190-4436-87ef-c3b05591e032 · outbound

This paper cites FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction,

Reference 68

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T15:48:05.921856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.592441Z digest=sha256:594c39a128888a6939d87218b9b752f27d190f7e205bdf8f7b19dd2b4682f699

Observation 55ad2281-72cf-4765-8890-eabf11a64b56 · outbound

This paper cites Available: https://github.com/ibm-granite/granite-3.

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding Available: https://github.com/ibm-granite/granite-3

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:48:07.630644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T15:48:05.414109Z digest=sha256:cfbec73e15b23275183131d2c4e1302e01a91e06ad7311a47dac73d07c60399a

Pith citing papers

No inbound Pith citation observations are available.