Pith. sign in

Paper Citation Record · LEDGER

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2505.23219.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23219 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:10.198594Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24ae3333-f6f1-45e2-97c8-e76e0b08e213 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Efficient memory management for large language model serving with pagedattention,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:06.403875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:06.403875Z digest=sha256:3549392c1c94354421e65d3c430e661cfde398a0ca03f437322e24b49e1f667d

Observation 91206771-a7bc-4a05-a2bb-0b33cdef305d · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and ver- ification,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Specinfer: Accelerating large language model serving with tree-based speculative inference and ver- ification,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:16.911548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:06.537216Z digest=sha256:f59b99d3f96320fba6e007d907342bb7be70ec80fb4b1ab4c01f81eff1b55f33

Observation 9fb6f691-5aac-4615-bf7b-817a44725dce · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:06.695859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:06.695859Z digest=sha256:1bb3cf4c76afd83c573b699568015312d01e695fed5a5746edcf51a8930c2a47

Observation 94627f43-b782-4c43-bcd0-91244600d07d · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Break the sequential dependency of LLM inference using lookahead decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:16.657708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:06.850459Z digest=sha256:e4ad10806cb9e8d06649b924fd648695fb17b809bc9d5abead6a93eb3a9de15f

Observation a371d8a0-6efc-4a5c-9ca1-5897462139c6 · outbound

This paper cites Codl: efficient CPU-GPU co-execution for deep learning inference on mobile devices,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Codl: efficient CPU-GPU co-execution for deep learning inference on mobile devices,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:06.977455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:06.977455Z digest=sha256:c9064420a9958e063a35b56cd6d8f81bcabeea010f28b74b465a312a7244ceb1

Observation 4e581d66-b47b-4ca8-b845-521ed67b90f7 · outbound

This paper cites Edgenn: Efficient neural network inference for cpu-gpu integrated edge devices,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Edgenn: Efficient neural network inference for cpu-gpu integrated edge devices,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:16.396266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.117354Z digest=sha256:a9847fa64d689b96c413f22fcc3bbc155b153835e9daa3d66443952cf4f32b90

Observation 0c9e5fbb-7dfa-4666-9e70-05434b1d00ee · outbound

This paper cites High-throughput cnn inference on embedded arm big. little multicore processors,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism High-throughput cnn inference on embedded arm big. little multicore processors,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:15.960000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.252057Z digest=sha256:17d2359c530af4ac5301bb206a0a7a3b2040289350e90a7771566ce49e368a6a

Observation 32fc7452-b213-4759-93c1-5844b13f9b00 · outbound

This paper cites Dopia: online parallelism management for integrated cpu/gpu archi- tectures,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Dopia: online parallelism management for integrated cpu/gpu archi- tectures,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:15.597486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.376808Z digest=sha256:4926ff37c50e89d362292c1ab8ff7a946edd553ac3a8ea59271ad62cd20959ac

Observation c2baf9dd-8904-47ae-9ed1-bba66802decb · outbound

This paper cites Apple m4,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Apple m4,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:15.312135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.475644Z digest=sha256:4b0ee32a3a6b6a9d08ecdef8f0de229b2c1629c8fbc25d061ab94204c2e1ae34

Observation a9dad62b-73f1-495f-8299-3ae6003b911f · outbound

This paper cites vllm github,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism vllm github,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.987334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.573662Z digest=sha256:72b20e2634cce283dcf0e215a7cab1bfb4a0795a331e0c9df5dace6fe9cba353

Observation 58b9bea6-9407-493a-97a4-732d593c3e7f · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:07.687263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:07.687263Z digest=sha256:52995ec7cdac7e66a54dfdd41c2f8704f19fe773e2e845336faf6bd5f0a805a8

Observation cefeb681-f9e5-4b0f-bc37-fde6ef58283c · outbound

This paper cites Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:07.770435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:07.770435Z digest=sha256:d1d328d54787ddfe6b80bf630efabf2c4dcd525a549d3fe48dccc68c2c958c47

Observation 7432d218-f45f-4fdd-8522-98a42ebbed96 · outbound

This paper cites Intel core ultra processor family,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Intel core ultra processor family,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.692481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.867375Z digest=sha256:1bb6c116018a6e4f1632ba37bb831863f5514c8339f82ce18953e0305bd781b1

Observation f3fc7d8d-41d0-46f8-9d9f-ff0a5f54c87e · outbound

This paper cites Meet jetson, the platform for ai at the edge,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Meet jetson, the platform for ai at the edge,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.344380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:07.993100Z digest=sha256:573a06d1e33eeed491316f663b4e634d8e4cdddf93645da7b1b2e140c6688149

Observation f5589f97-86ef-487d-8f49-38482ae3dc32 · outbound

This paper cites Communication effi- cient distributed machine learning with the parameter server,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Communication effi- cient distributed machine learning with the parameter server,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.024404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.088083Z digest=sha256:45ebe2b68680c5a78ebc09daef773eb08cc6f0759eaf2b5d5d51d77ed3b0e194

Observation f016866a-0884-484b-9494-5f9d5be016f1 · outbound

This paper cites Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:13.716687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.196608Z digest=sha256:a48c14ce2911d4e574be78dbc4f2326e6c7ae85365a0e373f724491abfc39c36

Observation 5a4f5d34-2367-49d8-981a-1361d9009a5b · outbound

This paper cites Codl: efficient cpu-gpu co-execution for deep learning inference on mobile devices.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Codl: efficient cpu-gpu co-execution for deep learning inference on mobile devices

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:13.472004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.363551Z digest=sha256:9d8cfe12cbdfaa5f11b1c5404792f22ed3bee7efe4335965684cba9528b74024

Observation 9c66d894-9c9a-4eb7-98b3-69b0b4c28a22 · outbound

This paper cites Attention is all you need,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Attention is all you need,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:08.445461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:08.445461Z digest=sha256:c1b880cb34259a221aecbc7d6453c619f23b1764198b5ba96c34784b324fb755

Observation 5bb6186d-b056-45b1-953a-ddd20e647f17 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:08.525545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:08.525545Z digest=sha256:fea72c54132afe477ce442e19e8ef6f07afc269b8f3a51ef78ff646e071d199b

Observation f6096911-0308-49be-a5c8-4507020568d7 · outbound

This paper cites EAGLE: speculative sampling requires rethinking feature uncertainty,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism EAGLE: speculative sampling requires rethinking feature uncertainty,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:13.169659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.621537Z digest=sha256:37e2be0d63d3a260d1f1b275f0ab4a36e6b26b4005405927c87939f3e7a7bfd3

Observation 9f73d127-3366-4a33-8495-f3a5a873b60d · outbound

This paper cites Apple a17,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Apple a17,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:12.863170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.693852Z digest=sha256:eb32c700d608c8c72952417d717c58013a14ce065725c00cc83fcde199b479d2

Observation d9580a79-5031-479f-be40-09730113aa80 · outbound

This paper cites Amd reveals next-gen desktop processors for extreme pc gam- ing and creator performance,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Amd reveals next-gen desktop processors for extreme pc gam- ing and creator performance,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:12.552778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.789304Z digest=sha256:1788f1438f26cd62c325eda2cadb4f6e65d4f4542e649908ed64a7f51e81b0d1

Observation 0a0d0a69-c945-4d6d-b371-6b32b9cb5ad6 · outbound

This paper cites Qualcomm snapdragon,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Qualcomm snapdragon,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:12.235562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:08.899326Z digest=sha256:9b28e68611aa2d5af0ae7329bcbef63caa524d15de05d16e8468fdee6cc26758

Observation 0ee5fff8-57fe-4f4d-8b57-00257fc7c042 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.016191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.016191Z digest=sha256:e419a634d777692f4ec6ffc7df148b2af8fe8b752a1ef7009d8bedbbd2e662bc

Observation 29985fa0-3612-487d-8f4e-3f1164b86d7e · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.920095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:09.116736Z digest=sha256:1b806b113b2998ae5ae9f88eed638803222f2062918679c13f1ffe85dc452986

Observation 9b6bcb50-3f38-4bdc-9901-6ba9cc0e51d5 · outbound

This paper cites Wave quantization,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Wave quantization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.668637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:09.228587Z digest=sha256:03bf8497cc57a206fdc6fdadb9ead7cf7d0f3f7a4af663360a33a666935b1553

Observation f69929ea-c697-4a01-94f8-042787a02f7c · outbound

This paper cites Nvidia fastertransformer: Transformer related optimization, including bert, gpt,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Nvidia fastertransformer: Transformer related optimization, including bert, gpt,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.430108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:09.390790Z digest=sha256:82a82ac2195454feb546a7f3aee026f570c7be6dedcb804886e8b4a5a4df0854

Observation 9f84d62a-3cc6-4af2-a63b-164c31ab2b39 · outbound

This paper cites Ctranslate2,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Ctranslate2,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.228226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:09.458408Z digest=sha256:efe486c1635e35d47bcb8fb4d4a5f796859b03cf49be02c0a485d02ca53155c0

Observation f060c808-1a7c-4988-9be3-720d8f49e024 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.524187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.524187Z digest=sha256:dfd5aeb504210c1546cec8586bf48b14e8f2e217d8b3629795110e325d7ccee3

Observation cbe957d3-64d0-4562-ac55-ab52abffefe9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.585286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.585286Z digest=sha256:5a8dd6dc38dd1fbf9eebc4176fe9c02c265d3f1295394b04fa378ea4627a1c00

Observation 985f7dad-d686-4f00-91bd-36f94c835bb3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Training Verifiers to Solve Math Word Problems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.637595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.637595Z digest=sha256:b973a96acb3c98ea2e3c6e570646f0353a6cb826dc7343712a0916c4c8859865

Observation b32f7167-717d-43fe-93a9-d99c0b8e39ff · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Evaluating Large Language Models Trained on Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.703036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.703036Z digest=sha256:a97b33f1806f7f15567dbf21e3cb3f79caacbdbb3f918e579c66f99fe374f49c

Observation eaae6fc4-7c10-4305-896f-fc8f7963e2d3 · outbound

This paper cites Program Synthesis with Large Language Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Program Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.762157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.762157Z digest=sha256:3f3dffc2c247dbe00eae9ebb485f810eccc96b1aa5c7a3902159809d31ac167e

Observation 0de30b68-523e-419a-a089-cada98f314b7 · outbound

This paper cites Edgenn: Efficient neural network inference for CPU-GPU integrated edge devices,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Edgenn: Efficient neural network inference for CPU-GPU integrated edge devices,

Reference 34

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:02:10.643834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:09.831626Z digest=sha256:0e792c5075628c5d5c12e064fabb6927a93cf4beeff7562b1e1e028a7064a1b4

Observation 5c1462e8-01ac-4aa8-b1ee-df2badc6a558 · outbound

This paper cites Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.067511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:09.895407Z digest=sha256:6d94f1f20704f7e42dbf9ee8957be3a4597bf3096c34d8cea1b15e1562db2a98

Observation c58228b5-db65-4718-9996-92174d32b9aa · outbound

This paper cites Communication-efficient model parallelism for distributed in-situ trans- former inference,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Communication-efficient model parallelism for distributed in-situ trans- former inference,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.955947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.955947Z digest=sha256:4ccf40f7302e77ba7d61cc378d130bdf69ef61fa1d86a5fed9eb3ff167d049e3

Observation 3d3743d6-38b0-495b-a278-4c1245722ebb · outbound

This paper cites Petals: Collaborative Inference and Fine-tuning of Large Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Petals: Collaborative Inference and Fine-tuning of Large Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.010089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.010089Z digest=sha256:2c584f60a6e3ccfd2c018963fa7a7bf61cd1a47039cfbb1b0e2b7930076612bf

Observation 0d670f62-08ce-4a51-8d72-5d4dec06a26e · outbound

This paper cites Asymo: scalable and efficient deep-learning inference on asymmetric mobile cpus,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Asymo: scalable and efficient deep-learning inference on asymmetric mobile cpus,

Reference 38

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:02:10.429864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:02:10.076767Z digest=sha256:5bec61045442a3a13fd5fd902c5bd650aa55c5790db1f2d0f6eb27300796b36a

Observation d0fb433f-8dc3-4c79-b65e-8b1912fdd96f · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.146184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.146184Z digest=sha256:bfdb96503e32cd0d04cb48d22a45356efe10a2f705ed1800731d038800886587

Observation 4012f9de-80bc-48e5-bbbf-2660dc0bc53f · outbound

This paper cites LLMCad: Fast and Scalable On-device Large Language Model Inference.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.198594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.198594Z digest=sha256:3898ee62d247ba48e877aa727cf0657eb130ba9263d6f356079f64624e3f9cea

Observation 84bb3486-6f11-49d1-900f-b3d522f463cb · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:08.290303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:08.290303Z digest=sha256:bda3f6081d804dc8701161d35ff7769af22dd18604473503bf9d2d990d5783cd

Pith citing papers

No inbound Pith citation observations are available.