Pith. sign in

Paper Citation Record · LEDGER

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2506.09554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09554 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:49:05.158865Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:07:58.959186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.580222Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8af69b5d-aa14-4aa8-a383-048688a79ac6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.534373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.534373Z digest=sha256:9f3deefe1ec77ac31c9bae148bd1bff0a6c226d343b3acf742f3966d5e56bd94

Observation 9f5b5f4c-0136-46ec-8aab-02981a2b5328 · outbound

This paper cites Plug in the safety chip: Enforcing constraints for llm-driven robot agents,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Plug in the safety chip: Enforcing constraints for llm-driven robot agents,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.996503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:03.625036Z digest=sha256:38f10553a4bfc17def65024dc852c30fa2967d87083de0dc881daad974ecd0a4

Observation 620bec5f-18e1-4327-a0c7-fcc3e9d3281c · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators A Survey on Efficient Inference for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.789514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.789514Z digest=sha256:21ac06a0ed284e16f8837dfc8000692bee2e5bb8161fdd11c4ae037b27fca5c3

Observation 89ecc77c-05b5-476d-8f20-896aca0967b4 · outbound

This paper cites LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.916941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.916941Z digest=sha256:34ec3243d4e8ae3ea85e8290ad71f51e9d56437f80e6055ac0d3a59ffa25c741

Observation b6633e4d-7b70-499f-83b9-4d4bc9029cc4 · outbound

This paper cites Characterizing the performance of accelerated jetson edge devices for training dnns,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Characterizing the performance of accelerated jetson edge devices for training dnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.859489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.079899Z digest=sha256:8556d52cd2d798efd6aebbad9c58d9fa9ac7e63bee68ede07f8da48475dc94a9

Observation 550cf011-0485-43ba-8ed4-d1302fef8095 · outbound

This paper cites Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:49:05.357455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.191036Z digest=sha256:ef1d3c3d3945a2ba6a6f792d513559d0a9ceac47fc976ad366b3f108d472e6a8

Observation 5952f333-940e-4bb0-8077-efa06d8ea389 · outbound

This paper cites A preliminary performance analysis of llm inference on edge accelerators,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators A preliminary performance analysis of llm inference on edge accelerators,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.675928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.296823Z digest=sha256:7f107bc9e62781e067ab6d6239253b9a13420e167599de93b246e29ca97eed5f

Observation 90bfe4cf-22eb-4e00-989f-2a49948496f8 · outbound

This paper cites Wikitext-2,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Wikitext-2,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.518047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.415311Z digest=sha256:f9c729a1211f1491209c1ed7fb115be09657dd8ea43a2027e7a91f1199ab76f9

Observation 32235912-1aff-4d2f-b2b7-0818a9672b08 · outbound

This paper cites Longbench,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Longbench,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.334059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.556604Z digest=sha256:82465afcb6874920aa1dc26dc243989288aecef962c039d21506d4442cce1d2a

Observation 37ed37ed-a5e6-4ed3-a6ea-ef8ce45a6286 · outbound

This paper cites Gpt3.int8(): 8-bit matrix multiplication for transformers at scale,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Gpt3.int8(): 8-bit matrix multiplication for transformers at scale,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.160849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.676505Z digest=sha256:866def5b62631066af24e05c98c5f1f84b45d137b648bc41570ac1dd1c618818

Observation 8315a6d7-71c0-4c55-ab46-ddabf1f332ec · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Splitwise: Efficient generative LLM inference using phase splitting

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:49:04.761986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:04.761986Z digest=sha256:a5248c575566cc7c822131e88460d3b9cd1b87a81a7762a0a662c90d08542465

Observation 0f1a1851-5edd-4ea6-b0f6-8c6664812145 · outbound

This paper cites When compared with INT4, INT8’s power savings range between 20% and 43% (median 32%).

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators When compared with INT4, INT8’s power savings range between 20% and 43% (median 32%)

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:49:06.022150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.852213Z digest=sha256:201d3aebc8fc482dcef25e3e15105a7a754786b14ad185624c539b7f8f42e169

Observation cc80d7e8-585a-4c02-b903-569a8cb2d6b3 · outbound

This paper cites Against INT4, INT8 consistently yields over 27% power savings.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Against INT4, INT8 consistently yields over 27% power savings

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:05.886948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:04.987410Z digest=sha256:395cca5ca3ab08585da4ba60ceea3d56781ea2df6893c063a157480ccf99f63b

Observation cb242ee6-3e20-4f56-bead-44e9b01bc9a0 · outbound

This paper cites •Energy Consumption: Similarly, INT8 achieves lower energy usage, with a median reduction of 24% compared to FP16 and 55% com- pared to INT4.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators •Energy Consumption: Similarly, INT8 achieves lower energy usage, with a median reduction of 24% compared to FP16 and 55% com- pared to INT4

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:05.737086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:05.068187Z digest=sha256:77938fcd957d1ff6c74564377ced13c49d3e24c8f67a20305b80f656bb64008d

Observation dfb5deec-c2a4-4ee6-b37a-dce61653e0fb · outbound

This paper cites an unresolved cited work.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:49:05.557627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T04:49:05.158865Z digest=sha256:ec80591545ae262c68d2607f989365250345a09052d3556049578b8e026d2f87

Pith citing papers

Observation 09e339dd-26dd-4613-8d3b-302929859ad7 · inbound

On the Sustainability of AI Inferences in the Edge cites this paper.

On the Sustainability of AI Inferences in the Edge Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:58.959186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:58.959186Z digest=sha256:331f4ee7b9860802a3567b08f8cbae5804e7c45aefb85dccd983853dc32073b0

Observation 056f7926-2373-437c-93e6-ce0d485aa266 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.582044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:cd0e29f0a7c5a96d355d21e52714f5a9a902f9fc0771d2d1b82587c17b7c6542

Observation bfb30276-44c1-4ed4-866f-de8350740453 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:1f4d7e3812243a33860c4ed8fe8eee891e883d8fd98d6f5571098e245d05a69b