Pith. sign in

REVIEW 1 cited by

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.02847 v1 pith:P74LQJEO submitted 2025-06-03 cs.AR cs.SYeess.SY

CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge

classification cs.AR cs.SYeess.SY
keywords edgeclonedevicesenergyllmsdeployinginferencemaintaining
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Deploying large language models (LLMs) on edge devices is crucial for delivering fast responses and ensuring data privacy. However, the limited storage, weight, and power of edge devices make it difficult to deploy LLM-powered applications. These devices must balance latency requirements with energy consumption and model accuracy. In this paper, we first quantify the challenges of deploying LLMs on off-the-shelf edge devices and then we present CLONE, an in-depth algorithm-hardware co-design at both the model- and system-level that intelligently integrates real-time, energy optimization while maintaining robust generality. In order to maximize the synergistic benefits of these algorithms in always-on and intermediate edge computing settings, we specialize in a 28nm scalable hardware accelerator system. We implement and extensively evaluate CLONE on two off-the-shelf edge platforms. Experiments show that CLONE effectively accelerates the inference process up to 11.92x, and saves energy up to 7.36x, while maintaining high-generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education

    cs.CY 2026-06 unverdicted novelty 4.0

    ELEVATE is a framework and prototype for deploying LLM-powered 3D avatar tutors locally on consumer hardware with a three-stratum design separating interaction, execution, and governance layers.