REVIEW 2 cited by
An Evaluation of Edge TPU Accelerators for Convolutional Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Edge TPUs are a domain of accelerators for low-power, edge devices and are widely used in various Google products such as Coral and Pixel devices. In this paper, we first discuss the major microarchitectural details of Edge TPUs. Then, we extensively evaluate three classes of Edge TPUs, covering different computing ecosystems, that are either currently deployed in Google products or are the product pipeline, across 423K unique convolutional neural networks. Building upon this extensive study, we discuss critical and interpretable microarchitectural insights about the studied classes of Edge TPUs. Mainly, we discuss how Edge TPU accelerators perform across convolutional neural networks with different structures. Finally, we present our ongoing efforts in developing high-accuracy learned machine learning models to estimate the major performance metrics of accelerators such as latency and energy consumption. These learned models enable significantly faster (in the order of milliseconds) evaluations of accelerators as an alternative to time-consuming cycle-accurate simulators and establish an exciting opportunity for rapid hard-ware/software co-design.
Forward citations
Cited by 2 Pith papers
-
Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput
Google proposes ML Productivity Goodput, a product of scheduling, runtime, and program goodputs, as a fleet-level metric for identifying and tracking efficiency improvements in large ML accelerator fleets.
-
A Unified Framework for Mapping and Synthesis of Approximate R-Blocks CGRAs
A CGRA design flow that maps neural network channels onto approximate DRUM multipliers and static voltage islands, reporting ~30% power reduction for MobileNetV2 with only output RMSE, not top-1 accuracy, as the quali...
Discussion (0). Continue with ORCID to comment.