Pith. sign in

REVIEW 4 cited by

Benchmarking TPU, GPU, and CPU Platforms for Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.10701 v4 pith:6ISI6BNW submitted 2019-07-24 cs.LG cs.PFstat.ML

classification cs.LGcs.PFstat.ML
keywords deeplearningmodelsplatformsbenchmarkperformancespecializedalong
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training deep learning models is compute-intensive and there is an industry-wide trend towards hardware specialization to improve performance. To systematically benchmark deep learning platforms, we introduce ParaDnn, a parameterized benchmark suite for deep learning that generates end-to-end models for fully connected (FC), convolutional (CNN), and recurrent (RNN) neural networks. Along with six real-world models, we benchmark Google's Cloud TPU v2/v3, NVIDIA's V100 GPU, and an Intel Skylake CPU platform. We take a deep dive into TPU architecture, reveal its bottlenecks, and highlight valuable lessons learned for future specialized system design. We also provide a thorough comparison of the platforms and find that each has unique strengths for some types of models. Finally, we quantify the rapid performance improvements that specialized software stacks provide for the TPU and GPU platforms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSTAR: Accelerating Diffusion Transformers via Spatial and Temporal Redundancy Reduction

    cs.AR 2026-07 conditional novelty 6.0 of 10

    DSTAR reports 7.33x latency speedup and 41.89x energy savings over an A100 GPU on seven diffusion transformers by quantizing differential activations to as few as 2 bits and reusing block-wise sparse attention scores.

  2. Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

    cs.SE 2025-01 conditional novelty 6.0 of 10

    A 55-criteria guideline and audit of 274 code benchmarks finds that most benchmarks skip data quality checks, prompting calls for more rigorous, reproducible benchmark construction.

  3. SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services

    cs.DC 2025-01 conditional novelty 6.0 of 10

    An SMDP-based dynamic batching policy minimizes a weighted sum of average response time and power consumption for batch service queues with size-dependent service times.

  4. DeepLL: Considering Linear Logic for the Analysis of Deep Learning Experiments

    cs.PL 2024-12 reject novelty 3.0 of 10

    A sketch showing how petri-net-style control flow and API resources of a deep learning training phase can be encoded as linear logic implications, with a proof of one reachability property.

Pith tools