Pith. sign in

REVIEW 14 cited by

MLPerf Inference Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.02549 v2 pith:HBVGSANQ submitted 2019-11-06 cs.LG cs.PFstat.ML

MLPerf Inference Benchmark

classification cs.LG cs.PFstat.ML
keywords inferencesystemshardwaremlperforganizationssoftwarebenchmarkbenchmarking
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Machine-learning (ML) hardware and software system demand is burgeoning. Driven by ML applications, the number of different ML inference systems has exploded. Over 100 organizations are building ML inference chips, and the systems that incorporate existing models span at least three orders of magnitude in power consumption and five orders of magnitude in performance; they range from embedded devices to data-center solutions. Fueling the hardware are a dozen or more software frameworks and libraries. The myriad combinations of ML hardware and ML software make assessing ML-system performance in an architecture-neutral, representative, and reproducible manner challenging. There is a clear need for industry-wide standard ML benchmarking and evaluation criteria. MLPerf Inference answers that call. In this paper, we present our benchmarking method for evaluating ML inference systems. Driven by more than 30 organizations as well as more than 200 ML engineers and practitioners, MLPerf prescribes a set of rules and best practices to ensure comparability across systems with wildly differing architectures. The first call for submissions garnered more than 600 reproducible inference-performance measurements from 14 organizations, representing over 30 systems that showcase a wide range of capabilities. The submissions attest to the benchmark's flexibility and adaptability.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FAIR+S: A validation study of a framework for sustainable research data and software

    cs.CY 2026-06 unverdicted novelty 7.0

    FAIR+S extends FAIR with sustainability metrics and is validated via expert survey confirming importance but revealing awareness gaps in green practices.

  2. Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining

    cs.CL 2026-04 unverdicted novelty 7.0

    Multilingual pretraining develops translation in two phases: early copying driven by surface similarities, followed by generalizing mechanisms while copying is refined.

  3. SLO-Guard: Crash-Aware, Budget-Consistent Autotuning for SLO-Constrained LLM Serving

    cs.LG 2026-04 conditional novelty 7.0

    SLO-Guard improves tuning budget consistency for SLO-constrained LLM serving by handling crashes explicitly and using a two-phase feasible-first exploration plus exploitation strategy.

  4. A Switch-Centric In-Network Architecture for Accelerating LLM Inference in Shared-Memory Network

    cs.AR 2026-03 unverdicted novelty 7.0

    SCIN uses an in-switch accelerator for direct memory access and 8-bit in-network quantization during All-Reduce, delivering up to 8.7x faster small-message reduction and 1.74x TTFT speedup on LLaMA-2 models.

  5. Edge-Inference Governors Need Memory-Clock State

    cs.PF 2026-06 accept novelty 6.0

    EMC-blind GPU-only latency fits miss 25–28% of tight deadlines on Jetson Orin; an EMC-aware two-cell refit holds misses ≤1.3% under a 2% QoS budget and selects a budget-feasible clock.

  6. The xPU-athalon: Quantifying the Competition of AI Acceleration

    cs.AR 2026-04 unverdicted novelty 6.0

    Quantitative benchmarks across recent AI accelerators reveal that optimal hardware choice varies with workload parameters and that several platforms incur substantially higher idle power than GPUs.

  7. Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures

    cs.DC 2026-04 unverdicted novelty 6.0

    Watt Counts supplies over 5,000 energy measurements across 50 LLMs and 10 GPUs and shows that hardware-aware selection can reduce server-scenario energy use by up to 70 percent with little effect on user experience.

  8. Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs

    cs.DC 2025-11 conditional novelty 6.0

    STAGE synthesizes high-fidelity Chakra-format execution graphs for distributed LLM workloads from symbolic tensor definitions, validated against real 128-GPU H100 traces and scaled to 32K GPUs.

  9. MatCreatioNN: Machine learning-guided computational discovery of photocatalysts for environmental applications

    cond-mat.mtrl-sci 2026-07 conditional novelty 5.0

    A ML funnel selects Zn- and Cr-based MOFs scoring 1.2-1.7x above PCN-224(Zr), but the comparison uses the same fitness score that selected them.

  10. Edge-Inference Governors Need Memory-Clock State

    cs.PF 2026-06 unverdicted novelty 5.0

    EMC state is required in latency models for edge inference governors; EMC-blind CPU/GPU fits miss 25-28% deadlines while EMC-aware refits limit misses to 1.3% and identify feasible energy points across vision and LLM ...

  11. HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation

    cs.SD 2026-04 unverdicted novelty 5.0

    A three-stage hierarchical AR model with dual-rate HuBERT/EnCodec tokens improves vocal-conditioned accompaniment generation, reaching FAD 1.71 and 51.5% preference vs ground truth on MUSDB18.

  12. HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation

    cs.SD 2026-04 unverdicted novelty 5.0

    HAFM uses a hierarchical autoregressive model with dual-rate HuBERT and EnCodec tokens to generate coherent instrumental music from vocals, achieving FAD 2.08 on MUSDB18 while matching prior systems with fewer parameters.

  13. Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference

    cs.PF 2026-05 unverdicted novelty 4.0

    Nvidia achieves 1.6x throughput with NVFP4 but hits a VRAM wall for 70B+ models, while Apple UMA enables linear scaling to 80B at 4-bit with up to 23x better energy efficiency.

  14. Copy First, Translate Later: Interpreting Translation Dynamics in Multilingual Pretraining

    cs.CL 2026-04 conditional novelty 4.0

    SLO-Guard, a crash-aware two-phase autotuner for vLLM serving, achieves no best-latency improvement over random search but demonstrates more consistent budget allocation across 150 trials on Qwen2-1.5B/A100.