Pith. sign in

REVIEW 2 cited by

Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13313 v1 pith:F46VGQI4 submitted 2024-09-20 cs.DC

classification cs.DC
keywords matrixperformancearchitecturesmultiplicationmultiplicationsaccuracyachievingapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study was aimed at simultaneously achieving sufficient accuracy and high performance for general matrix multiplications. Recent architectures, such as NVIDIA GPUs, feature high-performance units designed for low-precision matrix multiplications in machine learning models, and next-generation architectures are expected to follow the same design principle. The key to achieving superior performance is to fully leverage such architectures. The Ozaki scheme, a highly accurate matrix multiplication algorithm using error-free transformations, enables higher-precision matrix multiplication to be performed through multiple lower-precision matrix multiplications and higher-precision matrix additions. Ootomo et al. implemented the Ozaki scheme on high-performance matrix multiplication units with the aim of achieving both sufficient accuracy and high performance. This paper proposes alternative approaches to improving performance by reducing the numbers of lower-precision matrix multiplications and higher-precision matrix additions. Numerical experiments demonstrate the accuracy of the results and conduct performance benchmarks of the proposed approaches. These approaches are expected to yield more efficient results in next-generation architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mixed-precision numerics in scientific applications: survey and perspectives

    cs.CE 2024-12 conditional novelty 3.0 of 10

    A survey of mixed-precision numerical methods across CFD, climate, chemistry, and genomics, reporting speedups up to 8x on benchmarks and recommending co-design to unlock them.

  2. Hardware Trends Impacting Floating-Point Computations In Scientific Applications

    math.NA 2024-11 unverdicted novelty 2.0 of 10

    A review of floating-point hardware evolution and current AI-driven trends in reduced precision, mixed precision, and emulation.

Pith tools