Pith. sign in

REVIEW 1 cited by

Multiplying Matrices Without Multiplying

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.10860 v1 pith:VYJAMUNX submitted 2021-06-21 cs.LG cs.ARcs.PFstat.ML

classification cs.LGcs.ARcs.PFstat.ML
keywords matrixmatricesmultiplyingbeenfasterlearningmachinemethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Multiplying matrices is among the most fundamental and compute-intensive operations in machine learning. Consequently, there has been significant work on efficiently approximating matrix multiplies. We introduce a learning-based algorithm for this task that greatly outperforms existing methods. Experiments using hundreds of matrices from diverse domains show that it often runs $100\times$ faster than exact matrix products and $10\times$ faster than current approximate methods. In the common case that one matrix is known ahead of time, our method also has the interesting property that it requires zero multiply-adds. These results suggest that a mixture of hashing, averaging, and byte shuffling$-$the core operations of our method$-$could be a more promising building block for machine learning than the sparsified, factorized, and/or scalar quantized matrix products that have recently been the focus of substantial research and hardware investment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lookup Table-based Multiplication-free All-digital DNN Accelerator Featuring Self-Synchronous Pipeline Accumulation

    cs.AR 2025-06 conditional novelty 6.0 of 10

    An all-digital MADDNESS DNN accelerator macro using a self-synchronous pipeline and 10T-SRAM lookup tables achieves 174 TOPS/W and 2.01 TOPS/mm2 in 22nm post-layout simulation.

Pith tools