REVIEW 11 cited by
The Computational Limits of Deep Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep learning's recent history has been one of achievement: from triumphing over humans in the game of Go to world-leading performance in image classification, voice recognition, translation, and other tasks. But this progress has come with a voracious appetite for computing power. This article catalogs the extent of this dependency, showing that progress across a wide variety of applications is strongly reliant on increases in computing power. Extrapolating forward this reliance reveals that progress along current lines is rapidly becoming economically, technically, and environmentally unsustainable. Thus, continued progress in these applications will require dramatically more computationally-efficient methods, which will either have to come from changes to deep learning or from moving to other machine learning methods.
Forward citations
Cited by 11 Pith papers
-
Koopman Model Dimension Reduction via Variational Bayesian Inference and Graph Search
Variational Bayesian inclusion-flag estimates, thresholded into a directed graph, select a smaller Koopman dictionary while leaving output influence paths intact.
-
Quantum optical neural networks using atom-cavity interactions to provide all-optical nonlinearity
A simulated neural network uses atom-cavity two-level neurons as all-optical nonlinear activations and reports ~95% accuracy on MNIST and SAT-6.
-
Progressive Depth Up-scaling via Optimal Transport
Optimal-transport alignment of adjacent layers gives a cheap and slightly better initialization for progressive depth up-scaling than copying, averaging, or a learned predictor.
-
MOSAIC-FL, a micro-service based privacy-preserving framework with application to genomics
A gRPC micro-service FL stack with t-out-of-N CKKS secure aggregation matches cleartext accuracy on EMNIST and TCGA BRCA subtyping at modest extra cost for large models.
-
Physical Analogue Kolmogorov-Arnold Networks based on Reconfigurable Nonlinear-Processing Units
A proposed analog KAN chip uses silicon RNPUs as physically programmable nonlinear edges, with estimated ~250 pJ per inference and ~10x smaller area than a digital MLP.
-
Real-Time Analysis of Unstructured Data with Machine Learning on Heterogeneous Architectures
A graph neural network (ETX4VELO) reconstructs LHCb VELO tracks with performance comparable to the production 'search by triplet' algorithm while running end to end in the GPU-based first-level trigger, with additiona...
-
From Propagator to Oscillator: The Dual Role of Symmetric Differential Equations in Neural Systems
The same symmetric differential equation system can act as a stable signal propagator or as a self-oscillating signal generator, with the mode controlled by a parameter or by inhibitory-loop topology.
-
The Generalist Brain Module: Module Repetition in Neural Networks in Light of the Minicolumn Hypothesis
A review arguing that repeating a single generalist neural module, inspired by cortical minicolumns, yields robustness, scalability, and generalization benefits compared to monolithic networks.
-
What Makes Local Updates Effective: The Role of Data Heterogeneity and Smoothness
Under bounded second-order heterogeneity, local updates are shown to achieve faster convergence than mini-batch SGD in several convex and non-convex regimes, with matching lower bounds.
-
Improve Underwater Object Detection through YOLOv12 Architecture and Physics-informed Augmentation
Applying YOLOv12 with physics-flavored augmentations yields high reported mAP on four underwater detection benchmarks, but the claims are weakened by missing code, variance, and inconsistent speed numbers.
-
A Layered Self-Supervised Knowledge Distillation Framework for Efficient Multimodal Learning on the Edge
LSSKD trains compact classifiers with auxiliary self-supervised branches at each stage and past-epoch soft labels as targets, claiming teacher-free accuracy gains on classification benchmarks.
Discussion (0). Sign in to comment.