Pith. sign in

REVIEW 2 cited by

The Computational Complexity of Training ReLU(s)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.04207 v2 pith:LPZ2QAYG submitted 2018-10-09 cs.CC

classification cs.CC
keywords reluserroreventrainingcasecomplexitycomputationaldepth-2
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We consider the computational complexity of training depth-2 neural networks composed of rectified linear units (ReLUs). We show that, even for the case of a single ReLU, finding a set of weights that minimizes the squared error (even approximately) for a given training set is NP-hard. We also show that for a simple network consisting of two ReLUs, the error minimization problem is NP-hard, even in the realizable case. We complement these hardness results by showing that, when the weights and samples belong to the unit ball, one can (agnostically) properly and reliably learn depth-2 ReLUs with $k$ units and error at most $\epsilon$ in time $2^{(k/\epsilon)^{O(1)}}n^{O(1)}$; this extends upon a previous work of Goel, Kanade, Klivans and Thaler (2017) which provided efficient improper learning algorithms for ReLUs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Equivalence of Coarse and Fine-Grained Models for Learning with Distribution Shift

    cs.DS 2026-05 unverdicted novelty 8.0 of 10

    PQ and TDS learning are equivalent in the distribution-free setting for Boolean classes, implying hardness for TDS halfspace learning but efficient algorithms with membership queries.

  2. Omnipredicting Single-Index Models with Multi-Index Models

    cs.LG 2024-11 conditional novelty 7.0 of 10

    A new analysis of the Isotron algorithm yields omnipredictors for single-index models with about ε^-4 samples (ε^-2 for bi-Lipschitz links), improving the previous ε^-10 construction.

Pith tools