Pith. sign in

REVIEW 1 cited by

Unwrapping The Black Box of Deep ReLU Networks: Interpretability, Diagnostics, and Simplification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.04041 v1 pith:X6VBEX3P submitted 2020-11-08 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords deepblackdiagnosticsinterpretabilitylinearlocalnetworknetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The deep neural networks (DNNs) have achieved great success in learning complex patterns with strong predictive power, but they are often thought of as "black box" models without a sufficient level of transparency and interpretability. It is important to demystify the DNNs with rigorous mathematics and practical tools, especially when they are used for mission-critical applications. This paper aims to unwrap the black box of deep ReLU networks through local linear representation, which utilizes the activation pattern and disentangles the complex network into an equivalent set of local linear models (LLMs). We develop a convenient LLM-based toolkit for interpretability, diagnostics, and simplification of a pre-trained deep ReLU network. We propose the local linear profile plot and other visualization methods for interpretation and diagnostics, and an effective merging strategy for network simplification. The proposed methods are demonstrated by simulation examples, benchmark datasets, and a real case study in home lending credit risk assessment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CRITS: Convolutional Rectifier for Interpretable Time Series Classification

    cs.LG 2025-05 conditional novelty 5.0 of 10

    CRITS is an intrinsically interpretable time series classifier whose local saliency maps are the exact per-sample weights of the model, obtained without gradients, perturbations, or upsampling.

Pith tools