Pith. sign in

REVIEW 2 cited by

Gradients as Features for Deep Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.05529 v1 pith:CSUTGYMO submitted 2020-04-12 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords modeldeepfeaturesnetworkdifferentefficientgradientgradients
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We address the challenging problem of deep representation learning--the efficient adaption of a pre-trained deep network to different tasks. Specifically, we propose to explore gradient-based features. These features are gradients of the model parameters with respect to a task-specific loss given an input sample. Our key innovation is the design of a linear model that incorporates both gradient and activation of the pre-trained network. We show that our model provides a local linear approximation to an underlying deep model, and discuss important theoretical insights. Moreover, we present an efficient algorithm for the training and inference of our model without computing the actual gradient. Our method is evaluated across a number of representation-learning tasks on several datasets and using different network architectures. Strong results are obtained in all settings, and are well-aligned with our theoretical insights.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gradient Short-Circuit: Efficient Out-of-Distribution Detection via Feature Intervention

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Gradient Short-Circuit masks the top-gradient feature coordinates, approximates the resulting logits with a first-order Taylor step, and reports large FPR95 improvements on standard OOD benchmarks.

  2. Maximally-Informative Retrieval for State Space Model Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RICO ranks documents by how much they reduce an SSM's question perplexity, using gradient-document inner products, and matches BM25 while often beating E5 on answer quality without finetuning.

Pith tools