Pith. sign in

REVIEW 1 cited by

Re-parameterizing Your Optimizers rather than Architectures

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15242 v4 pith:PYVWDLPE submitted 2022-05-30 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords modelmodelsoptimizersrepoptimizersstructureextraknowledgemodel-specific
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The well-designed structures in neural networks reflect the prior knowledge incorporated into the models. However, though different models have various priors, we are used to training them with model-agnostic optimizers such as SGD. In this paper, we propose to incorporate model-specific prior knowledge into optimizers by modifying the gradients according to a set of model-specific hyper-parameters. Such a methodology is referred to as Gradient Re-parameterization, and the optimizers are named RepOptimizers. For the extreme simplicity of model structure, we focus on a VGG-style plain model and showcase that such a simple model trained with a RepOptimizer, which is referred to as RepOpt-VGG, performs on par with or better than the recent well-designed models. From a practical perspective, RepOpt-VGG is a favorable base model because of its simple structure, high inference speed and training efficiency. Compared to Structural Re-parameterization, which adds priors into models via constructing extra training-time structures, RepOptimizers require no extra forward/backward computations and solve the problem of quantization. We hope to spark further research beyond the realms of model structure design. Code and models \url{https://github.com/DingXiaoH/RepOptimizers}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What is YOLOv6? A Deep Insight into the Object Detection Model

    cs.CV 2024-12 unverdicted novelty 1.0 of 10

    A review-style paper that restates YOLOv6's architecture and benchmark tables from the YOLOv6 paper without adding new experiments or analysis.

Pith tools