Pith. sign in

REVIEW 4 cited by

Gradient-based Bi-level Optimization for Deep Learning: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.11719 v4 pith:RCE5YPHZ submitted 2022-07-24 cs.LG math.OC

classification cs.LGmath.OC
keywords optimizationbi-levelupdateformulationgradient-basedsurveycategorydata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Bi-level optimization, especially the gradient-based category, has been widely used in the deep learning community including hyperparameter optimization and meta-knowledge extraction. Bi-level optimization embeds one problem within another and the gradient-based category solves the outer-level task by computing the hypergradient, which is much more efficient than classical methods such as the evolutionary algorithm. In this survey, we first give a formal definition of the gradient-based bi-level optimization. Next, we delineate criteria to determine if a research problem is apt for bi-level optimization and provide a practical guide on structuring such problems into a bi-level optimization framework, a feature particularly beneficial for those new to this domain. More specifically, there are two formulations: the single-task formulation to optimize hyperparameters such as regularization parameters and the distilled data, and the multi-task formulation to extract meta-knowledge such as the model initialization. With a bi-level formulation, we then discuss four bi-level optimization solvers to update the outer variable including explicit gradient update, proxy update, implicit function update, and closed-form update. Finally, we wrap up the survey by highlighting two prospective future directions: (1) Effective Data Optimization for Science examined through the lens of task formulation. (2) Accurate Explicit Proxy Update analyzed from an optimization standpoint.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synergistic Localization and Sensing in MIMO-OFDM Systems via Mixed-Integer Bilevel Learning

    cs.NI 2025-07 reject novelty 6.0 of 10

    Jointly training localization and sensing models with a learned subcarrier-selection mask reduces localization and sensing error in MIMO-OFDM settings.

  2. AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ADADEDUP adaptively prunes object detection datasets by combining semantic clustering with proxy-model loss feedback, matching full-data mAP at 20% pruning.

  3. AffinityFlow: Guided Flows for Antibody Affinity Maturation

    cs.LG 2025-02 reject novelty 5.0 of 10

    AffinityFlow guides AlphaFlow structure generation toward low Rosetta binding energy, then inverse-folds the structures to propose antibody mutations, and reports top scores on a computational affinity maturation benchmark.

  4. Cellular Traffic Prediction via Byzantine-robust Asynchronous Federated Learning

    cs.LG 2025-05 reject novelty 4.0 of 10

    The paper proposes BAFDP, an asynchronous Byzantine-robust federated learning algorithm with local differential privacy, and reports superior traffic prediction accuracy over eight baselines on three datasets.

Pith tools