REVIEW 4 major objections 5 minor 68 references
FX-DARTS: Designing Topology-unconstrained Architectures with Differentiable Architecture Search and Entropy-based Super-network Shrinking
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read FX-DARTS shows neural architectures can be searched without cell-sharing or two-operator priors, matching constrained NAS at 76.4% top-1 on ImageNet.
desk verdict FX-DARTS is a credible attempt to remove DARTS priors, with a concrete mechanism and honest evaluation, but the paper never isolates whether entropy shrinking is what makes the searched architectures good. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the node-wise softmax-normalized architectural weight $a^o_{(i,j)}$ from Eq. (10), which sums over all incoming edges and operators and thereby removes the 'two operators from two distinct nodes' rule. The shrinking mechanism is the cell-level sparsity entropy $H^{(k)}_{\mathrm{cell}}(\alpha)=\sum_{2<j<N} H^{(k,j)}_{\mathrm{node}}(\alpha)$, added to the cross-entropy loss as $\sum_k \lambda_k H^{(k)}_{\mathrm{cell}}$; the $\lambda_k$ are adaptively increased or decreased by a feedback rule depending on whether the actual entropy reduction meets the expected $\Delta E$. Cyclic reinitialization of the model parameters and threshold-based discretization (Algorithm 2) complete the pipeline. Theorem 1 guarantees monotone decrease of the sparsity entropy under a sufficiently large $\lambda$, assuming bounded gradients of the cross-entropy term.
What would settle it
Retrain each threshold-pruned FX-DARTS architecture and the un-pruned super-network under identical budgets on the same validation split; if the pruned models do not consistently match or beat the super-network's accuracy, the entropy-shrinking route to discrete architectures fails.
Extended reading notes
Core claim
The paper's central claim is that DARTS can be made topology-unconstrained: every cell may have a unique structure, and a node may retain any number of operators from any previous nodes, without sacrificing search stability. The Entropy-based Super-network Shrinking (ESS) framework is what makes this possible, by normalizing architectural parameters node-wise, adding a cell-level sparsity entropy to the loss, adaptively adjusting the entropy coefficients, and periodically reinitializing model parameters to decouple architecture optimization from super-network training. The resulting architectures reach 78.32% top-1 on CIFAR-100 and 76.4% top-1 on ImageNet-1K with search costs between 1.6 and 4.3 GPU-hours. The paper itself notes that the advantage over constrained-space methods is not overwhelming.
Load-bearing premise
The load-bearing premise is that forcing each node's operator weights to become sparse, then deleting everything below a 0.02 threshold, yields a discrete architecture whose final accuracy tracks what the super-network promised.
Editorial extensions
If this is right
- A single FX-DARTS run can return several discrete architectures spanning different parameter and FLOP counts, so practitioners can trade accuracy for cost without rerunning the search.
- Searched networks are no longer limited to the shared normal/reduction cell pattern, expanding the set of candidate architectures available for a given task.
- The search cost of 1.6 to 4.3 GPU-hours keeps topology-unconstrained search in the same practical range as DARTS itself.
- Architectures found on TinyImageNet transfer to ImageNet-1K at 76.0 to 76.4% top-1 accuracy, indicating the flexible search is not tied to the CIFAR proxy.
Reading between the lines
- If every constrained architecture is contained in the unconstrained space, the constrained-space accuracy is a lower bound on what a perfect unconstrained search could achieve; the moderate gains reported here likely reflect optimizer limits rather than a ceiling of the space itself.
- The entropy-shrinking feedback loop is a generic sparsification controller: replacing sparsity entropy with latency, memory, or energy proxies could yield hardware-diverse architectures from the same search framework.
- A direct transfer test would run the same ESS shrinking on a non-cell-based differentiable search space, such as transformer depth or mixed-precision configurations, to see whether the stability benefits generalize.
- The reported sensitivity to the hand-set coefficients $c_1,c_2$ and to the expected entropy reduction $\Delta E$ suggests that automatic tuning of the shrinking schedule could close the remaining gap to constrained-space methods.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FX-DARTS proposes to relax two standard DARTS priors: the requirement that cells of the same type share one topology and the rule that each intermediate node keeps exactly two operators from two distinct source nodes. The method uses node-wise normalized architectural parameters, an entropy-based sparsity penalty with feedback-adapted per-cell coefficients, a warm-up/architecture-optimization split, cyclic reinitialization of model parameters, and threshold-based pruning of low-weight operators. The authors report competitive CIFAR-10/100 results, multi-task results on TinyImageNet/SVHN/Flowers102, and ImageNet-1K top-1 accuracies around 76% at search costs of a few GPU-hours, and they claim that the entropy-based shrinking framework makes this feasible in the enlarged, less constrained space.
Significance. If the empirical claims hold, the paper addresses a question of genuine interest to the NAS community: whether differentiable search can be extended to a substantially less constrained cell topology space without suffering instability and without losing accuracy relative to strong DARTS-style priors. The paper is transparent about its main limitation in the conclusion, provides ablations for several components, and evaluates across multiple datasets and operator spaces. The theoretical contribution is modest: Theorem 1 and Corollary 2 establish that the sparsity entropy can be driven down by gradient descent under a sufficiently large adaptive coefficient, which is a statement about the optimization objective, not about the quality of the pruned architecture. The central performance claim is therefore empirical, and the missing attribution controls described below are the main obstacle to accepting that claim at face value.
major comments (4)
- [§II-A, §IV-D, Table I] The paper never isolates the effect of the ESS search signal. Section II-A itself cites Li and Talwalkar [40], who show that random search is competitive in the original DARTS space. Since FX-DARTS enlarges the space and also changes the evaluation protocol (17 cells, no drop-path, adjusted channels), a random-search or random-thresholding baseline in the same unconstrained space is needed to attribute the reported gains to ESS rather than to the enlarged search space and the evaluation setup. Without such a control, the central claim that entropy-based shrinking produces better architectures than comparably sized arbitrary subgraphs in the same space is unsupported.
- [§II-B, Tables I, III, V] The related-work section explicitly discusses GOLD-NAS and DNAD as methods that also operate in unconstrained or flexible spaces, but neither appears in any experimental comparison. Because the paper's contribution is specifically about searching in a topology-unconstrained space, omitting the two most directly comparable unconstrained-search methods prevents the reader from assessing whether FX-DARTS advances the state of the art for flexible architectures. At minimum, the authors should add these comparisons or justify their absence with concrete experimental or computational reasons.
- [Algorithm 1 line 18, Algorithm 2, Eq. (10)] Algorithm 2 is executed inside the architecture-optimization loop (Algorithm 1, line 18), but the paper does not specify what happens to the normalization in Eq. (10) after an operator is pruned. If the operator is simply removed from the softmax denominator, the remaining weights change value; if it is masked but kept in the denominator, the entropy loss and gradient computations are inconsistent with the actual discrete architecture. This matters because the same thresholding operation is used both as a search-time regularizer and as the final discretization. The authors should state the exact bookkeeping rule for pruned operators and report whether the super-network accuracy before and after final thresholding differs, since Section IV-B(g) asserts that pruning at epsilon=0.02 has minimal impact but no such numbers are given.
- [§IV-D-1-b, Tables III and IV] The experimental protocol creates a potential selection bias. The baselines are re-evaluated with 17 cells and without drop-path under a protocol chosen by the authors, yet there is no evidence that this protocol is neutral or that the baselines' hyperparameters were re-tuned under it. In addition, the hyperparameters of FX-DARTS (search epochs, shrinking coefficients, expected entropy reduction) appear to be selected using the same multi-task benchmark that is later reported as the headline multi-task result in Table III, with Table IV reporting ablations on that same benchmark. The authors should either use a held-out validation split for hyperparameter selection or explicitly report the selection procedure, so that the reported margins (e.g., 0.86% on TinyImageNet over P-DARTS) are not the result of tuning on the test tasks.
minor comments (5)
- [§II-B, last sentence] The section ends with the incomplete sentence 'There are' immediately before Section III; this appears to be a missing passage or a formatting error and should be completed or removed.
- [Eq. (14)] The text says lambda_k for k = 1,...,N, but the sum is over L cells; the index set should be k = 1,...,L.
- [Corollary 2] The formula for lambda can be negative when the angle theta between the two gradients is small and the performance-gradient term dominates; if lambda is intended to be a positive regularization coefficient, the authors should state the required conditions and how the adaptive mechanism handles this case.
- [Table V] The top-5 accuracy of FX-DARTS-O1 (Tiny, 64E) is listed as 93.7%, which is higher than that of the stronger top-1 configuration FX-DARTS-O3 (Tiny, 64E) at 93.0%; this inconsistency is not discussed and should be explained or corrected.
- [§IV-D-1-a and Fig. 2] The naming convention 'FX-DARTS-O1 (48E)' is defined in the text as 3 x Tsearch, but the figure and table captions do not repeat this definition; adding a brief parenthetical in the table caption would improve readability.
Circularity Check
No significant circularity: the central empirical claims rest on held-out evaluations and ablations, while the entropy-decrease theorem is a stated consequence of the entropy regularizer rather than a derivation of architecture accuracy.
full rationale
FX-DARTS's main claim is empirical: that a topology-unconstrained differentiable search with entropy-based shrinking finds competitive architectures (78.32% CIFAR-100, 76.4% ImageNet top-1) at low search cost. These results are obtained by searching on CIFAR-10 or TinyImageNet training splits and reporting test-set accuracy, with three runs per architecture and component ablations in Table IV. The entropy loss in Eqs. (12)-(14) is an optimization objective, and Theorem 1 in Section IV-C only shows that adding a sufficiently large λH term to the loss makes the entropy decrease under a first-order gradient step. That is a direct consequence of the loss definition, and the paper does not use it to derive architecture accuracy; it is explicitly framed as 'convergence analysis of the sparsity entropy.' The adaptive λ feedback and the expected entropy reduction in Eq. (15) are hyperparameters of the search procedure, not fitted to the reported test results. The self-citations (CR-LSO, DNAD) appear in related-work discussion only and are not load-bearing; no uniqueness theorem or prior author result is invoked to force the choice of ESS. The honest limitation in Section V, admitting that the experiments have not demonstrated overwhelming advantages over constrained-space NAS, further weighs against any claim that the outcome was forced by the method's construction. The absence of a random-search control and the unspecified denominator recomputation after pruning in Algorithm 2 are correctness/attribution concerns, not evidence that a prediction reduces by construction to its inputs.
Assumptions & free parameters
free parameters (4)
- adaptive entropy coefficient initial value and update rates =
lambda_init = 1e-4; c1 = 1.05; c2 = 0.95
- discretization threshold epsilon =
0.02
- search schedule parameters =
Tsearch = 16, Twarm = Tsearch/2, Rinit = 5
- expected entropy reduction budget Delta E =
Initial total entropy divided by L * len(train loader) * Tsearch * Rinit
assumptions (5)
- domain assumption The DARTS continuous relaxation (Eq. 3-4) is a valid differentiable proxy for discrete architecture selection.
- domain assumption Super-network performance predicts final discrete architecture performance after threshold pruning.
- domain assumption Assumption 1 (bounded gradient of CE loss) holds, so Theorem 1 applies.
- domain assumption Node-wise normalization (Eq. 10) creates a meaningful competition across both input edges and operators.
- domain assumption Baselines re-evaluated under the authors' protocol (17 cells, no drop-path) are representative of the original methods' true performance.
Cite this review
Pith. "Pith review of FX-DARTS: Designing Topology-unconstrained Architectures with Differentiable Architecture Search and Entropy-based Super-network Shrinking." pith.science (2026). https://pith.science/paper/2LFVY4G7
@misc{pith2026250420079,
author = {Pith},
title = {Pith review of: FX-DARTS: Designing Topology-unconstrained Architectures with Differentiable Architecture Search and Entropy-based Super-network Shrinking},
year = {2026},
howpublished = {\url{https://pith.science/paper/2LFVY4G7}},
note = {Machine review of arXiv:2504.20079}
}
read the original abstract
Strong priors are imposed on the search space of Differentiable Architecture Search (DARTS), such that cells of the same type share the same topological structure and each intermediate node retains two operators from distinct nodes. While these priors reduce optimization difficulties and improve the applicability of searched architectures, they hinder the subsequent development of automated machine learning (Auto-ML) and prevent the optimization algorithm from exploring more powerful neural networks through improved architectural flexibility. This paper aims to reduce these prior constraints by eliminating restrictions on cell topology and modifying the discretization mechanism for super-networks. Specifically, the Flexible DARTS (FX-DARTS) method, which leverages an Entropy-based Super-Network Shrinking (ESS) framework, is presented to address the challenges arising from the elimination of prior constraints. Notably, FX-DARTS enables the derivation of neural architectures without strict prior rules while maintaining the stability in the enlarged search space. Experimental results on image classification benchmarks demonstrate that FX-DARTS is capable of exploring a set of neural architectures with competitive trade-offs between performance and computational complexity within a single search procedure.
Figures
Reference graph
Works this paper leans on
-
[40]
Random search and reproducibility for neural architecture search,
L. Li and A. Talwalkar, “Random search and reproducibility for neural architecture search,” in Proceedings of the 35th Conference on Uncer- tainty in Artificial Intelligence , Tel Aviv, Israel, Jul. 2019, pp. 367–377
work page 2019
-
[1]
Going deeper with convolutions,
C. Szegedy, W. Liu, Y . Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V . Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” inProceedings of the 28th IEEE Conference on Computer Vision and Pattern Recognition, 2015, Boston, MA, USA, Jun. 2015, pp. 1–9
work page 2015
-
[2]
A structure constraint matrix factorization framework for human behavior segmentation,
H. Gao, C. Lv, T. Zhang, H. Zhao, L. Jiang, J. Zhou, Y . Liu, Y . Huang, and C. Han, “A structure constraint matrix factorization framework for human behavior segmentation,” IEEE Transactions on Cybernetics , vol. 52, no. 12, pp. 12 978–12 988, 2021
work page 2021
-
[3]
BERT: pre- training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: pre- training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, Jun. 2019, pp. 4171–4186
work page 2019
-
[4]
A survey of the usages of deep learning for natural language processing,
D. W. Otter, J. R. Medina, and J. K. Kalita, “A survey of the usages of deep learning for natural language processing,” IEEE Transactions on Neural Networks and Learning Systems , vol. 32, no. 2, pp. 604–624, Apr. 2020
work page 2020
-
[5]
Optimal control via neural networks: A convex approach,
Y . Chen, Y . Shi, and B. Zhang, “Optimal control via neural networks: A convex approach,” arXiv preprint, arXiv:1805.11835 , 2018
arXiv 2018
-
[6]
Hierarchical deep reinforcement learning for continuous action control,
Z. Yang, K. Merrick, L. Jin, and H. A. Abbass, “Hierarchical deep reinforcement learning for continuous action control,” IEEE transactions on neural networks and learning systems, vol. 29, no. 11, pp. 5174–5184, Nov. 2018
work page 2018
-
[7]
LSTM-MSNet: Lever- aging forecasts on sets of related time series with multiple seasonal patterns,
K. Bandara, C. Bergmeir, and H. Hewamalage, “LSTM-MSNet: Lever- aging forecasts on sets of related time series with multiple seasonal patterns,” IEEE Transactions on Neural Networks and Learning Systems, vol. 32, no. 4, pp. 1586–1599, Apr. 2020
work page 2020
Show all 68 references
-
[8]
Object classification using CNN-based fusion of vision and lidar in autonomous vehicle environment,
H. Gao, B. Cheng, J. Wang, K. Li, J. Zhao, and D. Li, “Object classification using CNN-based fusion of vision and lidar in autonomous vehicle environment,” IEEE Transactions on Industrial Informatics , vol. 14, no. 9, pp. 4224–4231, 2018
2018
-
[9]
An interacting multiple model for trajectory prediction of intelligent vehicles in typical road traffic scenario,
H. Gao, Y . Qin, C. Hu, Y . Liu, and K. Li, “An interacting multiple model for trajectory prediction of intelligent vehicles in typical road traffic scenario,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 9, pp. 6468–6479, 2021
2021
-
[10]
Trajectory prediction of cyclist based on dynamic bayesian network and long short-term memory model at unsignalized intersections,
H. Gao, H. Su, Y . Cai, R. Wu, Z. Hao, Y . Xu, W. Wu, J. Wang, Z. Li, and Z. Kan, “Trajectory prediction of cyclist based on dynamic bayesian network and long short-term memory model at unsignalized intersections,” Science China Information Sciences , vol. 64, no. 7, p. 172207, 2021
2021
-
[11]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[12]
Motif discoveries in unaligned molecular sequences using self-organizing neural networks,
D. Liu, X. Xiong, B. DasGupta, and H. Zhang, “Motif discoveries in unaligned molecular sequences using self-organizing neural networks,” IEEE Transactions on Neural Networks , vol. 17, no. 4, pp. 919–928, 2006
2006
-
[13]
Deep neural network for structural prediction and lane detection in traffic scene,
J. Li, X. Mei, D. Prokhorov, and D. Tao, “Deep neural network for structural prediction and lane detection in traffic scene,” IEEE Transactions on Neural Networks and Learning Systems , vol. 28, no. 3, pp. 690–703, Mar. 2016
2016
-
[14]
Xception: Deep learning with depthwise separable convolu- tions,
F. Chollet, “Xception: Deep learning with depthwise separable convolu- tions,” in Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, Jul. 2017, pp. 1251–1258
2017
-
[15]
Multi-scale context aggregation by dilated convolutions,
F. Yu and V . Koltun, “Multi-scale context aggregation by dilated convolutions,” in Proceedings of the 4th International Conference on Learning Representations, San Juan, Puerto Rico, May 2016, 12 pages
2016
-
[16]
AutoML: A survey of the state-of-the-art,
X. He, K. Zhao, and X. Chu, “AutoML: A survey of the state-of-the-art,” Knowledge-Based Systems, vol. 212, Art. no. 106622, 2021
2021
-
[17]
Designing network design spaces,
I. Radosavovic, R. P. Kosaraju, R. Girshick, K. He, and P. Doll ´ar, “Designing network design spaces,” in Proceedings of the 33rd IEEE/CVFConference on Computer Vision and Pattern Recognition , Seattle, W A, USA, Jun. 2020, pp. 10 428–10 436
2020
-
[18]
DARTS: Differentiable architecture search,
H. Liu, K. Simonyan, and Y . Yang, “DARTS: Differentiable architecture search,” in Proceedings of the 7th International Conference on Learning Representations, New Orleans, LA, USA, May 2019, 12 pages
2019
-
[19]
Learning transferable architectures for scalable image recognition,
B. Zoph, V . Vasudevan, J. Shlens, and Q. V . Le, “Learning transferable architectures for scalable image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 8697–8710
2018
-
[20]
Designing neural network architectures using reinforcement learning,
B. Baker, O. Gupta, N. Naik, and R. Raskar, “Designing neural network architectures using reinforcement learning,” in Proceedings of the 5th International Conference on Learning Representations , Toulon, France, Apr. 2017
2017
-
[21]
Efficient neural srchitecture search via parameter sharing,
H. Pham, M. Y . Guan, B. Zoph, Q. V . Le, and J. Dean, “Efficient neural srchitecture search via parameter sharing,” in Proceedings of the 35th International Conference on Machine Learning , Stockholmsm ¨assan, Sweden, Jul. 2018, pp. 4092–4101
2018
-
[22]
Hierarchical representations for efficient architecture search,
H. Liu, K. Simonyan, O. Vinyals, C. Fernando, and K. Kavukcuoglu, “Hierarchical representations for efficient architecture search,” in Pro- ceedings of the 6th International Conference on Learning Representa- tions, Vancouver, Canada, Apr. 2018
2018
-
[24]
Neural architecture optimization,
R. Luo, F. Tian, T. Qin, E. Chen, and T. Liu, “Neural architecture optimization,” in Proceedings of the 32nd International Conference on Neural Information Processing Systems , Montr ´eal, Canada, Dec. 2018, pp. 7827–7838
2018
-
[25]
Does unsupervised architecture representation learning help neural architecture search?
S. Yan, Y . Zheng, W. Ao, X. Zeng, and M. Zhang, “Does unsupervised architecture representation learning help neural architecture search?” in Proceedings of the 34th International Conference on Neural Information Processing Systems, Virtual Event, Dec. 2020, pp. 12 486–12 498
2020
-
[26]
D-V AE: A variational autoencoder for directed acyclic graphs,
M. Zhang, S. Jiang, Z. Cui, R. Garnett, and Y . Chen, “D-V AE: A variational autoencoder for directed acyclic graphs,” in Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, Dec. 2019, pp. 1586–1598
2019
-
[27]
CR-LSO: Convex neural archi- tecture optimization in the latent space of graph variational autoencoder with input convex neural networks,
X. Rao, B. Zhao, X. Yi, and D. Liu, “CR-LSO: Convex neural archi- tecture optimization in the latent space of graph variational autoencoder with input convex neural networks,” arXiv preprint, arXiv:2211.05950 , 2022
2022 arXiv
-
[28]
SNAS: stochastic neural architecture search,
S. Xie, H. Zheng, C. Liu, and L. Lin, “SNAS: stochastic neural architecture search,” in Proceedings of the 7th International Conference on Learning Representations , New Orleans, LA, USA, May 2019, 12 pages
2019
-
[29]
Progressive DARTS: Bridging the optimization gap for NAS in the wild,
X. Chen, L. Xie, J. Wu, and Q. Tian, “Progressive DARTS: Bridging the optimization gap for NAS in the wild,” International Journal of Computer Vision, vol. 129, no. 3, pp. 638–655, Mar. 2021
2021
-
[30]
Gold-NAS: Gradual, one-level, differentiable,
K. Bi, L. Xie, X. Chen, L. Wei, and Q. Tian, “Gold-NAS: Gradual, one-level, differentiable,” arXiv preprint, arXiv:2007.03331 , 2020
2007 arXiv
-
[31]
Discretization-aware architecture search,
Y . Tian, C. Liu, L. Xie, Q. Ye et al., “Discretization-aware architecture search,” Pattern Recognition, vol. 120, Dec. 2021, Art. no. 108186
2021
-
[32]
KL-DNAS: Knowl- edge distillation-based latency aware-differentiable architecture search for video motion magnification,
J. Singh, S. Murala, and G. S. R. Kosuru, “KL-DNAS: Knowl- edge distillation-based latency aware-differentiable architecture search for video motion magnification,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–11, 2024, Early Access, DOI:10.1109/TNNLS.2023...
2024
-
[33]
An architecture entropy regularizer for differentiable neural architecture search,
K. Jing, L. Chen, and J. Xu, “An architecture entropy regularizer for differentiable neural architecture search,” Neural Networks , vol. 158, pp. 111–120, 2023
2023
-
[34]
Improved differentiable architecture search with multi-stage progressive partial channel connections,
Y . Xue, C. Lu, F. Neri, and J. Qin, “Improved differentiable architecture search with multi-stage progressive partial channel connections,” IEEE Transactions on Emerging Topics in Computational Intelligence , vol. 8, no. 1, pp. 32–43, 2024
2024
-
[35]
MNGNAS: distilling adaptive combination of multiple searched networks for one- shot neural architecture search,
Z. Chen, G. Qiu, P. Li, L. Zhu, X. Yang, and B. Sheng, “MNGNAS: distilling adaptive combination of multiple searched networks for one- shot neural architecture search,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 11, pp. 13 489–13 508, 2023
2023
-
[36]
Exploiting operation importance for differentiable neural architecture search,
Y . Zhou, X. Xie, and S.-Y . Kung, “Exploiting operation importance for differentiable neural architecture search,” IEEE Transactions on Neural Networks and Learning Systems , vol. 33, no. 11, pp. 6235–6248, 2022
2022
-
[37]
Fair DARTS: Eliminating unfair advantages in differentiable architecture search,
X. Chu, T. Zhou, B. Zhang, and J. Li, “Fair DARTS: Eliminating unfair advantages in differentiable architecture search,” in Proceedings of the 16th European Conference on Computer Vision , Glasgow, UK, Aug. 2020, pp. 465–480
2020
-
[38]
Darts-: Robustly stepping out of performance collapse without indicators,
X. Chu, X. Wang, B. Zhang, S. Lu, X. Wei, and J. Yan, “Darts-: Robustly stepping out of performance collapse without indicators,” Proceedings of the 9th International Conference on Learning Representations , Virtual Event, May 2021, 12 pages
2021
-
[39]
Noisy differentiable architecture search,
X. Chu and B. Zhang, “Noisy differentiable architecture search,” in Proceedings of the 32nd British Machine Vision Conference , Virtual Event, Nov. 2021, pp. 217–231
2021
-
[41]
NAS evaluation is frustratingly hard,
A. Yang, P. M. Esperanc ¸a, and F. M. Carlucci, “NAS evaluation is frustratingly hard,” in Proceedings of the 8th International Conference on Learning Representations. OpenReview.net, Addis Ababa, Ethiopia, Apr. 2020
2020
-
[42]
Single-path NAS: Designing hardware- efficient convnets in less than 4 hours,
D. Stamoulis, R. Ding, D. Wang, D. Lymberopoulos, B. Priyantha, J. Liu, and D. Marculescu, “Single-path NAS: Designing hardware- efficient convnets in less than 4 hours,” in Proceedings of Joint Eu- ropean Conference on Machine Learning and Knowledge Discovery in Databases, W ...
2019
-
[43]
ProxylessNAS: Direct neural architecture search on target task and hardware,
H. Cai, L. Zhu, and S. Han, “ProxylessNAS: Direct neural architecture search on target task and hardware,” in Proceedings of International Conference on Learning Representations , New Orleans, Louisiana, United States, May 2019, pp. 1–9
2019
-
[44]
K- shot NAS: Learnable weight-sharing for nas with k-shot supernets,
X. Su, S. You, M. Zheng, F. Wang, C. Qian, C. Zhang, and C. Xu, “K- shot NAS: Learnable weight-sharing for nas with k-shot supernets,” in Proceedings of International Conference on Machine Learning , Virtual Event, Jul. 2021, pp. 9880–9890
2021
-
[45]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the 29th IEEE Conference on Computer Vision and Pattern Recognition , Las Vegas, NV , USA, Jun. 2016, pp. 770–778
2016
-
[46]
ShuffleNet: An extremely efficient convolutional neural network for mobile devices,
X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An extremely efficient convolutional neural network for mobile devices,” in Proceed- ings of the 31st IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, Jun. 2018, pp. 6848–6856
2018
-
[47]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proceedings of the 9th International...
2021
-
[48]
Operation and topology aware fast differentiable architecture search,
S. Siddiqui, C. Kyrkou, and T. Theocharides, “Operation and topology aware fast differentiable architecture search,” in Proceedings of the 25th International Conference on Pattern Recognition, Milan, Italy, Jan. 2021, pp. 9666–9673
2021
-
[49]
DOTS: decoupling operation and topology in differentiable architecture search,
Y . Gu, L. Wang, Y . Liu, Y . Yang, Y . Wu, S. Lu, and M. Cheng, “DOTS: decoupling operation and topology in differentiable architecture search,” in Proceedings of the 34th IEEE Conference on Computer Vision and Pattern Recognition, Virtual Event, Jun. 2021, pp. 12 311–12 320
2021
-
[50]
Evolving connec- tions in group of neurons for robust learning,
J. Liu, M. Gong, L. Xiao, W. Zhang, and F. Liu, “Evolving connec- tions in group of neurons for robust learning,” IEEE Transactions on Cybernetics, vol. 52, no. 5, pp. 3069–3082, May 2022
2022
-
[51]
Exploring randomly wired neural networks for image recognition,
S. Xie, A. Kirillov, R. B. Girshick, and K. He, “Exploring randomly wired neural networks for image recognition,” in Proceedings of the 17th IEEE/CVF International Conference on Computer Vision, , Seoul, Korea, Oct. 2019, pp. 1284–1293
2019
-
[52]
DNAD: Differentiable neural architecture distillation,
X. Rao, B. Zhao, and D. Liu, “DNAD: Differentiable neural architecture distillation,” TechRxiv, 2022
2022
-
[53]
Partially-connected neural architecture search for reduced computational redundancy,
Y . Xu, L. Xie, W. Dai, X. Zhang, X. Chen, G.-J. Qi, H. Xiong, and Q. Tian, “Partially-connected neural architecture search for reduced computational redundancy,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 9, pp. 2953–2970, Sep. 2021
2021
-
[54]
Relativenas: Relative neural architecture search via slow-fast learning,
H. Tan, R. Cheng, S. Huang, C. He, C. Qiu, F. Yang, and P. Luo, “Relativenas: Relative neural architecture search via slow-fast learning,” IEEE Transactions on Neural Networks and Learning Systems , Jan. 2021
2021
-
[55]
XNAS: Neural architecture search with expert advice,
N. Nayman, A. Noy, T. Ridnik, I. Friedman, R. Jin, and L. Zelnik-Manor, “XNAS: Neural architecture search with expert advice,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, Dec. 2019, pp. 1975–1985
2019
-
[56]
β-darts: Beta-decay regularization for differentiable architecture search,
P. Ye, B. Li, Y . Li, T. Chen, J. Fan, and W. Ouyang, “β-darts: Beta-decay regularization for differentiable architecture search,” in Proceedings of the 35th IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, Jun. 2022, pp. 10 864–10 873
2022
-
[57]
Cyclic differentiable architecture search,
H. Yu, H. Peng, Y . Huang, J. Fu, H. Du, L. Wang, and H. Ling, “Cyclic differentiable architecture search,”IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 1, pp. 211–228, Jan. 2023
2023
-
[58]
Regularized evolution for image classifier architecture search,
E. Real, A. Aggarwal, Y . Huang, and Q. V . Le, “Regularized evolution for image classifier architecture search,” inProceedings of the 33rd AAAI Conference on Artificial Intelligence, Honolulu, Hawaii, USA, Jan. 2019, pp. 4780–4789
2019
-
[59]
ShuffleNet V2: Practical guidelines for efficient CNN architecture design,
N. Ma, X. Zhang, H. Zheng, and J. Sun, “ShuffleNet V2: Practical guidelines for efficient CNN architecture design,” in Proceedings of the 15th European Conference on Computer Vision , vol. 11218, Munich, Germany, Sep. 2018, pp. 122–138
2018
-
[60]
MobileNets: Efficient convo- lutional neural networks for mobile vision applications,
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, “MobileNets: Efficient convo- lutional neural networks for mobile vision applications,” arXiv preprint, arXiv:1704.04861, 2017
2017 arXiv
-
[61]
Mo- bileNetV2: Inverted residuals and linear bottlenecks,
M. Sandler, A. G. Howard, M. Zhu, A. Zhmoginov, and L. Chen, “Mo- bileNetV2: Inverted residuals and linear bottlenecks,” in Proceedings of the 31st IEEE Conference on Computer Vision and Pattern Recognition , Salt Lake City, UT, USA, Jun. 2018, pp. 4510–4520
2018
-
[62]
Searching for a robust neural architecture in four gpu hours,
X. Dong and Y . Yang, “Searching for a robust neural architecture in four gpu hours,” in Proceedings of the 32nd IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, Jun. 2019, pp. 1761–1770
2019
-
[63]
ℓ-darts: Light- weight differentiable architecture search with robustness enhancement strategy,
L. Hu, Z. Wang, H. Li, P. Wu, J. Mao, and N. Zeng, “ ℓ-darts: Light- weight differentiable architecture search with robustness enhancement strategy,” Knowledge-Based Systems, vol. 288, Art. no. 111466, 2024
2024
-
[64]
Self-adaptive weight based on dual- attention for differentiable neural architecture search,
Y . Xue, X. Han, and Z. Wang, “Self-adaptive weight based on dual- attention for differentiable neural architecture search,” IEEE Transac- tions on Industrial Informatics , 2024, doi:10.1109/TII.2023.3348843
2024
-
[65]
STO-DARTS: Stochastic bilevel optimization for differentiable neural architecture search,
Z. Cai, L. Chen, T. Ling, and H.-L. Liu, “STO-DARTS: Stochastic bilevel optimization for differentiable neural architecture search,” IEEE Transactions on Emerging Topics in Computational Intelligence , 2024, doi:10.1109/TETCI.2024.3359046
2024
-
[66]
SaDENAS: A self-adaptive differential evolution algorithm for neural architecture search,
X. Han, Y . Xue, Z. Wang, Y . Zhang, A. Muravev, and M. Gabbouj, “SaDENAS: A self-adaptive differential evolution algorithm for neural architecture search,” Swarm and Evolutionary Computation, vol. 91, Art. no. 101736, 2024
2024
-
[67]
EG-NAS: Neural architecture search with fast evolutionary exploration,
Z. Cai, L. Chen, P. Liu, T. Ling, and Y . Lai, “EG-NAS: Neural architecture search with fast evolutionary exploration,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 10, 2024, pp. 11 159–11 167
2024
-
[68]
Operation-level early stopping for robustifying differentiable NAS,
S. Jiang, Z. Ji, G. Zhu, C. Yuan, and Y . Huang, “Operation-level early stopping for robustifying differentiable NAS,” Advances in Neural Information Processing Systems , vol. 36, pp. 70 983–71 007, 2024
2024
-
[69]
DARTS-PT-CORE: Collaborative and regularized perturbation-based architecture selection for differentiable nas,
W. Xie, H. Li, X. Fang, and S. Li, “DARTS-PT-CORE: Collaborative and regularized perturbation-based architecture selection for differentiable nas,” Neurocomputing, vol. 580, Art. no. 127522, 2024
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.