{"id":"14a88275-f6a8-4786-beba-82c1ff312e89","arxiv_id":"1908.00709","paper_version":6,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes AutoML into a four-stage pipeline and reviews neural architecture search methods, their performance, and open problems.","lead":"This paper surveys automated machine learning (AutoML), covering data preparation, feature engineering, hyperparameter optimization, and neural architecture search (NAS), with comparison tables for representative NAS algorithms. It is a synthesis of existing work rather than a new method, useful as an entry point for researchers entering the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tables 3 and 4 contain probable transcription errors and mix incomparable GPU types, so the survey's performance snapshot may mislead readers despite the authors' caveat in Section 6.1.2.","rationale":"The reader's weakest_assumption correctly identifies the compiled comparison tables as the least secure part of the survey. My independent reading found concrete, checkable transcription errors and a hardware-mixing problem in the GPU Days metric, so the concern is not merely a general worry about fairness. The paper's own Section 6.1.2 caveat partially mitigates the issue, but the tables are still presented as a usable comparison and the surrounding text draws conclusions from them. Since the central claim is breadth of coverage and that breadth is genuinely delivered, the appropriate outcome is not rejection but conditional acceptance: correct or clearly qualify the comparison tables. This changes the reader's UNVERDICTED verdict to CONDITIONAL because there is a specific, fixable defect in the main practical contribution.","tokens_in":58088,"tokens_out":2724,"duration_ms":30401,"concrete_test":"Cross-check every row in Tables 3 and 4 against the cited source papers, starting with the three most suspicious entries: DARTS on ImageNet (Table 4: 73.3/81.3 vs. the original DARTS reported top-5 of roughly 91.3), RENASNet on CIFAR-10 (Table 3: 91.12 vs. the likely 97.12), and the NASNet-A publication venue in Table 4 (ICLR17 vs. CVPR18). After correcting any errors, recompute whether the Section 6.1 statements about gradient-descent efficiency, random-search competitiveness, and general performance ordering still hold. If corrected entries change the ordering or the qualitative conclusions, the paper should be accepted only with corrected tables and an explicit statement that GPU Days are not normalized across hardware generations.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's distinctive value is a single-document overview of the whole AutoML pipeline, and Section 6 explicitly presents Tables 3 and 4 as a global performance comparison of NAS methods. That comparison is load-bearing for the paper's practical utility: a reader could reasonably use these tables to rank methods. The weakest point is the reliability of those tables. Two concrete problems are visible in the manuscript itself. First, Table 4 lists DARTS (searched on CIFAR-10) with ImageNet top-1/top-5 accuracy 73.3/81.3; the original DARTS paper reports top-5 accuracy near 91.3, so 81.3 is very likely a transcription error. Second, Table 3 lists RENASNet at 91.12 top-1 accuracy on CIFAR-10, while neighboring CIFAR-10 rows are in the 96-98 range; 91.12 is probably a misprint for 97.12. Additionally, GPU Days as defined in Eq. 9 counts only number of GPUs times days, yet the tables mix K40, P100, 1080Ti, Titan, and V100 hardware; a day on an 800-GPU K40 cluster is not comparable to a day on one 2080Ti. The authors explicitly concede in Section 6.1.2 that 'the comparison is not quite fair,' but the tables are still presented without a normalization scheme or error flags. Because Sections 6.1 draws qualitative conclusions from these entries (e.g., that gradient-descent methods reduce resource use and reach SOTA, and that random search is comparable), the transcription and hardware-mixing issues can distort the very rankings the survey offers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of automated machine learning organized around the ML pipeline: data preparation, feature engineering, model generation, and model evaluation. The largest part is devoted to neural architecture search, covering search spaces, architecture optimization methods, model-evaluation acceleration, and a comparative summary of representative NAS algorithms on CIFAR-10 and ImageNet. It also discusses one/two-stage NAS, one-shot NAS, joint hyperparameter and architecture optimization, resource-aware NAS, and several open problems. The authors' stated contribution is breadth relative to earlier surveys, which they support with a comparison table of existing surveys.","tokens_in":58351,"tokens_out":4614,"duration_ms":47107,"significance":"If the factual content is corrected, the survey would be a useful single-document entry point: it compiles a wide set of references across all pipeline stages, reproduces several central equations (DARTS relaxation, Eq. 4; bilevel objective, Eq. 5; Gumbel-Softmax, Eq. 8; MnasNet objective, Eq. 12) faithfully, and organizes NAS topics in a way that is helpful for newcomers. The main added value beyond earlier NAS-focused surveys is the pipeline-level coverage and the compiled comparison tables in Tables 3 and 4. The usefulness of those tables, however, is the weakest point: several entries appear to be transcription errors, and the resource metric is hardware-incomparable. These issues are fixable, so they do not invalidate the survey's organizational contribution, but they do require substantive revision.","major_comments":[{"comment":"Table 4 lists DARTS (searched on CIFAR-10) as achieving ImageNet top-1/top-5 accuracy of 73.3/81.3. The original DARTS paper (Liu et al., ICLR 2019) reports a top-5 accuracy of about 91.3% in this setting, so the 81.3 entry is almost certainly a transcription error. Because Tables 3 and 4 are presented as a global performance comparison and feed the qualitative conclusions in Section 6.1, this error should be corrected and the tables should be checked against the primary sources.","section":"Table 4, DARTS row"},{"comment":"Table 3 lists RENASNet+c/o at 91.12% top-1 accuracy on CIFAR-10, while every neighboring method is in the 96–98% range. This is almost certainly a misprint for 97.12; as printed, it would place an EA+RL method far below its actual performance and distort the method-class comparison in Section 6.1.","section":"Table 3, RENASNet row"},{"comment":"GPU Days defined in Eq. (9) as N×D is not comparable across rows because Tables 3 and 4 mix K40, P100, 1080Ti, Titan, Titan Xp, V100, and 2080Ti hardware with no normalization, and many entries omit the GPU type entirely. The caveat in Section 6.1.2 that 'the comparison is not quite fair' is useful, but it appears only after the tables, and the qualitative claims in Section 6.1 about gradient-descent methods reducing resource use and random search being comparable rely precisely on this resource measure. Please add per-row hardware flags, a normalization footnote, or restrict the resource-efficiency conclusions to entries with comparable hardware.","section":"Section 6.1, Eq. (9)"}],"minor_comments":[{"comment":"The flow chart contains the typo 'exsiting' in the decision box; it should read 'existing'.","section":"Figure 2"},{"comment":"The discussion of Tables 3 and 4 and the comparison caveat appears under the heading 'NAS-Bench Dataset' (Section 6.1.2) before the actual NAS-Bench discussion begins; consider restructuring so the caveat and table caveats precede the benchmark discussion.","section":"Section 6.1.1/6.1.2 numbering"},{"comment":"The notation 'MB' and 'ML' in Eqs. (2) and (3) should be typeset as M^B and M^L to avoid ambiguity about whether the exponent applies to the multiplication or to M.","section":"Eqs. (2) and (3)"},{"comment":"The name 'Auto-WEAK' should be 'Auto-WEKA', matching the referenced system by Thornton et al.","section":"Section 7.7"},{"comment":"The entry 'Federate Learning' should be 'Federated Learning'.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper's practical value hinges on the reliability of the comparison tables, yet at least two entries are transparent transcription errors and the resource metric mixes incomparable hardware. An audit of Tables 3 and 4 against primary sources should be requested before acceptance; this is a correctness issue, not merely a presentation issue. The organizational contribution and the survey breadth are otherwise solid."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a survey, and it does what a survey should: it organizes AutoML along the pipeline—data preparation, feature engineering, HPO, NAS—and covers enough ground that a newcomer could get a genuine roadmap from one document. The discussion of one/two-stage, one-shot, and joint architecture-hyperparameter search is clear, and the key equations (DARTS relaxation, GPU Days, MnasNet objective) are reproduced faithfully. The claim of broader coverage than earlier surveys holds up. That is the real value here: consolidation, not novelty.\n\nThe soft spots are concentrated in the performance comparison. Tables 3 and 4 have at least two probable transcription errors: DARTS on ImageNet listed as 73.3/81.3 top-1/top-5, while the original paper reports top-5 near 91.3, and RENASNet at 91.12 on CIFAR-10, which sits far below every neighboring row and is almost certainly 97.12. The GPU Days metric also mixes K40, P100, 1080Ti, and V100 without normalization, so a day on an 800-GPU K40 cluster is not comparable to a day on one 2080Ti. The authors do concede in Section 6.1.2 that the comparison is not quite fair, but the caveat comes after the tables have already been presented as a global comparison, and the qualitative conclusions in Section 6.1 (gradient methods reduce cost, random search is competitive) lean on those same entries. A reader using the tables to rank methods could be misled.\n\nNone of this sinks the survey. The descriptive content is generally accurate, the organization is useful, and the open-problems section is thoughtful. These are fixable problems—verify every table entry against the source, add a hardware note to the GPU Days definition, and flag known typos. But as published, the tables are the weakest link, and they happen to be the part most likely to be lifted.\n\nFor a serious referee: yes, this deserves peer review. The survey has real pedagogical value, and the flaws are emendable rather than fatal. I would ask for table verification and a more prominent disclaimer before acceptance. For your own work: I would not cite it as a source for specific accuracy numbers, but it is a reasonable survey to point a student toward for the bird's-eye view.","headline":"A broad, readable AutoML survey whose comparative NAS tables are the main value but also the main liability: treat the tables as a rough map, not a ranking.","tokens_in":58873,"tokens_out":1277,"would_cite":false,"duration_ms":17159,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that a single document can organize the entire AutoML pipeline, from data preparation through model evaluation, and that this full-pipeline view is what earlier surveys lacked.","keywords":["automated machine learning","neural architecture search","hyperparameter optimization","data preparation","feature engineering","model evaluation","deep learning pipeline","survey"],"falsifier":"Run a representative set of the Table 3 and Table 4 methods, such as DARTS, ENAS, RandomNAS, AmoebaNet-B, and P-DARTS, on identical GPUs, epochs, batch sizes, and augmentation settings, then compare the rank order by accuracy and GPU days; if the reported order changes materially, the survey's comparative claims fail.","tokens_in":57859,"feed_emoji":"🤖","tokens_out":6654,"duration_ms":60811,"temperature":0.7,"pith_summary":"This paper is a survey that tries to establish a broad, current map of automated machine learning (AutoML) organized by the stages of the deep-learning pipeline: data preparation, feature engineering, model generation, and model evaluation. Its distinctive claim is coverage: unlike prior surveys that treat only neural architecture search (NAS) or only hyperparameter optimization, it covers all four stages in one document, with the most space given to NAS as the currently dominant sub-topic. The paper also compiles the reported accuracy and GPU-day cost of representative NAS algorithms on CIFAR-10 and ImageNet, and uses that snapshot to discuss one- versus two-stage NAS, one-shot NAS, joint hyperparameter and architecture optimization, and resource-aware search. A sympathetic reader would value this as an entry point that supplies the field's vocabulary, taxonomy, and open problems in one pass. The survey's own caveat is that the compiled comparisons are not perfectly fair because methods were evaluated under different settings.","feed_headline":"Maps the whole AutoML pipeline in a single survey","feed_subtitle":"Covers data prep, feature engineering, hyperparameter tuning, and neural architecture search in one organized guide.","key_machinery":"The load-bearing device is the AutoML pipeline itself: a four-stage diagram (data preparation, feature engineering, model generation, model evaluation) that the survey uses as its table of contents. Within model generation, the paper's NAS taxonomy, which places every algorithm along three dimensions (search space, architecture optimization method, and model evaluation method), does the analytical work of organizing a large literature. The comparison tables use the GPU-day metric, defined as $N \\times D$ where $N$ is the number of GPUs and $D$ the number of days, as the standard unit of search cost.","core_discovery":"The paper's central claim is that it covers a broader range of AutoML methods than earlier surveys [10, 45, 46, 9, 8], because it follows the complete pipeline rather than a single sub-topic. On its own terms, it establishes an organizational scheme: data preparation (collection, cleaning, augmentation), feature engineering (selection, construction, extraction), model generation (search space plus optimization), and model evaluation (low fidelity, weight sharing, surrogates, early stopping). Within NAS it reports that search spaces can be entire-structured, cell-based, hierarchical, or morphism-based, and that architecture optimization methods divide into evolutionary algorithms, reinforcement learning, gradient descent, surrogate-model-based optimization, grid and random search, and hybrids. It further argues that gradient-based and random-search methods now achieve results comparable to earlier RL and EA methods at a fraction of the cost, and that one-shot and weight-sharing NAS exhibit a bias that decoupled optimization tries to fix. The paper concludes by listing open problems: flexible search spaces, interpretability, reproducibility, robustness, joint hyperparameter and architecture optimization, complete pipeline systems, and lifelong learning.","pith_inferences":["If the breadth claim holds, then AutoML is best understood as a pipeline-integration problem rather than a single algorithm family, which suggests that progress in one stage, such as data augmentation, can shift the value of methods in another stage.","The paper's own fairness caveat implies that Tables 3 and 4 should not be used to rank methods; a reader should treat them as existence proofs and cost anecdotes until controlled benchmarks settle the order.","A natural testable extension of the survey's logic is to build a single benchmark suite that varies training settings and augmentation across all four pipeline stages, extending the NAS-Bench idea to data preparation and feature engineering.","The competitive performance of random search reported for NAS suggests that the field's real advantage may lie in search-space design and evaluation fidelity rather than in the optimizer itself, a hypothesis the survey does not fully pursue."],"forward_implications":["A newcomer can acquire from one document an organized picture of AutoML's complete pipeline rather than stitching together separate surveys for NAS, hyperparameter optimization, and feature engineering.","The compiled tables give a rough cost-versus-accuracy landscape: early reinforcement-learning and evolutionary searches cost thousands of GPU days, while differentiable and random-search approaches complete in under a day with competitive accuracy.","The one-shot and weight-sharing distinction implies that NAS results should be evaluated for rank correlation between search and evaluation stages, not only final accuracy, and that the Kendall Tau metric is a meaningful diagnostic.","The open-problems list frames concrete research targets: flexible search spaces free of human bias, interpretable search decisions, reproducible NAS benchmarks, robustness to noisy and adversarial data, and joint optimization of hyperparameters and architectures."],"supporting_citations":[{"why":"The NAS-focused survey whose narrower scope this paper's complete-pipeline coverage is defined against.","marker":"[10]"},{"why":"Another NAS-only survey used in Table 1 to show that prior work omitted the other pipeline stages.","marker":"[45]"},{"why":"A NAS challenges survey that also covers only the NAS sub-topic, supporting the breadth comparison.","marker":"[46]"},{"why":"An AutoML survey that covers feature engineering and hyperparameter optimization but little NAS, highlighting the gap this paper fills.","marker":"[9]"},{"why":"An AutoML benchmark survey that covers data preparation, feature engineering, and hyperparameter optimization but no NAS.","marker":"[8]"},{"why":"The foundational NAS work by reinforcement learning that frames the entire NAS section.","marker":"[12]"},{"why":"DARTS, the gradient-descent baseline that anchors the discussion of differentiable architecture search and its costs.","marker":"[17]"},{"why":"The random-search study that supports the survey's claim that simple search can be a competitive NAS baseline.","marker":"[180]"}],"fun_headline_variants":["AutoML survey covers the full pipeline","From data prep to NAS: the AutoML survey","One survey, all AutoML steps: data to search","Survey maps AutoML pipeline, zeroing in on NAS","AutoML state-of-the-art, pipeline-wide"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's comparative snapshot of NAS methods assumes that accuracy and GPU-day figures collected from different papers, using different hardware, training budgets, and augmentation schemes, are comparable enough to support the trends it draws.","fun_headline_variants_meta":{"raw":{"variants":["AutoML survey covers the full pipeline","From data prep to NAS: the AutoML survey","One survey, all AutoML steps: data to search","Survey maps AutoML pipeline, zeroing in on NAS","AutoML state-of-the-art, pipeline-wide"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000138,"raw_usage":{"total_tokens":1163,"prompt_tokens":961,"completion_tokens":202,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":128}},"tokens_in":577,"tokens_out":202,"duration_ms":2927,"temperature":1.0,"reasoning_tokens":128,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:35:15.410999+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a representative set of the Table 3 and Table 4 methods, such as DARTS, ENAS, RandomNAS, AmoebaNet-B, and P-DARTS, on identical GPUs, epochs, batch sizes, and augmentation settings, then compare the rank order by accuracy and GPU days; if the reported order changes materially, the survey's comparative claims fail.","supporting_citations":[],"review_version":1}