{"id":"a12bb27b-e3d0-4566-befe-7a51a1b7b5d1","arxiv_id":"2501.09655","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of LLM applications in electronic design automation, organized by design stage and adaptation technique.","lead":"This paper surveys how large language models are being applied across chip design, from system specifications to physical layout. It organizes dozens of recent works and highlights trends in model size, fine-tuning, and agent-based workflows.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The meta-study's corpus is unauditable: 'all publications in selected venues' cannot be checked because the venues and search protocol are never stated, so the survey's 'comprehensive' central claim rests on an unverifiable selection.","rationale":"Read in good faith, the paper's contribution is an organized map of LLM-for-EDA work plus a Section 3 meta-study claimed to be exhaustive. To support that, the inclusion rules need to be reproducible; currently they are not, and the reported trends cannot be audited. This is the only concern I found that, if it lands, directly undermines the strongest central claim. Other issues (off-topic Table 1 entry, Claude3 attribution, ACM template leftovers) are quality defects that do not change the survey's utility but do support the reader's conditional verdict. The reader's weakest assumption already names the unstated venues and loose inclusion criteria; I agree with that characterization. Therefore I would not move the verdict: keep it conditional on the authors either specifying the venues and protocol or softening the 'all publications' and 'comprehensive' language.","tokens_in":19293,"tokens_out":4300,"duration_ms":45204,"concrete_test":"Ask the authors to release the exact venue list, search queries, and date range, then independently reconstruct the corpus by querying IEEE Xplore and ACM DL for title/abstract terms ('large language model' OR 'LLM' OR 'ChatGPT' OR 'GPT-4') AND ('Verilog' OR 'RTL' OR 'logic synthesis' OR 'physical design' OR 'EDA' OR 'CAD') across DAC, ICCAD, DATE, ASP-DAC, ISPD, and MLCAD for 2022-2024, plus a Google Scholar sweep. Compute recall and precision of the survey's Section 3 citations against this reconstructed set. If recall is materially below 90%, or if many included works violate the stated 'selected venues'/regular-paper criterion, the 'all publications' and 'comprehensive' claims must be relaxed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 opens by saying it 'reviews all publications in the selected venues' from 2022 to 2024, and the survey's central claim is a 'comprehensive understanding of the current state' of LLM/EDA research. But the selected venues are never named and no search protocol or inclusion/exclusion log is given; the only criteria are negative ('excludes... works that use LLMs as an application', 'Only regular papers are considered while invited papers are excluded'). The bibliography of Section 3 includes arXiv preprints, workshop papers, and at least one off-topic entry (Table 1 lists [53], a study of JavaScript unit-test generation), so the de facto corpus does not transparently match the stated criteria. Because the meta-study's trend statements ('which areas are well-explored / unexplored', 'which techniques have been proven effective', 'observable trends') are all derived from this unlisted corpus, the trend claims are unfalsifiable. This is a load-bearing gap: a different venue choice or search string would likely change the reported trends. Independently, the factual slip 'Perplexity's Claude3' (Section 2) and template artifacts (author block 'Trovato et al.', ACM reference dated 2018, placeholder key words) do not bear directly on the central argument but reinforce the need for a cleanup pass.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys applications of large language models in electronic design automation. It organizes the literature by design stage (system-level design, RTL design, logic synthesis and physical design, and analog circuit design), then discusses LLM methodology (model architecture and size, customization techniques, and multimodal feature representation), and closes with academic infrastructure, application bottlenecks, and ethics/security/efficiency considerations. The paper claims to provide a comprehensive understanding of the current state and future potential of LLM applications in EDA, and Section 3 presents itself as a meta-study reviewing all LLM-for-CAD publications in selected venues from 2022 to 2024.","tokens_in":19545,"tokens_out":4457,"duration_ms":47817,"significance":"If the claims were fully supported, the survey would be a useful organizational reference for a fast-moving area: it collects a broad set of recent results, provides comparative tables (Tables 1 and 2) and a dataset overview (Table 3), and identifies concrete gaps such as the small size of VerilogEval and the limited exploration of multi-modal circuit representations. The survey makes no machine-checked derivations and ships no code, so its value is inherently organizational; for that kind of contribution, auditability of the corpus selection and accuracy of the attributed claims are the central quality criteria. The paper has no obvious internal circularity, since its own prior works [8,9] are cited in the body, but the central comprehensiveness claim does not depend on those citations.","major_comments":[{"comment":"Section 3 states that the meta-study 'reviews all publications in the selected venues' for 2022-2024, but it never names those venues and gives no search protocol, database, or inclusion/exclusion log. The only stated criteria are negative ('excludes... works that use LLMs as an application' and 'Only regular papers are considered while invited papers are excluded'). The de facto corpus in Tables 1-3 includes arXiv preprints, workshop papers, and at least one non-CAD item ([53], JavaScript unit-test generation), so the corpus does not transparently match the stated criteria. Because the trend claims in Section 3 are derived from this unlisted corpus, they are currently unfalsifiable; the authors should name the venues, give the search and screening protocol, and report the screening decisions, or revise the claim to describe a curated selection rather than 'all publications' in those venues.","section":"Section 3"},{"comment":"Section 4.1 asserts, without citation, that 'encoder-only models are shown to be more effective compared to encoder-decoder ones in recent research advancement.' This is unsupported and disconnected from the EDA evidence presented elsewhere in the paper: the survey's own tables and text focus on decoder-only and encoder-decoder-style generation models, and no EDA-specific encoder-only comparison is given. This claim is load-bearing for the paper's model-selection guidance, so it should either be removed or replaced with a cited, EDA-relevant comparison.","section":"Section 4.1"},{"comment":"The text uses Figure 2 to support the central trend claim that 'State-of-the-art customization methods have outperformed the best of prompt engineering methods with GPT-4', but the figure and its surrounding discussion do not state the exact evaluation configuration used (e.g., VerilogEval-machine versus VerilogEval-Human, pass@1 versus pass@k, temperature, or number of samples). Without this information, the quantitative comparison among RTLCoder, BetterV, VerilogEval, and prompt-engineering baselines cannot be checked or reproduced. Please specify the evaluation setup and the version of each benchmark, and state which curves in the figure support each qualitative conclusion.","section":"Section 3.2, Figure 2"}],"minor_comments":[{"comment":"The manuscript contains clear template artifacts that must be removed: the running header repeatedly reads 'Trovato et al.', the ACM Reference Format line gives a 2018 date and placeholder DOI, the Additional Key Words are the placeholder sentence 'Do, Not, Us, This, Code, Put, the, Correct, Terms, for, Your, Paper', and the footer says 'Manuscript submitted to ACM' on every page.","section":"Throughout"},{"comment":"The text identifies 'Perplexity's Claude3' and cites [3], but reference [3] is Anthropic's Model Card for Claude 3; the model is Anthropic's, not Perplexity's, and the attribution should be corrected.","section":"Section 2"},{"comment":"The discussion of LCDA reports a '25x speedup compared to state-of-the-art methods' without specifying the baseline method or the exact design point; please state the comparison target used in [65].","section":"Section 3.1"},{"comment":"There are several wording and typographical issues in this section, including 'close-sourced' for 'closed-source' and the sentence 'Encoder-only model [16] is primarily used for language prediction task, e.g., sentiment prediction', which misdescribes BERT-style models and should be rewritten.","section":"Section 4.1"},{"comment":"The tables would be more actionable if they included a unified column for quantitative evaluation metrics (e.g., pass rate or version of the benchmark); currently the same project can appear with different descriptions in different tables and the metric information is scattered in the text.","section":"Tables 1-3"}],"recommendation":"major_revision","confidential_remarks":"The template artifacts (repeated 'Trovato et al.' header, 2018 ACM reference, placeholder keywords) suggest the authors may not have seen the actual compiled submission; I would ask the editor to confirm that the revision process starts from a properly formatted manuscript. The self-citations to [8] and [9] are to related work by the same group and are clearly identified in the text, so I see no novelty-disclosure concern. The main technical gap is the unauditable meta-study in Section 3; if that is fixed, the survey could become a useful contribution, but in its current form the comprehensiveness claim is not verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's real value is organizational: it gives newcomers a structured map of LLM work across system-level design, RTL design, logic synthesis, physical design, and analog circuits, and it connects that work to methodology choices (prompting, fine-tuning, agents, RAG, multi-modal features). The tables comparing projects and the Figure 2 meta-analysis of VerilogEval/RTLLM numbers are genuinely handy — they condense a lot of scattered results into one place. For a survey, that is the job, and it does the job reasonably well.\n\nThe soft spots are real but mostly fixable. The template artifacts are embarrassing — the running header “Trovato et al.,” the ACM reference dated 2018, the placeholder keywords — and they signal a missing cleanup pass. The factual slip calling Claude3 “Perplexity’s Claude3” (Section 2) is a straightforward attribution error; the cited source is Anthropic. The claim that encoder-only models are “shown to be more effective compared to encoder-decoder ones” appears without a citation, which is exactly the kind of comparative statement a survey should not leave unsupported. These do not sink the paper, but they need correction.\n\nThe more serious issue is the meta-study in Section 3. The stress-test note is on target: the survey says it reviews “all publications in the selected venues” but never names the venues, never gives the search protocol, and gives only negative inclusion criteria. The de facto corpus includes arXiv preprints, workshop papers, and — as the off-topic Table 1 entry for JavaScript unit-test generation shows — at least one paper that does not match the stated EDA/HDL framing. Because the trend claims (“which areas are well-explored,” “which techniques have been proven effective”) are derived from this unlisted corpus, those claims are not independently checkable. That is a load-bearing gap for a survey that bills itself as comprehensive. A different venue choice or search string would plausibly change the trends.\n\nThat said, the central organizing function of the survey does not depend on the meta-study being exhaustive. Someone new to LLM-for-EDA will still get a fair picture of the main methods, benchmarks, and open problems from Sections 3 and 4. The paper is not a definitive quantitative account, but it is a reasonable entry point.\n\nBottom line: this deserves a serious referee, not a desk reject. A referee should ask the authors to name the venues and inclusion criteria, clean up the factual and template errors, and either cite or retract the architectural claim. With that revision, it would be a solid survey for newcomers and for researchers wanting a quick landscape overview.","headline":"A useful but sloppy survey of LLM-for-EDA that deserves peer review mainly so the authors can fix the factual errors, template artifacts, and the unverifiable meta-study corpus.","tokens_in":20074,"tokens_out":1149,"would_cite":false,"duration_ms":14080,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims LLMs now touch every major stage of chip design, with RTL generation the most developed application, while physical design and analog automation remain comparatively open.","keywords":["large language models","electronic design automation","hardware design","RTL code generation","Verilog","domain adaptation","autonomous agents","survey"],"falsifier":"Re-run the meta-study with the venue list and inclusion criteria made explicit and count regular papers per design stage; if the resulting distribution differs materially from the survey's reported trends — for instance, if physical design or analog papers are substantially undercounted — the survey's central map is falsified.","tokens_in":19100,"feed_emoji":"🔌","tokens_out":7147,"duration_ms":69361,"temperature":0.7,"pith_summary":"This paper is a survey of how large language models have been applied to electronic design automation (EDA) between 2022 and 2024. It claims to provide a comprehensive map of the field, organized by design stage — system-level design, RTL design, logic synthesis and physical design, and analog circuit design — and by the LLM technique used, such as prompt engineering, fine-tuning, retrieval-augmented generation, and autonomous agents. On the evidence it assembles, RTL code generation is the most active and most mature area, while backend physical design and analog design are comparatively open. The survey also identifies where progress comes from: domain-adapted and quality-weighted fine-tuning now outperform prompt engineering in RTL generation, and closed-source models such as GPT-4 currently lead autonomous agent frameworks.","feed_headline":"Survey maps LLMs' uneven progress across chip-design stages","feed_subtitle":"RTL generation is the mature hotspot; physical design and analog layout remain the open frontier.","key_machinery":"The organising machinery of the survey is a two-axis taxonomy: the EDA design-stage pipeline (system-level design, RTL design, synthesis and physical design, analog design) crossed with a set of LLM techniques (prompt engineering, domain-adaptive pre-training and supervised fine-tuning, retrieval-augmented generation, and autonomous agent frameworks). This taxonomy is what lets the paper turn a list of papers into claims about which stages are well-explored, which techniques work, and where the gaps are. The quantitative anchor is a small set of shared benchmarks — VerilogEval, RTLLM, and VerilogEval-Human — that let the survey compare functional correctness across models and methods in RTL generation.","core_discovery":"On the paper's own terms, the central result is an organized account of a young field: LLMs are being deployed across the whole EDA pipeline, but very unevenly. The RTL design stage dominates the literature, and within it three technique families have emerged — prompt engineering for commercial models, domain adaptation through supervised fine-tuning and custom datasets, and autonomous agent loops that let the model call tools and react to compiler or simulation feedback. The paper reports two quantitative patterns from its meta-study and benchmark comparisons: naive supervised fine-tuning improves with model size, while state-of-the-art customization methods that score or discriminatively guide generation (RTLCoder, BetterV) beat both naive fine-tuning and GPT-4 prompt engineering on VerilogEval; and autonomous agent methods raise VerilogEval-Human pass rates substantially, with GPT-4-based agents outperforming open-source ones. For the remaining stages, the paper concludes that script generation and documentation Q&A are established LLM niches, multi-modal and graph-based circuit representations are largely unexplored, and analog design work is only beginning.","pith_inferences":["Editorial extension: because the meta-study never names its selected venues, the survey's 'well-explored versus unexplored' map is only as trustworthy as an unstated venue choice; a re-run with explicitly defined venues could shift the balance between RTL and physical design.","Editorial extension: the paper's finding that naive SFT tracks model size, while quality-weighted methods outperform at smaller sizes, suggests that data quality weighting and controlled decoding may be a cheaper route to EDA-specific models than raw parameter scaling.","Editorial extension: a testable prediction of the survey's own outlook is that the next wave of LLM-EDA work will focus on graph- and image-based circuit representations; if that does not materialize, the survey's identified gap may not have been the binding constraint.","Editorial extension: the paper's emphasis on trust, verification, and proprietary-data leakage implies that industrial adoption of LLM-generated RTL will hinge on verification tooling and privacy-preserving training at least as much as on pass-rate improvements."],"forward_implications":["RTL code generation is the near-term payoff area: quality-weighted fine-tuning and controlled generation, not larger base models alone, are what currently push functional correctness past GPT-4 prompt engineering.","Backend physical design and analog layout remain the biggest open spaces; the survey names netlist generation, layout feature encoding, and timing estimation as specific targets for future LLM work.","Scaling LLMs for EDA is bottlenecked by data, not architecture: with VerilogEval at roughly 8,000 samples, expanding benchmark datasets is a precondition for benefiting from larger models.","Closed-source models like GPT-4 are currently the stronger choice for autonomous agent frameworks, but the gap is not fixed; open models improve when given planning and tool-integration scaffolds.","Multi-modal input — graph encoders for netlists and images for layouts — is the survey's main stated next frontier, and the field has barely started it."],"supporting_citations":[{"why":"Supplies the system-level co-design example (LCDA) and the 25x speedup over state-of-the-art that anchors Section 3.1.","marker":"[65]"},{"why":"Provides the VerilogEval benchmark used throughout the survey to compare functional correctness of RTL generation and to argue datasets are too small for scaling.","marker":"[34]"},{"why":"RTLCoder shows a small open-source fine-tuned model with scored Verilog data outperforming GPT-3.5, supporting the claim that customization beats prompt engineering.","marker":"[35]"},{"why":"BetterV uses a generative discriminator to guide CodeLlama generation and achieves state-of-the-art on VerilogEval-machine, evidence for controlled generation.","marker":"[44]"},{"why":"ChipNeMo demonstrates domain-adaptive pretraining, custom tokenizers, SFT, and retrieval for chip design, the basis of the domain-adaptation discussion.","marker":"[33]"},{"why":"ChatEDA supplies the LLM-as-agent instruction-tuning and EDA script generation example used in the synthesis and physical design section.","marker":"[24]"},{"why":"VerilogCoder contributes the quantified agent-framework result: 94.2% versus 60.3% pass rate with and without the agent, grounding the closed-source-leading claim.","marker":"[26]"},{"why":"Chip-Chat provides the end-to-end conversational system-level design example, from specification toward tapeout.","marker":"[5]"},{"why":"RTLLM is the prompt-engineering benchmark and self-planning method used to compare RTL generation quality.","marker":"[36]"},{"why":"RAG-EDA supplies the customized retriever and reranker for EDA documentation Q&A, backing the retrieval-augmented generation discussion.","marker":"[47]"}],"fun_headline_variants":["LLMs dominate RTL design but lag in analog layout","EDA survey: LLMs surge in RTL, stall in physical design","Chip design LLMs: RTL mature, analog still frontier","Survey: LLMs transform RTL, but rest of EDA lags","Uneven LLM adoption across chip design stages revealed"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's comprehensiveness depends on its unstated choice of which venues and papers its 2022–2024 meta-study covers; if that selection is narrower or biased, the reported distribution of work across design stages will not reflect the field as a whole.","fun_headline_variants_meta":{"raw":{"variants":["LLMs dominate RTL design but lag in analog layout","EDA survey: LLMs surge in RTL, stall in physical design","Chip design LLMs: RTL mature, analog still frontier","Survey: LLMs transform RTL, but rest of EDA lags","Uneven LLM adoption across chip design stages revealed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00051,"raw_usage":{"total_tokens":2458,"prompt_tokens":896,"completion_tokens":1562,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1473}},"tokens_in":512,"tokens_out":1562,"duration_ms":12386,"temperature":1.0,"reasoning_tokens":1473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:46:09.388411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the meta-study with the venue list and inclusion criteria made explicit and count regular papers per design stage; if the resulting distribution differs materially from the survey's reported trends — for instance, if physical design or analog papers are substantially undercounted — the survey's central map is falsified.","supporting_citations":[{"cited_title":"On the Viability of using LLMs for SW/HW Co-Design: An Example in Designing CiM DNN Accelerators","cited_arxiv_id":"2306.06923","evidence_quote":"Supplies the system-level co-design example (LCDA) and the 25x speedup over state-of-the-art that anchors Section 3.1."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ChatEDA supplies the LLM-as-agent instruction-tuning and EDA script generation example used in the synthesis and physical design section."}],"review_version":1}