{"id":"0b87eaf5-3708-43bc-b84a-bd5e5da42adb","arxiv_id":"2507.19525","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new multimodal benchmark of 3,614 circuit QA pairs shows that large language models perform worst on back-end layout and computation tasks, and that current models generally underperform on circuit design questions.","lead":"This paper introduces MMCircuitEval, a benchmark of 3,614 circuit-design questions that tests how well AI models handle text and image tasks across chip design stages. It finds that most models score poorly on back-end layout questions and numerical circuit computations, pointing to a gap in training data and model capability.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The answer-scoring function in §III-C is validated only by an informal 100-question-per-stage review in §IV-B; if it is miscalibrated across stages or model families, the reported performance gaps and model rankings could be artifacts of the evaluator rather than of circuit competence.","rationale":"The reader identified the evaluator validity as the weakest assumption, and my independent reading converges on the same point. The benchmark construction and data curation are described in reasonable detail, and the dataset itself is likely reusable, but the empirical conclusions are entirely mediated by a scoring function whose validation is informal and whose same-family bias is acknowledged but untested. The proposed exact-match check on the substantial closed-form subset is feasible with the released data and would settle whether the reported rankings and gaps are artifacts of the scorer. This does not push the verdict below conditional; it reinforces the condition that the evaluator must be validated before the performance claims are accepted. A rejected or unverdictable stance would be too strong given that the dataset and taxonomy have independent value and the broad trends are plausible.","tokens_in":16247,"tokens_out":2645,"duration_ms":34552,"concrete_test":"Re-score all 1,220 closed-form questions (738 single-answer choice, 86 multi-answer choice, 396 fill-in-the-blank) with exact-match correctness, using set equality for multi-answer items. Then compare model-by-model and stage-by-stage rankings and gap magnitudes between exact-match accuracy and the paper's weighted metric. If the Spearman rank correlation across models is below ~0.85, or if the back-end versus front-end gap or the computation versus knowledge gap reverses direction, the evaluator bias concern is confirmed and the empirical conclusions are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claims—back-end design is hardest, computation is hardest, GPT-family models lead, ChipExpert excels—are all computed through a weighted similarity/GPT-preference score, not against ground-truth correctness. Section III-C assigns GPT-4-turbo preference double weight alongside BLEU, ROUGE, and embedding cosine similarity. Section IV-B reports only that testers were 'generally positive' about the aggregate scores; no inter-annotator agreement, no error analysis, and no exact-match comparison are provided. This matters because 1,220 of the 3,614 items (33.7%) are closed-form: 738 single-answer choice, 86 multi-answer choice, and 396 fill-in-the-blank, for which exact-match scoring is well-defined and objective. Without such an anchor, a systematic bias toward verbose, fluent, or GPT-styled outputs could reorder models or exaggerate the stage and ability gaps that the paper highlights. The limitation section explicitly concedes that the evaluator 'may favor models in the same family,' but this admitted risk is never quantified. Since every headline result depends on this scorer, validating it against exact correctness on the closed-form subset is the single most load-bearing missing check.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MMCircuitEval, a benchmark of 3,614 multimodal question-answer pairs for evaluating LLMs in circuit design, organized by EDA stage (general knowledge, specification, front-end, back-end), circuit type, tested ability, question type, and difficulty. The authors describe an expert-reviewed curation pipeline, evaluate a wide range of proprietary and open-source LLMs using a weighted combination of BLEU, ROUGE, embedding cosine similarity, and GPT-4-turbo preference, and report that models perform worst on back-end design and computation tasks, with GPT-family models and ChipExpert leading. The paper also reports CoT prompting experiments and discusses data collection and limitations.","tokens_in":16517,"tokens_out":2899,"duration_ms":32180,"significance":"If the benchmark and its evaluation protocol are sound, MMCircuitEval would be a useful and much-needed resource: it is the first multimodal circuit-focused benchmark spanning multiple EDA stages, it has a substantial and categorically rich question set, and it is publicly released. The strengths of the paper are the scale and diversity of the dataset, the expert-review curation procedure, the fine-grained metadata, and the broad coverage of tested models. However, the central empirical claims — that back-end design is hardest, that computation is the weakest ability, and that certain model families lead — are all computed through a scorer whose validation is informal. The evaluator validation, model comparisons without uncertainty estimates, and the admitted family bias of the GPT-based scorer are load-bearing gaps that must be addressed before the reported performance gaps and rankings can be accepted.","major_comments":[{"comment":"The proposed evaluator is not adequately validated for the claims built on it. The manual check in §IV-B samples 100 questions per stage and reports only that testers were \"generally positive\" about aggregate scores; there is no inter-annotator agreement, no error analysis, and no comparison against exact-match correctness. This matters because 1,220 of the 3,614 items (738 single-answer choice, 86 multi-answer choice, and 396 fill-in-the-blank) have well-defined ground-truth answers for which exact-match scoring is objective. Since all headline results — the stage ordering, the ability gaps, and the model rankings — are computed through the weighted similarity/GPT-preference score, a systematic bias in the scorer could reorder models or exaggerate gaps. Please validate the evaluator on the closed-form subset against exact-match accuracy, report agreement statistics and error patterns, and show that rankings are stable when the GPT-preference component is ablated.","section":"§III-C, §IV-B"},{"comment":"The double weight assigned to GPT-4-turbo preference creates a circularity risk that is acknowledged but never quantified. The paper states in Limitations that the evaluator \"may favor models in the same family (e.g., models in the GPT series)\", yet GPT-family models are among the top performers in Table III. Because the weight is a free parameter, the reported superiority of GPT-4v/GPT-4o over other models could be partly an artifact of the scorer rather than of circuit competence. Please report rank correlations or rank changes when GPT preference is down-weighted or removed, and ideally calibrate the metric weights against human correctness labels rather than setting GPT preference to 2 by construction.","section":"§III-C, §V"},{"comment":"All reported accuracies are single point estimates without confidence intervals, error bars, or significance tests. Differences such as GPT-4v at 69.4% versus GPT-4o at 68.0%, or the stage-level declines of 12.0–21.8% reported in §IV-C, may be within sampling noise. This is especially important for the fine-grained conclusions about stage ordering and ability ordering, which are based on a single run over 3,614 items. Please report bootstrap confidence intervals or per-item variance, and for the CoT experiment in Table V (100 random questions per stage) provide uncertainty estimates or a paired test before claiming improvement.","section":"Table III, Table IV, Table V"},{"comment":"The claim that GPT-generated questions \"exhibit minimal overlap with the training corpora of existing foundation models\" is asserted without any supporting check. This is load-bearing for fair horizontal comparison because GPT-generated questions are used to evaluate GPT-family models; if overlap is substantial, memorization could inflate their scores. Please describe the method used to test overlap (e.g., n-gram or embedding-based inspection, training-data cutoff analysis) or soften the claim to a statement of intent rather than a demonstrated property.","section":"§III-B, §III-D"}],"minor_comments":[{"comment":"The affiliation \"School of Intergrated Circuits, Southeast University\" should read \"Integrated\"; also, \"LlaMa\" is spelled inconsistently (e.g., LlaMa3.2 vs Llama 3) across the text and tables.","section":"Author affiliations"},{"comment":"The description of the evaluator validation should specify how the 100 questions per stage were selected, how many testers were involved, their domain expertise, and what instruction they received; \"generally positive\" is too vague to support the conclusion that the metric is effective.","section":"§IV-B"},{"comment":"The reference for LlaMa3.2-Vision-Instruct-90B is cited as [33], which is the Llama 3 herd paper; please provide the specific vision-model reference or clarify that the model is a derivative of that release.","section":"Table III"},{"comment":"The source column lists \"Synthesis\" for the MMCircuitEval row, which is not defined; if it refers to paraphrase-based augmentation of existing questions, the terminology should be explained in the text.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of the journal and the benchmark resource itself is commendable. The main risk is not the data construction but the evaluation protocol: the scorer validation, the absence of uncertainty estimates, and the unquantified family bias of the GPT-based component are all fixable within a revision, so I do not see a reason to reject. I would ask the editor to ensure that the revision explicitly addresses the exact-match anchor and the rank-stability analysis, since these are the preconditions for trusting the headline empirical claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline is that MMCircuitEval is a useful, reusable resource: 3,614 multimodal QA pairs spanning general knowledge, specification, front-end, and back-end design, labeled by circuit type, ability, and difficulty, with expert review and a broad model evaluation. That part is genuinely new and worth building on. It covers ground prior benchmarks don't – RTL comprehension, datasheet/spec questions, and back-end layout questions – and the authors release the benchmark, which matters.\n\nWhat the paper does well: the dataset construction is thoughtful, the categories are sensible, and the evaluation across 40-plus models is a lot of work. The fine-grained labels for ability and difficulty are a real asset. The paper also honestly acknowledges in its Limitations section that the GPT-based evaluator may favor same-family models, which is more than many benchmark papers do.\n\nThe soft spot is the evaluator. The reported numbers are not exact-match accuracies; they come from a weighted score where GPT-4-turbo preference counts double, alongside BLEU, ROUGE, and embedding cosine similarity. The validation in Section IV-B is a small informal check – 100 questions per stage, no inter-annotator agreement, no error analysis. Since 1,220 of 3,614 items are closed-form (single-answer, multi-answer, fill-in-the-blank), exact-match scoring would be a straightforward anchor. Its absence means the stage and ability gaps could partly reflect evaluator bias. The contamination claim – that GPT-generated questions have minimal overlap with training corpora – is asserted without evidence.\n\nThese are real issues, but they don't sink the benchmark. The broad trends (back-end hardest, computation hardest) are plausible and consistent with prior EDA-specific findings. The ranking of top models could shift under a stricter scorer, but the dataset itself remains a valid measurement instrument for the community.\n\nThis paper deserves serious peer review. A good referee should ask for exact-match scoring on the closed-form subset, error bars or significance tests, and a contamination check. As is, I'd treat the headline numbers as provisional; the dataset is the contribution, not the precise rankings. I'd bring it to a reading group for people working on LLM/EDA evaluation.","headline":"A genuinely useful multimodal circuit QA benchmark with a plausible but under-validated scoring function; the dataset deserves adoption, the rankings should be read with caution.","tokens_in":17092,"tokens_out":1472,"would_cite":true,"duration_ms":16923,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MMCircuitEval is a 3,614-question multimodal benchmark spanning digital and analog circuit design, and its results show current models are weakest at back-end design and computation-heavy questions.","keywords":["multimodal large language models","benchmark","electronic design automation","circuit design","back-end design","question answering","chain-of-thought","model evaluation"],"falsifier":"Have several circuit engineers independently score a random sample of model answers as correct or incorrect, then compare their human rankings against the benchmark's composite score. If the human ordering of models differs substantially from the composite ordering, or if a masked re-scoring by GPT-4-turbo changes when model identity is revealed, the paper's reported model rankings would not be reproducible.","tokens_in":16089,"feed_emoji":"🔌","tokens_out":6635,"duration_ms":69392,"temperature":0.7,"pith_summary":"The paper presents MMCircuitEval, a benchmark of 3,614 expert-reviewed question-answer pairs for testing multimodal large language models on circuit design. The authors claim it is the first such benchmark to span the major stages of the electronic design automation workflow, covering general knowledge, design specification, front-end design, and back-end design across both digital and analog circuits. Every question is labeled by design stage, circuit type, tested ability (knowledge, comprehension, reasoning, computation), and difficulty, so model performance can be diagnosed along each axis. Evaluations of a broad set of models show that current systems are weakest on back-end design questions and on computation-heavy problems, which the paper takes as evidence that circuit-specific training data and image-processing strategies are the main bottlenecks. If the benchmark measures what it claims, it gives the field a reusable instrument for tracking progress in AI-assisted chip design.","feed_headline":"Benchmark finds LLMs weakest at chip back-end design","feed_subtitle":"MMCircuitEval's 3,614 expert-reviewed questions show where AI chip assistants fail, from layout knowledge to numerical computation.","key_machinery":"The load-bearing object is the benchmark itself: a curated set of 3,614 question-answer pairs organized along four axes—design stage (general knowledge, specification, front-end, back-end), circuit type (digital versus analog), tested ability (knowledge, comprehension, reasoning, computation), and difficulty (easy, medium, hard). Answers are scored by a weighted composite that doubles the weight of a GPT-4-turbo preference rating and singly weights BLEU-4, average ROUGE, and embedding cosine similarity, with models required to provide explanations. That composite is the machinery that turns free-form model outputs into the accuracy numbers underlying every comparison in the paper.","core_discovery":"The paper's central claim is that MMCircuitEval is a valid, reusable instrument for measuring how well multimodal large language models handle circuit-design questions across the full EDA workflow, and that the measurements it reports reveal a consistent weakness profile in current models. Across 3,614 questions, the strongest tested model answers about 69 percent correctly, most open-weights models stay below 50 percent, back-end design questions trail other stages by roughly 12 to 22 percentage points, and computation questions are the hardest category for nearly every model family. The paper also reports that models with image encoders often score lower on multimodal questions than text-only models that receive automatically generated image captions, and that a small circuit-specialized model can outperform much larger general-purpose ones. Taken together, these results support the paper's conclusion that circuit-specific training data and image-processing strategies, rather than raw model scale, are the main levers for progress.","pith_inferences":["A direct validity check would mask model identity in the GPT preference scorer; if scores change depending on whether the evaluated model comes from the same family, the evaluator has a self-preference that needs correction.","The ability and difficulty labels suggest a natural curriculum: circuit-specialized training could order data from knowledge to computation and easy to hard.","The benchmark's format could be extended to interactive or multi-turn design scenarios where the model must request missing datasheet or netlist information before answering.","Because the dataset contains real-world datasheets and netlists, it could be used to synthesize additional training pairs, attacking the data scarcity that the paper identifies as the main bottleneck."],"forward_implications":["Adding high-quality back-end design and layout data to training corpora becomes the direct route to raising model scores, since back-end accuracy lags every other stage by 12 to 22 points.","A circuit-specialized small model beating much larger general-purpose models implies that domain-specific fine-tuning can offset raw scale in EDA tasks.","Because captioned text-only models can outperform models with native visual encoders on multimodal questions, circuit images currently hurt most multimodal models more than they help.","Chain-of-thought prompting helps most on front-end reasoning and little on knowledge retrieval or computation, so test-time compute is best spent on reasoning-heavy questions.","The per-question labels for ability and difficulty convert the benchmark from a single leaderboard into a diagnostic tool for per-skill model comparison."],"supporting_citations":[{"why":"Motivates domain-adapted LLMs for chip design, the use case MMCircuitEval evaluates.","marker":"[1]"},{"why":"Semiconductor-industry LLM used as motivation for circuit-focused evaluation.","marker":"[2]"},{"why":"Existing EDA benchmark focused on the OpenROAD toolchain, used as a comparison baseline.","marker":"[8]"},{"why":"Existing EDA tool-documentation QA benchmark, used as a comparison baseline.","marker":"[9]"},{"why":"Source of the ChipExpert circuit-focused baseline and its results in the comparisons.","marker":"[25]"},{"why":"GPT-4 technical report backing the GPT-4/GPT-4v baselines and the GPT-4-turbo preference scorer.","marker":"[26]"},{"why":"GPT-4o system card; GPT-4o is used for ability-label assignment and answer correctness evaluation.","marker":"[37]"},{"why":"Llama 3 family provides the backbone for several baselines, including the text-only circuit-specialized model.","marker":"[33]"},{"why":"Cited as the chain-of-thought technique used in the test-time improvement experiments.","marker":"[54]"}],"fun_headline_variants":["Circuit test exposes LLMs' biggest gap: back-end design","Multimodal LLMs lose to caption-only models on circuits","Small chip-focused model outperforms giant LLMs in EDA","Open-source LLMs fail half of circuit design questions","New benchmark: image input hurts LLM circuit reasoning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes the weighted mix of text-similarity scores and GPT-4-turbo preference faithfully reflects whether an answer is correct for circuit problems, with validation limited to a manual check of 100 questions per stage and no inter-annotator agreement or exact-match comparison.","fun_headline_variants_meta":{"raw":{"variants":["Circuit test exposes LLMs' biggest gap: back-end design","Multimodal LLMs lose to caption-only models on circuits","Small chip-focused model outperforms giant LLMs in EDA","Open-source LLMs fail half of circuit design questions","New benchmark: image input hurts LLM circuit reasoning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1861,"prompt_tokens":977,"completion_tokens":884,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":802}},"tokens_in":593,"tokens_out":884,"duration_ms":11013,"temperature":1.0,"reasoning_tokens":802,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:46:14.114144+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several circuit engineers independently score a random sample of model answers as correct or incorrect, then compare their human rankings against the benchmark's composite score. If the human ordering of models differs substantially from the composite ordering, or if a masked re-scoring by GPT-4-turbo changes when model identity is revealed, the paper's reported model rankings would not be reproducible.","supporting_citations":[{"cited_title":"SemiKong: Curating, Training, and Evaluating A Semiconductor Industry-Specific Large Language Model","cited_arxiv_id":"2411.13802","evidence_quote":"Semiconductor-industry LLM used as motivation for circuit-focused evaluation."},{"cited_title":"EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD","cited_arxiv_id":"2405.06676","evidence_quote":"Existing EDA benchmark focused on the OpenROAD toolchain, used as a comparison baseline."},{"cited_title":"Language models are few-shot learners,","cited_arxiv_id":null,"evidence_quote":"Cited as the chain-of-thought technique used in the test-time improvement experiments."}],"review_version":1}