Pith. sign in

REVIEW 4 cited by

LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.04535 v2 pith:YKDYWC3Z submitted 2023-10-06 cs.LG cs.AR

classification cs.LGcs.AR
keywords hardwarellmsdesigntestautomatedgenerationstimuliconditions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Hardware design verification (DV) is a process that checks the functional equivalence of a hardware design against its specifications, improving hardware reliability and robustness. A key task in the DV process is the test stimuli generation, which creates a set of conditions or inputs for testing. These test conditions are often complex and specific to the given hardware design, requiring substantial human engineering effort to optimize. We seek a solution of automated and efficient testing for arbitrary hardware designs that takes advantage of large language models (LLMs). LLMs have already shown promising results for improving hardware design automation, but remain under-explored for hardware DV. In this paper, we propose an open-source benchmarking framework named LLM4DV that efficiently orchestrates LLMs for automated hardware test stimuli generation. Our analysis evaluates six different LLMs involving six prompting improvements over eight hardware designs and provides insight for future work on LLMs development for efficient automated DV.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnalogTester: A Large Language Model-Based Framework for Automatic Testbench Generation in Analog Circuit Design

    cs.MA 2025-07 conditional novelty 6.0 of 10

    An LLM multi-agent framework automatically generates TED testbenches for op-amps, bandgap references, and low-dropout regulators from research papers, with reported task success rates above 80 percent.

  2. Wit-HW: Bug Localization in Hardware Design Code via Witness Test Case Generation

    cs.AR 2025-08 unverdicted novelty 5.0 of 10

    Wit-HW generates witness test cases via mutation and uses spectrum-based comparison of passing and failing traces to rank buggy statements, reporting 49%, 73%, and 88% localization at Top-1, Top-5, and Top-10 across 41 bugs.

  3. SV-LLM: An Agentic Approach for SoC Security Verification using Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    SV-LLM automates SoC security verification with six cooperating LLM agents, reaching 84.8% vulnerability detection accuracy and 82% to 89% bug validation rates on benchmarks the paper does not disclose.

  4. VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation

    cs.AR 2025-07 conditional novelty 4.0 of 10

    A new pipeline and dataset of 20,392 synthesis-checked Verilog modules for LLM fine-tuning is presented, claimed to be the largest high-quality dataset of its kind.

Pith tools