Pith. sign in

REVIEW 3 cited by

OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.15112 v1 pith:O32XPLZY submitted 2025-03-19 cs.AR

classification cs.AR
keywords designbenchmarkgenerationdatasetcodedatasamplestraining
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The automated generation of design RTL based on large language model (LLM) and natural language instructions has demonstrated great potential in agile circuit design. However, the lack of datasets and benchmarks in the public domain prevents the development and fair evaluation of LLM solutions. This paper highlights our latest advances in open datasets and benchmarks from three perspectives: (1) RTLLM 2.0, an updated benchmark assessing LLM's capability in design RTL generation. The benchmark is augmented to 50 hand-crafted designs. Each design provides the design description, test cases, and a correct RTL code. (2) AssertEval, an open-source benchmark assessing the LLM's assertion generation capabilities for RTL verification. The benchmark includes 18 designs, each providing specification, signal definition, and correct RTL code. (3) RTLCoder-Data, an extended open-source dataset with 80K instruction-code data samples. Moreover, we propose a new verification-based method to verify the functionality correctness of training data samples. Based on this technique, we further release a dataset with 7K verified high-quality samples. These three studies are integrated into one framework, providing off-the-shelf support for the development and evaluation of LLMs for RTL code generation and verification. Finally, extensive experiments indicate that LLM performance can be boosted by enlarging the training dataset, improving data quality, and improving the training scheme.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing

    cs.AR 2026-08 conditional novelty 6.0 of 10

    A mutation-testing audit of RTL benchmark testbenches finds RTLLM v2.0 mostly below a 95% fault-kill floor and shows that an output-token cap can reorder a leaderboard.

  2. Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

    cs.LG 2025-06 conditional novelty 6.0 of 10

    CVDP is a new 783-problem benchmark for LLM-based RTL design and verification, and it finds that state-of-the-art models pass no more than 34% of code generation tasks.

  3. Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback

    cs.AR 2025-04 conditional novelty 6.0 of 10

    VeriPrefer improves Verilog generation LLMs by automatically building testbenches, using them to construct preference pairs, and training with direct preference optimization to increase functional correctness.

Pith tools