Pith. sign in

Paper Citation Record · LEDGER

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2506.14074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14074 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:45:24.830628Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.523834Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4fefa38d-41cd-4acc-9549-bf39075cadd6 · inbound

Revolution or Hype? Seeking the Limits of Large Models in Hardware Design cites this paper.

Revolution or Hype? Seeking the Limits of Large Models in Hardware Design Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:29.819033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:29.819033Z digest=sha256:b2f10350e92c29025c08d235cba6316c8512d5cc700100ff469c8d1f3b08c672

Observation b4554590-ea58-42ce-b493-aeecd25816e8 · inbound

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks cites this paper.

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.925789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:27:11.257817Z digest=sha256:517bb6867acfe1c4c12a2808c61f2c78b8c6fbab83182b1a20be1f4875b269de

Observation 6f2649cf-7615-4134-8db7-818d406a1eb7 · inbound

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement cites this paper.

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:49:56.192583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T10:48:11.646826Z digest=sha256:5b82ac120cf274b1cffeeeda7c4e5fc192d7ff1d286ce4ee28d6d66d45731f77

Observation 05a66429-c6ca-49f4-a6b7-c14695a4b5dc · inbound

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs cites this paper.

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.542242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:14:18.075112Z digest=sha256:13321928ad17c24e0e5cad2ad39726a0365d7f40845a1334d221a455b2b8a38c

Observation 3d8b9413-7a79-4d95-a73a-12b631e29fea · inbound

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs cites this paper.

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:24.133526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T10:18:18.439819Z digest=sha256:6e97c3af8a01a4836aa139d1e5be19be65448367a499865384e4427d39c19eec

Observation e9be5ce3-29dd-49ba-a076-528c46bea872 · inbound

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification cites this paper.

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.037753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:00:22.126447Z digest=sha256:1d9ad8802db191fcb028edd2e9d44daf9020bb37a76299a37003066c023afcd0

Observation 66379d30-45c0-42e4-963b-2152a4175cc1 · inbound

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration cites this paper.

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.318679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:24:08.563327Z digest=sha256:71ad45f759f392f846855d82e06e0903a3fb2bf289db4a2ea24c1565012a425a

Observation c4b63f7b-e5d4-4c85-a536-402b82b2bc90 · inbound

SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation cites this paper.

SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:26.586732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T10:16:07.200458Z digest=sha256:8c5bbea76defdfc30c4e9bb57030809aeae995d8816f476be682e6afcec8cd0d

Observation e755e713-d726-41ca-9ac0-5d652deb0f22 · inbound

RuC: HDL-Agnostic Rule Completion Benchmark Generation cites this paper.

RuC: HDL-Agnostic Rule Completion Benchmark Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:29.100776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T08:16:15.514823Z digest=sha256:85e5f564c83fb0229f5ea8bdb1e244aee9a69790a8d257c5bc92fe381a6d2d39

Observation 70ad44aa-d7cd-4b57-b824-2abbb68a642a · inbound

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation cites this paper.

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:05.892499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:59:20.163621Z digest=sha256:97e386201947fdbee8fc189c5e846567eef662e4e92e757cd67d5e349bf47483

Observation 98ed2132-5527-46c7-aa76-26bd1d1860bf · inbound

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision cites this paper.

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:42:37.291864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T14:42:16.802186Z digest=sha256:d124510a1b29dbba9d4cf82cea15db5906662380ae674c7dfc31604b20b532cd

Observation 1eca5840-96c5-4243-80c5-a90980d25c3c · inbound

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications cites this paper.

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:49.883285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T16:15:54.347613Z digest=sha256:7baa39fb434473b365bfc06fa11b5418aab9c0923189db32c98f585f27d854bd

Observation 6731d922-2cd0-4f9f-a721-5779b433838a · inbound

CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs cites this paper.

CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:57:07.337550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T23:06:44.339563Z digest=sha256:a645069a6e3ce7d456e03db8b7144d3d07c296ccfdec239f01da6e414b8e52fb

Observation ad382723-37ab-41e8-a3aa-e1d3cccd64ac · inbound

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models cites this paper.

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.492161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:01:24.292776Z digest=sha256:961e5feb1ac8c721ba81eba862a7fb6df452a15781ac14ab6fdec4293ae85354

Observation 2d10777a-2324-4932-a571-b82ca9fbc2ac · inbound

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation cites this paper.

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.897580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:50:01.901413Z digest=sha256:45e0ec3e49cf44e76c9568e8d477d76b83af2fd37edaac99d98a0d50872df395

Observation 7449a3fe-d2c1-4234-9a82-afca4826b75a · inbound

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement cites this paper.

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:59.398936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T00:26:54.438203Z digest=sha256:e17c6b4b16670a13178d25d4fc5748f150ee00758cd6b5765fa7abd82587e90a

Observation 222552d4-52ea-4ec0-abce-269bc2e165ce · inbound

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement cites this paper.

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T10:45:43.149976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:45:43.149976Z digest=sha256:eea5c028aa2a813eedf237a28d30bda72c5504d164280451c39e944557cc3e94

Observation 90f9680b-adf9-41a8-8e44-a8343be03cc5 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.525727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:48:58.638094Z digest=sha256:a3b496c8d3689c38a47cc3a2d395a930a62b9226d73689a6410e826797a6da0c

Observation d767d608-007f-4b11-897a-32f4e6770163 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:39.759592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:15:10.471751Z digest=sha256:aa9e9ca54f3ba5b9b1fdbffb65c7226b6d032d252457c6834eb06f4d6ccdbff2

Observation 71f95b0d-fe5c-455d-bc95-6b9a60d62bd0 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T10:03:40.395732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:03:40.395732Z digest=sha256:09d0459845557c052350bb12819467a579158aadcdc0e8bc0a1046b0899ca0de

Observation 08b2b0cb-1415-4b92-9a54-412c598df5aa · inbound

Agentic Hardware Design as Repository-Level Code Evolution cites this paper.

Agentic Hardware Design as Repository-Level Code Evolution Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:45:58.335609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T01:49:07.838897Z digest=sha256:c7aca3d5890498740198a8540b23508bd2fdfa149b413b41b222d49bac248da0

Observation c54ee891-db8a-4e77-9d1d-edc26fd2d2d5 · inbound

ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation cites this paper.

ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:12:17.834017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:12:17.834017Z digest=sha256:463e00bccfc6cac1e09c5f5cfd483f0142a2e5f8477ff49a3a8cde275fdf6c3a

Observation c960fe16-40d7-4547-8447-a1ef52a6030a · inbound

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering cites this paper.

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T19:12:13.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:12:13.427128Z digest=sha256:7d6251596e88266bacd611eca24f585c83971a34b6f38f22b92198c6338167a5

Observation 1c791175-4f57-49b6-a500-c8c73ff05263 · inbound

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows cites this paper.

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:46:04.086561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:46:04.086561Z digest=sha256:7108e2cfc4b4ffcd0cf578984bd7e368431302f2e2347a532f2009eebb77cf12

Observation 60de5fe6-2dfe-447c-b64e-78344d342b27 · inbound

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? cites this paper.

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:45:24.830628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:45:24.830628Z digest=sha256:b466d562bda111f69352631b3b8432388ac09210a4a893fc1aeeb049199eafee