Pith. sign in

Paper Citation Record · LEDGER

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

As of 17 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 26 inbound Pith citation observations for arXiv:2506.14074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14074 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:24.526196Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:08:22.502422Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.523834Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a7838e3-f16a-4832-9b13-90fe09896501 · outbound

This paper cites Cursor: The ai code editor, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Cursor: The ai code editor, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:25.118478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.394670Z digest=sha256:c9af14cd80afcbe6f846dda1aba182398395a41cb652b87b0430a9d38bdf8c8f

Observation becc77b3-2f9e-4970-9033-2b67cf9f6757 · outbound

This paper cites Large Language Model for Verilog Generation with Code-Structure-Guided Reinforcement Learning.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Large Language Model for Verilog Generation with Code-Structure-Guided Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.400299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.400299Z digest=sha256:bca718b4aa80c1de003f2511dc3f5803acbc997dbebd369bb4c6da912124bd4e

Observation fdc6561f-6e08-4e3c-ac6a-a7224fda8cc2 · outbound

This paper cites Craft RTL : High-quality synthetic data generation for verilog code models with correct-by-construction non-textual representations and targeted code repair.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Craft RTL : High-quality synthetic data generation for verilog code models with correct-by-construction non-textual representations and targeted code repair

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:25.099964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.406265Z digest=sha256:9b3a1a23734d8ef64e64c491198b630269edfecafb11464a34904313fd6ba0ea

Observation 4f091f46-d4e1-4384-98cd-aa0f2890d5a8 · outbound

This paper cites VerilogEval : Evaluating large language models for Verilog code generation, 2023.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification VerilogEval : Evaluating large language models for Verilog code generation, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:25.083406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.411209Z digest=sha256:22a0adfbe0c65913d178dede828f5153416ae81705c33151503abeea36503beb

Observation 59e82c01-446d-4df1-9992-4b23ef37607a · outbound

This paper cites VerilogCoder: Autonomous Verilog Coding Agents with Graph-based Planning and Abstract Syntax Tree (AST)-based Waveform Tracing Tool.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification VerilogCoder: Autonomous Verilog Coding Agents with Graph-based Planning and Abstract Syntax Tree (AST)-based Waveform Tracing Tool

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.418817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.418817Z digest=sha256:e758ecb57359492013ea75f3c4c8c9e18e0552c3981af4673a483db772565b3c

Observation 9236c40c-e7bd-4c83-842d-0cefb2f1798c · outbound

This paper cites Rtllm: An open-source benchmark for design rtl generation with large language model, 2023.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Rtllm: An open-source benchmark for design rtl generation with large language model, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.424937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.424937Z digest=sha256:7a06f0b5075bc12dc6501b875bbc7c35541d5feae43eff68006c21b38ebcc284

Observation ea25378d-a37b-4a9a-98e1-80ba1852acef · outbound

This paper cites OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.431514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.431514Z digest=sha256:d9715d1a8e693894a64597c81a82baef47757ae827ec260e68ce9e2f0c6220f9

Observation a7326f1d-2bd8-4b07-8af0-ca65a504e6e0 · outbound

This paper cites Revisiting verilogeval: A year of improvements in large-language models for hardware code generation.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Revisiting verilogeval: A year of improvements in large-language models for hardware code generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.437022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.437022Z digest=sha256:98b6c9a340c1daeedc9879aa3ae133e4c4c2aea208cb867aff7ef23d1c15a1cc

Observation 2148fa78-28b0-49ea-9d80-1ddf8b5fc72b · outbound

This paper cites Rtl-repo: A benchmark for evaluating llms on large-scale rtl design projects, 2024.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Rtl-repo: A benchmark for evaluating llms on large-scale rtl design projects, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:25.055507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.443021Z digest=sha256:45cb87f51440d9dd5e36113e9ad93df8ab2e0d56f1d74f400b78dd75102f1ab0

Observation 91b7475d-605b-48fa-88ea-a52c599c26dd · outbound

This paper cites LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.447859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.447859Z digest=sha256:4901e27bea56056b2734dd686fe932c0f993c452f9da7db053132e4f460cc8c6

Observation a7c348e9-3467-4a16-977a-d8ca47d0aea8 · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.453844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.453844Z digest=sha256:830ba64ba6ec7cd8f5ef396c39245d54408d0157d6489f3ad95e70f3e98a0641

Observation fa957e42-0ff2-4692-bc52-0e672e2e646a · outbound

This paper cites Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.458814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.458814Z digest=sha256:790f777f0cc86a33bdb1e56f8c365842b25165df074d311477190477479780a8

Observation ad66fe9e-216c-4ec2-ab35-6428c64f3bf2 · outbound

This paper cites Cocotb: Coroutine-based cosimulation testbench for vhdl and systemverilog, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Cocotb: Coroutine-based cosimulation testbench for vhdl and systemverilog, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:25.026688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.464206Z digest=sha256:244b7c488edee90d3242ad6e351a93c5fae7fb10fea229c4b53df2cc5d4573ad

Observation ea19c1ab-1376-47dc-8b02-40a897753ab7 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Bleu: a method for automatic evaluation of machine translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:24.469383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:59:24.469383Z digest=sha256:b78fc55a4781ff490c9919dbf57a64307a92a0cb866e5448f1575802bcb66158

Observation 3b7dd254-d76b-4adc-ae35-3da22695ee7b · outbound

This paper cites Icarus verilog, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Icarus verilog, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:25.010624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.474509Z digest=sha256:dc8d83e4bb7e2d6bb7c1573e9dffb96c13a29dafe28c6f372dbbfb372dadc32b

Observation d4e3ab9a-785f-43d6-8b08-ea640f9ba7ba · outbound

This paper cites Yosys open synthesis suite, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Yosys open synthesis suite, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.994180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.479128Z digest=sha256:983260239ceb476e0696e4ee16fb4755c71e8dde0ac4a79214e44ff29f5c9de0

Observation 96b2e548-3ed9-4b6e-a3a9-3d1f3912dec3 · outbound

This paper cites Verilator: Open-source systemverilog simulator, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Verilator: Open-source systemverilog simulator, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.976406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.484129Z digest=sha256:54a73a9f11cd695f77ce54cb26168f482bb5015347018c24b5b21e88e9296501

Observation d8356655-b1d4-49b9-87ab-4155e3cd5c20 · outbound

This paper cites Xcelium logic simulator, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Xcelium logic simulator, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.961616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.489216Z digest=sha256:92e790c541691f982c41dcf47beee685ae25a7888e7f6eb2126e166df380113c

Observation fe8cbc2b-7f8e-4bbe-bcae-ab30a9f10cbb · outbound

This paper cites Claude 3.7 sonnet, 2025.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Claude 3.7 sonnet, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.944192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.493833Z digest=sha256:301c20a73f1cc5aeff3bcf2d21cfd056f75c6b0f773c2c3db6a8bb604ab2b815

Observation 824be2fd-d8ef-4049-86c5-89e563adc852 · outbound

This paper cites Gpt-4.1, 2025 a.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Gpt-4.1, 2025 a

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.928185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.498684Z digest=sha256:2ae59d33f01092c5ba91f852cc4c61dc0cc0b7a150c67caad2e2bfdfc38152f3

Observation 73c66bb6-35f6-4703-a187-c15f7211bc08 · outbound

This paper cites Openai o1, 2024.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Openai o1, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.905761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.503453Z digest=sha256:6b0ef3bdbf791fcc27668d40d2b5be0d6d9e7f00022622cb10a6faafabd20cca

Observation 6553af7d-8cfd-4f66-808e-a4ae49435b65 · outbound

This paper cites Openai o4-mini, 2025 b.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Openai o4-mini, 2025 b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.875567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.507947Z digest=sha256:c634dd30992170d920fc590c49d32e23229429c6db0301a5adebe76dab611fd0

Observation 3f5bf0cf-6cf5-41f7-84d6-d8f274f2d92f · outbound

This paper cites Meta llama 3.1 405b, 2024 a.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Meta llama 3.1 405b, 2024 a

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.853477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.512350Z digest=sha256:49b805f6dfcc29f42c7a1886d9671bccec7f767527503bd992bebe34b7b4405c

Observation 914e6c69-9e97-4b42-b37d-40ea6b424762 · outbound

This paper cites Meta llama 3.1 70b, 2024 b.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Meta llama 3.1 70b, 2024 b

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.833259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.517088Z digest=sha256:813d0d057cd925b8d3ed35f69990c4fccfb6dfcd606b9397f554d1ec85e680ae

Observation b77c1f11-36b5-403b-8c70-3bc98e952442 · outbound

This paper cites Unsupervised k-means clustering algorithm.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Unsupervised k-means clustering algorithm

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.815184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.521566Z digest=sha256:30ae1437bdfdb69b95d6eb6ff3caaa9f1ff16ad7c9000235f1825b646a62f7fa

Observation 8865192a-cfcb-4134-8fb2-33ccc13adb85 · outbound

This paper cites Understanding how dimension reduction tools work: An empirical approach to deciphering t-sne, umap, trimap, and pacmap for data visualization.

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification Understanding how dimension reduction tools work: An empirical approach to deciphering t-sne, umap, trimap, and pacmap for data visualization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:59:24.797777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T19:59:24.526196Z digest=sha256:e230f2a4c5325788486183c2ba5bfc131f7b80d0960eb79c4530396dab9a5861

Pith citing papers

Observation 4fefa38d-41cd-4acc-9549-bf39075cadd6 · inbound

Revolution or Hype? Seeking the Limits of Large Models in Hardware Design cites this paper.

Revolution or Hype? Seeking the Limits of Large Models in Hardware Design Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:29.819033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:29.819033Z digest=sha256:bf1f85701e6d3cbe5ddbfc25b4252f5e262d62afa25dd1e66ad770b92d392241

Observation b4554590-ea58-42ce-b493-aeecd25816e8 · inbound

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks cites this paper.

HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.925789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T11:27:11.257817Z digest=sha256:86c0ca1e2fc3ba7649085f236bdd17026651b82122216654e45640898877a3de

Observation 6f2649cf-7615-4134-8db7-818d406a1eb7 · inbound

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement cites this paper.

Dr. RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:49:56.192583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T10:48:11.646826Z digest=sha256:4df54b17b078fed7c73daa2fd804e511e2b06265b9df3704d049821cf5ab1ec7

Observation 05a66429-c6ca-49f4-a6b7-c14695a4b5dc · inbound

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs cites this paper.

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:17:37.542242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:14:18.075112Z digest=sha256:7f5b0a6f53f54cb174601aaea59415a133655942f365d69c053b1c2148c0eecd

Observation 3d8b9413-7a79-4d95-a73a-12b631e29fea · inbound

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs cites this paper.

Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:24.133526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T10:18:18.439819Z digest=sha256:58268152e6dc6efab011dffe0a793cc7732067f2d05f6ecc98313f73275d52ee

Observation e9be5ce3-29dd-49ba-a076-528c46bea872 · inbound

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification cites this paper.

Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.037753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:00:22.126447Z digest=sha256:b14408b493edd0a0d4f07ff841a4389dd6aee71ee2f1fb8150b448a6cea3c1e0

Observation 66379d30-45c0-42e4-963b-2152a4175cc1 · inbound

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration cites this paper.

ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:36:08.318679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T01:24:08.563327Z digest=sha256:a11e5b55ada4f2b52bb6950bf1bb9220df8e47bb1b0744daf28132bb3b67814f

Observation c4b63f7b-e5d4-4c85-a536-402b82b2bc90 · inbound

SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation cites this paper.

SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:36:26.586732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T10:16:07.200458Z digest=sha256:639b5a002609f5b8058f815ce621e1eee510ff6dcb00e10887cc92844e276847

Observation e755e713-d726-41ca-9ac0-5d652deb0f22 · inbound

RuC: HDL-Agnostic Rule Completion Benchmark Generation cites this paper.

RuC: HDL-Agnostic Rule Completion Benchmark Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:29.100776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T08:16:15.514823Z digest=sha256:3ddfcda4b493552795b02c44e9bdce8f2e9bb37dbf072a1db32c5f52bdb56105

Observation 70ad44aa-d7cd-4b57-b824-2abbb68a642a · inbound

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation cites this paper.

ChipMATE: Multi-Agent Training via Reinforcement Learning for Enhanced RTL Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:05:05.892499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T21:59:20.163621Z digest=sha256:96db0d75f4d7f14f3c8b53025c855e9413860478ce65cd99aa545c4351982d95

Observation 98ed2132-5527-46c7-aa76-26bd1d1860bf · inbound

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision cites this paper.

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:42:37.291864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T14:42:16.802186Z digest=sha256:a9038716bec12ee6f6d330448e476227fb5717f54084eef30e7bc1cc90e02714

Observation 1eca5840-96c5-4243-80c5-a90980d25c3c · inbound

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications cites this paper.

AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:49.883285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T16:15:54.347613Z digest=sha256:35c1a091343d12ec54436e707297a57260c1b737cd88380395aee171c376f5e9

Observation 6731d922-2cd0-4f9f-a721-5779b433838a · inbound

CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs cites this paper.

CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:57:07.337550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T23:06:44.339563Z digest=sha256:537894fd61ef42ecac30d463e7e7d46c4b4e9bfee416581d312f2dafdde98220

Observation ad382723-37ab-41e8-a3aa-e1d3cccd64ac · inbound

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models cites this paper.

RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.492161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T17:01:24.292776Z digest=sha256:c9bc8e599b9fd493be1a7df25fe76200ccfb3dc1ac223c421bccb699ff9f5050

Observation 2d10777a-2324-4932-a571-b82ca9fbc2ac · inbound

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation cites this paper.

Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:58:32.897580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T06:50:01.901413Z digest=sha256:b7b9a08df4a10516a05bcbff9419c2035ea3978acd58e98984ace47d7c5c08f1

Observation 7449a3fe-d2c1-4234-9a82-afca4826b75a · inbound

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement cites this paper.

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:59.398936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T00:26:54.438203Z digest=sha256:452dd3190ebc9f3622c4ff1ec08fa68de455dbfee3c2c48cf2816de04b33e108

Observation 222552d4-52ea-4ec0-abce-269bc2e165ce · inbound

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement cites this paper.

Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T10:45:43.149976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:45:43.149976Z digest=sha256:bbc55c0930fd010153a1175355b1898a114e8ec6246ff3bed96ae7c77917a225

Observation 90f9680b-adf9-41a8-8e44-a8343be03cc5 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.525727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T01:48:58.638094Z digest=sha256:6c182dd84b56e90bb7e8263227f375840d011b0cc980031432669ea26933af31

Observation d767d608-007f-4b11-897a-32f4e6770163 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:39.759592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T10:15:10.471751Z digest=sha256:5515f621db0598adcfda7e55bb2ebd72952b5abbd28e284d7934a432e8bb0eaf

Observation 71f95b0d-fe5c-455d-bc95-6b9a60d62bd0 · inbound

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research cites this paper.

CHIA: An open-source framework for principled, agentic AI-driven hardware/software co-design research Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T10:03:40.395732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:03:40.395732Z digest=sha256:bbbd8a594a9758107594808baa538d7aa533cf338c1d766eefacd0440ab42ee7

Observation 08b2b0cb-1415-4b92-9a54-412c598df5aa · inbound

Agentic Hardware Design as Repository-Level Code Evolution cites this paper.

Agentic Hardware Design as Repository-Level Code Evolution Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:45:58.335609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T01:49:07.838897Z digest=sha256:5a791e7471cfbe8d14f8df200e7a161a2affc02fcaaccb679496a4cc371587c3

Observation c54ee891-db8a-4e77-9d1d-edc26fd2d2d5 · inbound

ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation cites this paper.

ChipVerilog: A Large-Scale OpenCores-Derived Benchmark for LLM-Based Verilog RTL Generation Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:12:17.834017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:12:17.834017Z digest=sha256:bc418432d78466101da66325d7d947b61738753a6304321c4b0cc2bfe1c154a6

Observation c960fe16-40d7-4547-8447-a1ef52a6030a · inbound

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering cites this paper.

When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T19:12:13.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:12:13.427128Z digest=sha256:d64a93aa45fe84ff653a8ad485f2acb33d30c8ea38e04d2335ae1e158dd93ccf

Observation 1c791175-4f57-49b6-a500-c8c73ff05263 · inbound

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows cites this paper.

Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T17:46:04.086561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:46:04.086561Z digest=sha256:4984dee2abd68c997fdd909a4be9c8f242ae537605786086e12d01d02a6e53e2

Observation 60de5fe6-2dfe-447c-b64e-78344d342b27 · inbound

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? cites this paper.

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing? Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:45:24.830628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:45:24.830628Z digest=sha256:9faa57f6c96825f1a5b84395635ca79f4d5cf7a76df045c8c00ff87a885d67c3

Observation c31426c2-2a29-4db9-a2d4-bdfdb26e00af · inbound

GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing cites this paper.

GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:08:22.502422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:08:22.502422Z digest=sha256:2c805396db9acf465f182b0b8aae8821f6f35f9d245bae8587d21865e6d9adbd