CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning
cs.AR, cs.CL
Submitted: 2026-09-15
Updated: 2026-09-19
Code: https://github.com/YosysHQ/mcy
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Design verification remains one of the most resource-intensive stages of hardware development, often consuming up to 70% of the total design effort.
Terminology
Abstract
Design verification remains one of the most resource-intensive stages of hardware development, often consuming up to 70% of the total design effort. While recent work has explored using Large Language Models (LLMs) to automate testbench generation, most existing approaches focus narrowly on functional correctness, overlooking the critical aspect of coverage quality. To bridge this gap, we present CovR, an agentic framework for automated testbench generation that combines self-reflection loops with simulation-based feedback to maximize coverage. Using this pipeline, we construct a large-scale dataset of 16,514 natural specification RTL reasoning testbench tuples with a strong teacher model, enabling coverage-aware supervision. Building on this, we propose a reinforcement learning (RL) framework tailored for coverage-driven testbench generation, leveraging tool-derived rewards from simulation and coverage feedback to optimize a student model. Experimental results show that the CovR finetuned model achieves 93.81% cov@10 on VerilogEval and RTLLM V2.0, and 87.76% cov@10 on CVDP, outperforming state-of-the-art approaches by 7.97% and 3.59%, respectively. Furthermore, deploying the finetuned model back into the agentic refinement pipeline further improves cov@10 to 94.27% on VerilogEval and RTLLM V2.0 and 91.39% on CVDP. Moreover, when integrated as a plug-in stimulus engine for full verification workflows, CovR improves coverage by 18.95% and mutation detection score by 1.19%, while revealing 4.46% undetected failures, highlighting the importance of optimizing for coverage in LLM-based hardware verification.
Sources
- Evaluating Large Language Models Trained on Code
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5-Coder Technical Report
- GRPO with State Mutations: Improving LLM-Based Hardware Test Plan Generation
- TB or Not TB: Coverage-Driven Direct Preference Optimization for Verilog Stimulus Generation
- Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
- How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
- Qwen3 Technical Report
- Insights from Verification: Training a Verilog Generation LLM with Reinforcement Learning with Testbench Feedback
- VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
- VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification
- LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- RTLSeek: Boosting the LLM-Based RTL Generation with Multi-Stage Diversity-Oriented Reinforcement Learning
- PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4