SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications
cs.AI, cs.SE
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/EDGAhab/specread
Terminology
Sources
- VerilogEval: Evaluating Large Language Models for Verilog Code Generation
- Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation
- RTLLM: An Open-Source Benchmark for Design RTL Generation with Large Language Model
- AssertLLM: Generating and Evaluating Hardware Verification Assertions from Design Specifications via Multi-LLMs
- AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications
- AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design
- RTL-BenchLS: A Large-Scale Benchmark for RTL Reasoning and Generation with Large Language Models
- Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
- VClare: Resolving Imperfect Specifications in LLM-Based Verilog Generation
- Understanding and Mitigating Errors of LLM-Generated RTL Code
- ChipNeMo: Domain-Adapted LLMs for Chip Design
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- DocNLI: A Large-scale Dataset for Document-level Natural Language Inference
- ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts
- SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization
- Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
- Ask-EDA: A Design Assistant Empowered by LLM, Hybrid RAG and Abbreviation De-hallucination
- Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA
- EDA Corpus: A Large Language Model Dataset for Enhanced Interaction with OpenROAD
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection