Automatic Harness Evolution for Hardware Design Verification: Can LLMs Consolidate Gains Across Discovered Harnesses?
cs.SE, cs.LG
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- ChipNeMo: Domain-Adapted LLMs for Chip Design
- Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation
- Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Automated Design of Agentic Systems
- AFlow: Automating Agentic Workflow Generation
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Self-Harness: Harnesses That Improve Themselves
- Recursive Harness Self-Improvement
- EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
- Rethinking the Evaluation of Harness Evolution for Agents
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
- Stop Comparing LLM Agents Without Disclosing the Harness
- Agentic Hardware Design as Repository-Level Code Evolution
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties