PerfReasoning: How Well Do LLMs Reason on Hardware Performance?
cs.AI, cs.PF
Submitted: 2026-09-03
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement.
Terminology
Abstract
Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.
Sources
- HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
- The Championship Simulator: Architectural Simulation for Education and Competition
- Bench4HLS: End-to-End Evaluation of LLMs in High-Level Synthesis Code Generation
- Data-Driven Offline Optimization For Architecting Hardware Accelerators
- PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM/VLM Agents for VLSI Physical Design
- SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
- VerilogEval: Evaluating Large Language Models for Verilog Code Generation
- OpenLLM-RTL: Open Dataset and Benchmark for LLM-Aided Design RTL Generation
- The gem5 Simulator: Version 20.0+
- RTLLM: An Open-Source Benchmark for Design RTL Generation with Large Language Model
- KernelBench: Can LLMs Write Efficient GPU Kernels?
- Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
- AIRCHITECT: Learning Custom Architecture Design and Mapping Space
- Computer Architecture's AlphaZero Moment: Automated Discovery in an Encircled World
- ArchEval: Measuring AI Agents as Computer Architects
- KernelBenchX: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
- MultiKernelBench: A Multi-Platform Benchmark for Kernel Generation
- ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection