HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning".
Jane: The paper was written by Qingyun Zou, Feng Yu, Hongshi Tan, Yao Chen, Bingsheng He et al. from National University of Singapore.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary: Tom: Okay, so we’ve established what HLS-Seek is trying to fix—the gap between easy-to-write software and hard-to-run hardware—and now we're looking at the paper's summary, focusing on how they actually achieved this using Reinforcement Learning.
Jane: If I understand correctly, the main innovation here is using a "Proxy Comparative Reward." Can you break down what that means in simpler terms for our listeners?
Lu: It’s clever because instead of needing a perfect, pre-existing simulation environment to grade every single code snippet—which is computationally impossible—they are using a comparative reward signal.
Meng: Wait, so they aren't training the model on the *absolute* best possible code, but rather teaching it to generate code that is better than some baseline or a comparison sample? That sounds like a massive data saving trick.
Lalam: Precisely, Meng; comparing against something known—a weak performance versus a strong performance—is much more tractable for an RL agent than trying to optimize against perfect reality all at once.
Tom: Right, so the RL framework is guiding the LLM's code generation process iteratively, improving the quality step-by-step based on these comparisons rather than just massive datasets of perfect examples.
Jane: That sounds like they’re letting the AI learn by comparison, like a student who learns best by seeing an A+ paper and then trying to write something that matches that structure and quality.
Lu: And the "Proxy" part is what makes it feasible; it’s abstracting those complex hardware metrics into a measurable score that the RL agent can actually use to adjust its policy during training.
Meng: From an engineering standpoint, implementing this comparative reward function itself must be complicated; how robust are these proxy metrics when applied across different target FPGA architectures?
Lalam: Because the field of hardware acceleration is moving so fast, having a framework like HLS-Seek that learns to optimize against measurable proxies means that the underlying cultural hurdle—the inability to easily test and validate performance at scale—gets significantly reduced.
Improvements: Tom: We’ve talked about the overall mechanism, and now we’re digging into what improvements "HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning" suggests over existing methods.
Jane: The paper really seems to emphasize moving beyond just syntactic correctness to achieving genuine *functional* and *resource* efficiency, which is a big step up from what we usually see in pure code generation models.
Lu: I think the key improvement they are highlighting is the shift in objective function; they aren't just optimizing for code completion probability, but for hardware-specific quality metrics like latency or area utilization.
Meng: So, if a standard LLM generates code that is readable and passes basic compilation checks, HLS-Seek actively pushes it to be *smaller* or *faster* on the target silicon? That changes the entire optimization loop.
Lalam: It implies that future AI tools shouldn't just be general-purpose code generators; they need to become specialized domain experts who inherently understand the physical constraints of the hardware they are targeting.
Tom: And it sounds like this approach makes the entire cycle—design, generation, synthesis, validation—much more tightly coupled and intelligent than before.
Jane: It’s like giving the AI not just a textbook to read from, but also a stopwatch and a power meter while it writes the answers.
Lu: Specifically, I think they are showing that by incorporating this comparative reward structure, they bypass some of the brittle assumptions that plague current end-to-end synthesis tools when trying to integrate generative models.
Meng: If this scales, it means we could design complex accelerators for niche scientific problems—like genomics or advanced radar processing—much faster than current manual expert cycles allow.
Lalam: The impact here extends beyond just speed; it democratizes hardware acceleration. Before, only large labs with massive compute clusters could afford to iterate on custom silicon designs, but this framework lowers the barrier dramatically.
Conclusion: Tom: Wow, we’ve covered a lot of ground discussing "HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning," and it really feels like a paradigm shift is coming in how we design custom hardware.
Jane: It’s amazing to think that AI can move from just writing code that *looks* right, to writing code that *physically performs* right on specialized silicon.
Lu: I’m incredibly excited about the possibilities for domain-specific architectures; imagine entire fields of science being able to prototype custom compute units overnight instead of waiting years for fabrication.
Meng: From my side, the practical hurdle I see getting cleared is the iteration time; if we can rapidly prototype performance improvements using this, development cycles shrink dramatically, which is a massive win for industry adoption.
Lalam: Looking at the cultural shift, this technology means that expertise in hardware architecture becomes more
Conclusion: Tom: So we’ve spent a lot of time dissecting how it works, but we need to wrap up the big picture for everyone listening about the core achievement of "HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning."
Jane: It’s truly a remarkable feat, Tom. We’ve moved from teaching a model to just write syntactically correct code to training it to actually design the physical hardware itself—that's what matters for our listeners.
Lu: I think the most important thing for me is that this signals a major paradigm shift in how we view AI; we’re no longer just generating text or images, we are becoming automated architects.
Meng: From an engineering standpoint, it means that instead of needing massive teams to manually optimize kernels over and over again, we now have a viable path toward truly autonomous hardware design.
Lalam: The biggest cultural impact I see is that this democratizes sophisticated silicon design, making high-performance accelerators available to smaller organizations who can’t afford traditional expert-level hardware engineers.
Tom: That's a huge vision, Lalam, and it brings us back to the practical success of HLS-Seek.
Jane: It feels like we've seen the end of an era where AI was just a fancy autocomplete tool for hardware design.
Lu: And it’s definitely pushing the boundaries of what this specific architecture allows for complex optimization patterns we once thought were impossible to teach an LLM.
Meng: We can actually look forward to deploying these tools in industrial settings, knowing the iterative process is far more efficient and cost-effective than previous manual methods.
Lalam: It’s a new level of collaboration between powerful AI and accelerating hardware itself, creating a future where performance is baked into the design intent.
Tom: Well, that’s our take on it for now, Jane. It's clear that HLS-Seek represents a significant leap in automated hardware engineering.
National University of Singapore
cs.LG, cs.AI
Submitted: 2026-05-13
Updated: 2026-09-04
Comments: Accepted at ICCAD 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: HLS-Seek is a novel framework designed to address a critical gap in existing Large Language Model (LLM) approaches to High-Level Synthesis (HLS).
Key concepts
- Proxy Comparative Reward
- This innovative reward signal allows the AI to train without needing a perfect simulation environment. Instead of optimizing against ideal reality, the model learns by comparison, generating code that is measurably better than a known baseline or sample.
- High-Level Synthesis (HLS)
- HLS is the process of converting high-level software code into optimized hardware descriptions for specialized silicon. The goal is to ensure the resulting code achieves genuine functional and resource efficiency, moving beyond simple syntactic correctness.
- Reinforcement Learning (RL) in Code Generation
- RL guides an AI model's code generation process by making iterative improvements. The model learns to optimize its output step-by-step based on measurable performance comparisons and rewards, rather than just relying on massive datasets of perfect examples.
Terminology
Summary
HLS-Seek is a novel framework designed to address a critical gap in existing Large Language Model (LLM) approaches to High-Level Synthesis (HLS). While previous LLM methods focused primarily on achieving functional correctness, HLS-Seek introduces Quality of Results (QoR)—a measure of hardware performance such as latency and resource utilization—into the training loop. This is critically important because QoR dictates the real-world efficiency and deployment viability of hardware accelerators, enabling HLS-Seek to achieve state-of-the-art results in both functional correctness and hardware efficiency using a highly efficient, proxy reward mechanism.
The Core Insight: Proxy Comparative Reward
The fundamental breakthrough of HLS-Seek is the realization that Reinforcement Learning (RL) for HLS does not require absolute synthesis results. Instead, the system leverages relative comparisons between candidates.
This allows the training process to be driven by a lightweight comparative reward model that predicts which of two designs achieves a better resource-latency trade-off. This approach bypasses the prohibitively expensive
synthesis-in-the-loop paradigm, which would require minutes or hours per evaluation. The model is trained on this relative advantage, achieving 99.53% Pareto-dominance accuracy
without ever invoking real synthesis for to determine the reward signal.
The Three-Stage Training Pipeline
HLS-Seek employs a comprehensive three-stage training pipeline to build robust, hardware-aware reasoning into the LLM:
-
Diversity-Oriented SFT: The model is initially fine-tuned on a broad HLS corpus, exposing it to
diverse pragma configurations
and mapping natural language descriptions to HLS implementations. This stage focuses on coverage over optimality. -
Cold-Start SFT: Using the Pareto-Proximal Quality-Diversity Sampling strategy, the the model is trained on high-quality data that includes
reasoning traces
generated by commercial models, instilling ahardware reasoning capability.
-
Group Relative Policy Optimization (GRPO): The final stage uses GRPO with a four-component reward system:
-
Reasoning Format Reward (r f): Ensures the output conforms to a structured format.
-
Compilation Reward (r comp): Verifies syntactic correctness via the C compiler.
-
Functional Correctness Reward (r c): Confirms semantic correctness against testbenches.
-
Predicted QoR Reward (r q): The core proxy reward, used for final optimization.
The Uncertainty-Aware Proxy Mechanism
To prevent reward hacking
as the LLM explores new pragma configurations, HLS-Seek integrates an uncertainty-aware Monte Carlo (MC) dropout switching mechanism. This system uses the MC dropout variance (u i) to estimate its own confidence in a candidate's QoR score. If u i exceeds a threshold, the model selectively falls back to real Vitis HLS synthesis.
The resulting true QoR is then used to online update the proxy,
creating a self-improving reward system that maintains fidelity while keeping synthesis overhead low.
Performance and Efficiency
HLS-Seek demonstrates superior performance compared to foundational models and other HLS-specific baselines. On the HLS-eval benchmark, it achieves 81.5% syntax correctness pass@1 and 81.4% functional correctness pass@5, outperforming GPT-5.1 with a substantial gain in both metrics. Furthermore, the model excels in hardware efficiency:
-
Latency: It achieves the
lowest latency on 16/30 kernels.
-
Pareto Dominance: It
Pareto-dominates HLS-specific baselines
on multiple benchmarks.
The efficiency of this approach is also notable, achieving an 8.5× faster training time compared to the full synthesis-in-the-loop RL methods, which require immense computational resources and time.
Improvements for AI systems
Based on a rigorous analysis of HLS-Seek (Qor-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning), the core innovations are not limited to High-Level Synthesis but represent transformative methodologies applicable to any AI system where the true objective function is costly, multi-objective, and requires continuous optimization.
Here are the specific improvements applicable to general AI systems:
The Improvement: Replace reliance on absolute, high-cost evaluation metrics (e.g, real-world latency or precise loss values) with a lightweight Siamese comparative reward model trained to predict the relative advantage between two candidates.
Specific Mechanism: Train a shared encoder (e.g., using a Transformer architecture) on pairs of outputs (A vs B) to calculate a preference probability P(A > B). This model uses techniques like Bradley-Terry modeling and is trained with both full dominance labels (score 1.0) and softer, relative superiority labels (score 0.5).
What the Improved AI System Can Do:
-
Achieve Massive Computational Scale: The system can perform vastly more training iterations because proxy inference takes milliseconds (ms), whereas real-world evaluation takes minutes or hours, achieving a massive speedup (e.g., 8.5x faster than real-reward RL).
-
Optimize Multi-Objective Tasks: It enables the simultaneous optimization of disparate objectives (e.g., maximizing throughput while minimizing power consumption) by ranking candidates based on their Pareto dominance relative to the other candidates in a specific design space, rather than requiring an artificial weighting of absolute values.
The Improvement: Integrate Monte Carlo (MC) dropout uncertainty estimation into the proxy reward loop, creating a self-corrective mechanism that determines when the proxy model is unreliable and requires costly real-world verification.
Specific Mechanism: The system performs multiple stochastic forward passes (M=10) through the encoder. If the variance (Var) of scores exceeds a predefined threshold (tau), the system triggers a fallback to real-world evaluation (e.g., running Vitis HLS or a physical simulation). The results are then fed back to update and fine-tune the proxy model online.
What the Improved AI System Can Do:
-
Prevent Reward Hacking: It robustly guards against the policy exploiting blind spots in a proxy model by grounding rewards at states where the proxy is inherently unreliable, ensuring fidelity without constant overhead.
-
Maintain Efficiency and Fidelity Simultaneously: It guarantees high reward accuracy (e.g., 99.53% dominance accuracy) while keeping the total cost of real-world evaluation below a predefined target (e.g., 5% trigger rate), making it applicable to high-stakes, expensive systems where continuous optimization is required.
The Improvement: Implement a layered training regimen that moves from broad coverage to specialized, reasoning-based optimization, enabling the sophisticated learning of complex trade-offs.
Specific Mechanism: The system utilizes three distinct phases:
-
Diversity SFT: Broad exposure to diverse inputs and outputs (basic syntax/mapping).
-
Cold-Start SFT: Training on curated, Pareto-proximal data paired with explicit Chain-of-Thought (CoT) reasoning traces, instilling structural awareness.
-
GRPO (Group Relative Policy Optimization): Applying a four-component reward (r format, r comp, r correctness, r QoR) within a group of candidates to drive fine-grained policy refinement.
What the Improved AI System Can Do:
-
Learn Deep, Structured Reasoning: By training with reasoning traces (CoT), the the system learns not just what to output, but why, enabling it to propose complex, non-obvious solutions (e.g., algorithmic restructuring) that are inaccessible to simple prompt-based approaches.
-
Achieve Global Optimization: It moves beyond local optima by using GRPO within a group, allowing the system to evaluate and select the best overall design among multiple candidates based on a comprehensive set of criteria (including syntactical compliance and hardware efficiency).
The Improvement: Formally define and optimize for holistic quality metrics beyond single points, focusing on the Pareto front.
Specific Mechanism: The system optimizes for Pareto dominance across multiple objective dimensions (e.g., latency, resource utilization). It also incorporates a soft reward structure (r QoR) that rewards designs with relative superiority even when they are incomparable (e.g., having the same latency but different resource trade-offs).
What the Improved AI System Can Do:
- Identify Optimal Trade-offs: The system can produce solutions that are not just fast, but holistically superior, achieving a balanced trade-off between competing hardware constraints, surpassing models that are either purely speed-oriented or purely resource-conservative.
Abstract
High-Level Synthesis (HLS) compiles algorithmic C/C++ descriptions into hardware, with Quality of Results (QoR)---latency and resource utilization---critically governed by pragma configurations and code structure. Existing natural-language-to-HLS (NL-to-HLS) training approaches prioritize functional correctness while largely ignoring QoR. We observe that reinforcement learning (RL) for HLS does not require absolute synthesis results---only relative comparisons between candidates. Based on this insight, we propose HLS-Seek, a QoR-aware NL-to-HLS framework that avoids full synthesis-in-the-loop RL via a comparative proxy reward model achieving 99.53% Pareto-dominance accuracy. To prevent reward hacking, we introduce uncertainty-aware Monte Carlo (MC) dropout switching that selectively invokes real Vitis HLS synthesis for low-confidence candidates and online updates the proxy, creating a self-improving reward system. HLS-Seek achieves 84.7% syntax correctness pass@1 and 81.4% functional correctness pass@5 on HLS-Eval with only 7B parameters, surpassing GPT-5.1 on functional pass@5, while achieving 8.5 times faster training than real-reward RL. On QoR evaluation, HLS-Seek achieves the lowest latency on 19/30 kernels and Pareto-dominates HLS-specific baselines on 9 kernels.
Sources
- Evaluating Large Language Models Trained on Code
- ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement Learning
- Qwen2.5-Coder Technical Report
- ForgeHLS: A Large-Scale, Open-Source Dataset for High-Level Synthesis
- Evaluating Large Language Models for Automatic Register Transfer Logic Generation via High-Level Synthesis
- SynthAI: A Multi Agent Generative AI Framework for Automated Modular HLS Design Generation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- HLSPilot: LLM-based High-Level Synthesis
- HLStrans: Dataset for C-to-HLS Hardware Code Synthesis
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks