SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization
Fangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu
Department of Computer Science and Engineering, The Chinese University of Hong Kong
cs.AR, cs.AI
Submitted: 2026-08-14
Updated: 2026-08-17
Comments: 12 pages, 8 figures
Code: https://github.com/sallyliu921/CB-EVO
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: SynAct is an adaptive closed-loop LLM reasoning–acting agent that iteratively optimizes PPA on a commercial synthesis tool, with particular emphasis on WNS.
Terminology
Summary
SynAct is an adaptive closed-loop LLM reasoning–acting agent that iteratively optimizes PPA on a commercial synthesis tool, with particular emphasis on WNS. At each iteration, SynAct conditions its decision on the latest timing reports and diagnostic outputs, mitigating information overload through a multilayer GraphRAG module that retrieves scenario-relevant command documentation on demand. To improve exploration efficiency, Bayesian optimization over a GrammarVAE latent space leverages historical optimization experience accumulated in earlier iterations. Built on ReAct for coupled reasoning and action and AutoGen-style multi-agent collaboration for role coordination, SynAct continuously interprets circuit feedback and produces each optimization command with an explicit rationale.
The main contributions are: (1) A closed-loop LLM agent that iteratively optimizes PPA by adapting decisions to the current circuit state. (2) A multi-layer GraphRAG module for scenario-relevant knowledge retrieval, mitigating information overload. (3) A BO-guided experience refinement mechanism over a GrammarVAE latent space for experience-guided exploration. (4) SynAct reduces average WNS to 27% of that from bootstrap synthesis while maintaining balanced area and power trade-offs.
The framework consists of three main components. The Analysis Agent examines current reports and produces a diagnosis through two stages: Preprobe, where it parses raw report logs and proactively constructs follow-up probes using analyze* commands, and Postprobe, where probe outputs are incorporated to refine the diagnosis. The Optimization Agent takes the structured diagnosis and user request as input and generates executable synthesis command candidates, supported by RAG-grounded candidates derived from scenario-relevant manual sections retrieved by GraphRAG, and BO-seeded candidates that adapt historically effective command patterns. Candidate Selection performs safety filtering, removing candidates whose WNS falls beyond a preset threshold or whose reward drops below 0.2 times the initial bootstrap reward, then selects the one with the highest score balancing reward, exploration, and redundancy avoidance.
The GraphRAG module organizes tool documentation into a three-layer graph of optimization scenarios, executable commands, and configurable variables. Document preprocessing categorizes content into applicable scenarios, executable commands, and configurable variables. Entity and relationship extraction uses an LLM to convert unstructured text into structured entities and identify intra-layer and inter-layer relationships. Scenario-driven retrieval uses the Analysis Agent's diagnosis as a query, generates a scenario description, embeds it with a domain-customized text embedding model fine-tuned on EDA tool documentation, selects top-k relevant scenarios, and expands to connected commands and variables.
The BO-guided experience refinement uses GrammarVAE to encode synthesis commands into a continuous latent space, where nearby points correspond to grammatically similar commands. A trust region centered at the current best command's latent point focuses optimization. An experience log stores latent encodings and rewards of executed commands. A reward-weighted RBF surrogate is refit from scratch on the full experience log each iteration. UCB acquisition selects the next latent point, which is decoded into a command recommendation for the Optimization Agent. All candidates are executed from the same parent checkpoint, their rewards recorded, and the trust region is recentered and resized based on performance.
Experiments on a commercial synthesis tool across 14 open-source RTL designs show that SynAct reduces average WNS to 27.03% of that from bootstrap synthesis, substantially outperforming ChatLS (71.73%) and CBTune (66.67%). SynAct achieves nonnegative WNS on two designs, whereas neither baseline does. For secondary metrics, SynAct reduces TNS to 17.86% of the bootstrap synthesis result, compared with 96.43% for ChatLS and 77.14% for CBTune. Area and power remain close to bootstrap levels, with 99.28% area, 98.92% dynamic power, and 98.64% static power.
Runtime analysis shows candidate evaluation dominates at 84.85% of SynAct's runtime, with bootstrap synthesis at 10.4%, Analysis Agent at 2.56%, Optimization Agent at 1.67%, and other overhead at 0.52%. SynAct's total runtime averages 988.62% of bootstrap time, less than half of CBTune's 2174.18% but higher than ChatLS's 289.83%. Token usage averages about 139k tokens per design, with 69% from the Optimization Agent and 31% from the Analysis Agent.
Generalizability across LLMs was tested by rerunning with GPT-5.2, which reduced WNS and TNS Ratio Avg. to 20.37% and 13.69% respectively, showing that a stronger LLM can exploit SynAct's mechanism more effectively. Validation of latent-space BO guidance showed that near command pairs have a mean reward difference of 0.188 versus 0.238 for far pairs, a 21.0% reduction, and BO-seeded candidates achieve a mean reward of 0.297 versus 0.289 for RAG-grounded candidates, obtaining the highest mean reward in 65.3% of benchmark iterations. Ablation studies showed that removing BO raises the WNS Ratio Avg. from 27.0% to 37.5%, vanilla RAG raises it to 28.6%, and removing RAG entirely raises it to 38.3%.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Adaptive closed-loop decision-making with state-conditioned actions: I can implement a system that continuously re-evaluates its next action based on the latest environment feedback (e.g., timing reports, diagnostic outputs) rather than following a fixed plan. This improves responsiveness to dynamic, real-world conditions.
-
Multi-layer GraphRAG for on-demand knowledge retrieval: I can build a hierarchical knowledge graph (scenarios → commands → variables) that retrieves only scenario-relevant documentation, reducing information overload. This enables the AI to access precise, contextual knowledge without processing irrelevant data, improving efficiency and accuracy in complex tool environments.
-
Bayesian optimization over a GrammarVAE latent space for experience-guided exploration: I can encode past successful actions into a continuous, grammar-aware latent space and use a trust-region Bayesian optimizer (with reward-weighted RBF surrogate and UCB acquisition) to propose new actions. This leverages historical experience to explore efficiently, avoiding random or redundant attempts.
-
Two-stage diagnostic probing (Preprobe/Postprobe): I can implement an agent that first parses raw logs, then proactively generates follow-up queries (e.g., analyze* commands) to gather missing information, and finally refines its diagnosis. This improves the quality of situational understanding before acting.
-
Safety-filtered candidate selection with multi-objective scoring: I can add a filtering layer that rejects actions with unacceptable risk (e.g., performance below a threshold) and then scores remaining candidates on a balance of reward, exploration novelty, and redundancy avoidance. This ensures safe, diverse, and effective action selection.
-
Domain-customized embedding fine-tuning for retrieval: I can fine-tune a text embedding model on domain-specific documentation (e.g., EDA tool manuals) to improve semantic retrieval accuracy, making the AI more effective in specialized fields.
-
Multi-agent role coordination (Analysis Agent + Optimization Agent): I can split complex tasks into specialized agents—one for diagnosis, one for action generation—with structured handoffs. This improves modularity, parallelization, and clarity of reasoning.
-
Experience log with latent encodings for continual learning: I can store all executed actions, their latent representations, and rewards, then refit the surrogate model from scratch each iteration. This enables the AI to adapt its exploration strategy over time without catastrophic forgetting.
What the improved AI system can do:
-
Optimize complex, multi-objective engineering workflows (e.g., chip design PPA) by iteratively refining actions based on real-time feedback.
-
Reduce performance degradation (e.g., worst-case slack) to 27% of baseline while maintaining balanced trade-offs across other metrics (area, power).
-
Achieve nonnegative outcomes (e.g., timing closure) on some designs where prior methods fail.
-
Operate with 3.4x lower runtime overhead than a comparable BO-based method and use 139k tokens per task, making it practical for real deployment.
-
Generalize across different underlying LLMs—stronger models yield even better results (e.g., 20% WNS ratio), showing the framework is model-agnostic and scalable.
-
Avoid redundant or unsafe actions via safety filtering, improving reliability in production settings.
-
Learn from past experience to propose better actions over time, with a 21% reduction in reward variance between similar commands, confirming meaningful latent-space structure.
Sources
- ChipNeMo: Domain-Adapted LLMs for Chip Design
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4