NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation
Jiarui Ma, Jianghan Wang, Yuheng Ma, Ziyi Zhuang, Xiaoguang Liu
Southern University of Science and Technology
eess.SY, cs.AI, cs.SY
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: accepted by MLCAD 2026
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 75/100
The gist: NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation Abstract Summary: Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their
Terminology
Summary
NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation
Abstract Summary:
Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning. Although netlists are textual, they encode structured circuit objects through topology and parameters. The paper presents NetlistBench, a structure-verified benchmark for SPICE netlist recognition and manipulation. NetlistBench contains 2,342 cases across 24 task families, covering parameter and connectivity recognition and edits, hierarchical operations, equivalence judgment, and long-horizon compound editing. Model outputs are evaluated by a deterministic structure-aware oracle. Across six non-thinking LLMs, performance varies substantially with operation-level structural complexity. Simple local edits reach 96%–100% accuracy, while device addition drops to 41%–83% and equivalence judgment to 49%–90%. Enabling reasoning substantially improves weaker models but does not eliminate structure-preservation failures, with performance still degrading sharply as the edit horizon increases. NetlistBench identifies netlist reliability as a distinct bottleneck for trustworthy LLM-based circuit design automation.
Introduction Summary:
Large language models are increasingly explored across the lifecycle of integrated circuit design, including domain-adapted chip-design assistance, analog circuit generation, simulation-driven optimization, and multimodal netlist extraction. Many of these workflows share SPICE netlists as a recurring representation layer. In generation-oriented settings, LLMs may synthesize netlists from specifications, schematics, or circuit images. In simulation-driven optimization loops, they often revise existing netlists according to performance or simulator feedback. In netlist-to-schematic or netlist-to-layout workflows, they may interpret connectivity, hierarchy, and device relationships. Errors at this representation layer can directly corrupt downstream simulation, optimization, or layout reasoning. Failures in LLM-based circuit workflows may originate either from high-level design reasoning or from low-level netlist corruption, yet existing evaluations rarely separate these two sources. Reliable netlist operation is therefore a prerequisite for trustworthy LLM-based circuit design workflows. General LLM-for-code benchmarks evaluate executable functional correctness through unit tests or repository test suites but do not capture the device-specific terminal semantics, shared-node connectivity, and ordered subcircuit interfaces of SPICE netlists. Existing evaluations of LLM-based circuit design typically focus on end-to-end outcomes, such as syntactic validity, simulation success, specification improvement, or downstream task completion. Recent circuit-oriented benchmarks mainly assess domain-level capabilities, including circuit interpretation, topology reasoning, schematic understanding, AMS-domain multimodal reasoning, or graph-structured reasoning. While these evaluations reveal important limitations, they do not isolate the elementary operations required to interpret and modify SPICE netlists correctly.
The main contributions of this work are:
-
Formulating SPICE netlist reliability as a representation-level evaluation problem, focusing on whether LLMs can correctly recognize and manipulate netlists as structured circuit artifacts.
-
Introducing NetlistBench, a structure-verified benchmark covering netlist structural-property recognition and natural-language-guided netlist manipulation, evaluating representative frontier, flash-class, and open-weight LLMs, showing that reliability varies sharply across operation type and task horizon.
-
Developing a structure-aware evaluation pipeline based on canonical circuit representations, enabling manipulation outputs to be verified beyond raw text matching or final simulation outcomes.
Background Summary:
SPICE netlists are simulator-facing circuit descriptions that encode devices, terminals, nodes, parameters, models, ports, and subcircuit hierarchies in a compact, positional textual format. Although a netlist appears as a sequence of text lines, its underlying semantics correspond to a structured circuit object. Each device statement typically begins with an instance prefix, followed by an ordered sequence of node names connected to specific device terminals, a model reference, and optional parameter assignments. For hierarchical circuits, subcircuit definitions (.subckt) establish ordered port interfaces, and each subcircuit instance binds external nodes to internal ports strictly according to their positional order. A defining characteristic is that electrical connectivity is encoded implicitly through node-name sharing rather than explicit terminal-to-terminal links. For structural analysis and manipulation, connections are not represented as explicit terminal-to-terminal links; the circuit topology must be reconstructed from terminal–node bindings across the entire netlist. These properties make netlist operations different from ordinary text editing. A correct edit must preserve terminal-role bindings, maintain consistent node identities, and avoid unintended changes to unrelated devices or subcircuit interfaces.
Existing studies have reported limitations of LLMs in processing circuit representations. At the circuit-reasoning level, benchmarks such as CIRCUIT and AMSbench show that LLMs can struggle with topology-heavy circuit interpretation and multi-step circuit reasoning. At the generation and adaptation level, systems such as SPICEPilot, SPICEAssistant, AnalogCoder, and Spice Wizard rely on simulation feedback, syntax checks, or tool-assisted repair loops. At the netlist-analysis level, SPICED studies LLM-aided detection and localization of syntactical bugs and analog Trojans. Complementary representation-oriented work, such as CircuitFormer, highlights the mismatch between standard language tokenization and the graph-structured semantics of circuits, while Image2Net evaluates diagram-to-netlist conversion using graph-structured netlist comparison. These studies evaluate netlists within broader generation, simulation, conversion, or reasoning pipelines rather than isolating representation-level netlist operations.
Benchmark Design Summary:
NetlistBench evaluates representation-level netlist reliability through two modalities: recognition, which extracts or compares circuit structure, and manipulation, which applies explicit natural-language edits to SPICE netlists. This design separates structure interpretation and structure-preserving transformation from high-level design reasoning, simulator behavior, and optimization.
The source corpus uses two complementary SPICE netlist sources. AnalogGenie provides flat CMOS analog netlists originally developed for topology discovery, with 2,838 flat netlists (492 Simple, 1,752 Medium, 594 Complex) totaling approximately 58,000 device instances. ALIGN provides 931 hierarchical analog netlists from a layout automation flow, containing 21 topology families with 3–4 subcircuits per netlist. The total corpus is 3,769 netlists.
The instance construction pipeline is deterministic and template-driven. Each instance is represented as a self-contained triplet: I i = (N src(i), T inst(i), Y target(i)), where N src denotes the source SPICE netlist, T inst denotes the explicit task instruction, and Y target denotes the task-specific target used for evaluation. Both T inst and Y target are produced by deterministic, family-specific templates. For manipulation tasks, Y target is the uniquely determined target SPICE netlist. For recognition tasks, it is the canonical JSON answer. For equivalence judgment tasks, it is the binary structural-equivalence verdict. All instances undergo automated construction-time validation.
The structure-aware evaluation oracle parses each model output and the reference target into a canonical IR—a normalized structure that lists every device by instance name with its device kind, ordered terminal nodes, and parameters, together with top-level directives and, for hierarchical circuits, each subcircuit’s port interface and internal devices. An output passes only if its IR matches the reference IR under semantics-preserving normalizations: same set of named devices, identical terminal-node bindings, parameters equal up to numeric normalization, identical top-level directives, symmetric two-terminal passives compared with unordered terminals, and subcircuit definitions agreeing on port interface, internal devices, and directives. This exact-match-up-to-normalization rule directly encodes name-preservation and locality constraints. For equivalence-judgment tasks, the ground-truth labels are validated using VF2 graph-isomorphism checks on labeled bipartite device–net graphs, but scoring preserves device and node names rather than equating circuits up to renaming.
NetlistBench contains 24 task families. The recognition modality contains 800 cases across eight families: seven evaluate structured extraction of device parameters, terminal connectivity, node incidence, subcircuit interfaces, and instance mappings, while the eighth evaluates structural equivalence between netlist pairs. The manipulation modality contains 1,542 cases across 16 families: six single-edit families cover connectivity editing, device addition, removal and replacement, parameter editing, and rename propagation; five flat compound families combine 3, 6, 9, 12, or 15 dependent operations; and five hierarchical families evaluate subcircuit expansion, interface modification, and multi-step internal editing.
Evaluation Summary:
The paper evaluates six single-shot non-thinking models spanning frontier, flash-class, and open-weight tiers: Claude Sonnet 4.6, GPT-4.1, Gemini 2.5 Flash, DeepSeek-V4-Flash, Qwen3.6-Flash, and Qwen3-30B-A3B. All are queried through official provider APIs with explicit reasoning modes disabled. As a reasoning reference, DeepSeek-V4-Flash is additionally evaluated with its native thinking mode enabled. Two controlled secondary analyses are run on a paired stratified subset: SPICE versus PySpice output representation, and direct prompting versus native thinking and CoT prompting. Each model is queried once per case with deterministic decoding, and retries are used only for transport failures. Responses are graded by the structure-aware oracle and reduced to binary pass/fail outcomes. The parser tolerates incidental code fences, but empty, unparsable, or structurally invalid outputs fail. Pass rates are reported with Wilson 95% confidence intervals.
Performance across task families shows substantial variation. Local operations that primarily modify explicit text are the most reliable: device removal and parameter editing reach 96%–100% across models. Reliability decreases for operations that require maintaining connectivity or introducing new structure, including connectivity editing, device replacement, device addition, subcircuit port swapping, and inline expansion. Recognition exhibits a similar distinction. Explicit attributes such as device parameters, ordered terminal lists, and subcircuit ports are extracted with high accuracy, whereas relational queries vary substantially across models. Node incidence ranges from 15% to 98%, semantic terminal connectivity from 6% to 99%, and terminal-neighbor incidence from 2% to 97%. Structural equivalence judgment remains challenging, with pass rates from 49% to 90%. Overall non-thinking pass rates range from 41% to 82%. Enabling reasoning raises DeepSeek-V4-Flash from 55% to 81%, but does not consistently surpass the strongest non-thinking model.
Long-horizon compound editing reveals substantial reliability degradation. Accuracy declines consistently as the edit horizon increases. Claude drops from 80% at 3 steps to 26% at 15 steps, while GPT-4.1 drops from 71% to 34%. The remaining models decline to near-zero accuracy at longer horizons, with Gemini decreasing from 70% to 6%, DeepSeek-V4-Flash from 44% to 0%, and both Qwen3.6-Flash and Qwen3-30B from 34%/28% to 0%. Reasoning does not eliminate this trend: DeepSeek-V4-Flash with reasoning enabled falls from 74% at 3 steps to 31% at 15 steps. The degradation is not simply a consequence of weak atomic editing. Long-horizon compound tasks require models to track multiple dependent edit intents, update intermediate circuit state, and preserve edit locality across an extended instruction sequence. Because the edits are mutually dependent, errors compound across the sequence. The same downward trend appears in the hierarchical compound family.
The mitigation analyses on the paired subset (n=718) show that reasoning gives the larger aggregate gains: native thinking raises both models by roughly 30–40 points (p < 10-6), while CoT also improves performance but less strongly. PySpice has a smaller and less consistent effect: it improves DeepSeek-V4-Flash overall (p < 0.001), but not Qwen3.6-Flash (p = 0.51). Reasoning is the stronger mitigation, but errors still concentrate on long-horizon compound edits, hierarchical operations, and relational structural queries; neither mitigation makes current LLMs sufficiently reliable for unverified netlist editing.
Discussion Summary:
NetlistBench shows that netlist reliability cannot be reduced to general circuit knowledge or output-format compliance. Models tend to perform better on localized edits, such as parameter changes and device removal, while showing reduced reliability on tasks involving structural attachment, ordered port handling, equivalence judgment, or multi-step edits. The observed failures are frequently structural rather than purely procedural: recognition outputs usually follow the required JSON schema but contain incorrect circuit facts, while manipulation failures involve omitted edits, duplicated edits, loss of locality, unintended terminal rebinding, and topology drift. This error pattern follows directly from the representation properties: local substitutions and deletions often require only limited changes to already explicit text, while attachment, hierarchy, equivalence, and compound editing require the model to maintain an implicit circuit graph across terminal roles, node identities, and subcircuit interfaces. The observed failures indicate a structure-preservation bottleneck: models can often produce syntactically plausible netlists, but still lose edit locality, perturb unrelated bindings, or fail to maintain consistent topology across multiple dependent operations. Current LLMs should not be treated as standalone, unverified netlist editors. More robust workflows may need to decompose complex edits, verify each intermediate netlist structurally, and provide feedback when unintended changes are detected.
Limitations Summary:
NetlistBench evaluates bounded, circuit-block-level netlists rather than industrial-scale post-layout decks. Although the compound tasks increase the number of dependent edits, all evaluated netlists fit within the tested models’ context windows. The benchmark does not assess long-context retrieval, hierarchical partitioning, or direct processing of extracted netlists containing millions of device and parasitic statements. The current release covers a restricted circuit and syntax domain, primarily flat and hierarchical analog CMOS blocks from AnalogGenie and ALIGN. It does not comprehensively cover symbolic.param expressions, complex.model cards, behavioral or controlled sources, include hierarchies, extracted parasitics, or simulator- and PDK-specific syntax. Results should not be assumed to transfer unchanged to RF, power, digital, or mixed-signal netlists. Task instructions are generated from deterministic templates to isolate structural capabilities and enable unambiguous grading; they do not capture the full linguistic variability or design intent of real designer–assistant interactions. Each model–case pair is evaluated once, so the results characterize the tested API snapshots rather than complete output distributions.
Conclusion Summary:
NetlistBench shows that netlist reliability is a distinct bottleneck for LLM-based circuit design. Current models often handle local substitutions and simple extraction, but remain fragile on connectivity-sensitive edits, hierarchy manipulation, structural equivalence, and long-horizon compound transformations. Reasoning improves performance, yet does not make LLMs reliable unverified editors of simulator-facing netlists. These findings motivate decomposed editing workflows, structure-aware verification after each edit, and circuit representations that expose topology more directly than raw SPICE text.
Improvements for AI systems
Improvements to AI Systems:
-
Add a structure-aware verification layer – Implement a post-generation parser that converts LLM output into a canonical circuit IR (device instances, terminal-node bindings, parameters, subcircuit interfaces) and validates it against the intended edit before returning results. The system can reject or flag outputs that violate terminal-role bindings, introduce unintended node renames, or alter unrelated devices.
-
Decompose long-horizon edits into atomic steps – Instead of asking the model to apply 15 dependent edits in one pass, break the instruction sequence into individual operations. After each step, run the structure oracle to verify the intermediate netlist; if verification fails, feed the error back to the model for correction before proceeding. This prevents error compounding and maintains edit locality.
-
Add explicit graph-state tracking during generation – Augment the model’s context with a machine-readable representation of the current circuit graph (nodes, device-terminal bindings, subcircuit ports) alongside the raw SPICE text. This reduces reliance on implicit topology reconstruction and improves performance on connectivity-sensitive tasks like device addition, terminal rebinding, and hierarchical port swapping.
-
Train or fine-tune on structure-preserving edit examples – Create a dataset of paired (source netlist, edit instruction, target netlist) instances from NetlistBench’s 1,542 manipulation cases. Fine-tune the model to minimize structural drift by penalizing outputs whose canonical IR differs from the target IR, rather than using token-level loss alone.
-
Implement a hierarchical edit planner – For compound tasks, have the model first identify the affected subcircuits and devices, then generate a plan (e.g., “modify port order in subckt X, then update instance Y’s node bindings”). Execute each plan step with independent verification, and roll back to the last verified state if a step fails.
-
Add a relational query module for recognition tasks – For tasks like node incidence, terminal connectivity, and terminal-neighbor queries, use a deterministic graph traversal on the parsed netlist IR instead of relying on the LLM’s memory. The LLM can handle extraction of explicit attributes (parameters, port lists), while relational queries are delegated to a symbolic solver.
-
Enable reasoning with structured feedback loops – When reasoning mode is enabled, prompt the model to output intermediate reasoning that explicitly lists the current node set, device-terminal bindings, and intended changes before producing the final netlist. Compare this reasoning against the oracle’s expected state to catch topology drift early.
-
Add a confidence-based fallback mechanism – For tasks with high structural complexity (e.g., equivalence judgment, long-horizon edits), have the model output a confidence score. If confidence is below a threshold, automatically route the task to a verification tool or request a second model pass with the oracle’s feedback.
What the Improved AI System Can Do:
-
Reliably edit SPICE netlists with near-perfect accuracy on local operations (parameter changes, device removal) and significantly improved accuracy on connectivity-sensitive operations (device addition, terminal rebinding, subcircuit port changes) by using the structure oracle for verification and correction.
-
Handle long-horizon compound edits (15+ dependent operations) without catastrophic accuracy drops, maintaining correctness above 70% instead of falling to near-zero, by decomposing tasks and verifying each intermediate state.
-
Answer relational circuit queries (node incidence, terminal connectivity, neighbor identification) with near-deterministic accuracy by combining LLM extraction with symbolic graph traversal, reducing error rates from up to 98% down to single digits.
-
Perform structural equivalence judgment with high reliability (above 90%) by using the oracle’s canonical IR comparison rather than relying on the LLM’s pattern matching alone.
-
Operate as a trustworthy unverified editor in design workflows, flagging its own low-confidence outputs and requesting verification, thus preventing silent corruption of downstream simulation, optimization, or layout tasks.
-
Maintain edit locality – ensuring that changes to one device or subcircuit do not unintentionally alter unrelated bindings or interfaces, even in hierarchical netlists with multiple nested subcircuits.
Abstract
Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist recognition and manipulation remains poorly understood and is rarely separated from high-level design reasoning. Although netlists are textual, they encode structured circuit objects through topology and parameters. We present NetlistBench, a structure-verified benchmark for SPICE netlist recognition and manipulation. NetlistBench contains 2,342 cases across 24 task families, covering parameter and connectivity recognition and edits, hierarchical operations, equivalence judgment, and long-horizon compound editing. Model outputs are evaluated by a deterministic structure-aware oracle. Across six non-thinking LLMs, performance varies substantially with operation-level structural complexity. Simple local edits reach 96% -- 100% accuracy, while device addition drops to 41% -- 83% and equivalence judgment to 49% -- 90%. Enabling reasoning substantially improves weaker models but does not eliminate structure-preservation failures, with performance still degrading sharply as the edit horizon increases. NetlistBench identifies netlist reliability as a distinct bottleneck for trustworthy LLM-based circuit design automation.
Sources
- Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI
- SPICED: Syntactical Bug and Trojan Pattern Identification in A/MS Circuits using LLM-Enhanced Detection
- Evaluating Large Language Models Trained on Code
- CircuitFormer: A Circuit Language Model for Analog Topology Design from Natural Language Prompt
- ChipNeMo: Domain-Adapted LLMs for Chip Design
- Schemato -- An LLM for Netlist-to-Schematic Conversion
- Evaluating LLM-based Workflows for Switched-Mode Power Supply Design
- AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
- SPICEPilot: Navigating SPICE Code Generation and Simulation with AI Guidance
- TopoSizing: An LLM-aided Framework of Topology-based Understanding and Sizing for AMS Circuits
- Image2Net: Datasets, Benchmark and Hybrid Framework to Convert Analog Circuit Diagrams into Netlists
Related papers
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- A Geometric Decision Procedure for STL Feasibility and Repair
- Submodular Multi-Agent Policy Learning for Online Distributed Task Allocation in Open Multi-Agent Systems
- Policy-Level Recursive Self-Improvement for Embodied AI with a Criticality World Model
- Minimal Experiments for Robust Stabilization: Information, Spectral Geometry, and Duration
- Decentralized Power-Optimal Coordination for Spacecraft Swarms Using Time-Varying Magnetorquer Actuation