SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs
cs.CL, cs.PL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted by EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generating formally verified programs from natural language remains challenging: existing approaches either produce code in a single pass without recourse when verification fails, or rely on
Terminology
Abstract
Generating formally verified programs from natural language remains challenging: existing approaches either produce code in a single pass without recourse when verification fails, or rely on open-ended agentic reasoning that is non-deterministic and opaque. We introduce SKILLFORGE, a framework that decomposes formal code synthesis into a library of atomic, reusable skills, each targeting a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or targeted repair, and defined by a prompt template, tool binding, and decidable success criterion. A verification-driven harness orchestrates these skills: it submits candidates to the Dafny verifier, diagnoses failures into structured categories, deterministically routes to the appropriate repair skill, and iterates until formal correctness is proved or a budget is exhausted. On a curated benchmark of natural language to Dafny specification pairs, SKILLFORGE substantially outperforms both state-of-the-art agentic approaches (including ReAct-style agents, MCTS-based repair, and RL-guided verification) and traditional iterative baselines, while requiring fewer tokens and lower latency. Ablation studies confirm that every skill contributes measurably, and the harness converges rapidly with the majority of programs verified on the first attempt.
Sources
- MINIF2F-DAFNY: LLM-Guided Mathematical Theorem Proving via Auto-Active Verification
- A benchmark for vericoding: formally verified program synthesis
- Veri-Sure: A Contract-Aware Multi-Agent Framework with Temporal Tracing and Formal Verification for Correct RTL Code Generation
- DafnyBench: A Benchmark for Formal Software Verification
- Adaptive Proof Refinement with LLM-Guided Strategy Selection
- Large Language Models Based Automatic Synthesis of Software Specifications
- Towards AI-Assisted Synthesis of Verified Dafny Methods
- Proof2Silicon: Prompt Repair for Verified Code and Hardware Generation via Reinforcement Learning
- Cobblestone: A Divide-and-Conquer Approach for Automating Formal Verification
- Dafny as Verification-Aware Intermediate Language for Code Generation
- Laurel: Unblocking Automated Verification with Large Language Models
- Tail Optimality and Performance Analysis of the Nudge*(M) Scheduling Algorithm
- Large Language Models for Code Generation: The Practitioners Perspective
- VerMCTS: Synthesizing Multi-Step Programs using a Verifier, a Large Language Model, and Tree Search
- Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
- Chain-Aware Encoding for Microservice Trace Anomaly Detection
- Memento-Skills: Let Agents Design Agents
- VERINA: Benchmarking Verifiable Code Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering