SOVER: Formal Certification of Optimization Reformulations via LLM-Assisted SMT Verification
cs.AI, cs.LG, math.OC
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Accepted to EMNLP 2026 Findings
Code: https://github.com/baranwa2/SOVER
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex mathematical optimization problems across modeling languages.
Terminology
Abstract
Large Language Models (LLMs) have shown remarkable promise in translating and reformulating complex mathematical optimization problems across modeling languages. However, validating such transformations through empirical solver executions alone is unreliable, as solver outcomes may be affected by local minima, structural timeouts, numerical artifacts, and subtle semantic divergence between formulations. We introduce SOVER, an LLM-assisted SMT framework that separates semantic mapping from formal certification: Z3 checks domain cross-feasibility and global objective-order preservation for mixed-integer linear formulations, while dReal provides tolerance-aware feasibility/range and ε-argmin checks for continuous nonlinear formulations. We also introduce NLEquiv-150, a public benchmark of 100 equivalent and 50 deliberately hard non-equivalent nonlinear reformulation pairs. With LLM-extracted mappings, SOVER classifies 149/150 pairs (99.33%) correctly, including all 50 hard negatives; the sole error is an incomplete mapping extraction.
Sources
- OptiMUS: Scalable Optimization Modeling with (MI)LP Solvers and Large Language Models
- Evaluating Large Language Models Trained on Code
- The Weisfeiler-Lehman Method and Graph Isomorphism Testing
- The Llama 3 Herd of Models
- A survey on combinatorial optimization
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- Teaching Networks to Solve Optimization Problems
- Large Language Model Safety: A Holistic Survey
- Gemini: A Family of Highly Capable Multimodal Models
- Large Language Models in Operations Research: Methods, Applications, and Challenges
- A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions
- Fact-checking AI-generated news reports: Can LLMs catch their own lies?
- EvoCut: Strengthening Integer Programs via Evolution-Guided Language Models
- EquivaMap: Leveraging LLMs for Automatic Equivalence Checking of Optimization Formulations
- A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving
- StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
- Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection