ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization".
Jane: The paper was written by Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin et al. from Northwestern University and Wenzhou Buyi Pharmacy Chain Co., Ltd. and Wenzhou University and City University of Hong Kong and National University of Singapore.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, folks. Today we're digging into a paper that's got a mouthful of a title: "ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization." Jane, what caught your eye first?
Jane: Tom, honestly, it was that word "behavioral verification." We've all seen LLMs write code that looks fine, runs fine, and then quietly gives you a wrong answer. This paper is about catching exactly that kind of silent failure.
Lu: And it's a big deal, Jane. The authors found a gap of ninety percentage points between code that runs and code that's actually correct on some problems. Ninety points. That's not a rounding error, that's a chasm.
Meng: Right, and as someone who has to ship software, that's terrifying. A solver says "optimal," but the model forgot a constraint. You'd never know unless you had ground truth to compare against, and in the real world, you rarely do.
Tom: So the title is promising a way to catch these silent failures without needing a perfect answer key. That's the hook for me.
Jane: Exactly. And the way they do it is clever. They don't just ask the LLM to double-check its own work, because we know that doesn't work. Instead, they poke the model with the solver itself.
Lu: They perturb parameters. Shrink a capacity limit to almost nothing, and if the objective doesn't change, the constraint probably isn't in the code. It's using the solver as an external brain, not the LLM's own flawed introspection.
Meng: That's the part I like. It's a test you can actually run. You don't need a human expert to stare at the math. You just change a number, re-solve, and look at the objective.
Tom: And that's the core promise of "ReLoop" — a loop that keeps testing and repairing until the behavior matches what the problem description demands.
Jane: The authors also built a new benchmark, RetailOpt-one hundred ninety with one hundred ninety retail inventory problems designed to trip up LLMs with interacting constraints. So they're not just claiming this works; they're giving the community a way to measure it.
Lu: And the results show real gains. On Claude Opus four point six, they went from about twenty-two point six percent accuracy to thirty-one point one percent on the strict metric. That's a big jump for a hard problem.
Meng: It's still not perfect, but the point is they're closing the gap between "runs" and "correct." That's the direction we need.
Tom: So we've got a method, a benchmark, and a measurable improvement. Next, we should talk about what's actually inside the pipeline. That's where the real meat is.
Summary: Jane: So we're back with "ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization." Tom, we teased the big idea, but let's get into what the pipeline actually does.
Tom: Right. The paper splits it into two big pieces. First, structured generation. Instead of asking the LLM to just write code, they force it through four stages: understand, formalize, synthesize, verify.
Lu: And that formalize stage is where they make the LLM explicitly state variable types. Is this a continuous variable or an integer? Can you order two point seven pallets? That kind of reasoning catches a whole class of errors before they ever reach the code.
Meng: I like that, but the second piece is the real innovation. They call it behavioral verification. It runs in two layers. The first layer just checks execution — syntax, solver status, that kind of thing. The second layer is the perturbation testing we mentioned.
Jane: And that's where they test whether the formulation responds correctly to changes. If you shrink a cost coefficient to almost zero and the objective doesn't budge, that cost term is probably missing from the objective function.
Lu: It's a beautiful trick because it sidesteps the whole problem of LLM self-critique. The LLM extracts what constraints should exist, but the solver does the actual detection. You're not asking the model to judge its own work; you're asking the math to prove it.
Tom: And if a test flags a problem, the system doesn't just give up. It generates a targeted repair, but it's careful. There's a regression guard — if the repaired code changes the objective by more than four percent, they roll it back.
Meng: That's the engineer's touch. You don't want the fix to be worse than the disease. They also have a safety check that blocks the repair LLM from fabricating data values, which is a real risk when you're asking a model to fix code.
Jane: The results across benchmarks show that the two mechanisms are complementary. Structured generation helps most on complex, compositional problems. Behavioral verification shines on simpler problems with localized defects, like a missing constraint.
Lu: On MAMO-ComplexLP, behavioral verification was the biggest single contributor for Claude, adding four point four percentage points. On RetailOpt-one hundred ninety structured generation drove the largest gains.
Tom: So it's not one magic bullet. It's a toolkit where each piece handles a different kind of failure. That's a mature way to think about the problem.
Meng: And the fact that they got execution rates to one hundred percent on Claude Opus four point six tells me the recovery mechanisms are working. The code always runs, even if it's not always right.
Jane: Which brings us to the question of what this means for the future. We've got the method, we've got the benchmark. What's the impact?
Improvements: Tom: We're back with "ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization." Jane, we've covered the what. Let's talk about the so what.
Jane: The biggest improvement this paper suggests is a shift in how we think about verification for LLM-generated code. We're moving from "does it run?" to "does it behave correctly?"
Lu: And that's a philosophical shift. The paper shows that execution success and semantic correctness are almost completely decoupled. DeepSeek hit ninety-one point one percent execution but only zero point five percent accuracy on RetailOpt-one hundred ninety. That's a stunning demonstration that "it runs" means almost nothing.
Meng: For me, the practical improvement is the diagnostic recovery. When the solver says infeasible, ReLoop computes the Irreducible Inconsistent Subsystem — the minimal set of conflicting constraints — and feeds that back to the LLM. That's not a generic "try again" message. That's a precise diagnosis.
Tom: So instead of the LLM guessing why it failed, it gets told exactly which constraints are fighting each other. That's like a mechanic telling you which bolt is loose instead of just saying "car won't start."
Jane: And the repair loop is smart about what it fixes. It only acts on high-confidence warnings. Ambiguous signals are logged but ignored, because a false-positive repair can introduce a regression.
Lu: There's also a documented limitation that's really honest. The chain-of-thought prompting actually breaks fine-tuned models. OptMATH, which was trained to output code directly, collapsed from fifty-six point two percent to thirty point zero percent accuracy when forced into the four-stage reasoning template.
Meng: That's a real-world deployment warning. If you've fine-tuned a model on a specific format, you can't just bolt on a generic reasoning prompt. The paper documents eighty-four crashes and sixty-five regressions on that model. That's a cautionary tale.
Tom: So the improvement isn't just "make it better." It's also "know when your approach will make things worse."
Jane: And that's valuable. The paper gives us a map of where these techniques work and where they don't. That's how you build reliable systems, by knowing the boundaries.
Lu: The authors also release RetailOpt-one hundred ninety as a benchmark, which is a gift to the community. We now have a stress test that specifically targets compositional constraint interactions, which is where LLMs fail most.
Meng: And the benchmark is designed with progressive difficulty. Eight scenario families, from core operations to omni-channel logistics. It's not just one hard problem; it's a spectrum.
Tom: So we've got a method, a benchmark, and a clear-eyed view of limitations. What does this mean for the world beyond the lab?
Conclusion: Tom: We're wrapping up our discussion of "ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization." Jane, give us the send-off.
Jane: This paper tackles a problem that's been lurking under the surface of LLM-based optimization. Code that runs and returns a feasible solution can still be completely wrong, and until now, we had no good way to catch that without a perfect answer key.
Lu: ReLoop changes that by using the solver itself as a truth-teller. Perturb a parameter, watch the objective, and you get a signal that doesn't depend on the LLM's own blind spots. That's a genuinely new idea.
Meng: And the engineering is careful. Regression guards, safety checks, diagnostic feedback — this is a system built by people who've actually deployed software, not just written papers about it.
Tom: The results speak for themselves. Execution rates hit one hundred percent on Claude Opus four point six, accuracy improved across benchmarks, and the authors were honest about where the approach breaks down, like with fine-tuned models.
Jane: And they gave us RetailOpt-one hundred ninety a benchmark that will let the whole field measure progress on the hardest part of this problem — compositional constraint reasoning.
Lu: The ninety-point gap between feasibility and correctness should be a wake-up call for anyone deploying LLM-generated optimization code. This paper is a step toward closing that gap.
Meng: It's not a silver bullet. Structural errors still slip through. But it's a meaningful step, and it gives us a framework to build on.
Tom: So we say goodbye to "ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization." A solid contribution to making LLMs trustworthy where it matters — when the answer has to be right, not just plausible.
Jane: Thanks for listening, folks. We'll be back with the next paper soon. Until then, keep questioning the answers.
Tom: And keep verifying the behavior, not just the execution. See you next time.
Northwestern University · Wenzhou Buyi Pharmacy Chain Co., Ltd. · Wenzhou University · City University of Hong Kong · National University of Singapore
cs.SE, cs.AI, cs.LG, math.OC
Submitted: 2026-02-17
Updated: 2026-09-26
Comments: Code and benchmark: https://github.com/junbolian/ReLoop
Code: https://github.com/junbolian/ReLoop
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 66/100
The gist: when the LLM translates a natural-language problem into a mathematical formulation—and persist because existing post-hoc checks cannot detect them." To address this, ReLoop employs two
Key concepts
- Behavioral Verification
- This is the process of testing whether an LLM's generated code behaves correctly according to problem requirements, moving beyond just checking if the code runs. It involves testing how the formulation responds to changes, such as shrinking a cost coefficient, to detect missing constraints.
- Structured Generation
- This is a four-stage process for LLMs: understand, formalize (where they explicitly state variable types), synthesize, and verify. The formalization stage helps catch errors early by forcing the model to reason about details like whether a variable is continuous or an integer.
- Perturbation Testing
- This technique involves testing the model by changing input parameters, such as shrinking a capacity limit to almost nothing. If the objective function does not change after this test, it suggests a constraint might be missing from the code.
Terminology
Summary
Summary
The paper introduces ReLoop, a framework designed to address silent failures
in LLM-generated optimization code. The authors define a silent failure as code that (i) executes without error, (ii) returns a solver-feasible solution, yet (iii) is not semantically correct.
They identify a feasibility–correctness gap reaching 90 percentage points on compositional problems,
where state-of-the-art models achieve up to 91.1% solver-feasibility but only 0.5% formulation correctness.
The paper's key insight is that Silent failures originate at the modeling stage—when the LLM translates a natural-language problem into a mathematical formulation—and persist because existing post-hoc checks cannot detect them.
To address this, ReLoop employs two complementary mechanisms: Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify), preventing formulation errors at their source,
and "Behavioral verification detects errors that survive generation by testing whether the formulation responds correctly to solver-based parameter perturbation—an external semantic signal that bypasses LLM self-review and requires no ground truth."
The structured generation process is formalized as a chain: x −−−−−−→ U −−−−−→ M −−−−−−→ Ĉ −−−−→ C.
Stage 1 (Understand) extracts the objective direction, decision variables, constraints, and parameters from x.
Stage 2 (Formalize) transforms this into a mathematical specification M = (I, P, V, C, f): index sets I, parameters P, decision variables V with explicit type reasoning, constraints C, and objective f.
Stage 3 (Synthesize) generates executable Gurobi code Ĉ from M, structured to access all data through data[
key] dictionary patterns rather than hardcoded literals.
Stage 4 (Verify Completeness) cross-checks the generated code Ĉ against the original problem x.
Behavioral verification operates through two layers. L1 (Execution Verification) is the only blocking layer: a FATAL result triggers regeneration with error feedback (up to N attempts).
For infeasible models, L1 computes the Irreducible Inconsistent Subsystem (IIS)—the minimal set of conflicting constraints—and feeds specific constraint names and bounds to the LLM.
L2 (Behavioral Testing) includes two sub-modules: "Constraint Presence Testing (CPT) targets missing constraints: for each candidate extracted by the LLM (annotated with physical type: capacity, demand, etc.), CPT applies extreme perturbation to the governing parameter (capacity ×0.001, demand ×100, other ×0.01). Objective Presence Testing (OPT) targets missing cost/revenue terms with analogous perturbation (cost ×0.001, revenue ×100, other ×0.01). Both measure
the objective change ratio r = z ′ − z ∗ / z ∗ and classify results:
r < τl = 5%: WARNING (likely missing)—triggers repair. τl ≤ r ≤ τh = 30%: I NFO (uncertain)—logged, no repair. r > τh or infeasibility: PASS (present)—confirmed active."
The paper introduces RetailOpt-190, a new benchmark of 190 compositional retail optimization scenarios
across eight scenario families (F1–F8) following a progressive composition principle.
The families include "F1: Core Operations, F2: Assortment & Substitution, F3: Resource Constraints, F4: Demand Dynamics, F5: Feasibility Stress, F6: Discrete Logistics, F7: Network & Multi-Echelon, F8: Omni-channel. Each of
38 archetypes is instantiated with 5 numerical variants (±15%, deterministic seeds), yielding 190 instances."
Main results on RetailOpt-190 show that for Claude Opus 4.6, Base
accuracy is 22.6%, CoT
reaches 31.1%, and ReLoop
achieves 31.1% at ϵ = 10−4, with execution improving from 72.1% to 100.0%. For DeepSeek-V3.2, CoT collapses execution (91.1%→53.2%) because the four-stage decomposition produces intermediate mathematical notation that fails to translate into valid Gurobi syntax—L1 fully recovers this regression.
DeepSeek's accuracy improves from 0.5% (Base) to 5.8% (ReLoop). The paper notes that "Claude’s accuracy plateaus at 31.1% because the remaining two-thirds are predominantly structural silent failures: fundamentally different (yet internally consistent) problem decompositions that perturbation cannot detect."
Cross-benchmark results on MAMO-ComplexLP show Claude improving from 70.4% (Base) to 79.8% (+ReLoop), DeepSeek from 60.1% to 62.6%, and Qwen3-32B from 40.4% to 46.3%. On IndustryOR, Claude improves from 66.0% to 68.0%, DeepSeek from 50.0% to 62.0%, and Qwen3 from 43.0% to 46.0%. The paper identifies a limitation with SFT models: "OptMATH on MAMO: Base accuracy (56.2%) collapses under CoT (30.0%) because the SFT model’s rigid 'problem→code' generation pattern cannot accommodate our four-stage reasoning template—84 instances crash and 65 previously correct solutions are destroyed."
Ablation studies reveal the complementarity of the two mechanisms: "structured generation drives the largest gains on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), while behavioral verification dominates on localized defects (+4.4pp on MAMO-ComplexLP, its largest contribution across benchmarks). On MAMO,
L2 is the largest single accuracy contributor for Claude (+4.4pp; 11 corrected, 2 regressed) and adds +2.0pp for DeepSeek (4 corrected, 0 regressed). On RetailOpt-190,
L2 contributes execution recovery (Claude: 99.5%→100.0%) and practical accuracy gains (DeepSeek: 10.5%→11.1% at ϵ = 10−2), but strict accuracy (ϵ = 10−4) is unchanged because errors are predominantly structural."
The paper concludes that "ReLoop preserves or improves every foundation-model metric combination through L2’s non-blocking, regression-guarded design, and the 90-point feasibility–correctness gap underscores that semantic verification is becoming essential infrastructure for LLM-based optimization. Limitations include that
Structured generation assumes format compatibility: CoT disrupts SFT models’ learned patterns, and
Three failure modes remain beyond scope: coefficient magnitude errors, formulation equivalence errors, and unrepresented problem structures."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement and what the improved AI system can do:
Implementation: Replace single-pass code generation with a four-stage chain-of-thought process:
-
Stage 1 (Understand): Extract objective direction, decision variables, constraints, and parameters from natural language.
-
Stage 2 (Formalize): Produce a mathematical specification with explicit variable-type reasoning (continuous vs. integer vs. binary based on physical context).
-
Stage 3 (Synthesize): Generate executable optimization code referencing a pre-extracted data dictionary.
-
Stage 4 (Verify): Self-check completeness against the original problem before output.
Resulting capability: The AI system can now decompose complex optimization problems (e.g., multi-period inventory with perishability, substitution, and capacity constraints) into structured mathematical models, reducing silent formulation errors by up to 8.5 percentage points on compositional problems.
Implementation: Add a post-generation verification layer that tests semantic correctness without ground truth:
-
Constraint Presence Testing (CPT): Perturb constraint parameters (e.g., capacity ×0.001, demand ×100) and measure objective change. If change < 5%, flag the constraint as likely missing.
-
Objective Presence Testing (OPT): Perturb cost coefficients (×0.001) and revenue coefficients (×100). If objective change < 5%, flag the term as omitted.
-
Classification: Objective change 30% or infeasibility → PASS.
Implementation: Integrate a repair loop with three safety mechanisms:
-
Safety check: Block data-variable reassignment and dangerous imports (e.g.,
os,subprocess). -
Regression guard: Rollback any repair that crashes, degrades solver status, or shifts objective by >4%.
-
Skip guard: Bypass repair when no high-confidence WARNING exists.
Implementation: For infeasible models, compute the Irreducible Inconsistent Subsystem (IIS) and feed specific conflicting constraint names to the LLM for regeneration. For unbounded models, report unbounded ray variables.
Implementation: Before code generation, extract all numerical parameters into a structured JSON dictionary. Generate code that references this dictionary via data["key"] patterns rather than hardcoded literals. Fall back to self-contained generation if extraction fails.
-
Generate reliable optimization code for complex, multi-constraint problems (e.g., 20+ interacting constraints across multi-period, multi-product, multi-location dimensions) with up to 100% executable code and 31.1% strict formulation accuracy on Claude Opus 4.6—a 22.6% improvement over direct generation.
-
Detect and repair silent failures—code that runs and returns solver-feasible solutions but is semantically wrong—using external solver-based signals rather than LLM self-review, which inherits the generating model's blind spots.
-
Recover from execution failures with specific diagnostic feedback (IIS constraints, unbounded rays, syntax tracebacks) rather than generic retries, improving execution rates by up to 44.2 percentage points on mid-tier models.
-
Maintain monotonic improvement across all foundation models and benchmarks through non-blocking verification (L2 never discards valid solutions) and regression-guarded repair (rollback if objective shifts >4%).
-
Generalize across problem domains—retail inventory, generic LP/MILP, and real-world industrial optimization—without domain-specific tuning, as demonstrated on MAMO-ComplexLP and IndustryOR benchmarks.
-
Handle both data-embedded and schema-based prompts, enabling deployment in production settings where data is pre-loaded separately from code generation.
-
Provide confidence signals via the severity matrix: FATAL (blocks output), WARNING (triggers repair), INFO (reference only), PASS (verified), enabling users to gauge reliability of generated solutions.
-
Achieve 100% executable code on frontier models (Claude Opus 4.6) and consistent accuracy improvements across all foundation models, with the largest gains on compositional problems (+8.5pp) and localized defects (+4.4pp).
Abstract
Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations -- a feasibility-correctness gap reaching 90 percentage points on compositional problems. We introduce ReLoop, which addresses this gap through two complementary mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify), preventing formulation errors at their source. Behavioral verification detects errors that survive generation by testing whether the formulation responds correctly to solver-based parameter perturbation -- an external semantic signal that bypasses LLM self-review and requires no ground truth. The two mechanisms are complementary by error structure: structured generation drives the largest gains on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), while behavioral verification dominates on localized defects (+4.4pp on MAMO-ComplexLP, its largest contribution across benchmarks). Combined with diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6 and consistently improves accuracy on chat-tuned foundation models across three benchmarks; we further identify a known limitation of narrowly-tuned SFT models, whose learned output formats are brittle to chain-of-thought prompts -- an interaction we document and analyze. We release RetailOpt-190, 190 compositional retail optimization scenarios targeting the multi-constraint interactions where LLMs most frequently fail.
Sources
- Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow
- LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
- DeepSeek-V3 Technical Report
- Qwen3 Technical Report
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties