Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
cs.AI
Submitted: 2026-08-16
Updated: 2026-09-25
Comments: Code and data are available at https://github.com/junbolian/AdmitOR
Code: https://github.com/junbolian/AdmitOR
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills.
Terminology
Abstract
Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library can be retrieved again and again, and on a stream of new problems there is no ground-truth answer to decide with. Existing learners admit trajectories by matching known optima or labels, and label-free substitutes such as execution success or agreement at one instance can admit wrong models. We introduce ADMITOR, a label-free admission gate. It generates models from three model families, runs each on the stated problem and on instances with resampled parameters, keeps the largest group of models whose optimal values agree on every instance across families, and applies a threshold fitted on solver-verified problems to accept, abstain, or escalate, with a finite-sample bound on the false-discovery rate among accepted values. Inside a state-of-the-art skill learner, ADMITOR raises candidate-level admission precision to 0.927, against 0.871 for majority vote over the host's own samples and 0.726 for execution success, and its library, the smallest of the four, reaches the highest macro accuracy over five public benchmarks, 58.4 against 54.8 for majority vote. An ablation on the same records shows that the gain comes from the accepted value being external to the learner and unanimous across families; on this stream, resampling never changed an accepted value and only reduced coverage. The false-discovery bound holds on the calibration set but not on the benchmark stream: an audit of every false certificate traces most of them to benchmark texts that omit or round the numbers needed to reproduce the labeled answer, and a label-free check of the extracted numbers against the text flags most of these cases.
Sources
- Universal Self-Consistency for Large Language Model Generation
- LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- OptArgus: A Multi-Agent System to Detect Hallucinations in LLM-based Optimization Modeling
- ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
- Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow
- From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
- Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verification
- LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
- OptMATH: A Scalable Bidirectional Data Synthesis Framework for Optimization Modeling
- VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification
- ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation
- OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling
- SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling
- More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection