Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority
cs.SE, cs.AI, cs.SY, eess.SY
Submitted: 2026-09-24
Updated: 2026-09-24
Terminology
Sources
- Small Language Models are the Future of Agentic AI
- Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs
- When Correct Isn't Usable: Improving Structured Output Reliability in Small Language Models
- ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
- Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings
- Language Models (Mostly) Know What They Know
- Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra
- LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
- Agents4PLC: Automating Closed-loop PLC Code Generation and Verification in Industrial Control Systems using LLM-based Agents
- A Deterministic Control Plane for LLM Coding Agents
- Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations
- Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction
- Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
- Qwen3 Technical Report
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties