Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
cs.AI
Submitted: 2026-08-18
Updated: 2026-09-25
Code: https://github.com/KCL-Planning/VAL
Terminology
Sources
- Teaching Large Language Models to Self-Debug
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- QLoRA: Efficient Finetuning of Quantized LLMs
- Reinforced Self-Training (ReST) for Language Modeling
- Distilling the Knowledge in a Neural Network
- LoRA: Low-Rank Adaptation of Large Language Models
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- LEVER: Learning to Verify Language-to-Code Generation with Execution
- Correct Is Not Governed: Provenance Integrity in Agentic Workflows
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
- On the Planning Abilities of Large Language Models : A Critical Investigation
- STaR: Bootstrapping Reasoning With Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection