Verification as an Architectural Layer for LLM Agents: A V-Model Design, and a Pilot Study of Its Deterministic Core
cs.SE, cs.AI
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- ReAct: Synergizing Reasoning and Acting in Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- ReWOO: Decoupling Reasoning from Observations for Efficient Augmented Language Models
- Reason-Plan-ReAct: A Reasoner-Planner Supervising a ReAct Executor for Complex Enterprise Tasks
- ADaPT: As-Needed Decomposition and Planning with Language Models
- Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Self-Refine: Iterative Refinement with Self-Feedback
- Constitutional AI: Harmlessness from AI Feedback
- Large Language Models Cannot Self-Correct Reasoning Yet
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Let's Verify Step by Step
- Teaching Large Language Models to Self-Debug
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
- QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
- Think Locally, Explain Globally: Graph-Guided LLM Investigations via Local Reasoning and Belief Propagation
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties