Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Grading the Graders".
Jane: Large language models (LLMs) are increasingly paired with "verifiers" to detect errors, but these verification schemes suffer from conflating five different concepts—granularity, concept abstraction, risk tier, system-stack layer,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, we’re looking at this paper, "Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning." Basically, they're trying to sort out all these different ways we check if an AI is making a mistake.
Jane: It seems like the problem is that people use the word "level" in verification literature to mean five totally different things: how detailed the check is, how abstract the concept is, what kind of risk it’s dealing with, which system layer you're looking at, and even where you got your starting information.
Lu: They propose this new thing called Verification Autonomy Levels or VAL. It’s a way to classify any verification scheme along one axis: where does the specification come from and what exactly does the verdict guarantee?
Meng: So, instead of all these separate concepts, they want one simple way to look at it. What’s the big deal if we can just put everything on this single axis?
Tom: It matters because they argue that the most important thing for trust in LLM reasoning is actually where the ground truth comes from. That's what they focus on, and it helps sort through all that confusion.
Jane: They call it focusing on the epistemic status of the ground truth, which they argue is key for people to trust what an LLM reasons out.
Lu: The core idea is this six-level taxonomy, L0 through L5. It classifies any verification scheme by its anchor source and what that pass actually commits to guaranteeing.
Meng: But how do we know where it’s going in those levels? Is it just a guess?
Tom: No, they have a deterministic decision procedure and a runnable classifier that lets you put any specific scheme into one of these six categories. It's structured, not random.
Jane: The levels are defined by two things: the anchor source—who or what supplies the ground truth for the verdict—and the guarantee that a PASS actually commits to making.
Lu: And there are three things that happen when you move between these levels. First, it’s always "the anchor, not the judge, determines the level," and second, "the guarantee degrades downward but not upward" when you move from one level to another.
Tom: That sounds like a very clear rule for how we should think about these checks. So what are those levels actually based on?
Paper summary: Jane: They use three decisive questions to sort things out. First, they ask who declares the verification condition—what the verdict is actually about.
Lu: The answer to that first question can be one of four things: either it’s the LLM under test saying "I checked it," or it’s a strict rule from code or problem text, or there's an objective source like a gold answer or execution result, or it's just a property in a decidable system.
Meng: That covers the different kinds of starting points for any verification process. So after figuring out that, what’s next?
Tom: Question two is about the guarantee. What exactly does a PASS promise to deliver?
Jane: It’s either correctness—meaning every single proposed candidate satisfies the condition—or completeness—meaning no candidate was missed at all.
Lu: The paper points out that most deployed verifiers offer correctness while they are read as completeness, which is a big distinction. That's the whole substance of section four there.
Meng: So if we look at question three, what does that change?
Tom: Question three is about scope. If you’re talking about completeness, it asks over what domain that guarantee holds: maybe just a single property, or a whole class of programs' memory safety, or claimed universality.
Jane: And the decision procedure maps those answers to the levels. For instance, if you have completeness and universal scope, that lands you at level five—but they show that level five is impossible in the unrestricted case.
Lu: They show that if you have completeness with a domain scope, it’s level four, and single-property scope puts you at level three.
Meng: So what about the correctness side? What happens when the guarantee is correctness instead of completeness?
Tom: If you have correctness and an anchor based on a known-correct reference, like a gold label, that’s level two. But they flag that as "usage-degraded," meaning it's not as good as other options.
Jane: If you use a silver anchor, which is just deterministic re-computation with no gold reference, it’s at the L1 level or on the boundary of L2.
Lu: And if you use a problem-derived rule to check correctness, that puts you in level one. Otherwise—if it’s an LLM declaring something without any anchor—that's level zero.
Meng: I see how it works, but what about the completeness blind spot they mention? That seems like a big theoretical hurdle for this whole setup.
Paper summary: Tom: They claim that substitution and sampling-based verification at level two can prove that proposed candidates hold, but they cannot prove that no candidate was missed. That’s a property of the paradigm itself, not just some tuning failure.
Jane: They argue completeness is only possible for properties with a decidable fragment—things like mathematical or syntactic rules—or it's not possible at all in the open world.
Lu: And they state that empirical open-world verification, like fact-checking or diagnosis, caps out at anchored correctness, which is level two. It reaches rulescoped completeness only over a formalized sub-fragment if you want to get higher than that.
Meng: That means for real world applications outside of pure math or code checking, we’re capped at level two. That’s a practical constraint I need to keep in mind when building things.
Tom: Exactly. The paper shows how different approaches actually perform in symbolic mathematics, where the L3 solution set verifier catches missed solutions that the L2 substitution verifier misses by construction.
Jane: In medical diagnosis, they showed an L3 upgrade involved swapping a hand-written heuristic for a validated clinical decision rule as the judge's anchor.
Lu: The paper says that an L2 judge adds reportability, not accuracy, and an L3 judge adds completeness within its own ODD—its Operational Design Domain—but nothing adds completeness outside of it.
Meng: So we have this whole structure now: VAL classifies the scheme based on where the spec comes from and what it guarantees. It’s a checklist before you trust any verification claim.
Tom: It really provides that pre-purchase checklist, asking where the specification comes from before trusting any claim about an LLM’s reasoning.
Jane: The main message is that we need to know exactly what we're getting when we ask for verification, and that L2 is a correctness probe, not a guarantee of completeness.
Lu: And they show how this framework separates the different "level" axes—like granularity and risk tier—from the anchor-axis, which they argue is actually the one that matters most for trust.
Meng: So for practical impact, it means we need to be honest about what we are expecting from an AI verification tool in a specific domain.
Tom: That’s what this paper on Grading the Graders is all about: giving us a way to grade these verifiers based on where they come from and what they actually guarantee.
Conclusion: Tom: So we're wrapping up this deep dive on Grading the Graders, which is all about putting these new Verification Autonomy Levels, or VALs, on LLM checking systems.
Jane: Right, so basically they've created this six-level system to sort out all the confusing ways we used to measure how good an AI’s verification process actually is.
Lu: They're classifying verification schemes by where the ground truth comes from and what a successful check actually promises to guarantee.
Meng: It seems like they’ve managed to untangle five different concepts—granularity, abstraction, risk, system layers—by just focusing on that anchor source and the guarantee itself.
Tom: And they’re arguing that focusing on where that ground truth originates is actually the most important thing for building trust in AI reasoning.
Jane: That means we can stop looking at all those separate checks and just look at this single axis to understand how reliable a verification scheme is.
Lu: They have a deterministic way to place any specific check into one of these six levels, which is really interesting because it makes the whole thing predictable, not just theoretical.
Meng: Predictability matters when you're trying to build something that actually runs reliably in the real world instead of just being a clever idea on paper.
Tom: And they show that if you look at how different verification methods perform across things like math or medicine, it helps show exactly where each level is useful.
Jane: They point out that an L2 verdict, for example, is really just a correctness probe, not a guarantee that nothing was missed in the entire system.
Lu: That’s the big caveat they bring up: you can prove something holds but you can't always prove you haven't missed something else.
Meng: So for someone who only listens to this show, what does this mean? It means before you trust an AI check, you need to know exactly what level of guarantee it’s actually making.
Tom: Exactly. We just saw how they map out all the possibilities from zero up to five, and that gives us a really solid way to assess these tools.
Yajie Yin
cs.CL
Submitted: 2026-08-19
Updated: 2026-10-03
Comments: v3: code and the full literature assessment released on Zenodo (DOI: 10.5281/zenodo.23120985); v2 added a reproducibility study (blind-LLM inter-rater kappa~0.8; human raters pending), an external-framework transfer test, and 15+ fixes. Writing was assisted by an AI language model; all experiments and research decisions are the author's own
Code: https://github.com/1549080929-debug/math_agent
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Large language models (LLMs) are increasingly paired with "verifiers" to detect errors, but these verification schemes suffer from conflating five different concepts—granularity, concept
Key concepts
- Verification Autonomy Levels (VAL)
- A six-level classification system (L0-L5) that categorizes verification methods based on where their ground truth comes from and what they promise to guarantee. It replaces confusing metrics by focusing on the anchor source and the guarantee made by the verification process.
- Anchor Source
- The entity or mechanism that supplies the ground truth upon which a verification verdict is based. This includes things like an LLM declaring it checked something, a deterministic rule from code, or an objective execution result. The paper argues this source determines the level of trust in the verification.
- Guarantee
- What a successful verification (a 'PASS') commits to achieving. The two main options are 'correctness'—ensuring every proposed answer is right—or 'completeness'—ensuring no possible answer was missed. The paper shows that guarantees degrade when moving between levels.
- Completeness Blind Spot
- The theoretical finding that substitution-based verification (L2) can prove candidates are correct but cannot prove no candidate was missed. This means completeness is only achievable by encoding the property into a decidable system (L3/L4), and empirical domains like fact-checking are capped at anchored correctness (L2).
Terminology
Summary
Large language models (LLMs) are increasingly paired with verifiers
to detect errors, but these verification schemes suffer from conflating five different concepts—granularity, concept abstraction, risk tier, system-stack layer, and epistemic source—which this paper addresses by proposing Verification Autonomy Levels (VAL), a meta-standard that classifies verification schemes along a single axis: where the verification spec comes from and what it guarantees. This framework aims to resolve the systematic conflation across 17 surveyed papers by focusing on the epistemic status of the ground truth, which is argued to be the most critical factor for trust in LLM reasoning<ref:2608.19009#pg6>.
How it works
The core contribution is Verification Autonomy Levels (VAL), a six-level epistemic taxonomy (L0–L5) classifying any verification scheme by its anchor source and guarantee, along with a deterministic decision procedure and a runnable classifier<ref:2608.19009#pg4>. The levels are defined by two properties: the anchor source (who or what supplies the ground truth the verdict rests on) and the guarantee (what a PASS commits to)<ref:2608.19009#pg8>. Three consequences follow this definition: first, the anchor, not the judge, determines the level,
and second, the guarantee degrades downward but not upward
when moving between levels<ref:2608.19009#pg8>.
The classification process is deterministic and follows three decisive questions in order<ref:2608.19009#pg10>:
-
Q1 (spec source): Who declares the verification condition—the thing the verdict is about? The answer is one of: (a) the LLM under test itself (
I checked it,
this is verified
); (b) a deterministic rule derived from the problem or code text (regex, parser, schema); (c) an objective source independent of both the problem and the generative mechanism under test (gold answer, measured value, execution result); or (d) a property encoded in a decidable system<ref:2608.19009#pg10>. -
Q2 (guarantee): What does a PASS commit to? Correctness—
every proposed candidate satisfies the condition
—or completeness—no candidate was missed.
The distinction is the entire substance of Sec. 4; most deployed verifiers offer the former while being read as the latter<ref:2608.19009#pg10>. -
Q3 (scope): If the guarantee is completeness, over what domain does it hold: a single property (this equation's solution set), a whole class (all programs' memory safety), or claimed universality?<ref:2608.19009#pg10>.
The decision procedure maps the answers to the level:
-
completeness + universal scope → L5 (rejected: impossible, unrestricted)<ref:2608.19009#pg4>.
-
completeness + domain scope → L4<ref:2608.19009#pg4>.
-
completeness + single-property scope → L3<ref:2608.19009#pg4>.
-
correctness + gold-label anchor (known-correct reference) → L2 (decidable anchor used only for correctness → flag
usage-degraded, upgradeable
)<ref:2608.19009#pg4>. -
correctness + silver anchor (deterministic re-computation, no gold reference) → L1/L2 boundary, classified L1 (objective about the computation, silent about the answer; VerifiAgent's tool)<ref:2608.19009#pg4>.
-
correctness + problem-derived rule → L1<ref:2608.19009#pg4>.
-
otherwise (LLM-declared / no anchor) → L0<ref:2608.19009#pg4>.
The Completeness Blind Spot
The central theoretical claim is that Substitution- and sampling-based verification (L2) can prove that proposed candidates hold; it cannot prove that no candidate was missed
<ref:2608.19009#pg5>. This blind spot is not a tuning failure but a property of the verification paradigm<ref:2608.19009#pg5>. Completeness is achievable only by re-encoding the property into a decidable system (L3/L4)—or not at all<ref:2608.19009#pg5>. This means the completeness blind spot is not confined to mathematics
and that empirical open-world propositions (fact-checking, diagnosis, legal analysis) have no such fragment and therefore cap at anchored correctness (L2)
<ref:2608.19009#pg5>.
Orthogonality of Axes
The paper disambiguates five confounded level
axes by showing that granularity, concept abstraction, risk, and system-stack are orthogonal to the VAL axis
<ref:2608.19009#pg4>. These axes answer what to check, how finely, at which layer, and what to do with the result<ref:2608.19009#pg4>. The anchor-axis is argued to be the one that matters for trust<ref:2608.19009#pg4>.
Empirical Findings Across Domains
The framework was exercised across four domains, reporting honest negative results where they occurred
<ref:2608.19009#pg6>. In symbolic mathematics, the L3 solution set verifier catches missed solutions that the L2 substitution verifier cannot see, by construction
<ref:2608.19009#pg6>. In medical diagnosis, an L3 upgrade involved replacing a hand-written heuristic with a validated clinical decision rule (Alvarado score) as the judge's anchor<ref:2608.19009#pg6>. The paper concludes that an L2 judge adds reportability, not accuracy; an L3 judge adds completeness within its ODD; nothing adds completeness outside it
<ref:2608.19009#pg6>.
Operational Design Domain (ODD)
An ODD must be a pre-declared syntactic predicate over inputs, not a post-hoc description of success
<ref:2608.19009#pg11>. Completeness is relative to an ODD; a verifier that is complete inside its ODD is, by construction, silent about anything outside it
<ref:2608.19009#pg11>. The engineering test of ODD honesty is a boundary probe where adversarial instances just outside the claimed ODD must break the completeness claim
<ref:2608.19009#pg14>.
Conclusion
The framework provides a pre-purchase checklist: ask where the spec comes from before trusting any verification claim, and know that an L2 verdict is a correctness probe, not a completeness guarantee
<ref:2608.19009#pg9>. The paper asserts that the highest grade any grader can earn is a guarantee it can actually deliver: correctness within its domain, completeness within its ODD, and honesty when it must abstain
<ref:2608.19009#pg19>.
REFERENCES
-
Han, J., Buntine, W., Shareghi, E. VerifiAgent: A Unified Verification Agent in Language Model Reasoning. EMNLP 2025. arXiv:2504.00406<ref:2608.19009#pg20>
-
Fang, J., Zhang, B., Wang, C., et al. Graph of Verification (GoV): Structured Verification of LLM Reasoning with Directed Acyclic Graphs. arXiv:2506.12509<ref:2608.19009#pg20>
-
Li, Y., Lin, Z., Zhang, S., et al. Making Large Language Models Better Reasoners with StepAware Verifier (DiVERSE). ACL 2023. arXiv:2206.02336<ref:2608.19009#pg20>
-
Liu, C., Yuan, Y., Yin, Y., et al. Safe: Enhancing Mathematical Reasoning in LLMs via Retrospective Step-aware Formal Verification. ACL 2025. arXiv:2506.04592<ref:2608.19009#pg20>
-
Miao, N., Teh, Y.W., Rainforth, T. SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning. ICLR 2024. arXiv:2308.00436<ref:2608.19009#pg20>
-
Juneja, G., Dutta, S., Chakraborty, T.
Improvements for AI systems
-
System architecture should adopt a canonical division of labor:
The LLM translates (decomposes, renders, declares), deterministic code judges within its ODD, and a human anchors the spec and audits the judge.
This ensures that roles are not crossed, preventing failures likewhen the LLM declares its own verification spec (L0).
-
Implement a three-part audit strategy to terminate trust recursion:
The chain terminates at the kernel of a decidable system (L4), at a definitional anchor (L3), or at a conventional primary standard (metrology, where calibration chains terminate at definitionally fixed units).
This prevents systems from endlessly questioning the verifier. -
Adopt an ODD-aware verification strategy:
An L3/L4 guarantee holds only within an ODD—here, the decidable domain in which the property can be encoded and the verdict computed.
Systems must explicitly state theirsyntactic predicate over inputs
(e.g.,solveset over 'polynomial equations of degree ≤ 4 with real coefficients'
). -
For open-world tasks, deploy an L0–L2 classification:
For empirical domains [fact-checking, diagnosis], L0–L2 are therefore not merely 'the lower three levels' but a complete classification in their own right.
This shifts the goal from unattainable universal completeness toanchored correctness,
which is more practically achievable. -
Integrate a decision procedure into the deployment pipeline:
The classification is deterministic: given the answers to Q1–Q3, the level is the first rule that fires.
This provides a clear, reproducible mapping for assigning VAL levels based on whether an anchor is LLM-declared (L0), problem-derived (L1), or objective truth (L2).
Abstract
Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors. Yet the verification literature uses the word "level" to mean at least five different things: verification granularity, concept abstraction, risk tier, system-stack layer, and the epistemic source of the ground truth. We propose Verification Autonomy Levels (VAL), a meta-standard that classifies any verification scheme along a single axis: where does the verification spec come from, and what does the verdict guarantee? VAL ranges from L0 (LLM self-declaration; no deterministic anchor) through L2 (objective ground truth; correctness only) to L3/L4 (decidable systems with single-property or domain-level completeness), with L5 impossible in the unrestricted case. Central to VAL is the completeness blind spot: substitution- and sampling-based verifiers can confirm that proposed candidates hold, but cannot prove that no candidate was missed. We further identify a dichotomy the literature has not stated: completeness is reachable only for formally specifiable properties, whereas empirical open-world verification (fact-checking, diagnosis) caps at anchored correctness (L2). We document this gap empirically across four domains (symbolic mathematics, behavior monitoring, medical diagnosis, code generation) and in the strongest formal-verification baseline in our survey. We show that granularity, concept hierarchy, risk, and system stack are orthogonal to VAL, resolving a conflation across 17 surveyed papers. Code and the full literature assessment are released on Zenodo (DOI: 10.5281/zenodo.23120985).
Sources
- VerifiAgent: a Unified Verification Agent in Language Model Reasoning
- Graph of Verification: Structured Verification of LLM Reasoning with Directed Acyclic Graphs
- Making Large Language Models Better Reasoners with Step-Aware Verifier
- Safe: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning
- $\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning
- Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
- Hierarchical Attention Generates Better Proofs
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
- LLM Output Drift: Cross-Provider Validation & Mitigation for Financial Workflows
- Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
- Prompting Frameworks for Large Language Models: A Survey
- Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
- Beyond Words: A Mathematical Framework for Interpreting Large Language Models
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering