Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency
cs.CL
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: 48 pages, 3 figures
Code: https://github.com/idavidrein/gpqa
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an
Terminology
Abstract
When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an interaction between a prover with unrestricted internal search and a resource-bounded human verifier. The verifier requests and checks supporting details without access to the LLM's internal state. Passed checks accumulate evidence toward an acceptance threshold. We prove anytime-valid soundness against adaptive provers: the probability of ever accepting a false claim is at most a chosen error level, provided the task supplies bounds on false passes and human checking errors that remain valid after every relevant history. A finite-horizon completeness bound additionally requires bounds on the adequacy of honest responses and sufficient diagnostic progress. Further checks can strengthen the evidence for acceptance, but each requires another adequate response and reliable human effort. Whether this tradeoff permits certification depends on the verifier's effort budget, cognitive load, expertise, and fatigue. We identify conditions under which the supplied bounds certify a specified sequence of local checks but not a specified global check under the same resource budgets.
Sources
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- Measuring Progress on Scalable Oversight for Large Language Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs
- When Should an AI Workflow Release? Always-Valid Inference for Black-Box Generate-Verify Systems
- Supervising strong learners by amplifying weak experts
- Training Verifiers to Solve Math Word Problems
- The Impossibility of Eliciting Latent Knowledge
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- AI safety via debate
- Why is plausibility surprisingly problematic as an XAI criterion?
- Prover-Verifier Games improve legibility of LLM outputs
- Debate Helps Supervise Unreliable Experts
- E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing
- Self-critiquing models for assisting human evaluators
- Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering