Argumentation for Explainable and Globally Contestable Decision Support with LLMs
cs.AI, cs.CL
Submitted: 2026-03-15
Updated: 2026-05-03
Comments: KR 2026, KR Meets Machine Learning and Explanation Track
Journal ref: Proceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning - KR meets Machine Learning and Explanation, 917-928. 2026
DOI: 10.24963/kr.2026/86
Code: https://github.com/adamdejl/argeval
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Large language models (LLMs) exhibit strong general capabilities, but their deployment in high-stakes domains is hindered by their opacity and unpredictability.
Terminology
Abstract
Large language models (LLMs) exhibit strong general capabilities, but their deployment in high-stakes domains is hindered by their opacity and unpredictability. Recent work has taken meaningful steps towards addressing these issues by augmenting LLMs with post-hoc reasoning based on computational argumentation, providing faithful explanations and enabling users to contest incorrect decisions. However, this paradigm is limited to pre-defined binary choices and only supports local contestation for specific instances, leaving the underlying decision logic unchanged and prone to repeated mistakes. In this paper, we introduce ArgEval, a framework that shifts from instance-specific reasoning to structured evaluation of general decision options. Rather than mining arguments solely for individual cases, ArgEval systematically maps task-specific decision spaces, builds corresponding option ontologies, and constructs general argumentation frameworks (AFs) for each option. These frameworks can then be instantiated to provide explainable recommendations for specific cases while still supporting global contestability through modification of the shared AFs. We investigate the effectiveness of ArgEval on treatment recommendation for glioblastoma, an aggressive brain tumour, and show that it can produce explainable guidance aligned with clinical practice.
Sources
- Docling Technical Report
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Advanced Layout Analysis Models for Docling
- Reasoning Models Don't Always Say What They Think
- Retrieval- and Argumentation-Enhanced Multi-Agent LLMs for Judgmental Forecasting (Extended Version with Supplementary Material)
- Contestability in Quantitative Argumentation
- gpt-oss-120b & gpt-oss-20b Model Card
- Wait, Wait, Wait... Why Do Reasoning Models Loop?
- ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar Argumentation
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection