Towards an Argumentative Foundation for Evaluative AI

arXiv:2608.07473 · cs.AI, cs.MA · Submitted 2026-04-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Towards an Argumentative Foundation for Evaluative AI".

Jane: The paper was written by Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago and Francesca Toni from Imperial College London and The University of Queensland and Cardiff University and King's College London.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the arXiv radio hour, everyone. I'm Tom, and today we're looking at a paper that's got a title that really makes you stop and think: "Towards an Argumentative Foundation for Evaluative AI."

Jane: And I'm Jane. Tom, I have to say, that title is a mouthful, but it's actually pointing at something really important. We've been talking a lot on this show about how AI just gives you answers, and this paper is basically saying, "What if AI didn't give you answers at all?"

Tom: Exactly. It's from a team spanning Imperial College London, the University of Queensland, Cardiff University, and King's College London. Authors include Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago, and Francesca Toni. These are serious names in the argumentation and explainable AI space.

Jane: And Tim Miller, he's the one who actually coined the term "Evaluative AI" in the first place, back in two thousand twenty-three. So this paper is very much a follow-up to his own vision, saying, "Okay, we said we want this, now here's how we actually build it."

Tom: Right. And the core idea in the title, "argumentative foundation," is the key. Instead of a neural network that spits out a single recommendation, you build a system that lays out arguments for and against different options, like a debate.

Jane: And that's such a shift. Normally, AI is like a judge delivering a verdict. This is more like AI being the courtroom itself, presenting both sides and letting the human be the jury.

Tom: And that's why the word "evaluative" matters. The AI isn't deciding; it's helping you evaluate. The paper is arguing that the best way to do that is with formal argumentation, which is a branch of logic that's been around for decades.

Jane: So it's not just a vibe or a metaphor. They're saying, "We have actual mathematical tools for representing arguments and attacks and supports, and we should use those."

Tom: Precisely. And the implications of that are huge. It means the AI's reasoning process isn't a black box. You can see the arguments, you can see the connections, and you can even challenge them.

Jane: Which is exactly what you want in high-stakes fields. Healthcare, law, finance. You don't want a system that just says, "Trust me." You want a system that says, "Here's why, and here's what speaks against it."

Tom: And that's the promise here. It's a position paper, so they're not claiming to have the final system built. They're laying out the roadmap and saying, "This is the direction we should go."

Jane: And I love that they're being so explicit about it. They're not just saying, "Argumentation is nice." They're saying, "Here's the formal problem, here's the framework, and here's why it's the right tool." It's a very bold, clear statement of intent.

Tom: Bold and clear. And the next segment, we're going to dig into what they actually mean by "Evaluative AI" and how this argumentation framework is supposed to work. Stick around.

Summary: Tom: So, Jane, we've got the title sorted. Now let's talk about what this paper actually says. "Towards an Argumentative Foundation for Evaluative AI" is essentially a position paper, but it's a position paper with a very concrete technical proposal.

Jane: Right. And the summary is basically this: they want to take this idea of Evaluative AI, where you present competing hypotheses with evidence for and against, and give it a proper formal backbone using something called weighted Quantitative Bipolar Argumentation Frameworks.

Tom: That's a mouthful. Let's break that down. "Bipolar" means you have two types of relations: support and attack. "Quantitative" means everything has a weight, a number between zero and one. So you're not just saying "this supports that," you're saying "this supports that with strength zero point eight."

Jane: And that's a huge step up from just having a list of pros and cons. You can actually compute a final strength for each hypothesis by combining all those weighted supports and attacks. It's like a mathematical tug-of-war.

Tom: Exactly. And the paper formalizes this as a "ranking-based EAI problem." The AI doesn't just pick one winner. It produces a ranked list of all the hypotheses, ordered by their computed strength.

Jane: And that ranking is not the final decision. It's the system's "preference," given the evidence. The human still makes the call. The AI is just saying, "Based on what I know, this one looks stronger, but here's the full picture."

Tom: And that's the key difference from traditional XAI. In the old model, the AI makes a decision and then tries to explain it. Here, the AI never makes a decision. It just lays out the evidence structure and lets the human evaluate.

Jane: And they're very clear that this isn't just a theoretical exercise. They want this to be computable. They want algorithms that can take a set of arguments and relations and actually produce that ranking.

Tom: Right. And they also talk about the properties that a good ranking should have. Things like monotonicity, which means if you add more supporting evidence for a hypothesis, its rank shouldn't go down.

Jane: That sounds obvious, but it's actually a really important guarantee. If the system is going to be trusted, it has to behave rationally. You can't have it saying, "We added more evidence for this, but it's now ranked lower." That would be insane.

Tom: And they list several other principles like that. Balance, equivalence, dominance. They're essentially saying, "If you're going to build a ranking system, these are the rules it should follow."

Jane: So the summary is really about taking a fuzzy idea, "AI that helps you evaluate," and turning it into a precise, computable problem with clear requirements. That's what a good position paper does.

Tom: And the implications are that this becomes a research agenda. Other researchers can pick up these principles and these frameworks and start building actual systems.

Jane: And that's what makes this exciting. It's not just a philosophical discussion. It's a blueprint for building something real. And speaking of real, the next segment is going to look at the specific improvements they're suggesting over existing methods.

Improvements: Tom: So Jane, we've covered the basics. Now let's talk about what this paper is actually improving on. "Towards an Argumentative Foundation for Evaluative AI" doesn't exist in a vacuum. There's a previous approach called Weight of Evidence, or WoE.

Jane: Right, and that was the original way to do Evaluative AI. It was proposed by Le et al., and it uses a statistical measure to quantify how much each piece of evidence supports or opposes a hypothesis. It's a solid idea, but the paper argues it has some real limitations.

Tom: And the big one is structure. WoE treats each piece of evidence as independent. But in the real world, evidence interacts. A symptom might support a diagnosis, but that diagnosis might then support a treatment. WoE can't capture that chain of reasoning.

Jane: And that's where the argumentation framework shines. It's a graph. You have nodes for evidence and hypotheses, and edges for support and attack. You can have multi-step reasoning, where a hypothesis becomes evidence for another hypothesis.

Tom: Exactly. In the paper's healthcare example, you have symptoms at the bottom, diagnoses in the middle, and treatments at the top. The diagnoses are both hypotheses relative to the symptoms and evidence relative to the treatments. That's a richer structure than WoE can handle.

Jane: And that's a huge improvement. It means the system can reason about intermediate concepts, not just directly from raw evidence to final answer. That's how human experts actually think.

Tom: Another improvement is explainability. With WoE, you get a number, but it's hard to see where that number came from. With an argumentation framework, you have a visual graph. You can literally see the path from a symptom to a treatment.

Jane: And you can use existing explanation methods for argumentation to say, "This piece of evidence contributed this much to this hypothesis." There are even attribution methods that tell you exactly which arguments mattered most.

Tom: And then there's contestability. This is a big one. The paper argues that because the whole structure is explicit, a human can challenge it. They can say, "I disagree with the weight you put on this symptom," or "I think this evidence shouldn't be connected to this hypothesis."

Jane: And that's something you can't do with a black-box neural network. You can't argue with a weight matrix. But you can argue with a graph. You can point at a node and say, "That's wrong."

Tom: And the paper even mentions counterfactual explanations. You can ask, "What would I need to change to make this hypothesis rank higher?" That's a powerful tool for human-AI collaboration.

Jane: So the improvements are really about moving from a flat, statistical approach to a structured, interactive, and contestable one. It's a fundamental upgrade in how we think about AI-assisted decision-making.

Tom: And it's not just about being fancier. It's about being more trustworthy. And that's what we're going to dig into in the next segment, looking at the actual first page of the paper and the motivating example they use.

First Page: Tom: Alright, Jane, let's get into the nitty-gritty of the first page of "Towards an Argumentative Foundation for Evaluative AI." The paper opens with this really compelling healthcare example, and there's a figure that just makes everything click.

Jane: And that figure is the skeleton of the whole approach. It's a graph with three layers. At the bottom, you have symptoms like high fever and cough. In the middle, you have diagnoses like viral pneumonia and asthma. At the top, you have treatments like anti-viral therapy and antibiotics.

Tom: And the edges between them are either green for support or red for attack. So a positive PCR test supports viral pneumonia, which then supports anti-viral therapy. But asthma might attack the idea of antibiotics, because antibiotics don't help asthma.

Jane: And that's the beauty of it. The diagnosis layer is doing double duty. It's a hypothesis with respect to the symptoms below, but it's also evidence with respect to the treatments above. That's the multi-step reasoning we talked about.

Tom: And the paper actually runs the numbers. They use a specific semantics called O-QuAD, and they give every argument an initial weight of zero point five and every edge a weight of one. And the result is that Anti-viral and Antibiotics are ranked equally at zero point four zero, while Bronchodilator comes in lower at zero point three one.

Jane: And that's a concrete, computable result. You can see exactly how the numbers come out. And that's what makes this more than just a diagram. It's a working example.

Tom: And the paper uses this to show how the framework handles conflict. The symptoms might support both viral and bacterial pneumonia, and those two diagnoses might attack each other. The semantics has to resolve that conflict and produce a final strength for each treatment.

Jane: And the way it does that is recursive. The strength of a treatment depends on the strength of the diagnoses that support or attack it, which in turn depends on the strength of the symptoms. It's a bottom-up computation.

Tom: Right. And the paper makes a point that this whole structure is interpretable. You can look at the graph and see why Anti-viral is ranked where it is. You can trace the path from high fever and cough through viral pneumonia to anti-viral therapy.

Jane: And that's the whole point of Evaluative AI. It's not about giving you the answer. It's about giving you the reasoning so you can make the decision yourself. And this first page really demonstrates that in action.

Tom: And it also sets up the contestability angle. If you're a doctor and you disagree with the initial weight of zero point five on a symptom, you can change it. You can say, "No, this patient's cough is more severe, let's bump that up to zero point eight." And the ranking will update accordingly.

Jane: And that's the human-in-the-loop promise. The AI isn't a black box that you just accept. It's a tool that you can poke and prod and adjust until it matches your understanding.

Tom: And that's a really powerful vision. And it's one that we're going to see extend beyond a single doctor. The next segment, we're going to talk about the multi-agent vision, where multiple AI systems can debate each other.

Conclusion: Tom: And that brings us to the final segment on "Towards an Argumentative Foundation for Evaluative AI." We've covered the title, the summary, the improvements, and the first page. Let's wrap this up.

Jane: So the big picture is that this paper takes a really promising idea, Evaluative AI, and gives it a formal, computational foundation. Instead of just saying "AI should present options," they say "Here's exactly how to represent those options as arguments, how to compute their strengths, and how to rank them."

Tom: And they do it with weighted Quantitative Bipolar Argumentation Frameworks. That's the technical heart of the paper. It gives you a graph structure, a way to compute strengths, and a way to explain the results.

Jane: And the implications are genuinely significant. In healthcare, a doctor could use this to see not just "what's the diagnosis?" but "why is this diagnosis ranked higher than that one, and what evidence would change my mind?"

Tom: And it extends to law, finance, any domain where decisions are high-stakes and need to be justified. The ability to contest the AI's reasoning is huge for building trust.

Jane: And the paper also opens up this multi-agent vision. Multiple AI systems, each with their own argumentation framework, could debate each other. They could share evidence, challenge each other's arguments, and converge on a more robust evaluation.

Tom: And that's the long-term research agenda. This paper is a position paper, so it's setting the stage. It's saying, "Here's the foundation, now let's build on it."

Jane: And I think the most exciting part is that it's all explainable and contestable by design. It's not an afterthought. The whole structure is built to be transparent and challengeable.

Tom: And that's what makes it human-centred. It puts the human in the loop, not as a rubber stamp, but as an active participant in the reasoning process.

Jane: So, farewell to "Towards an Argumentative Foundation for Evaluative AI." It's a paper that gives us a roadmap for building AI systems that don't just tell us what to do, but help us think better.

Tom: And that's a future worth getting excited about. Thanks for listening, everyone. We'll be back with the next paper soon.

Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago, Francesca Toni

Imperial College London · The University of Queensland · Cardiff University · King's College London

cs.AI, cs.MA

Submitted: 2026-04-25

Updated: 2026-08-11

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 61/100

The gist: The paper proposes a long-term research agenda for a formal, computational foundation for Evaluative AI (EAI) based on computational argumentation, specifically weighted Quantitative Bipolar

Key concepts

Evaluative AI
This concept describes an AI system that does not make a final judgment. Instead, it lays out all competing hypotheses along with the evidence for and against them, allowing the human user to act as the ultimate decision-maker or jury.
Argumentative Foundation
The core idea of representing reasoning using formal logic. Arguments are structured as nodes in a graph, and their relationships—whether they support or attack other claims—are represented by edges between them.
Weighted Quantitative Bipolar Argumentation Framework
This is a specific technical tool used to model arguments. It assigns numerical weights (quantification) to every piece of evidence, defining relations as either 'support' or 'attack' (bipolar), allowing for a computable ranking of hypotheses.
Contestability and Explainability
Because the AI's reasoning is presented in a transparent graph, users can see exactly how a conclusion was reached. This structure allows humans to challenge specific weights or connections, making the system highly trustworthy.

Terminology

Summary

The paper proposes a long-term research agenda for a formal, computational foundation for Evaluative AI (EAI) based on computational argumentation, specifically weighted Quantitative Bipolar Argumentation Frameworks (wQBAFs). The authors argue that EAI, which supports human decision-making by presenting competing hypotheses with evidence for and against each rather than a single recommendation, can be formally understood as a ranking-based problem. They advocate that wQBAFs provide a principled paradigm for realizing this ranking-based EAI in an explainable and contestable manner, and can empower a multi-agent vision for distributed and human-centred EAI systems.

The paper's contributions are threefold: (1) EAI can be formally understood as a ranking-based problem; (2) wQBAFs provide a principled paradigm for realising ranking-based EAI in an explainable and contestable manner suitable for human-centred EAI systems; (3) wQBAFs can empower agent-level EAI within a multi-agent vision for EAI, in which diverse evaluative models interact and jointly deliberate.

The paper formally defines the ranking-based EAI problem. Given a finite set of hypotheses H and a finite set of evidence E, an initial weight function τ: (E ∪ H) → [0,1] maps each element to its initial weight, capturing plausibility before considering dependencies. Pro and con relations are defined as disjoint binary relations: pro, con ⊆ (E × (E ∪ H)) ∪ (H × H), where evidence can relate to other evidence or hypotheses, and hypotheses can relate to other hypotheses. A relation weight function w: (pro ∪ con) → [0,1] quantifies the strength of influence of each relation. The ranking-based EAI problem is then defined as: Given E = ⟨E, H, pro, con, τ, w⟩, determine a total preorder ≽E on H, representing the relative plausibility of hypotheses given the weighted pro and con relations and initial weights.

The argumentative solution maps an EAI problem to a wQBAF Q = ⟨A, R−, R+, τ, w⟩ where A = E ∪ H, R− = con, and R+ = pro. Solutions are identified in two steps: first, compute the strengths of arguments using quantitative evaluation methods that update initial weights considering support and attack relations; second, derive a ranking over hypotheses by comparing their strengths. The wQBAF solution is naturally explainable: qualitatively, the graphical structure shows relationships between arguments and reasoning paths from evidence to hypotheses; quantitatively, final strengths are explainable via attribution methods and counterfactual explanations, affording contestability by humans who can modify weights to change hypothesis strengths.

The paper introduces seven principles for EAI that serve as design guidelines and evaluation criteria. Point-wise principles include: Principle 1 (Monotonicity), requiring that rankings change consistently with changes in underlying factors, such as adding pro evidence never causing a rank to fall; Principle 2 (Balance), requiring rankings remain unchanged when pro and con evidence are balanced. Pair-wise principles include: Principle 3 (Equivalence), stating two hypotheses must be ranked equally if they share the same pro and con evidence with the same initial weight; Principle 4 (Dominance), requiring a hypothesis be ranked no lower than another if, given equivalent con evidence, its pro evidence set is a superset or its aggregated pro weight is greater. General principles include: Principle 5 (Robustness), requiring stability of the overall ranking under minor perturbations; Principle 6 (Explainability), requiring transparent logic and clear explanations for rankings; Principle 7 (Contestability), ensuring rankings can be challenged by users regarding evidence veracity, initial weights, and the final ranking itself.

The multi-agent vision views each wQBAF as an autonomous argumentative agent. Agents may engage in argumentative communication, exchanging arguments or relations to surface new evidence, reconcile inconsistencies, or converge through principled exchange protocols. Multiple wQBAFs may be fused into group-level evaluations via semantic alignment of arguments, clustering of evidence structures, or ensemble-style aggregation over hypothesis rankings. This positions argumentation-based EAI as a foundation for distributed, resilient, and genuinely deliberative EAI, where interacting agents construct evaluations that are less biased, more robust, explainable, and contestable.

The paper discusses limitations: the framework assumes wQBAFs are correctly constructed, whereas reliable extraction of argumentative structure remains a bottleneck (though LLM-based argument mining may help); contestability may introduce risks of strategic or unintentional misleading evidence, requiring mechanisms like permissioned contribution and trust-aware weighting; multi-agent EAI systems may exhibit deep and sometimes irreducible disagreement, suggesting the need for managing pluralism and supporting structured disagreement rather than enforcing consensus.

Improvements for AI systems

Based on the paper, I can improve AI systems in the following concrete ways:

Improvement: Instead of an AI that outputs a single recommendation and then justifies it, the AI will generate a ranked list of all plausible hypotheses, each accompanied by structured pro and con evidence.

What the improved system can do: In a medical diagnosis setting, instead of saying You have bacterial pneumonia, here's why, it will present: "Ranked hypotheses: (1) Bacterial pneumonia (strength 0.40), (2) Viral pneumonia (strength 0.38), (3) Asthma (strength 0.31). For each, here are the supporting and opposing symptoms with their individual influence weights." This prevents cognitive fixation and preserves human agency.

Abstract

Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EAI that are explainable and contestable, setting the ground for a long-term research agenda towards distributed and human-centred EAI systems.

Sources

Related papers