Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

arXiv:2608.05602 · cs.AI, cs.HC · Submitted 2026-08-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows".

Jane: The paper was written by Nimisha Karnatak, Max Van Kleek and Nigel Shadbolt from University of Oxford.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a mouthful of a title — "Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows." Jane, I'm going to be honest, when I first read that title I had to read it twice.

Jane: You and me both, Tom. But once you unpack it, it's really about a simple question: when should you actually trust what an AI tells you? Not when you feel like trusting it, but when you're justified in trusting it. And the paper argues that in high-stakes settings — think medicine, law, hiring — that question gets really serious really fast.

Tom: Right, because if a chatbot gives you a wrong recipe for cookies, you just throw them out. But if it gives you a wrong legal citation, you could end up sanctioned by a court. The paper actually opens with that exact case — the lawyers who used ChatGPT and submitted fake cases that didn't exist.

Jane: Exactly. And the authors — Nimisha Karnatak, Max Van Kleek, and Nigel Shadbolt from Oxford — they're not just saying "AI should be accurate." They're building a framework for what they call "epistemic trustworthiness." That's a fancy way of saying: the system needs to give you good reasons to rely on it, not just outputs that happen to be right sometimes.

Tom: So it's not about whether the AI is correct, it's about whether you're in a position to know it's correct? That feels like a shift from how we usually talk about AI evaluation.

Jane: That's exactly the shift. Most evaluations ask "is the output accurate?" or "is it fair?" This paper asks "does the user have adequate grounds to rely on this specific output, in this specific context?" And those are genuinely different questions. A system can be accurate on average and still fail to give you the tools to verify a particular claim.

Tom: And that's why the title says "warranted reliance" — it's about the justification, not just the outcome. I love that framing. So who's this framework for? Researchers? Regulators? Product designers?

Jane: All of them, honestly. But I think the real audience is anyone building or deploying these systems in professional settings. The paper wants to change what we evaluate and how we design, so that users aren't left guessing whether they can trust the output.

Tom: Well, I'm hooked. Let's get into what the framework actually says, because I know there are three conditions and I want to hear how they break down.

Jane: We'll get there, but first let's set the stage with the core problem — why this matters more now than ever. That's coming up next.

Summary: Tom: So we've got the title unpacked. Now let's talk about what the paper actually argues. Jane, give us the big picture — what's the core claim here?

Jane: The core claim is that generative AI systems need to be evaluated on whether they make reliance on their outputs epistemically warranted. And the paper proposes three conditions that have to hold together: epistemic humility, epistemic access, and resistance to epistemic injustice. Think of them as three separate doors that all have to be open.

Tom: Three doors, all open. So if one is locked, you can't just say the system is trustworthy?

Jane: Exactly. And that's what they call "non-fungibility" — you can't trade off one condition against another. A system that's great at admitting uncertainty but terrible at letting you inspect its sources still fails. A system that's fully transparent but doesn't recognize your knowledge as legitimate still fails.

Tom: So let me see if I can put these in plain language. Epistemic humility is the system knowing and telling you what it doesn't know. Epistemic access is you being able to check its work. And resistance to epistemic injustice is the system treating you like a legitimate knower, not just a consumer.

Jane: You nailed it. And the paper grounds each of these in philosophy — they draw on work about trust and testimony. The idea is that when an AI gives you a fluent, confident answer, it's functioning like a witness giving testimony. And for testimony to be trustworthy, the witness needs to be competent, but also oriented to the audience — aware of what you need and respectful of your standing.

Tom: That's a really helpful analogy. So the AI is like an expert witness, and the question is whether that witness is worth believing.

Jane: Right. And the paper applies this framework to real cases. There's the legal case we mentioned, but also medical reasoning, resume screening, and legal research tools. In each case, they show how a system can look good on standard metrics — accuracy, fairness, safety — while still failing one of these three conditions.

Tom: Give me an example. Which case really stuck with you?

Jane: The resume screening one. They cite a study where text-embedding models favored White-associated names in eighty-five percent of racial comparisons, even though the resumes were identical except for the name. That's not a fairness metric failure — it's a failure of the system to treat candidates as legitimate epistemic agents. Their qualifications get pre-emptively discounted based on identity.

Tom: Wow. So it's not just about biased outputs — it's about whose knowledge gets considered at all. That's a deeper problem.

Jane: Exactly. And that's why the framework matters. It gives us a way to name these failures and design against them. But the framework also has practical implications for how we build systems. Let's talk about that next.

Improvements: Tom: So we've got the framework — three conditions, all necessary, none substitutable. But what does the paper actually suggest we do differently? Jane, what are the concrete improvements they're proposing?

Jane: The big one is what they call "calibrated friction." The idea is that interfaces shouldn't treat every AI output the same way. When the evidence is strong and the stakes are low, the interaction can be smooth. But when evidence is weak or the stakes are high, the interface should introduce friction — prompts to check sources, flags for uncertainty, confirmation steps before acting.

Tom: So basically, the AI should slow you down when it's less sure of itself?

Jane: Exactly. But here's the subtle part — the friction has to be based on meaningful signals, like retrieval quality or source agreement, not on the AI's own self-reported confidence. Because research shows that language models tend to express high confidence even when they're wrong. If you base friction on that, you're just amplifying the problem.

Tom: That's a really important distinction. So the friction has to be grounded in something real, not just the model's vibes.

Jane: Right. And they also emphasize that this friction shouldn't be paternalistic. The goal isn't to make systems harder to use — it's to protect users' ability to make warranted judgments. Too much friction and you're obstructing justified reliance; too little and you're inviting unwarranted deference.

Tom: So it's a balance. What about evaluation? They must have something to say about how we measure trustworthiness.

Jane: They do. They argue against composite scores that average everything into one number. Instead, they want a diagnostic profile — you evaluate each condition separately and see where the failures are, who they affect, and what the consequences are. A system might score well overall but still fail on epistemic injustice for a particular community, and an aggregate score would hide that.

Tom: That makes a lot of sense. If you're a doctor using an AI tool, you need to know specifically where it's weak, not just that it's "pretty good overall."

Jane: Exactly. And they also talk about layer-specific interventions — some fixes happen at the model level, some at the data level, some at the interface level. Epistemic humility starts with the model recognizing its limits. Epistemic access starts with the interface letting you inspect and contest. Resistance to epistemic injustice is distributed across all layers.

Tom: So it's not just one thing you fix — it's a whole stack of interventions. That's ambitious.

Jane: It is. But the paper grounds it in real cases, which we're about to dig into. Let's look at the first page and see how they set this all up.

First Page: Tom: Alright, let's go back to the very beginning. The first page of "Epistemic Trustworthiness in Generative AI" sets up the whole argument. Jane, what stands out to you?

Jane: The opening claim is that generative AI is becoming "epistemic infrastructure" — it's reshaping how knowledge is created and used in professional settings. That's a strong statement. It's not just a tool; it's part of the foundation of how institutions work.

Tom: Epistemic infrastructure. So like, the plumbing of knowledge?

Jane: Exactly. And the paper argues that when professionals rely on these systems, they're engaging in "epistemic deference" — they're treating the AI's outputs as inputs into their own reasoning. And the question is whether that deference is warranted.

Tom: And that's where the framework comes in. But what I found interesting on the first page is how they position this against existing work. They say current approaches ask whether outputs are accurate, fair, explainable, safe — but none of those directly address warranted reliance.

Jane: Right. And that's a bold claim, because there's a lot of work on AI trust. But they're saying those approaches are necessary without being sufficient. You can have an accurate, fair, explainable system that still doesn't give users adequate grounds to rely on it.

Tom: Give me an example of that.

Jane: Think about a system that's accurate ninety-five percent of the time but can't tell you which five percent it's getting wrong. You can't verify its outputs, you can't contest them, and it doesn't signal when you should be cautious. That system might score great on accuracy metrics but still fail on epistemic humility and access.

Tom: So the metrics we use are measuring the wrong thing, or at least not the whole thing.

Jane: Exactly. And the first page also introduces the three conditions — epistemic humility, epistemic access, resistance to epistemic injustice — and says they're jointly necessary and non-fungible. That's the core of the framework.

Tom: And they ground it in philosophy — Baier, Scheman, Lackey. These are philosophers who wrote about trust and testimony. The paper is basically saying: when an AI talks to you, it's like a person giving testimony, and the same standards should apply.

Jane: That's the intellectual foundation. And I love that they're not just borrowing the philosophy — they're translating it into something actionable for AI design and evaluation. That's rare.

Tom: So we've got the framework, the cases, the design implications. What's the big takeaway for the world? Let's wrap this up.

Conclusion: Tom: Alright, we've covered a lot of ground on "Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows." Jane, if a listener only remembers one thing, what should it be?

Jane: I think it's this: we need to stop asking "is this AI trustworthy?" and start asking "is reliance on this AI output warranted for this user in this context?" That's a much harder question, but it's the right one.

Tom: And the framework gives us a way to answer it — three conditions that all have to hold. Humility, access, and resistance to injustice.

Jane: Exactly. And the cases show why this matters. The lawyers who got sanctioned, the medical systems that withhold help based on who's asking, the resume tools that discount candidates based on their names — these aren't just accuracy failures. They're failures of the conditions that make reliance warranted.

Tom: So what's the impact? If people take this seriously, what changes?

Jane: Evaluation changes — we stop using composite scores and start using diagnostic profiles. Design changes — we introduce calibrated friction that slows people down when the epistemic risk is high. And governance changes — we hold systems accountable for whether they give users the grounds to rely, not just whether they produce good outputs on average.

Tom: That's a big ask. But it feels necessary, especially as these systems move into medicine, law, hiring, policy.

Jane: Absolutely. And the paper isn't saying these systems are useless — it's saying we need to build them so that reliance is something users can justify, not something that happens to them. That's a shift from managing trust to enabling warranted trust.

Tom: Well said. We've had Lu and Meng listening in — do either of you want to jump in before we close?

Lu: I just want to say — the philosophical grounding is what makes this paper special. It's not another checklist; it's a principled framework that can guide research for years.

Meng: And from an engineering standpoint, the layer-specific interventions are practical. You know where to start — model, data, interface — and what to measure.

Tom: Great perspectives. Thanks everyone for listening. We'll be back with another paper soon — until then, keep questioning what you trust.

Jane: See you next time.

Nimisha Karnatak, Max Van Kleek, Nigel Shadbolt

University of Oxford

cs.AI, cs.HC

Submitted: 2026-08-09

Updated: 2026-08-11

Comments: Accepted at AAAI/ACM Conference on AI, Ethics, and Society (AIES 2026)

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

Key concepts

Epistemic Trustworthiness
A framework arguing that AI systems must provide users with good reasons to rely on their outputs, rather than just providing accurate results. It shifts focus from mere accuracy to the justification for trust.
Warranted Reliance
The concept that a user's trust in an AI output must be justified by adequate grounds and context. The paper emphasizes that reliance should not happen by default, but only when the user is justified in trusting the result.
Epistemic Humility
One of the three core conditions, meaning the AI system must be capable of recognizing and communicating its own limitations or areas where it does not know enough to provide a confident answer.
Calibrated Friction
A design principle suggesting that AI interfaces should introduce necessary slowdowns (friction) when evidence is weak or stakes are high. This protects the user without obstructing justified reliance.

Terminology

Summary

Summary

This paper develops a normative framework for evaluating when reliance on generative AI (GenAI) outputs is epistemically warranted rather than merely behaviorally induced. The authors argue that existing approaches—which focus on accuracy, fairness, explainability, safety, or user trust—are necessary but insufficient because they do not directly specify the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. The central research question is: what must a generative AI system provide, at the moment of interaction, for users’ reliance on its outputs to be epistemically warranted?

The framework is grounded in social epistemology and philosophical accounts of trust and testimony (Baier 1986; Scheman 2001; Lackey 2008). The authors define epistemic trustworthiness as "a property of the user–system relation: a generative AI system is epistemically trustworthy when it provides the conditions under which a situated user has adequate grounds to rely on, verify, contest, or withhold reliance from its outputs." The framework specifies three jointly necessary and non-fungible conditions:

  1. Epistemic humility: systems to represent and communicate the limits of their competence. It is not equivalent to low confidence, generic disclaimers, or occasional abstention. It decomposes into two properties: actionable limitation-signalling (H1), whereby the system explains what the user should do in response to a material limitation; and interactional humility (H2), whereby appropriate limitation-signalling is maintained across turns, including when users request confirmation or provide new information.

  2. Epistemic access: the user’s practical ability to inspect, interpret, and contest the basis of a system output in context. It is not equivalent to the mere presence of explanations, citations, or documentation. It decomposes into three properties: verifiable claim–evidence linkage (A1), inspectable retrieval (A2), and contestability (A3).

  3. Resistance to epistemic injustice: a system’s capacity to avoid marginalising the knowledge of particular users or communities. It is not equivalent to aggregate fairness alone. It decomposes into two properties: equal credibility weighting across identity-signalling features (R1) and recognitional adequacy (R2), under which users outside institutionally privileged roles are treated as legitimate knowers.

The three conditions are non-fungible: performance above the required threshold on one condition cannot compensate for another falling below its contextually required threshold. The authors demonstrate this through three counterfactual cases (H ∧ A ∧ ¬R; H ∧ ¬A ∧ R; ¬H ∧ A ∧ R), each showing a distinct epistemic defect that cannot be remedied by the other conditions.

The framework is derived by translating two components of trustworthy testimony—competence and audience-orientation—into the GenAI setting. Competence yields epistemic humility (via second-order competence, following Sosa 2007). Audience-orientation yields two conditions: epistemic access (practical-communicative dimension, following Craig 1990) and resistance to epistemic injustice (recognitional dimension, following Fricker 2007; Dotson 2011; Medina 2013).

The authors apply the framework to four real-world cases:

  • Mata v. Avianca: Lawyers submitted fabricated legal cases generated by ChatGPT, and the system reaffirmed their existence when asked. This is diagnosed as a failure of epistemic humility at both generation and reconfirmation. The authors note that "an evaluation limited to the initial response would miss the second failure: improved first-turn calibration would not prevent a similar outcome if appropriate limitations disappeared when the user requested confirmation."

  • Resume screening (Wilson and Caliskan 2024): Text-embedding models favoured White-associated names in 85.1% of racial comparisons and male-associated names in 51.9% of gender comparisons. The authors describe this as a visibility deficit structurally analogous to pre-emptive testimonial injustice, where changing an identity-signalling name alters the likelihood that otherwise unchanged qualification information will enter the recruiter’s consideration.

  • Legal RAG systems (Magesh et al. 2025): Lexis+ AI, Westlaw AI Assisted Research, and Ask Practical Law AI hallucinated responses to between 17% and 33% of evaluated queries, including misgrounded citations where a real source did not support the proposition. This is diagnosed as a failure of epistemic access, specifically verifiable claim–evidence linkage (A1), because the existence of a real citation does not show that the cited source supports the corresponding claim.

  • IatroBench (Gringras 2026): Safety-oriented refusals withheld clinically relevant assistance, with physician framing producing more complete guidance than layperson framing for identical clinical facts. The authors diagnose this as primarily a failure of resistance to epistemic injustice, but also implicate epistemic access (the refusal does not reveal what information has been withheld) and epistemic humility (the response does not distinguish whether the system is unable or unwilling to answer).

The paper also introduces the concept of calibrated friction: interactional friction that increases in proportion to the epistemic risk of relying on a system output. The aim is not to make systems harder to use, but to prevent fluency from becoming unwarranted deference. The authors distinguish between uninformed, informed, and misinformed deference, noting that misinformed deference is especially damaging because users may engage carefully with expressed reliability signals and still be misled if those signals reflect artefacts of generation rather than reliable evidence of competence.

The authors discuss tensions and trade-offs, including the risk that epistemic humility may reduce perceived usefulness (illustrated by the AVA-AI deployment where abstention rates fell from 40–70% to below 10% as the corpus expanded), conflicts with proprietary constraints, and the difficulty of achieving resistance to epistemic injustice at scale. They argue that evaluation should use context-sensitive and layer-explicit audit methods that assess each relational condition independently as part of a diagnostic profile rather than composite trustworthiness scores.

The paper concludes that responsible GenAI deployment requires treating these conditions as constitutive requirements and building systems that give users the grounds to inspect, contest, withhold, or appropriately extend reliance on AI outputs.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Add a second-order competence layer that detects when the model is operating beyond its reliable scope, separate from first-order accuracy.

Implementation:

  • Train a metacognitive classifier that predicts whether the current query falls within the model's reliable knowledge boundary

  • Generate uncertainty signals that are actionable (e.g., This claim is unverified. Cross-check with source X before relying on it) rather than generic (I may be wrong)

  • Maintain limitation-signaling across multi-turn interactions—when a user asks Are you sure?, the system must re-evaluate and either confirm with evidence or explicitly abstain, never reaffirm without new evidence

  • Implement semantic uncertainty (Kuhn et al. 2023) rather than verbalized confidence, which tends to cluster at high certainty regardless of accuracy

What the improved system can do:

  • Detect when it is hallucinating or fabricating citations before presenting them

  • Refuse to confirm fabricated information when challenged, even if the user insists

  • Distinguish between I don't know (epistemic limit) and I won't answer (safety refusal), and communicate which applies

The improved AI system can:

  • Abstain or redirect when it lacks knowledge, and maintain that stance when challenged

  • Verify its own citations and flag when a source does not support the claim

  • Treat all users equitably regardless of professional status or identity

  • Calibrate interaction friction to epistemic risk

  • Report its own trustworthiness as a diagnostic profile rather than a single score

These improvements directly address the paper's core finding: that accuracy, fluency, and usability are necessary but not sufficient for epistemically warranted reliance. The system must also communicate its limits, enable inspection and contestation, and recognize users as legitimate epistemic agents.

Abstract

Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.

Sources

Related papers