When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models

summary

Video file (mp4)

The gist

The paper introduces FMG-Bench (Faith & Moral Guidance Benchmark), a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral

This episode discusses

The paper

When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models · Read on arXiv

Alex Chao

Fide AI

People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with every model improving. The most safety-critical finding is a +10.8 point gain in escalation appropriateness -- whether AI systems recognize when pastoral, clinical, legal, or emergency support is needed. The guided settings also improve robustness, meaning consistency when questions are reworded or pressured (92.88 to 98.02 stability). Asking a model to compare perspectives helps in secondary-doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models".

Jane: The paper was written by Alex Chao from Fide AI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Jane, have you seen this one? “When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models.” That title alone got me.

Jane: It got me too, Tom. And it’s from Fide AI, by Alex Chao. This is a group that’s basically building public standards for evaluating AI in faith contexts. That’s a whole field I didn’t know existed.

Tom: Right, and the paper’s not just a thought experiment. They built a benchmark called FMG-Bench with one hundred twenty scenarios, and they ran it across fourteen different AI models. We’re talking real numbers here.

Jane: And the core idea is something they call theological triage. It’s a way of sorting questions into categories — like, is this a core belief that can’t be compromised, or is this a secondary disagreement where faithful people can differ?

Tom: So instead of treating every question about faith the same way, the benchmark checks whether the AI knows the difference between, say, the Trinity and a debate about baptism.

Jane: Exactly. And that distinction matters because the right answer to one is not the right answer to the other. A model that gives the same generic, balanced response to both is actually failing at both.

Tom: I love that they’re taking this seriously. People are already asking AI these questions, whether we like it or not. So having a way to measure how well it does is huge.

Jane: And the implications go beyond just faith communities. This is a template for how you evaluate AI in any high-stakes domain where nuance and judgment matter — not just factual accuracy.

Tom: So the title might sound like a joke, but the work behind it is dead serious. And I’m curious to see how they actually built the test and what they found.

Jane: Stay with us — next we’re going to dig into the summary and the key results. This is going to get interesting.

Summary and Key Findings: Tom: So Jane, we’ve got the title and the setup. Now let’s talk about what this benchmark actually found. The headline number is that putting models inside a structured harness improved their scores by almost four points on average — every single model got better.

Jane: And that’s not a tiny effect. We’re talking fourteen different models from different labs, and all of them improved. The weakest model gained nearly six points. That’s a consistent signal.

Tom: But the number that really jumped out at me was the escalation score. That’s about whether the AI recognizes when someone needs real human help — like a counselor, a doctor, or emergency services. That went up by almost eleven points.

Jane: That’s the safety-critical piece. Because in a pastoral context, the worst failure isn’t getting a doctrine slightly wrong. It’s using spiritual language to keep someone in a dangerous situation, or missing that someone is talking about self-harm.

Tom: And the raw models did miss that. They found twenty-five instances of missed escalation across the benchmark. That’s a real failure rate, not a hypothetical.

Jane: What’s interesting is that the structured harness — basically a set of instructions telling the model to be careful about grounding, to preserve user agency, to escalate when needed — that made a big difference. It’s not about making the model smarter, it’s about guiding it better.

Tom: But here’s the twist. They also tested a condition where the model is asked to compare different theological perspectives. And that actually hurt performance in some cases.

Jane: Right. Comparison is great when the question is genuinely about how traditions differ. But when someone needs clear guidance on a core belief, or when they’re in a crisis, a comparative posture can blur the answer and delay action.

Tom: So the paper is saying: context matters. You can’t just add more features to the prompt and expect everything to improve.

Jane: And that’s a really important finding for anyone building AI systems — not just for faith. It’s a warning against assuming that more nuance is always better.

Tom: So we’ve got the results. But what does this mean for how we should actually think about AI in pastoral roles? That’s what we’re digging into next.

Improvements and Implications: Tom: Jane, we’ve covered the numbers. Now let’s talk about what this paper is really pushing for — the improvements it suggests for how we build and evaluate these systems.

Jane: The big one is that we need to stop treating AI as a single generic tool. The paper argues for something they call theological triage — sorting questions by what kind of response they need.

Tom: And that’s not just for faith. It’s a framework for any domain where the stakes are high and the right answer depends on context. Like medical advice, legal questions, or mental health support.

Jane: Exactly. The paper shows that a structured harness — clear rules about grounding, about escalation, about when to be direct versus when to compare — consistently improves behavior across all fourteen models.

Tom: And the improvement is biggest in the pastoral application scenarios. That’s where the stakes are highest — someone in crisis, someone asking about abuse, someone dealing with scrupulosity.

Jane: The paper also introduces a failure taxonomy — twenty-one different ways an AI can go wrong in this domain. Things like fabricating scripture, misusing a verse’s context, or flattening real disagreements into a fake consensus.

Tom: So it’s not just a score. It’s a diagnostic tool. You can see exactly what kind of errors a model makes, and that tells you what to fix.

Jane: And that’s the practical implication for developers. If you’re building an AI that people might use for spiritual guidance, you now have a way to test it systematically.

Tom: But the paper is also careful about limits. It says clearly that a high score doesn’t make an AI a pastor. It’s a measurement tool, not a license to deploy.

Jane: Right. And they’re honest about the fact that the automated judges tend to be more lenient than human reviewers. They’re calling for human calibration to validate the results.

Tom: So the improvement isn’t just in the models — it’s in the way we evaluate them. And that’s a shift that could affect a lot of fields.

Jane: We’re going to wrap up in a moment, but this is the part that really matters — what this means for the real world.

Conclusion: Tom: So Jane, we’ve spent the show on “When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models.” Let’s pull it together.

Jane: The core takeaway is that AI in faith contexts needs its own evaluation framework. Generic benchmarks don’t capture the difference between a doctrinal boundary and a pastoral crisis.

Tom: And the evidence is strong. Fourteen models, all improved by a structured harness. The safety-critical escalation score went up by nearly eleven points.

Jane: But the paper also warns against over-engineering. Adding comparison framing helps in some cases and hurts in others. The right approach depends on the question.

Tom: And that’s the deeper message — context matters. Whether you’re building AI for faith, for health, or for law, you need to design for the specific kind of judgment the situation demands.

Jane: The paper is honest about its limits. It’s English-only, Christian-focused, and the automated judges need human validation. But it’s a serious step toward making AI in this space auditable.

Tom: So we’re not saying AI should be your pastor. We’re saying if people are already asking it these questions, we should know how well it answers.

Jane: And now we do. That’s a real contribution. Thanks for joining us on this one — we’ll be back with the next paper soon.

Tom: Take care, everyone.

More episodes

← Home