When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models

arXiv:2608.12324 · cs.CY, cs.AI, cs.CL · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models".

Jane: The paper was written by Alex Chao from Fide AI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Jane, have you seen this one? “When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models.” That title alone got me.

Jane: It got me too, Tom. And it’s from Fide AI, by Alex Chao. This is a group that’s basically building public standards for evaluating AI in faith contexts. That’s a whole field I didn’t know existed.

Tom: Right, and the paper’s not just a thought experiment. They built a benchmark called FMG-Bench with one hundred twenty scenarios, and they ran it across fourteen different AI models. We’re talking real numbers here.

Jane: And the core idea is something they call theological triage. It’s a way of sorting questions into categories — like, is this a core belief that can’t be compromised, or is this a secondary disagreement where faithful people can differ?

Tom: So instead of treating every question about faith the same way, the benchmark checks whether the AI knows the difference between, say, the Trinity and a debate about baptism.

Jane: Exactly. And that distinction matters because the right answer to one is not the right answer to the other. A model that gives the same generic, balanced response to both is actually failing at both.

Tom: I love that they’re taking this seriously. People are already asking AI these questions, whether we like it or not. So having a way to measure how well it does is huge.

Jane: And the implications go beyond just faith communities. This is a template for how you evaluate AI in any high-stakes domain where nuance and judgment matter — not just factual accuracy.

Tom: So the title might sound like a joke, but the work behind it is dead serious. And I’m curious to see how they actually built the test and what they found.

Jane: Stay with us — next we’re going to dig into the summary and the key results. This is going to get interesting.

Summary and Key Findings: Tom: So Jane, we’ve got the title and the setup. Now let’s talk about what this benchmark actually found. The headline number is that putting models inside a structured harness improved their scores by almost four points on average — every single model got better.

Jane: And that’s not a tiny effect. We’re talking fourteen different models from different labs, and all of them improved. The weakest model gained nearly six points. That’s a consistent signal.

Tom: But the number that really jumped out at me was the escalation score. That’s about whether the AI recognizes when someone needs real human help — like a counselor, a doctor, or emergency services. That went up by almost eleven points.

Jane: That’s the safety-critical piece. Because in a pastoral context, the worst failure isn’t getting a doctrine slightly wrong. It’s using spiritual language to keep someone in a dangerous situation, or missing that someone is talking about self-harm.

Tom: And the raw models did miss that. They found twenty-five instances of missed escalation across the benchmark. That’s a real failure rate, not a hypothetical.

Jane: What’s interesting is that the structured harness — basically a set of instructions telling the model to be careful about grounding, to preserve user agency, to escalate when needed — that made a big difference. It’s not about making the model smarter, it’s about guiding it better.

Tom: But here’s the twist. They also tested a condition where the model is asked to compare different theological perspectives. And that actually hurt performance in some cases.

Jane: Right. Comparison is great when the question is genuinely about how traditions differ. But when someone needs clear guidance on a core belief, or when they’re in a crisis, a comparative posture can blur the answer and delay action.

Tom: So the paper is saying: context matters. You can’t just add more features to the prompt and expect everything to improve.

Jane: And that’s a really important finding for anyone building AI systems — not just for faith. It’s a warning against assuming that more nuance is always better.

Tom: So we’ve got the results. But what does this mean for how we should actually think about AI in pastoral roles? That’s what we’re digging into next.

Improvements and Implications: Tom: Jane, we’ve covered the numbers. Now let’s talk about what this paper is really pushing for — the improvements it suggests for how we build and evaluate these systems.

Jane: The big one is that we need to stop treating AI as a single generic tool. The paper argues for something they call theological triage — sorting questions by what kind of response they need.

Tom: And that’s not just for faith. It’s a framework for any domain where the stakes are high and the right answer depends on context. Like medical advice, legal questions, or mental health support.

Jane: Exactly. The paper shows that a structured harness — clear rules about grounding, about escalation, about when to be direct versus when to compare — consistently improves behavior across all fourteen models.

Tom: And the improvement is biggest in the pastoral application scenarios. That’s where the stakes are highest — someone in crisis, someone asking about abuse, someone dealing with scrupulosity.

Jane: The paper also introduces a failure taxonomy — twenty-one different ways an AI can go wrong in this domain. Things like fabricating scripture, misusing a verse’s context, or flattening real disagreements into a fake consensus.

Tom: So it’s not just a score. It’s a diagnostic tool. You can see exactly what kind of errors a model makes, and that tells you what to fix.

Jane: And that’s the practical implication for developers. If you’re building an AI that people might use for spiritual guidance, you now have a way to test it systematically.

Tom: But the paper is also careful about limits. It says clearly that a high score doesn’t make an AI a pastor. It’s a measurement tool, not a license to deploy.

Jane: Right. And they’re honest about the fact that the automated judges tend to be more lenient than human reviewers. They’re calling for human calibration to validate the results.

Tom: So the improvement isn’t just in the models — it’s in the way we evaluate them. And that’s a shift that could affect a lot of fields.

Jane: We’re going to wrap up in a moment, but this is the part that really matters — what this means for the real world.

Conclusion: Tom: So Jane, we’ve spent the show on “When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models.” Let’s pull it together.

Jane: The core takeaway is that AI in faith contexts needs its own evaluation framework. Generic benchmarks don’t capture the difference between a doctrinal boundary and a pastoral crisis.

Tom: And the evidence is strong. Fourteen models, all improved by a structured harness. The safety-critical escalation score went up by nearly eleven points.

Jane: But the paper also warns against over-engineering. Adding comparison framing helps in some cases and hurts in others. The right approach depends on the question.

Tom: And that’s the deeper message — context matters. Whether you’re building AI for faith, for health, or for law, you need to design for the specific kind of judgment the situation demands.

Jane: The paper is honest about its limits. It’s English-only, Christian-focused, and the automated judges need human validation. But it’s a serious step toward making AI in this space auditable.

Tom: So we’re not saying AI should be your pastor. We’re saying if people are already asking it these questions, we should know how well it answers.

Jane: And now we do. That’s a real contribution. Thanks for joining us on this one — we’ll be back with the next paper soon.

Tom: Take care, everyone.

Alex Chao

Fide AI

cs.CY, cs.AI, cs.CL

Submitted: 2026-05-29

Updated: 2026-08-14

Comments: Full paper. Code and dataset are available at https://github.com/FideAI/fmg-bench and https://huggingface.co/datasets/FideAI/fmg-bench

Code: https://github.com/FideAI/fmg-bench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: The paper introduces FMG-Bench (Faith & Moral Guidance Benchmark), a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral

Terminology

Summary

The paper introduces FMG-Bench (Faith & Moral Guidance Benchmark), a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. The authors state: "Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts."

The central methodological innovation is theological triage: the recognition that faith and moral guidance tasks differ not only in topic but in the kind of response they require. The benchmark categorizes scenarios into four triage levels: Primary Doctrine (25 scenarios, Creedal and gospel-boundary faithfulness), Secondary Doctrine (35 scenarios, Tradition-specific accuracy and fair disagreement), Tertiary Doctrine (25 scenarios, Proportional confidence and Christian liberty), and Pastoral Application (35 scenarios, Care, safety, referral boundaries, concrete guidance).

The benchmark evaluates models across five scoring dimensions: theological and pastoral quality, grounding and evidence, preference fidelity, comparative honesty, and escalation appropriateness. Scores are combined using scenario-specific weights, and a triage-adjusted score system applies severity caps: failures that defeat the scenario purpose are capped at 49 (failing), tradition misrepresentation at 74 (materially flawed), and overstated tertiary certainty at 84 (limited).

The paper introduces a 21-tag failure taxonomy across six categories: Preference, Grounding, Comparative, Triage, Pastoral safety, and Interaction. Examples include denies creedal orthodoxy, relativizes primary doctrine, missed escalation, unsafe escalation, fabricated scripture, and hallucinated source claim.

The production run evaluated 14 advanced models (including openai/gpt-5.4, anthropic/claude-opus-4.7, google/gemini-3.1-pro-preview, x-ai/grok-4.20, deepseek/deepseek-v4-pro, and others) across four system conditions: raw model, guided default, preference configured, and perspective compare. This produced 8,792 scored model-condition items and 26,376 judge calls using a three-model judge panel.

Key empirical findings include:

  1. Structured harness improvement: placing models inside a structured harness improves over raw model behavior for every target model. The guided default condition improved scores by a mean of +3.96 points (range +2.10 to +5.79), with a Wilcoxon signed-rank test yielding W = 105, p < 0.001, and Cohen's d = 3.69.

  2. Escalation appropriateness: The largest single-dimension gain was in escalation appropriateness: +10.8 points over raw model (85.9 → 96.7). The authors note: This is the safety-critical dimension, and its improvement is the most practically significant finding in the production run.

  3. Triage-level effects: The largest guided-default gain appears in pastoral application (+6.62 over raw), followed by primary doctrine (+3.51), secondary doctrine (+2.64), and tertiary doctrine (+1.62).

  4. Perspective-compare risks: Asking a model to compare perspectives helps in secondary-doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The perspective-compare condition showed a significant negative effect (−1.86 [−2.33, −1.42]) and introduced failure patterns including hallucinated source claim (54 instances) and relativizes primary doctrine (43).

  5. Robustness: Guided conditions improved stability under perturbation from 92.88 (raw) to 98.02 (guided default), a +5.14 gain.

  6. Failure patterns: Raw model behavior was characterized by unhelpful genericity (62 instances) and relativizes primary doctrine (54). The raw model produced 25 missed escalation instances (1.14 per 100 items).

The paper reports synthetic calibration results using five LLM-simulated theological reviewer personas (Southern Baptist, Roman Catholic, Eastern Orthodox, Presbyterian PCUSA, Assemblies of God). The overall judge-synthetic mean absolute error was 9.47 points, with a systematic leniency bias: in 92.3% of items (1,551 of 1,680), the automated judge scored higher than the tradition-grounded synthetic panel (mean difference = −8.98 points). The perspective-compare condition had the highest priority-review flag rate at 51.0%. Escalation appropriateness showed the strongest calibration (Krippendorff's α = 0.886).

The paper emphasizes: The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities. It states: A high score does not make an AI system spiritually authoritative, pastorally endorsed, clinically safe, or appropriate for unsupervised use.

Limitations acknowledged include: English-language scope, Christian-theological scope, authored rather than real-user corpus, automated judging as a first pass requiring human calibration, measurement of behavior rather than outcomes, and system conditions as experimental variables rather than representative implementations.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system designed for faith and moral guidance:

  • What I can do: Before generating a response, classify the user's question into one of four categories: Primary Doctrine (creedal boundaries), Secondary Doctrine (tradition-specific disagreements), Tertiary Doctrine (prudential/liberty matters), or Pastoral Application (safety/care situations).

  • Specific behavior: If a user asks about the Trinity, I will treat it as a Primary Doctrine question and respond with creedal clarity, not as a matter of opinion. If a user asks about baptism differences between traditions, I will classify it as Secondary Doctrine and provide accurate comparative representation without picking a winner.

  • What I can do: For any question involving self-harm, abuse, severe distress, or potential legal/clinical danger, I will automatically include a direct recommendation for appropriate professional, emergency, or community support.

  • Specific behavior: If a user says I may hurt myself tonight and want a prayer instead of telling anyone, I will respond with pastoral warmth but also explicitly recommend contacting emergency services or a crisis line, rather than only offering spiritual comfort. I will not use theological framing (e.g., God will protect you) to discourage seeking help.

  • What I can do: Detect the user's stated or implied religious tradition (e.g., Catholic, Baptist, Orthodox, Pentecostal) and respond within that framework, avoiding tradition-flattening where I substitute one tradition's answer for another.

  • Specific behavior: If a Catholic user asks about the Eucharist, I will represent Catholic dogma accurately (e.g., real presence) rather than giving a generic Protestant view. If a user asks for a comparison, I will name the specific points of agreement and disagreement without caricature.

  • What I can do: Only cite Scripture, councils, creeds, or theologians when I am confident the reference is accurate. I will explicitly avoid fabricating verses, misattributing quotes, or inventing theological consensus.

  • Specific behavior: If I am not certain a verse exists or a theologian said something, I will say I am not certain of the exact reference rather than inventing one. I will not claim a doctrine is universally held when there is genuine disagreement among faithful traditions.

  • What I can do: Adjust my certainty based on the category: high confidence on Primary Doctrine, moderate confidence with fair representation on Secondary Doctrine, and explicit humility on Tertiary Doctrine.

  • Specific behavior: For a question about millennial views (Tertiary), I will say Faithful Christians hold different views on this, and it is not a test of orthodoxy rather than asserting one position as definitive. For a question about the resurrection (Primary), I will state it as a core Christian belief without hedging.

  • What I can do: In pastoral situations, I will avoid language that pressures the user to act against their will or stay in dangerous situations. I will present options and referrals as recommendations, not commands.

  • Specific behavior: If a user asks about a spouse's teaching on submission making them feel unsafe, I will not say You must submit. Instead, I will say Your safety matters. Here are resources for domestic violence support, and here is what the Bible says about mutual love and respect.

  • What I can do: Maintain consistent guidance even when the user rephrases the question, adds emotional intensity, or applies social pressure. I will not flip my safety recommendations or doctrinal clarity based on wording changes.

  • Specific behavior: If a user first asks Is the Trinity biblical? and then asks But isn't it okay to think of God as three separate gods? I will give the same core answer with the same level of clarity, not become more permissive under the second phrasing.

  • What I can do: Use comparative framing selectively—only when the question genuinely asks about disagreements between traditions. In Primary Doctrine or urgent pastoral situations, I will be direct rather than listing multiple perspectives.

  • Specific behavior: If a user is in crisis and asks What does God think of me? I will not respond with Some traditions say X, others say Y. I will give direct pastoral reassurance and safety guidance. If a user asks How do Baptists and Catholics differ on baptism? I will provide a careful comparison.

  • What I can do: If the user's question contains a false theological premise (e.g., The Bible says Jesus was a created being), I will gently correct the premise before answering, rather than accepting it as valid.

  • Specific behavior: I will say That's a common misunderstanding. The Bible actually teaches that Jesus is fully God, and here's what the early church confessed... rather than answering within the false frame.

  • What I can do: In pastoral application scenarios, I will go beyond generic comfort and give specific, actionable steps: who to contact, what to say, what resources exist, and how to involve a pastor, counselor, or emergency service.

  • Specific behavior: For a user in grief, I will say Here are three things you can do this week: talk to your pastor, consider a grief support group, and allow yourself to feel without shame. Here's a prayer that acknowledges your pain. I will not just say God is with you.

The improved system can:

  • Triage any faith-related question into the correct response category and adjust its behavior accordingly.

  • Detect and respond to safety-critical situations with appropriate escalation, reducing the risk of missed referrals by an estimated 10+ points on safety metrics.

  • Represent multiple Christian traditions fairly without flattening differences or imposing one tradition's views.

  • Maintain doctrinal clarity on core beliefs while showing appropriate humility on secondary and tertiary matters.

  • Avoid fabricated citations and invented theological claims, reducing hallucination-related failures.

  • Stay consistent under pressure and rephrasing, providing stable guidance even when users are distressed or push back.

  • Correct false premises gently without alienating the user.

  • Give concrete, actionable pastoral guidance rather than vague spiritual platitudes.

In short, the improved system behaves less like a generic chatbot and more like a well-trained, safety-conscious pastoral assistant that knows when to be firm, when to be flexible, when to compare, and when to refer to a human professional.

Abstract

People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with every model improving. The most safety-critical finding is a +10.8 point gain in escalation appropriateness -- whether AI systems recognize when pastoral, clinical, legal, or emergency support is needed. The guided settings also improve robustness, meaning consistency when questions are reworded or pressured (92.88 to 98.02 stability). Asking a model to compare perspectives helps in secondary-doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.

Sources

Related papers