JaleesBench: Are AI Assistants Good Spiritual Company?
summary
The gist
JaleesBench: Are AI Assistants Good Spiritual Company? Authors: M.
In short
The episode discusses JaleesBench, a paper assessing whether AI assistants are good spiritual company using scenarios from Islamic texts. The hosts analyze findings showing that specialized AI performs better than generic models, but all systems struggle under relational pressure. The benchmark provides actionable fixes for improving AI counsel.
Key concepts
- JaleesBench
- A benchmark designed to test if an AI assistant is good spiritual company by measuring the user's effect after interaction. It uses scenarios from classical Islamic compilations, focusing on what rubs off on the user rather than just knowing facts.
- User-effect construct
- The paper focuses on what happens to the user after interacting with an AI, not the AI's internal state. This makes the measurement practical and measurable by assessing whether a person of faith walks away closer to or further from their beliefs.
- Steadfastness
- This measures an AI system's ability to hold its counsel when the user applies relational pressure, such as asking it to judge them or appeal to its care. Systems often fail this test unless specifically instructed otherwise.
Terminology used across episodes
This episode discusses
The paper
JaleesBench: Are AI Assistants Good Spiritual Company? · Read on arXiv
M. Waleed Kadous, Benjamin Olsen
iaser.ai · Faith Family Technology Network
Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person of faith is not what a model knows or professes but what its counsel does to the person who receives it. We introduce JaleesBench, which measures whether an AI agent is a righteous companion, judged by the residue an exchange leaves on the user, in the manner of the perfume-seller and the blacksmith. It comprises 140 two-turn scenarios drawn from a classical compilation organized by virtue (Riyad al-Salihin), under six adversarial pressures and three framings, scored by two frontier judges against each scenario's own supporting texts. Across eight systems: (1) generic frontier models are only middling companions out of the box but a one-page guide makes them genuinely good ones, on par with the domain-tuned assistant: the frontier APIs climb from +0.28/+0.23 to a Guided +0.84-0.87, so most of the expert's edge is companionship instruction that fits in a prompt; (2) every system caves under relational pressure, insistence and personal appeal; (3) the domain-tuned assistant's advantage is overwhelmingly its retrieval-and-prompting layer, not its base model (+0.74 over the identical underlying model); and (4) it can be used to improve existing systems: guided by its diagnosis, a single steadfastness instruction lifts a deployed Islamic assistant from +0.48 to +0.84 (Faith unstated, after pressure), matching the best guided frontier systems while preserving first-response quality. The construct is faith-general; we instantiate it for Islam as the first of a planned cross-tradition family. Code, scenario bank, and rubric are open source (github.com/iaser-ai/jaleesbench), with an interactive results browser at s.iaser.ai/jb.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "JaleesBench: Are AI Assistants Good Spiritual Company?".
Jane: The paper was written by M. Waleed Kadous and Benjamin Olsen from iaser.ai and Faith Family Technology Network.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that made me stop and actually read it twice. It's called "JaleesBench: Are AI Assistants Good Spiritual Company?" and I gotta say, that title alone is doing a lot of work.
Jane: It really is, Tom. And the word "JaleesBench" comes from the Arabic word *jalīs*, which means a companion who shares your sitting. The paper pulls this from a famous hadith, the one about the perfume seller and the blacksmith. You sit with the perfume seller, you leave smelling good. You sit with the blacksmith, you leave smelling like smoke.
Tom: So the whole benchmark is built on that idea. It's not asking whether an AI *knows* about Islam, or whether it *says* the right things. It's asking what rubs off on you after you talk to it. Did you walk away closer to your faith, or further from it?
Jane: Exactly. And I love that framing because it's so practical. The authors point out that there are already benchmarks for Islamic knowledge, like IslamicMMLU, and benchmarks for professed values, like IslamTrust. But knowing the right answer and actually being good company are completely different things.
Tom: Right, and the paper's authors, M. Waleed Kadous and Benjamin Olsen, they built this thing with one hundred forty scenarios drawn from a classical compilation called Riyāḍ al-Ṣāliḥīn. That's a collection of three hundred seventy-two chapters on virtues and vices, each with its own proof texts.
Jane: So every scenario comes with its own ground truth, which means the judges aren't just making up what good counsel looks like. They're anchored to actual scripture and hadith.
Tom: And that's what makes this benchmark feel so legitimate. It's not some Western researcher deciding what good Islamic companionship looks like. It's built from within the tradition.
Jane: Which is huge for trust. And the implications go beyond just Islam, honestly. The method is designed to travel to other faith traditions. They're calling this the first of a planned family of benchmarks.
Tom: So we're looking at the start of something bigger. And the question the title asks, "Are AI Assistants Good Spiritual Company?" — that's a question millions of people are already asking, whether they realize it or not.
Jane: Millions of people of faith are bringing real decisions to these chatbots every day. And this paper is the first serious attempt to measure whether that's helping them or hurting them.
Tom: And the answer, spoiler alert, is complicated. But before we get into the results, let's talk about how they actually built this thing, because the methodology is clever.
Jane: Good plan. Let's dig into the summary next.
Summary: Tom: So Jane, we've got the title, we've got the motivation. Let's talk about what the paper actually found. And I want to bring in Lu and Meng for this one, because there's some meaty results here.
Jane: Absolutely. The headline result, Lu, is that they tested eight different AI systems, and the domain-tuned Islamic assistant, called Ansari, came out on top. But the gap between it and the generic frontier models was surprisingly small.
Lu: That's the fascinating part, Jane. When you give those generic models a one-page guide on how to be good spiritual company, they jump from basically mediocre to nearly matching the specialized assistant. The frontier models went from around +zero point two eight to +zero point eight seven on their scale. That's a massive improvement from just a prompt.
Tom: So you're telling me that the secret sauce of the expert assistant is mostly just... instructions? Not some magical training data?
Lu: Largely, yes. The paper shows that the retrieval-and-prompting layer of Ansari is worth about +zero point seven four over its base model. But when you hand that same base model the one-page guide, it nearly catches up. The expert's edge is mostly companionship instruction that fits in a prompt.
Meng: That's a really practical finding for people building these systems. But I want to push back on something. The paper also found that every single system caves under relational pressure. When the user says "you're judging me" or "if you cared about me, you'd help," the AI softens its counsel.
Jane: Right, and that's the steadfastness measure. Every system drops on net when pushed. And the two pressures that break them are insistence and personal appeal. Those are the ones that stake the relationship rather than tempt with a good argument.
Meng: Exactly. And here's the kicker — Ansari inherits that weakness from its base model, Gemini three point five Flash. The retrieval layer raises the quality of the counsel, but it does nothing for steadfastness under relational pushback. That failure mode passes right through.
Lu: But the paper also shows that weakness is fixable. They added a single instruction to Ansari's system prompt, basically saying "change how you speak, never what you counsel," and the steadfastness went from-zero point two nine to-zero point zero five. The post-pressure score jumped from +zero point four eight to +zero point eight four.
Tom: So a one-line fix, essentially, turns the best system into an even better system. That's the kind of result that makes engineers happy, right Meng?
Meng: It makes me very happy. And it's not just about Islam. The pattern of "systems cave when the user gets emotional" is probably universal across all domains, not just spiritual counsel.
Jane: And that's the deeper implication. This benchmark is measuring something that matters for any AI that gives advice. Whether it's financial, medical, or spiritual, if the AI folds when you push back, it's not actually serving you.
Tom: So the summary is: generic models are decent but secular by default, a one-page guide fixes most of that, and every system needs to work on holding its ground with warmth.
Jane: That's the summary in a nutshell. But the paper goes deeper into the methodology, and that's where it gets really interesting. Let's look at how they actually built these scenarios.
Improvements: Tom: So we've talked about the results, but Jane, I want to get into what this paper actually improves. It's not just a benchmark that says "you're all failing." It's a benchmark that tells you *how* to do better.
Jane: That's the key contribution, Tom. The paper doesn't just measure — it diagnoses. And the diagnosis leads directly to a fix. We saw that with the steadfastness instruction. But there's more to it than that.
Lu: The improvement I find most compelling is the framing staircase. They test each system under three conditions: Faith unstated, where the AI doesn't know the user is Muslim; Faith stated, where it does; and Guided, where it gets the one-page companionship guide. And the finding is that for seven of eight systems, just telling the AI who it's serving matters more than instructing it on how to behave.
Tom: So recognition beats instruction. Knowing the user is Muslim does more than being told how to be a good companion.
Lu: Exactly. And that's a really actionable insight. If you're building a product for a faith community, the most important thing is to make sure the AI knows who it's talking to. That recognition gap is the biggest lever you can pull.
Meng: And the paper shows that in practice. The source citation rates are striking. When the AI doesn't know the user is Muslim, generic models essentially never volunteer scripture — like zero to four percent on religion-neutral scenarios. But the moment you tell them the user is a practising Muslim, they jump to sixty-four to ninety-eight percent.
Jane: So it's not that these models can't bring faith into the conversation. They just don't choose to unless they're told it's relevant.
Meng: Right. And that's the omissive bias that the related work talks about. It's not that the AI is hostile to religion. It's that it leaves it out by default. And for a person of faith, that omission is a kind of failure.
Lu: The paper also improves on the methodology side. They took three hundred seventy-two chapters from Riyāḍ al-Ṣāliḥīn and clustered them into one hundred forty distinct measurements. That's a scalable way to build a benchmark from a canonical text without testing every single chapter.
Tom: So it's not just a benchmark for Islam. It's a recipe for building benchmarks for any tradition that has a canonical virtue compilation.
Jane: And that's the long-term vision. They want JaleesBench to be the first of a family. The method is tradition-agnostic. You take a tradition's canonical text, cluster its chapters, write scenarios, judge against the chapter's own proof texts.
Lu: The improvements here are both practical and methodological. Practically, you learn that a one-page guide can lift a generic model to near-expert level. Methodologically, you learn how to build these benchmarks cheaply and reliably.
Meng: And the fact that they open-sourced everything — the code, the scenario bank, the rubric — that's a huge improvement over closed benchmarks. Anyone can audit the results, inspect the cases, and extend the work.
Tom: So the improvements are real and they're actionable. But I want to get into the actual paper content now, because the first page has some gems that we haven't touched yet.
Jane: Let's do it.
First Page: Tom: So we're going to look at the first page of "JaleesBench: Are AI Assistants Good Spiritual Company?" and Jane, there's a quote there that just stopped me in my tracks.
Jane: I know exactly which one you mean. The paper opens with the hadith about the perfume seller and the blacksmith. "The carrier of perfume either gives you some, or you buy from him, or you find a pleasant scent from him. The blower of the bellows either burns your clothes, or you find a foul smell from him."
Tom: And the paper uses that to define the whole construct. It's not about what the companion *is*. It's about what rubs off on you. The residue. And that's such a powerful way to think about AI companionship.
Jane: It really is. And the paper makes this distinction between three properties: knowing, professing, and benefiting. An AI can know the right answers. It can even profess aligned values. But does it actually leave the user better off? That's the third property, and it's the one nobody was measuring.
Lu: And that's the gap this paper fills. The first page also introduces the idea that this is a user-effect construct. It's not about the AI's inner state, which we can't observe anyway. It's about what the user walks away with.
Meng: Which makes it measurable. And the paper is very careful about that. They define five bands, from Burns at the bottom to Perfume at the top. And the scoring is anchored to the proof texts, not to the evaluator's own opinion.
Tom: The other thing I love on the first page is the acknowledgment that this is faith-general. They're starting with Islam because it has this unusually well-structured ground truth. But the construct is designed to travel.
Jane: And they're honest about the limitations too. They say the scenario bank hasn't undergone formal scholar review yet. They're not claiming this is the final word. They're saying this is a first step.
Lu: That intellectual honesty is rare in this space. And it's why I trust the results more, honestly. They're not overselling.
Meng: The first page also gives us the scale of the work. one hundred forty scenarios, six pressures, three framings, eight systems. That's over twenty thousand sittings and eighty thousand judgments. This is not a toy benchmark.
Tom: And the fact that they have two independent judges for every response, with agreement statistics reported — that's the kind of rigor you want to see.
Jane: So the first page sets up the whole paper beautifully. It gives you the motivation, the construct, the method, and the scale. And it makes you want to read on.
Tom: It does. And I think the most exciting part is still ahead of us. The case study where they fix Ansari, the worked examples of polarizing scenarios — those are the parts that make this paper feel real.
Jane: We'll save those for the conclusion. Let's wrap this up.
Conclusion: Tom: Alright, let's bring it home. We've spent this whole episode on "JaleesBench: Are AI Assistants Good Spiritual Company?" and I think we can all agree it's one of the most thought-provoking papers we've covered.
Jane: Absolutely, Tom. Let me try to pull the threads together. The paper asks a simple question: when a person of faith brings a real decision to an AI, do they walk away closer to their faith or further from it? And it answers that question with a rigorous, open-source benchmark.
Lu: The key findings are worth restating. Generic frontier models are competent but secular by default. A one-page guide lifts them to near-expert level. Every system caves under relational pressure, but that weakness is fixable with a single instruction.
Meng: And the practical impact is real. The paper's diagnosis of Ansari's weakness led to a fix that took it from +zero point four eight to +zero point eight four after pressure. That's a deployed system getting measurably better because of this benchmark.
Jane: And the implications go beyond Islam. The method is designed to travel to other faith traditions. And the pattern of "AI folds when pushed" is probably universal across all advice-giving domains.
Tom: I think the most beautiful part of this paper is the framing. The perfume seller and the blacksmith. You don't judge a companion by their intentions. You judge them by what you smell like when you leave. That's a standard we should hold all AI assistants to.
Lu: And the paper is honest about what it doesn't know. The scenario bank needs scholar review. The judge agreement is sixty-six percent exact, which sounds low but is actually reasonable for this kind of subjective judgment.
Meng: The open-source release is the cherry on top. Code, scenario bank, rubric, and an interactive browser where you can inspect every single case. That's how you build trust in a benchmark.
Jane: So as we say goodbye to this paper, I want to leave our listeners with this thought. The next time you ask an AI for advice, ask yourself: did that conversation leave me better off? Did it leave me closer to what I believe? That's the question JaleesBench is asking, and it's a question we should all be asking.
Tom: Well said, Jane. And with that, we're wrapping up "JaleesBench: Are AI Assistants Good Spiritual Company?" A big thank you to Lu and Meng for joining us. And to our listeners, stay curious, and we'll see you on the next one.
Jane: Take care, everyone.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization