HoosierHelp: Benchmarking LLM Agents for Social Service Navigation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "HoosierHelp: Benchmarking LLM Agents for Social Service Navigation".
Jane: The paper was written by Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang et al. from University of Notre Dame.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. We’ve got a fascinating paper on the table today, and it’s called "HoosierHelp: Benchmarking LLM Agents for Social Service Navigation." Jane, I have to say, the title alone got me curious—this isn’t your typical AI benchmark.
Jane: Absolutely, Tom. And that’s exactly why I wanted to dig into it. When we hear "benchmark," we usually think of math problems or coding challenges. But this one is about something much more human—helping people find food, housing, utility assistance, that kind of thing. It’s built on real resources from Indiana, over three thousand nine hundred of them, covering twenty-one different service categories.
Tom: Right, and that’s what makes it so different. The authors took the Indiana two hundred eleven directory—that’s the free hotline people call to get connected with local help—and turned it into a structured database. Then they built an interactive test where an AI agent has to chat with a simulated person in need, ask the right questions, search the database, and recommend the right resource.
Jane: And the clever part is how they made the simulated users feel real. They didn’t just have everyone answer politely and directly. Some users are impatient, some ramble, some make requests the agent can’t actually do, like "can you just pay my bill for me?" And some even contradict themselves. That’s where the benchmark gets really interesting.
Tom: So it’s not just about whether the AI can call a tool correctly. It’s about whether it can handle a messy, real-world conversation where the person on the other end isn’t always cooperative.
Jane: Exactly. And the results show that even the best models, like GPT-five point four and Qwen3 point 7-Max, really struggle. The best model only got about sixty-seven percent of the final recommendations exactly right. And when users were self-contradictory, that number dropped to near zero for some models. It’s a sobering reminder that we’re still far from having reliable AI assistants for this kind of high-stakes navigation.
Tom: So the title "HoosierHelp" isn’t just a catchy name—it’s a real test of whether AI can be trusted with something as important as connecting people to basic needs.
Jane: And that’s the big implication. If we want AI to help with social services, we need to know where it fails. This paper gives us a map of exactly that. And I think the next thing we need to talk about is how they actually built this benchmark and what the agents had to do.
Tom: Good point. Let’s get into the details of the benchmark construction next.
Summary: Tom: So we’ve set the stage with the title, but now let’s get into the meat of "HoosierHelp: Benchmarking LLM Agents for Social Service Navigation." Jane, how did they actually set up these conversations?
Jane: So each episode starts with a hidden profile. The agent doesn’t see it, but it contains the user’s real need, their constraints, and the target resource they should end up with. The agent has to figure all of that out through conversation. And there are two tools the agent can use—one to search the database, and one to make a final recommendation.
Tom: And the search tool is pretty detailed, right? It’s not just "find me food." It has fields for service category, location, schedule, intake method, documents, and eligibility.
Jane: Exactly. And that’s where the grounding problem shows up. The paper found a big gap between what the agents searched for and what they finally recommended. The best model had a tool exact-match score of only thirty-eight point seven five percent, but a resource exact-match score of sixty-seven point five percent. That means the agents were often picking the right resource without faithfully preserving all the user’s constraints in the search.
Tom: That sounds like a recipe for trouble. If the agent doesn’t search with the right schedule or the right documents, it might recommend something the user can’t actually use.
Jane: Right. And the paper shows this clearly when they split the results by constraint satisfiability. When the user’s first choice directly matched a resource, models did pretty well. But when the first choice had no match and the user had to fall back to an alternative, performance dropped dramatically. For example, Qwen3 point 6-35B-A3B went from eighty-five percent down to thirty-eight percent on resource accuracy.
Tom: So the agents are bad at handling the "no results" case. They either give up, or they broaden the search too much and lose track of what the user actually said.
Jane: Precisely. And the user behavior split makes it even worse. Self-contradictory users were the hardest condition by far. Tool exact match fell below fifteen percent for every single model. That tells us the agents aren’t recognizing a contradiction as something they need to clarify before searching.
Tom: So they just take the last thing the user said and run with it?
Jane: Often, yes. Or they take the first thing and ignore the correction. Either way, they’re not treating the conversation as a source of facts that need to be reconciled.
Tom: And what about the interaction length? Did longer conversations help?
Jane: Not necessarily. The paper found that more turns didn’t reliably improve accuracy. Some models, like Qwen3 point 6-35B-A3B, averaged only four turns and had low tool accuracy. Others, like GPT-five point four, took over six turns and did better. But even the longer conversations didn’t close the gap on the hard cases.
Tom: So it’s not about talking more—it’s about asking the right questions and keeping track of the answers.
Jane: Exactly. And that brings us to what the paper suggests we should do about it.
Improvements: Tom: Alright, so we know the problem. What does "HoosierHelp: Benchmarking LLM Agents for Social Service Navigation" suggest we do about it? Jane, what are the key improvements the authors are pushing for?
Jane: The main call is for agents that are more robust to complex and non-ideal user interactions. Specifically, they highlight the need for better information elicitation and grounded tool use. That means agents need to actively ask for missing constraints, not guess them, and they need to keep those constraints consistent across the whole conversation.
Tom: So it’s about memory and attention. The agent has to remember what the user said in turn two when it’s making a search in turn six.
Jane: Right. And the paper shows that this is exactly where models fail. For example, in one of the diagnostic cases they shared, a user gave two needs—public safety and food assistance—with different schedules. The agent handled the first need perfectly but completely forgot about the food assistance. It just finalized with one recommendation.
Tom: So it’s not just about handling one complex conversation—it’s about handling multiple threads at once.
Jane: Exactly. And another improvement they suggest is better fallback reasoning. When a search returns no results, the agent should ask the user which constraint can be relaxed, rather than just broadening everything on its own. In one case, the agent expanded the location from a specific ZIP code to a whole county without asking, and ended up recommending the wrong resource.
Tom: That’s a classic case of the agent trying to be helpful but actually making things worse.
Jane: Yes. And the third improvement is about recognizing non-actionable requests. When a user says "can you just pay my bill," the agent needs to redirect that into the real need—utility assistance—without getting distracted. Some models handled this okay, but many lost the thread and recommended something irrelevant.
Tom: So the improvements are really about three things: asking the right questions, handling empty results intelligently, and staying focused on the real need.
Jane: And the paper suggests that current agents lack robust mechanisms for all three. They’re good at pattern-matching from the conversation to a plausible resource, but they’re not good at maintaining a faithful requirement state.
Tom: So the fix isn’t just a better prompt or a bigger model—it’s a different architecture or training approach that emphasizes grounding and consistency.
Jane: Exactly. And I think that’s a really important message for the field. We can’t just throw more compute at this problem. We need agents that are designed to be careful, not just clever.
Tom: That’s a great point. Let’s bring in Lu and Meng to get their take on the practical side of these improvements.
Conclusion: Tom: We’ve covered a lot of ground on "HoosierHelp: Benchmarking LLM Agents for Social Service Navigation." Jane, before we wrap up, let’s pull it all together.
Jane: Sure. The paper gives us a realistic, interactive benchmark for a high-stakes domain. It shows that current LLM agents are unreliable for social service navigation, especially when users are impatient, contradictory, or have constraints that don’t match anything at first. The best model only got about two-thirds of the final recommendations right, and that’s not good enough for something this important.
Tom: And the key takeaway is that we need agents that are better at eliciting information, handling fallback scenarios, and staying grounded in what the user actually said.
Jane: Right. It’s not about making the agent faster or more conversational. It’s about making it more careful and more consistent. The paper’s diagnostic failures show that even small lapses in memory or attention can lead to wrong referrals.
Tom: So what’s the bigger picture here? What does this mean for the world?
Jane: It means we’re not ready to deploy AI agents for social service navigation without human oversight. But it also gives us a clear roadmap for what needs to improve. If we can build agents that pass benchmarks like this, we could help millions of people access the resources they need, around the clock, without overburdening human specialists.
Tom: That’s the promise. And this paper is a crucial step toward making it real.
Jane: Absolutely. It’s a benchmark that measures something that actually matters. And for that, I think we should give the authors a lot of credit.
Tom: Agreed. So we’ll say goodbye to "HoosierHelp" and get ready to look at the next paper on the arXiv. Thanks for listening, everyone.
Jane: See you next time.
Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang, Yanfang Ye
University of Notre Dame
cs.HC, cs.AI
Submitted: 2026-07-03
Updated: 2026-08-12
Project page: https://www.211.org
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 63/100
The gist: "Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints.
Terminology
Summary
Summary
The paper introduces HoosierHelp, an interactive benchmark for evaluating LLM agents on social service navigation. The task requires agents to connect help-seeking individuals to resources that satisfy their needs and specific constraints. The benchmark is grounded in 3,971 Indiana public social service resources covering 21 service categories.
The paper states: "Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not capture the interaction complexity and constraint-grounding demands of this setting."
The benchmark construction involves three key dimensions of user realism: (i) need structure (single vs. composite needs), (ii) constraint satisfiability (direct match vs. fallback required), and (iii) interaction behavior patterns (normal, impatient, rambling, unsupported request, and self-contradictory). The paper explains: HoosierHelp enhances the realism of simulated users by varying their need structure, constraint satisfiability, and behavior patterns, including impatience, rambling, unsupported requests, and self-contradiction.
The agent interacts with LLM-simulated users conditioned on structured hidden profiles, issues structured resource-search calls via a search resources tool, and terminates with a final recommendation tool call selecting resource IDs from previously returned search results. The paper notes: Separating search from final recommendation also allows HoosierHelp to distinguish search errors from selection errors.
Experiments were conducted on 240 samples across seven LLMs: Qwen3.7-Max, Qwen3.6-27B, Qwen3.6-35B-A3B, GPT-5.4, GPT-OSS-120B, GPT-4.1-mini, and DeepSeek-V4-Flash. The main results show a consistent gap between tool grounding and final referral accuracy. The paper reports: Resource EM is substantially higher than Tool EM for every model. The best Resource EM is 67.50% for Qwen3.7-Max, while the best Tool EM is 38.75% for GPT-5.4.
Key findings include:
-
Fallback reasoning is the clearest stress test:
Every model drops substantially on fallback-required cases. For example, Qwen3.6-35B-A3B falls from 85.00% Resource EM on direct-match cases to 38.33% on fallback-required cases.
-
Self-contradictory users are the hardest condition:
Tool EM falls below 15% for every model, and Resource EM often collapses as well, indicating that agents rarely recognize contradictions as unresolved facts that require clarification before search.
-
Normal conversations are not solved:
The best normal Resource EM is 85.42%, and the best normal Tool EM is 66.67%.
-
Interaction length reflects strategy, not quality:
Qwen3.6-35B-A3B averages only 4.00 agent turns and frequently reaches a final recommendation, but its Tool EM is low. In contrast, GPT-5.4 averages 6.28 turns and obtains the highest Tool EM.
-
Field-level failures concentrate in schedule and intake errors:
Schedule and intake errors are common across models, reflecting the difficulty of mapping conversational availability and access preferences into structured filters.
The paper concludes: "current LLM agents remain unreliable for social-service navigation. The main failure is not simply retrieving a relevant resource, but maintaining a faithful requirement state through noisy interaction, tool feedback, fallback revision, and grounded final selection. The authors highlight
the need for agents with stronger information elicitation and grounded tool use."
Improvements for AI systems
Based on the paper's findings, here are specific improvements to AI systems for social service navigation:
Improvement: Add an explicit structured memory that tracks each user constraint (location, schedule, intake, documents, eligibility) as a verified, unverified, or contradictory state. The system must not finalize recommendations until all required fields are either verified or explicitly confirmed as unconstrained.
What it does: Prevents the observed failure where agents select plausible resources while omitting critical constraints. The system maintains a checklist per need and refuses to search until the constraint state is complete, reducing the gap between Tool EM (38.75%) and Resource EM (67.50%).
Improvement: When search resources returns zero results, automatically trigger a structured clarification flow: (1) inform the user of no match, (2) ask specifically which constraint (location, schedule, or intake) can be relaxed, (3) preserve all other verified constraints, (4) re-search with only the changed field.
Improvement: Implement a contradiction detector that flags when a user provides conflicting information within a turn or across turns. When detected, the system must issue a targeted clarification question before proceeding to search. The detector should look for direct denials of previously stated facts (e.g., Thursday works
followed by cannot do Thursday
).
Improvement: Add a preprocessing layer that separates actionable search constraints from conversational noise. This encoder should: (a) extract only facts matching the tool schema fields, (b) ignore off-topic background (e.g., phone has been buzzing
), (c) recognize unsupported requests (e.g., pay my bill
) as non-actionable and redirect to the underlying service need.
Improvement: Implement adaptive questioning that prioritizes missing constraints based on their impact on search results. The system should ask for schedule and intake before documents and eligibility, and should not repeat questions after receiving a no preference
answer. It should also recognize when additional turns won't improve accuracy and finalize with grounded recommendations.
Improvement: Add a post-search validation step that checks: (a) every selected resource id appeared in a previous tool result, (b) the selected resources satisfy all verified constraints, (c) for composite needs, each need has exactly one selected resource. If validation fails, the system must re-search or re-elicit rather than finalize.
Improvement: Detect user behavior patterns (impatience, rambling, contradiction) from early turns and adjust the interaction strategy: for impatient users, batch questions into fewer turns; for rambling users, explicitly confirm extracted facts before searching; for contradictory users, prioritize clarification before any tool call.
Capability summary: The improved system can reliably navigate social service conversations by maintaining faithful constraint state, systematically recovering from empty searches, detecting and resolving contradictions, filtering noise from rambling users, adapting to impatience, and ensuring all recommendations are grounded in actual database results. It would achieve significantly higher Tool EM (targeting >60% vs. current 38.75%) and Resource EM (targeting >80% vs. current 67.50%) across all benchmark conditions.
Sources
- Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction
- GLEN-Bench: A Graph-Language based Benchmark for Nutritional Health
- EvolveRouter: Co-Evolving Routing and Prompt for Multi-Agent Question Answering
- LongDA: Benchmarking LLM Agents for Long-Document Data Analysis
- SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment
- Non-Collaborative User Simulators for Tool Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- LLMs4All: A Review of Large Language Models Across Academic Disciplines
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support