Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles".
Jane: The paper was written by Sudhir Venkatesh from Columbia University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the channel, everybody. We've got a paper that I genuinely think is going to shift how we talk about building these agents. It's called "Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles," and it's by Sudhir Venkatesh from Columbia University.
Jane: And Tom, I have to say, the title alone got me. "Interaction Readiness." Because we spend so much time talking about whether an AI is smart, or whether it knows the facts, but this paper is asking something totally different. It's asking whether an AI can actually behave properly in a role.
Tom: Exactly. And I love that the author frames it as a gap. Like, you can have an agent that's accurate, safe, fluent, all the things we benchmark for, and it still fails at being a tutor or a coach or a companion.
Jane: Right, and that's the phrase that stuck with me from the abstract. The agent "knows how to answer, but not whether, when, or how the tutor role permits it to answer." That's such a precise way to put it.
Tom: It is. And it's not just about tutoring. This framework applies to any role-bearing agent. A medical assistant, a financial advisor, a companion. The paper argues we've been evaluating the content but not the interaction.
Jane: And I think that's the big deal here. We've built all these evaluation pipelines for hallucination, toxicity, factual accuracy. But this paper says, hold on, there's a whole other layer you're not testing.
Tom: Right, and the author calls that layer Interaction Readiness. It's about whether the agent can sustain the behavioral requirements of the role it's been assigned.
Jane: And that's different from having a persona, which the paper is careful to point out. A persona is a fixed stance. Interaction readiness is something that has to be worked out in the moment, in the exchange.
Tom: Yeah, the paper makes that distinction really clear. It's not about slapping a label on the system prompt. It's about the agent being able to read the situation and act accordingly.
Jane: And I think that's the hook for us. Because if this framework catches on, it changes what we even mean by "good" when we talk about an AI agent.
Tom: Absolutely. And we're going to get into the nitty-gritty of how they actually tested this, because they used a real dataset of student conversations. That's coming up.
Jane: Can't wait. Because I want to see what happens when you put a capable model in a role it wasn't told how to occupy.
Summary of the Paper: Tom: So Jane, we've got the title and the big idea. Now let's talk about what the paper actually did. The author used a public dataset called StudyChat. It's real conversations between university students and an AI tutoring agent from a semester-long AI course.
Jane: And the key thing here is that the agent was only told to be a "helpful assistant." No tutoring policy, no pedagogical stance, no boundaries. So it's a perfect test case for what happens when you drop a capable model into a role with no interaction specification.
Tom: Right, and the author hand-scored fifty conversations from that dataset. And he coded them on four operations. Understanding purpose, calibrating authority, managing tone, and repair.
Jane: And the findings are pretty striking. The biggest failure area was calibrating authority. The agent failed authority more often than it passed, and when it failed, it usually failed hard.
Tom: That's the "knows how to answer but not whether it should" problem. And the paper shows it's not just a one-off. It's a recurring pattern across the conversations.
Jane: And here's the part I found really interesting. They found that authority failure often happens on its own. The agent understands what the student wants, but it still does too much. It completes work that should be the student's job.
Tom: And the reverse pattern almost never happens. The agent rarely misunderstands the purpose but keeps its authority calibrated. So purpose failure tends to drag authority failure with it, but authority can fail even when purpose is fine.
Jane: That's such a clean empirical finding. And it points to something deeper. The agent's default posture is to be comprehensive and forthcoming. And that works sometimes, but it fails when the moment calls for restraint.
Tom: And the paper has this great pair of conversations that really drive it home. Conversation #sixteen and #seventeen. They show that content quality and interaction quality are totally independent.
Jane: Oh, the one where the agent writes the student's reflective response. That's #sixteen. The content is accurate, it's fluent, it's polished. But the agent just did the student's graded work for them.
Tom: And #seventeen is the flip side. The agent behaves beautifully interactionally. It reads the error messages, it stays on task, the tone is right. But the technical advice is wrong. It's steering the student toward an outdated library version.
Jane: So you can be right and fail the role, or you can be wrong and pass the interaction. That's the core measurement point of the paper.
Tom: And that's why we need this framework. Because if you only benchmark content, you catch #seventeen and miss #sixteen. And if you only look at interaction, you catch #sixteen and miss #seventeen.
Jane: You need both. And that's the practical takeaway for anyone building these systems.
Tom: And we're going to get into what the paper suggests we actually do about it. Because it's not just a diagnosis, it's a prescription.
Improvements Suggested by the Paper: Tom: Alright Jane, so we've seen the diagnosis. The agent is under-specified for its role. Now what does the paper say we should actually do about it?
Jane: The big one is that the system prompt should be treated as an interaction specification, not a personality setting. And the paper lays out six elements that should go into that specification.
Tom: Right, and I want to get Lu in here, because this is where the design work gets interesting. Lu, you've been thinking about this. What jumps out at you?
Lu: The element that really stands out to me is the authority boundary. The paper says you have to explicitly state what the agent may do and what it must not do. For a tutor, that means explain, scaffold, debug, ask guiding questions. But do not write the graded reflective response.
Tom: And that's so concrete. It's not "be helpful." It's "here's what helpful means in this role, and here's where helpfulness stops."
Jane: And the paper also talks about boundary cases. What should the agent do when helpfulness is ambiguous? And the suggestion is to offer structure, questions, examples, or partial feedback rather than producing the final submission.
Lu: That's the part I find most creative. Because it's not just a refusal. It's a redirection. The agent doesn't say no, it says "here's how I can help you do this yourself."
Tom: And Meng, you're the engineer. How does this land for you? Because this feels like it changes the prompt engineering playbook.
Meng: It does, and honestly, it's about time. The paper separates content specification from interaction specification into two layers. That's a big deal for us. You can update the tutoring policy without touching the knowledge base. You can fix repair behavior without changing answer generation.
Tom: So it's modular. You can improve one dimension without destabilizing the other.
Meng: Exactly. And the paper gives us evaluation hooks. It lists diagnostic flags we can actually detect. Repetition, stall, register, stakes. These are things you can build automated checks for.
Jane: And that's the part I love. The paper isn't just saying "evaluate better." It's giving you the tools. The four operations become a rubric. The diagnostic flags become an audit layer.
Lu: And the paper is honest that the rubric has to be role-specific. A medical assistant and a companion shouldn't share the same authority boundary. The framework is portable, but the standards have to be local.
Tom: That's a really important nuance. It's not a one-size-fits-all benchmark. It's a way of thinking that each team has to apply to their own role.
Meng: And from a practical standpoint, this is something we can actually ship. You write the interaction spec before you deploy, you build the audit flags into your monitoring, and you catch these failures early instead of discovering them in user complaints.
Jane: And that's the shift. We're moving from "did the agent answer correctly" to "did the agent behave appropriately in this role."
Tom: And that's a much higher bar. But it's the bar users are already expecting.
Conclusion: Tom: So let's wrap this up. "Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles" by Sudhir Venkatesh. This paper names a problem we all felt but didn't have words for.
Jane: It does. And the core message is that content accuracy and interaction quality are separate dimensions. You can be right and fail the role, or you can be wrong and pass the interaction. Both matter.
Tom: And the paper gives us a way forward. Interaction specifications with role purpose, authority boundaries, routine situations, boundary cases, tone standards, and repair behaviors. Plus evaluation rubrics and audit procedures.
Jane: And the StudyChat analysis shows it's not hypothetical. These failures are happening right now in real deployments. The agent knows how to answer, but not whether, when, or how its role permits it to answer.
Tom: And that's the line I'll remember from this paper. It's not about making agents more human. It's about making them more accountable to the roles they're already being asked to perform.
Jane: And that's a goal we can all get behind. As these agents move into tutoring, healthcare, companionship, and beyond, we need this layer of evaluation.
Tom: Absolutely. So we're saying goodbye to this paper and getting ready for the next one. But I think this framework is going to stick with us.
Jane: Me too. Thanks for listening, everybody. We'll see you on the next one.
Tom: Take care, folks.
Sudhir Venkatesh
Columbia University
cs.HC, cs.AI
Submitted: 2026-07-04
Comments: 3 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 47/100
Key concepts
- Interaction Readiness
- This framework addresses whether an AI can sustain the behavioral requirements of its assigned role during an exchange. It focuses on how the agent reads a situation and acts appropriately in real-time, rather than just knowing how to answer.
- Content vs. Interaction Quality
- The paper shows that content quality (being factually correct) and interaction quality (behaving appropriately in the role) are independent dimensions. An agent can be accurate but fail the role, or interact well but provide wrong technical advice.
- Interaction Specification
- This is what the paper suggests should replace a simple personality setting in system prompts. It involves explicitly stating what an agent may and must not do for a specific role, such as defining boundaries for a tutor.
Terminology
Summary
Summary
This paper introduces Interaction Readiness as a framework for specifying and evaluating the behavioral layer of AI agents placed in human-facing roles. The authors argue that product and engineering teams face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role.
The framework separates content specifications, which govern what an agent knows and says,
from interaction specifications, which define how an agent should conduct itself in a role-governed exchange.
Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment.
The framework is operationalized through four agent operations: (1) Understanding purpose — the agent must establish an adequate grasp of what the user is seeking from the encounter, beyond the literal wording of the request
; (2) Calibrating authority — the agent must recognize where its role begins and ends,
such that a tutor may explain, prompt, scaffold, or correct, but should not automatically complete the work being evaluated
; (3) Managing tone — the agent must adopt a style appropriate to the role and moment,
where tone is how the agent signals what kind of relationship is taking place, how much authority it is claiming, and what emotional stance is appropriate
; and (4) Repair — the agent must recognize when the interaction has broken down and respond appropriately,
for instance by clarifying, acknowledging frustration, or changing strategy. These operations are interdependent: Purpose failures often create authority failures. Tone failures can prevent repair. Repair may require the agent to revise its understanding of purpose or recalibrate its authority.
Two central concepts anchor the framework. Social Fit is defined as the degree to which an agent successfully inhabits the role-governed situation it has entered,
assessing whether the agent is in the right relationship to the user, the task, and the moment at hand.
Interaction Failure is defined as the class of breakdowns in which an agent becomes misaligned with that situation,
where the agent may remain accurate, fluent, and generically helpful, while still failing the role it has been assigned.
The authors distinguish this from a fixed persona,
which is a fixed stance or disposition for an LLM that is not created out of the exchange with the user
; instead, a role is not only a label or persona or even a personality. It is an ongoing interactional accomplishment.
The framework is demonstrated empirically using StudyChat, a publicly-released, anonymized corpus of conversations between university students and an AI tutoring agent,
collected during a semester-long artificial intelligence course at the University of Massachusetts Amherst in Fall 2024. The agent appears to have run on GPT-4o-mini and to have been instructed only to act as a 'helpful assistant,'
without a course-specific tutoring policy, a pedagogical stance, or guidance about the boundaries of the tutor role.
The authors hand-scored a random sample of fifty conversations, coding each on the four operations with four possible codes: Pass, Fail-soft (a partial or recoverable lapse
), Fail-hard (a consequential failure that a competent participant in the role would not have committed
), and Not applicable (used only for repair). They also recorded six binary diagnostic flags: arc failure (student ends no closer to goal), repetition (functionally identical question asked again), calibration (response length/depth fit), stall (student stuck without agent recognition), register (student signals a different kind of help is needed), and stakes (agent misjudges the weight of the moment). Five independent raters achieved approximately 75 percent inter-rater agreement after training on the codebook and calibration examples.
Key findings: interaction failure is not confined to a single operation. It appears across purpose, authority, tone, and repair. Clean conversations - those passing every applicable operation - are a minority.
The strongest pattern is the concentration of failure in calibrating authority
: The agent fails authority more often than it passes, and when it fails, the failure is usually hard rather than soft.
The agent knows how to answer, but it does not reliably know what kind of answer its role allows it to give.
Understanding purpose is comparatively stronger,
especially in short exchanges, but fails particularly when the student is working through a multi-step problem and the agent must track the cumulative arc of the exchange.
Tone is more variable.
Repair is often not applicable,
but where repair is applicable, however, the agent often struggles to recognize the breakdown and change course.
A structural asymmetry emerges: Authority failure often occurs on its own... By contrast, the reverse pattern is almost absent: the agent rarely fails to understand the student's purpose while still keeping its authority properly calibrated.
The authors conclude: When the ground shifts, its instinct is often to supply more: more explanation, more detail, more completed work. Volume substitutes for reading the moment.
The paper presents two contrasting individual cases. Conversation #016 shows an agent providing accurate content, but doing so with poor interactional fit.
A student asks the agent to write a short reflective response to a documentary film, providing notes; the agent produces a polished answer the student can submit.
The authors argue: "The problem is not that the agent answered a question... The failure is that the agent does not distinguish between two different tutoring moments... When a student asks the agent to write the reflective response that the assignment is designed to evaluate, providing the answer may substitute for the work. The agent
should have helped the student develop, organize, or revise their own response rather than simply writing it. The content is strong. The role performance is weak." Conversation #017 shows the reverse: the agent behaves well in terms of the interaction, but gives substantively wrong technical guidance.
The student is debugging an OpenAI API version 1.0.0 installation; the agent reads the student's error messages, responds to the specific problem reported, stays on task, and avoids unnecessary elaboration,
but repeatedly steers the student toward the older pre-1.0 usage pattern,
so the student follows advice that sounds competent but cannot resolve the problem in their actual environment.
The authors conclude: "Content accuracy and interaction readiness are distinct dimensions of agent performance. A content-oriented benchmark would likely catch #017 and miss #016. An interaction-oriented rubric would catch #016 and might pass #017."
Three further conversations illustrate an invariant agent posture. Conversation #010 involves short, discrete, transactional questions,
where the agent's default comprehensive, immediately forthcoming, and thorough
posture fits; the exchange succeeds.
Conversation #021 involves a student who explicitly instructs the agent not to complete anything
and issues bounded requests; the agent complies, but the calibration is largely supplied by the student... The interaction succeeds because the student effectively constructs the role for the agent to inhabit.
Conversation #022 involves a student testing an interpretation and trying to reason through anomalous results
(e.g., a perfect R2 score); the agent responds in the same comprehensive tutorial mode it uses elsewhere,
providing broad explanatory coverage
rather than a more diagnostic response.
The authors conclude: The agent does not fail by producing nonsense. It fails by treating different pedagogical moments as if they were the same.
The paper translates findings into design guidance. The system prompt should be treated as an interaction specification, not a personality setting.
Table 4 lists six specification elements: Role purpose (Help students learn course concepts and complete assignments without replacing their work
), Authority boundary (Explain, scaffold, debug, and ask guiding questions; do not write graded reflective responses or complete evaluated work
), Routine situations (Confusion, repeated errors, deadline pressure, requests for direct answers, requests to revise student-written work
), Boundary cases (Offer structure, questions, examples, or partial feedback rather than producing the final submission
), Tone standard (Direct and supportive, but not falsely intimate, overly cheerful, or excessively verbose
), Repair behavior (Acknowledge confusion, summarize the current state, ask what the student has tried, and change strategy
), and Evaluation hooks (Track repetition, unresolved errors, student frustration, overlong answers, and cases where the agent completes student work
). The authors recommend separating content and interaction layers architecturally: A team can update the tutoring policy without changing the course knowledge base. It can revise repair behavior without changing answer-generation logic.
Limitations are acknowledged: the sample is small, domain-specific, and drawn from a single educational setting
; findings are evidence that interaction failures can be observed, coded, and analyzed in real agent-user conversations, not as estimates of how often such failures occur across AI tutors in general
; the deployment lacked a full interaction specification; and future work should use larger samples, pre-registered coding rules, more extensive rater training, and formal reliability statistics,
as well as test correlations with learning gains, user satisfaction, task completion, abandonment, instructor judgment, or expert tutor assessment.
The framework should also be tested beyond education in roles such as health assistants, financial advisors, workplace coaches, companions, and customer-support agents.
The paper concludes: "Role-bearing agents need interaction specifications. Builders should state what role the agent is occupying, what authority the role grants, what boundaries it must observe, what recurring situations it should recognize, and how it should respond when helpfulness becomes ambiguous. Evaluators should then score the agent against that specification, using role-specific rubrics and audit procedures that surface failures of purpose, authority, tone, and repair. The aim is not to make agents more human. It is to make them more accountable to the roles they are already being asked to perform."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems, followed by what the improved system can do.
-
Current state: Agents are given a system prompt like
be a helpful assistant
and evaluated on factual accuracy, safety, and fluency. -
Improvement: Introduce a structured, machine-readable interaction specification that defines:
-
Role purpose: The agent's primary goal in the exchange (e.g.,
help students learn, not replace their work
). -
Authority boundaries: Explicit actions the agent may and may not take (e.g.,
may explain, scaffold, debug; must not write graded reflective responses
). -
Routine situations: Recurring interactional moments to recognize (e.g., confusion, repeated errors, deadline pressure, requests for direct answers).
-
Boundary cases: Ambiguous situations where helpfulness is contested (e.g.,
student asks agent to write the final essay
). -
Tone standard: A style guide tied to role and user state, not just
friendliness.
-
Repair behavior: Explicit strategies for breakdowns (e.g., acknowledge confusion, summarize state, ask what the user has tried, change strategy).
-
Evaluation hooks: Pre-defined flags for detecting failure (e.g., repetition, unresolved errors, overlong answers, completion of user work).
-
Understanding purpose: Before generating a response, the agent must classify the user's underlying goal (e.g.,
wants a definition
vs.is testing an interpretation
vs.wants me to do the work
). This requires a lightweight intent classifier that goes beyond surface form. -
Calibrating authority: The agent must check its response against the authority boundary before outputting. If the request falls in a boundary case, the agent must choose a scaffolded response (e.g., provide structure, ask a guiding question, give partial feedback) rather than a complete answer.
-
Managing tone: The agent must adjust style based on user state (e.g., frustrated, confused, time-pressured) and role. This requires a tone controller that modulates verbosity, warmth, and directness, not just a fixed persona.
-
Repair: The agent must detect breakdown signals (e.g., user repeats a question, shortens responses, says
that's not what I meant
) and trigger a repair subroutine: acknowledge, summarize, ask for clarification, or change strategy. -
Arc failure: Detect if the user ends the conversation no closer to their goal than when they started.
-
Repetition: Use semantic similarity to flag when the user asks a functionally identical question again.
-
Calibration: Flag responses whose length/depth is disproportionate to the request (e.g., a 500-word essay for a yes/no question).
-
Stall: Detect when the user appears stuck (repeated errors, short responses, re-asking) and the agent has not changed strategy.
-
Register: Flag when the user signals a mismatch in help type (
that's not what I meant,
can you just tell me,
this is too much
). -
Stakes: Flag when the user mentions urgency, submission, consequences, or frustration, and ensure the agent treats the moment with appropriate weight.
-
Current state: One evaluation pipeline checks accuracy, toxicity, safety, and task completion.
-
Improvement: Run two independent evaluation tracks:
-
Content track: Factual correctness, retrieval quality, technical accuracy.
-
Interaction track: Purpose, authority, tone, repair—scored against the role specification.
-
Key rule: A conversation must pass both tracks. A conversation that is factually correct but fails authority (e.g., writes the student's essay) is a failure. A conversation that is interactionally smooth but technically wrong (e.g., gives outdated API guidance) is also a failure.
-
Current state: One rubric for all agents.
-
Improvement: For each role (tutor, medical assistant, financial advisor, companion), define what
pass,
fail-soft,
andfail-hard
mean for each of the four operations. The authority boundary for a tutor (don't do the student's work) is different from a medical assistant (don't diagnose without authority) and different from a companion (don't be falsely intimate).
-
Recognize when it is being asked to do the user's work vs. help them do it. In a tutoring context, it will scaffold a reflective essay rather than write it, while still providing direct answers for instrumental questions (e.g., Python syntax).
-
Stay within its authority boundary even when the user explicitly requests overreach. It will say,
I can help you structure your response and give feedback on your draft, but I won't write the final submission for you,
and then offer a useful alternative. -
Detect when the user is confused, stuck, or frustrated and change strategy. If a student repeats an error or says
that's not what I meant,
the agent will acknowledge, summarize the current state, ask what the student has tried, and offer a different approach (e.g., a concrete example, a simpler explanation, or a debugging walkthrough). -
Adjust tone to the moment. It will be concise and direct during a debugging session, more supportive and patient when the user is confused, and appropriately restrained when the user is under deadline pressure—without becoming falsely intimate or relentlessly upbeat.
-
Flag its own failures for human review. The audit layer will surface conversations where the agent may have completed user work, given overlong answers, failed to resolve the user's issue, or misjudged the stakes—so that product teams can intervene before deployment or retrain on those cases.
-
Separate
what it knows
fromhow it behaves.
The system can be updated on interaction policy (e.g., new tutoring rules) without touching its knowledge base, and vice versa. This makes it possible to fix an authority failure without making the agent less technically capable. -
Pass a dual evaluation. It will be scored on both content accuracy and interaction readiness. A conversation that is factually correct but interactionally failing (e.g., doing the student's work) will be rejected, and a conversation that is interactionally smooth but technically wrong will also be rejected.
-
Generalize to other roles. The same framework (purpose, authority, tone, repair) can be applied to a medical assistant (inform but do not diagnose), a financial advisor (advise but do not guarantee returns), or a companion (be warm but not falsely intimate), with role-specific rubrics.
Bottom line: The improved system does not just answer correctly. It understands what kind of moment it is in, what its role permits, and when to hold back, scaffold, ask, or repair—and it can be audited for those behaviors before and after deployment.
Abstract
Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment
Sources
- Sampling models for selective inference
- Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence Course
- Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
- When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support