DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow".
Jane: The paper was written by Tao Long, Xuanming Zhang, Sitong Wang, Zhou Yu and Lydia B. Chilton from Columbia University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: We're starting today with a really interesting paper called "DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow."
Jane: That's quite a mouthful, Tom, but it basically explores how we can make AI understand the subtle social rules we follow every day.
Tom: Right, and this research comes from a brilliant team at Columbia University including Tao Long and Xuanming Zhang.
Jane: They're looking at how to move past simple commands and toward an agent that understands context, like when you're organizing a seminar series.
Lu: This is such a massive leap for the field because most AI today lives in a vacuum without any social awareness. If an agent can navigate something as delicate as human relationships, we're looking at a whole new category of digital companions that actually understand us.
Meng: I'm going to be the skeptic here and ask if they can actually handle the messiness of real people in these workflows. Social rules change constantly, and a system that doesn't account for sarcasm or sudden changes in plans is going to fail miserably when it hits real production environments.
Lalam: But Meng, that's exactly where alignment comes in to bridge that gap between logic and sociality. We aren't just teaching them commands; we are helping them adopt our unique cultural styles and ways of communicating.
Tom: So how do they actually achieve that without us having to explain every single preference from scratch?
Jane: They use a very clever framework grounded in distributed cognition, which we should explore next.
Title: Tom: The team uses a concept called distributed cognition to share the mental load between the human and the AI.
Jane: It's like having a partner that tracks all the boring details while you make the important decisions.
Tom: They've built this three-part system with a coordination agent, an interactive dashboard, and a policy module.
Jane: The dashboard is especially cool because it makes the agent's reasoning visible so you aren't just guessing what it's thinking.
Lu: I think the simulation mode is the real game-changer in this setup for testing social outcomes. You can run a rehearsal with simulated people to see how your policies work before any real social stakes are involved.
Meng: That makes perfect sense from an engineering standpoint because it provides a safe sandbox for testing everything. You can see if your logic holds up when a simulated speaker asks for something completely unexpected or breaks a rule.
Lalam: And that rehearsal process builds a beautiful layer of trust within our communities. When people see that the AI is transparent and follows established norms, they'll feel much more comfortable delegating tasks to it.
Tom: Does this system actually get better as you use it, or do you have to keep fixing it manually?
Jane: It actually learns from you through those reusable artifacts they mentioned.
Title: Tom: The research really highlights how the system gets smarter by turning your manual edits into permanent policies and templates.
Jane: It's not just about fixing one mistake, but actually teaching the agent your long-term preferences.
Tom: They even use these things called "stop hooks" to make sure the AI pauses and asks for help whenever it hits something uncertain.
Jane: That prevents those awkward moments where an AI might say something totally inappropriate to a professor or a client.
Tom: Their technical evaluations were also quite telling, especially regarding how the agent processes information.
Jane: They found that providing a high-level summary of the task state is much more effective than just dumping a massive amount of raw data into the prompt.
Tom: It's like giving a person a briefing instead of just handing them a stack of unorganized papers.
Jane: Exactly, and it helps the agent select the right policy much more accurately for the current situation.
Tom: They also looked at how well the system identifies potential problems, or edge cases.
Jane: It turns out that giving the model access to existing policies makes its detection much more reliable.
Lu: I find those stop hooks to be one of the most profound parts of this whole architecture. They create a moment of intentional pause for human guidance, which keeps us in the driver's seat.
Meng: I agree with Lu, and from an implementation standpoint, it's a very robust way to handle errors. Instead of the system failing silently, it flags the issue and waits for us to define the new rule.
Lalam: This creates such a beautiful cycle where every correction helps refine our digital coexistence. We aren't just fixing bugs; we are actually helping the AI understand our cultural nuances over time.
Tom: It really bridges that gap between machine logic and human sociality.
Jane: And it does so in a way that feels very safe and predictable for the user.
Lu: I can't wait to see how these policies evolve as they are shared across different communities, Tom.
Meng: That evolution is exactly what makes this so useful for engineers, because it means the system learns from its own context.
Tom: : "--- CONCLUSION ---"
Conclusion: Tom: We've really seen how "DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow" provides a roadmap for better AI partners.
Jane: It's about moving from simple automation to true social coordination that we can actually trust.
Tom: The combination of simulation and explicit policies is such a practical way to build that confidence.
Jane: I think it shows that we don't have to give up control to get the benefits of delegation.
Tom: It's a massive shift in how we think about the relationship between human intent and machine execution.
Jane: Exactly, because instead of fighting the AI, we are actually teaching it our specific ways of working.
Lu: I can imagine these agents becoming part of the very fabric of how we build communities online. They could manage everything from academic seminars to neighborhood organizing with total social grace.
Meng: If we keep prioritizing this kind of safety and testing, real-world deployment will be much smoother for everyone involved. I'm looking forward to seeing how these "stop hooks" become a standard feature in every agentic system.
Lalam: It's a vital step toward ensuring technology respects our culture and our need for agency. We are building tools that don't just work, but actually belong in our lives.
Tom: That is such a perfect way to end on, Lalam.
Jane: I think we've all learned a lot about how much control we can actually afford to hand over when the system is transparent and predictable.
Tom: We definitely have, but we're running out of time for this session, and our next paper is a complete departure from everything we just discussed.
Jane: It's a total curveball, so don't go anywhere!
Columbia University
cs.HC, cs.AI, cs.CY, cs.ET
Submitted: 2025-09-16
Updated: 2026-09-14
Comments: 22 pages, 6 figures. ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: This paper introduces DoubleAgents, an interactive system designed for "human–agent alignment in coordination tasks" that are "socially embedded." It addresses the challenge of delegating complex
Key concepts
- Distributed Cognition
- A concept used to share the mental load between a human and an AI agent. Instead of the human doing everything, the system acts as a partner that tracks details while the human makes important decisions, allowing for more effective collaboration in complex workflows.
- Stop Hooks
- These are mechanisms that cause an AI to pause and ask for human guidance whenever it encounters uncertainty. This prevents inappropriate actions and allows users to define new rules, ensuring the system remains safe, predictable, and under human control during social interactions.
- Simulation Mode
- A feature that provides a safe sandbox for testing social outcomes. Users can run rehearsals with simulated people to see how different policies work and how the AI handles unexpected requests or rule-breaking before applying them in real-world, high-stakes social environments.
Terminology
Summary
This paper introduces DoubleAgents, an interactive system designed for human–agent alignment in coordination tasks
that are socially embedded.
It addresses the challenge of delegating complex workflows where user preferences are often implicit, evolving, and difficult to specify upfront,
providing a framework that enables comfortable offloading
while ensuring users maintain necessary control.
The Distributed Cognition Framework
DoubleAgents is grounded in a distributed cognition approach, distributing the cognitive load of coordination across three integrated components:
-
A coordination agent that
maintains state and proposes plans and actions,
serving as a reasoning partner and memory for the user. -
A dashboard visualization that makes the agent’s reasoning
legible for user evaluation
through temporal, social, and procedural views. -
A policy module that transforms user edits into
reusable alignment artifacts,
which include coordination policies, email templates, and stop hooks.
Iterative Alignment via Simulation
The system utilizes a dual-mode approach involving simulation for rapid iteration and live deployment for real-world coordination.
By using LLM-based simulated respondent agents, the system allows users to compress weeks into minutes
before actual social stakes are involved. This simulation mode enables users to:
-
Review proposed plans and approve actions in a safe environment.
-
Iteratively refine policies and templates based on how the agent responds to different simulated personas.
-
Build confidence in the system's robustness by observing how it handles
potential edge cases
before moving to live deployment.
Technical Implementation and Evaluation
The system's performance is driven by two core technical mechanisms: policy selection and edge case detection. Technical evaluations conducted using various large language models revealed:
-
Policy selection is most effective when utilizing
summary-based context,
which provides astructured, high-level snapshot of the system
rather than raw, unabstracted data. -
Edge case detection—the ability to identify situations that fall
outside the scope of existing coordination policies
—is most accurate when employingfew-shot prompting with policy context.
This allows the agent to reliably determine whether a scenario requires escalation or can be resolved via existing rules.
Empirical Results and User Experience
A two-day lab study and three real-world deployments demonstrated that participants' comfort in offloading tasks and reliance on DoubleAgents both increased over time.
The research identified several key outcomes of the alignment process:
-
The use of policies and templates helps users
encode their values,
which significantly reduces the effort required for subsequent tasks. -
Visualizations support
confident evaluation
by allowing users to verify agent reasoning at a glance. -
Stop hooks are essential for
maintaining safe delegation boundaries,
as they allow the system to pause and ask for guidance during points of uncertainty.
Ultimately, the study shows that alignment established in simulation transfers effectively
to live coordination, allowing users to manage complex tasks like organizing seminar series with significantly reduced cognitive overhead.
Improvements for AI systems
1. Implementation of a Policy-Driven Distributed Cognition Architecture
-
Improvement: Replace monolithic instruction-following prompts with a modular system that separates State Tracking, Policy Retrieval, and Action Execution. This uses a dedicated
Progress Summary
module to transform raw, noisy logs into high-abstraction, structured JSON snapshots. -
Capabilities: The AI can maintain high reasoning accuracy in long-horizon tasks by avoiding context saturation. It will select specific behavioral rules (e.g.,
prioritize faculty over students
) from a library of natural language policies rather than attempting to infer complex social norms from raw data alone.
2. Integration of a High-Fidelity Persona-Based Simulation Layer (Rehearsal Mode)
-
Improvement: Integrate an LLM-based multi-agent simulation engine that uses diverse, persona-driven respondent agents (e.g.,
the busy professor,
the non-responsive student
) to runstress tests
on the agent's current policy set before real-world deployment. -
Capabilities: The AI can proactively surface edge cases and potential social faux pas in a zero-stakes environment. This allows the system to identify gaps in its reasoning—such as how to handle a request for remote participation—before any actual communication is sent to real stakeholders.
3. Transformation of User Interventions into Reusable Alignment Artifacts
-
Improvement: Implement a
Policy Module
that automatically translates user-initiated edits (to email tone, plan structure, or scheduling logic) into three distinct, persistent artifacts: Coordination Policies, Communication Templates, and Stop Hooks. -
Capabilities: The AI achieves
long-term social intelligence.
Instead of merely correcting an error once, the system updates its internal rulebook. If a user corrects an email tone once, the system learns the specific style and applies it to all future interactions; if a user resolves an edge case (e.g.,Zoom is okay
), that decision is codified into a policy that prevents future unnecessary escalations.
4. Deployment of Uncertainty-Triggered Stop Hooks
for Safe Delegation
-
Improvement: Develop an Edge Case Detector that compares incoming information against the current active policy library to quantify
policy coverage.
When a scenario is detected as falling outside existing coverage, the system triggers a mandatory execution pause (a Stop Hook). -
Capabilities: This prevents
hallucinated autonomy
and social errors. The AI knows exactly when it is unqualified to act, transitioning from an autonomous agent to a collaborative partner by presenting the user with a concise, actionable question (e.g.,Speaker X requested a change; should I accommodate or decline?
) rather than making an unaligned decision.
5. Semantic State Summarization for Enhanced Policy Selection
-
Improvement: Implement a specialized summarization pipeline that converts temporal, social, and procedural data into a
Summary-Based Context
specifically optimized for policy retrieval. -
Capabilities: The system can perform highly accurate policy selection (as evidenced by the paper's F1 score improvements), ensuring the agent always applies the correct behavioral logic to the current state of a complex, multi-turn workflow.
Abstract
Aligning agentic AI with user intent is critical for delegating complex, socially embedded tasks, yet user preferences are often implicit, evolving, and difficult to specify upfront. We present DoubleAgents, a system for human-agent alignment in coordination tasks, grounded in distributed cognition. DoubleAgents integrates three components: (1) a coordination agent that maintains state and proposes plans and actions, (2) a dashboard visualization that makes the agent's reasoning legible for user evaluation, and (3) a policy module that transforms user edits into reusable alignment artifacts, including coordination policies, email templates, and stop hooks, which improve system behavior over time. We evaluate DoubleAgents through a two-day in-lab interactive simulation study (n=10), three real-world deployments, and a technical evaluation. Participants' comfort in offloading tasks and reliance on DoubleAgents both increased over time, correlating with the three distributed cognition components. Participants still required control at points of uncertainty - edge-case flagging and context-dependent actions. We contribute a distributed cognition approach to human-agent alignment in socially embedded tasks. We further introduce interactive simulation as a methodological testbed for rapid iteration and alignment testing of agentic systems.
Sources
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
- Constitutional AI: Harmlessness from AI Feedback
- Blind Judgement: Agent-Based Supreme Court Modelling With GPT
- Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?
- Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
- Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
- Position: Towards Bidirectional Human-AI Alignment
- Interactive AI Alignment: Specification, Process, and Evaluation Alignment
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- TravelPlanner: A Benchmark for Real-World Planning with Language Agents
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- Towards Aligning Personalized Conversational Recommendation Agents with Users' Privacy Preferences
- NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support