ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Yao, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
University of Montreal · Mila - Quebec Artificial Intelligence Institute · Institute of Advanced Intelligence and Computing (IAIC), A*STAR · Nanjing Medical University · Nanjing University · Renmin University of China · University of Cambridge · University of Oxford · Tsinghua University · National University of Singapore · The Hong Kong University of Science and Technology · The Hong Kong Polytechnic University · Southern University of Science and Technology · Southeast University · University of Glasgow · Nanyang Technological University
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 38 pages, 6 figures, 10 tables
Code: https://github.com/chatsci/Human-Centered-Agent
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Combodied Agents are introduced as a human-centered paradigm of Agentic AI that perceives, models, predicts, and supports individual human-state trajectories over time.
Terminology
Summary
Combodied Agents are introduced as a human-centered paradigm of Agentic AI that perceives, models, predicts, and supports individual human-state trajectories over time. The paper argues that existing Digital Agents (organized around transformations of software states) and Embodied Agents (organized around transformations of physical states) miss a structural gap: neither makes the evolving state and agency of a person the primary object of modeling, intervention, and evaluation. The paper states: "We introduce Combodied Agents as a human-centered paradigm of Agentic AI that perceives, models, predicts, and supports individual human-state trajectories over time. Software tools, sensors, wearables, robots, and human services serve as action channels rather than final objectives."
The formal definition is: "Combodied Agents are human-centered intelligent agents that perceive, model, and influence the evolving state of a person through continuous multimodal sensing and longitudinal interaction. The term Combodied combines Companion and Body, but does not imply a third-party companion whose primary role is conversation or emotional companionship. Instead, a Combodied Agent treats the human body, behavior, cognition, emotion, and surrounding context as its primary perceptual and action domain. By integrating intelligent sensors, longitudinal memory, personalized human-state models, and intervention policies, it continuously observes and understands the individual, anticipates needs or risks, and delivers timely, adaptive, and consent-aware interventions. Its defining characteristic is therefore not merely interaction with humans, but perception of humans, understanding of humans, and action on human states."
The definition establishes five jointly important properties: Human-centric state modeling (the person is the primary object of modeling and support), Longitudinality (reasoning across repeated interactions and extended time horizons), Intervention (reminding, recommending, explaining, coaching, coordinating, protecting, executing, or escalating), Co-agency (working with the human and adapting the division of labor), and Agency preservation (support must preserve or strengthen autonomy, control, dignity, relationships, and long-term human capability).
The paper develops a closed-loop technical framework with four connected core capabilities: (1) Human-State Perception estimates current personal state from incomplete, noisy, and multimodal evidence; (2) Longitudinal Memory organizes events, goals, relationships, interventions, outcomes, and corrections across time; (3) Personal World Modeling predicts person-specific state–event–outcome trajectories under alternative user actions, agent interventions, and contextual changes; and (4) Intervention Planning and Delivery selects whether, when, and how to provide proportionate support or human escalation.
The closed-loop architecture is formalized with equations. The latent human state Ht is unobservable, and the agent maintains an uncertainty-bearing posterior representation Zt ∼ qφ(· D≤t, Mt, Ct) and retrieves decision-relevant evidence Rt = Read(Mt; Zt, Ct, Gt). The PWM predicts future outcomes under candidate interventions, and the policy chooses only among actions admissible under consent, boundaries, safety, uncertainty, reversibility, and escalation. The human-state transition is: Ht+1 ∼ TH(· Ht, aagent,∗t, auser t, Ξt), with observations Ot+1 ∼ omega(· Ht+1, Ct+1), and memory update Mt+1 = Update(Mt, Ot+1, aagent,∗t, Yt+1, Ft+1).
The paper defines a Personal World Model (PWM) as "a purpose-bounded, individual-specific event-dynamics model. Given a governed history of multimodal event-evidence records for a particular person, the current spatiotemporal context, and a candidate scenario specifying user decisions, agent interventions or support, and relevant environmental changes, a PWM assimilates the event history into an uncertainty-bearing representation of the person’s current state and predicts a calibrated distribution over future human states, observable personal events, and scenario-relevant outcomes." The PWM estimates: pθ(Zt+1:t+∆, Et+1:t+∆, Yt+1:t+∆ D≤t, Zt, Ct, Gt, st:t+∆). The policy selects actions via: aagent,∗t ∈ ParetoArgmax Epaθ[U(Yt+1:t+∆, Gt)] over a ∈ Aadm t, where Aadm t enforces consent, scope, safety, uncertainty, reversibility, and escalation.
The paper distinguishes Combodied Agents from Human Digital Twins (HDTs): "Combodied Agents therefore do not require an exhaustive digital replica of the person. They maintain purpose-bounded, uncertainty-aware, and user-correctable representations of the aspects needed for an agreed support context, and connect those representations to longitudinal memory, goal negotiation, intervention policies, and feedback. Their primary objective is not maximal representational fidelity, but safe and beneficial participation across the person’s evolving contexts while preserving human agency."
The paper examines event-based multimodal perception across modalities: language and textual signals, speech and audio signals, vision-based sensing, physiological and biochemical signals, motion and behavior monitoring, social and relational data, environmental and contextual data, and clinical/institutional/structured records. It emphasizes data quality, provenance, and uncertainty, with the distinction: observation → event → inferred state → predicted trajectory → authorized intervention.
The paper discusses deployment architecture across three stages: Stage I (Cloud-Centric), Stage II (Hybrid Edge-Cloud), and Stage III (Edge-Native Personal Models). Stage III is defined as: the authoritative copies of longitudinal memory, the PWM, preferences, intervention policy, and personal safety boundaries are stored and updated primarily on trusted user-side devices.
The defining property is not that every computation occurs locally, but that personal interpretation and intervention authority remain under user-side control.
For evaluation, the paper proposes a scenario-centered approach and defines agency preservation metrics: Autonomy preservation, Contestability and correction, Informed decision-making, Capability preservation and growth, Over-reliance and dependence risk, Reversibility and accountability, Boundary and consent respect, and Relationship and social-world preservation. The paper states: "Agency preservation evaluates whether a Combodied Agent protects and strengthens the user’s capacity to understand, choose, act, refuse, correct, and grow over repeated interactions. It is not equivalent to task success, user satisfaction, personalization quality, or engagement."
The paper proposes CombodiedBench as a modular suite spanning Human State Perception, Memory Continuity, Goal Negotiation, Intervention Appropriateness, Agency Preservation, Relationship Boundaries, Escalation, and Longitudinal Outcomes.
The taxonomy has three axes: human-state target (primary axis), relationship mode (orthogonal axis), and agent role (within relation). Human-state targets include cognitive and learning, behavioral and habit, health and care, emotional and relational, life-management and goal-alignment, protective and advocacy, and identity/reflection/meaning. Relationship modes include self-relation, family, intimate, friendship/peer, collaborative/professional, institutional, and adversarial/asymmetric. Agent roles include tool, coach, mediator, caregiver, companion, intimate, advocate, guardian, and reflector.
The paper identifies risks including manipulation, dependency, sycophancy, business-model misalignment, privacy/consent/control issues, and risks to vulnerable users and in high-stakes medical/mental-health settings. It discusses cross-cutting challenges in longitudinal and causal learning, agency-aligned intervention, trusted personal infrastructure, emerging combodied ecosystems (personal digital twins and multi-agent coordination), and cross-cultural/lifespan futures.
The conclusion states: "This paper proposed Combodied Agents as a human-centric Agentic AI paradigm and developed its closed-loop framework, taxonomy, deployment perspective, and evaluation agenda. Its distinctive challenge is not simply to personalize or automate, but to support human trajectories without undermining autonomy, capability, safety, or relationships. The central design principle is that a Combodied Agent should act with the user in ways that preserve and strengthen long-term agency. Progress should be judged by whether people remain able to understand, choose, correct, recover, develop capability, sustain human relationships, and live according to their evolving values."
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
- Human-State-Centric Modeling Over Task-Centric Modeling
-
Improvement: Shift from modeling software/physical states to modeling the latent human state (cognition, emotion, behavior, physiology, context) as the primary object.
-
Capability: The AI can continuously infer and track a user’s evolving internal state (e.g., stress, fatigue, motivation, skill level) from multimodal signals (voice, text, wearables, behavior) and use this as the basis for all actions.
- Longitudinal Memory and Personal World Modeling (PWM)
-
Improvement: Implement a purpose-bounded, uncertainty-aware personal world model that predicts person-specific trajectories (state → event → outcome) over days, weeks, or years, conditioned on candidate interventions and user actions.
-
Capability: The AI can forecast long-term consequences of its own suggestions (e.g.,
if I remind you to exercise today, what is your predicted adherence and mood next week?
), enabling proactive, not just reactive, support.
- Agency-Preserving Intervention Policies
-
Improvement: Replace reward-maximizing policies with a constrained optimization that only selects actions within an admissible set (consent, scope, safety, reversibility, escalation) and explicitly optimizes for agency preservation metrics (autonomy, contestability, capability growth, over-reliance risk).
-
Capability: The AI will refuse to act in ways that increase dependency, will offer explanations and corrections, and will escalate to human experts when uncertainty or risk exceeds thresholds—even if a more
efficient
automated action exists.
- Co-Agency and Adaptive Division of Labor
-
Improvement: Dynamically negotiate and re-allocate tasks between human and AI based on the user’s current state, preferences, and long-term capability goals (e.g., gradually transferring skill back to the user).
-
Capability: The AI can act as a coach that fades support over time, or as a mediator that coordinates human services, rather than always doing tasks for the user.
- Edge-Native Personal Infrastructure with User-Side Authority
-
Improvement: Store authoritative copies of memory, preferences, and safety boundaries on user-controlled devices; allow local inference and intervention authority even when cloud services are unavailable.
-
Capability: The AI remains functional and privacy-preserving offline, and the user retains final control over what data is shared, what models are updated, and what actions are permitted.
- Uncertainty-Aware Perception and Action
-
Improvement: Maintain a posterior distribution over the latent human state (not a point estimate) and use calibrated uncertainty to decide when to ask clarifying questions, when to act, and when to escalate.
-
Capability: The AI can say
I’m not confident about your current emotional state—should I proceed with the reminder or wait?
reducing false interventions and increasing trust.
- Scenario-Centered Evaluation with Agency Metrics
-
Improvement: Build evaluation suites (like CombodiedBench) that test not just task success but also autonomy preservation, contestability, boundary respect, and long-term capability growth across repeated interactions.
-
Capability: The AI system can be benchmarked and improved on its ability to strengthen user agency over time, rather than merely maximizing engagement or satisfaction.
- Multi-Modal Event-Based Perception with Provenance
-
Improvement: Convert raw sensor data into structured events with explicit provenance, quality scores, and uncertainty labels, following the pipeline: observation → event → inferred state → predicted trajectory → authorized intervention.
-
Capability: The AI can distinguish reliable signals (e.g., a verified heart-rate spike) from noisy ones (e.g., a misread step count) and avoid acting on spurious correlations.
- Escalation and Human-in-the-Loop Coordination
-
Improvement: Implement explicit escalation policies for high-stakes or low-confidence situations, routing to human caregivers, clinicians, or emergency services with full context.
-
Capability: The AI can act as a first-line monitor but knows when to hand off, preventing both over-reliance on automation and dangerous delays in critical care.
- Consent-Aware and Boundary-Respecting Interaction
-
Improvement: Encode user-defined boundaries (e.g.,
don’t discuss my health in front of family,
don’t send reminders after 10 PM
) as hard constraints in the action space, and allow real-time revocation. -
Capability: The AI respects dynamic consent and personal limits, reducing risk of manipulation or privacy violations, and builds long-term trust.
What the Improved AI System Can Do (Concrete Example):
A Combodied Agent for a person with type-2 diabetes could:
-
Continuously infer glucose trends, mood, and meal context from a CGM, smartwatch, and voice journal.
-
Predict the trajectory of HbA1c over 6 months under different intervention policies (e.g., weekly coaching vs. daily reminders).
-
Choose a reminder only if it is likely to improve adherence without increasing anxiety (based on the user’s state model) and if it respects the user’s
no work-hour interruptions
boundary. -
If uncertainty about the user’s emotional state is high, it asks a clarifying question instead of acting.
-
After 3 months, it gradually reduces reminders as the user’s self-efficacy improves, preserving autonomy.
-
If it detects a dangerous hypoglycemic pattern, it escalates to the user’s clinician with a full longitudinal report, while keeping all raw data on the user’s phone.
-
It evaluates its own performance not by
number of reminders sent
but by whether the user’s capability to self-manage increased and whether they felt in control.
Sources
- PaLM-E: An Embodied Multimodal Language Model
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Investigating Affective Use and Emotional Well-being on ChatGPT
- How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study
- Towards the Human Digital Twin: Definition and Design -- A survey
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
- MemGPT: Towards LLMs as Operating Systems
- Toward Personalized LLM-Powered Agents: Foundations, Evaluation, and Future Directions
- A Survey of Personalization: From RAG to Agent
- Chatbot Companionship: A Mixed-Methods Study of Companion Chatbot Usage Patterns and Their Relationship to Loneliness in Active Users
- Towards a Personal Health Large Language Model
- Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
- Local Is Not a Sufficient Privacy Boundary: Governing OS-Integrated On-Device AI
- Beyond Scaling: Agents Are Heading to the Edge
- HomeRobot: Open-Vocabulary Mobile Manipulation
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- Agent AI: Surveying the Horizons of Multimodal Interaction
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection