Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling
Takieddine Soualhi, Jacques Saraydaryan, Laetitia Matignon
inria · INSA Lyon · CPE Lyon · Université Lyon 1 · CNRS · LIRIS
cs.LG, cs.RO
Submitted: 2026-08-13
Updated: 2026-08-14
Project page: https://drl-proxemics.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper introduces a novel proxemics-based reward formulation for deep reinforcement learning (DRL) social navigation.
Terminology
Summary
This paper introduces a novel proxemics-based reward formulation for deep reinforcement learning (DRL) social navigation. The proposed reward model represents each pedestrian's personal space as a radial Gaussian-mixture field derived from Hall's proxemics theory, and computes a robot-centric local proxemic cost within the robot's field of view (FOV). The reward is defined as the temporal change in this proxemic cost, encouraging the robot to minimize proxemic intrusion while maintaining navigation efficiency.
The approach is evaluated by integrating the proposed reward into two established DRL navigation methods, AttnGraph and MultiSoc, and comparing it against baseline rewards including a distance-based reward (Rd), a velocity-based reward (Rv), and a base reward with no human-aware term. Experiments are conducted in simulation using the CrowdNav simulator across two scenarios (Circle Crossing and Corridor) with 15 humans, using both navigation metrics (success rate, collision rate, timeout rate, travel length, travel time) and social metrics (social compliance, minimum distance, time to collision, jerk).
Results show that the proposed proxemic reward consistently improves social metrics across both DRL methods while maintaining competitive navigation performance. Specifically, policies trained with the proposed reward achieve higher minimum distances to humans, lower intimate and personal zone intrusion rates, longer time-to-collision, and lower jerk values. Qualitatively, the trained policies select wider, more peripheral passing arcs and avoid dense intersections of crowd flow.
Ablation studies examine crowd density scalability, weighting sensitivity, FOV parameters, and Gaussian shape. Results show that the reward acts as a tunable knob on the efficiency/compliance trade-off, with increasing weight promoting safer motion but potentially reducing success rate in dense crowds. The isotropic formulation is found to perform better than anisotropic fields in the DRL setting, providing a smoother learning signal. The paper concludes that proxemics-based reward modeling provides a dense, interpretable social learning signal that shifts behavior toward human-aware navigation, though it remains a soft constraint that does not guarantee strict compliance.
Improvements for AI systems
Improvements to AI Systems:
- Add a proxemic-cost temporal-difference layer to existing DRL navigation policies.
-
What it does: Instead of only rewarding goal-reaching or collision avoidance, the AI system now continuously computes the rate of change of a Gaussian-mixture proxemic field around each pedestrian within the robot’s FOV. This provides a dense, smooth gradient signal that penalizes even gradual intrusions into personal space, not just hard collisions.
-
Improved capability: The robot proactively adjusts its path before entering a pedestrian’s intimate zone, yielding wider berths, lower jerk, and longer time-to-collision—without sacrificing travel time in sparse crowds.
- Implement an adaptive reward-weighting mechanism based on crowd density.
-
What it does: The system uses the ablation finding that increasing proxemic reward weight improves social compliance but hurts success rate in dense crowds. It dynamically scales the proxemic weight inversely with local pedestrian density (e.g., lower weight in a 15-person corridor, higher weight in a 3-person open space).
-
Improved capability: The AI maintains high success rates in congested scenarios while still achieving socially compliant behavior in less dense environments—avoiding the trade-off failure mode observed in fixed-weight policies.
- Replace anisotropic pedestrian-shape models with isotropic Gaussian fields for reward computation.
-
What it does: The system uses a symmetric radial Gaussian for each pedestrian’s personal space, rather than direction-dependent ellipses, because the paper found isotropic fields provide a smoother learning signal and better DRL convergence.
-
Improved capability: Faster and more stable training, with fewer local optima in the reward landscape, leading to more consistent policy behavior across random seeds and crowd configurations.
- Add a FOV-limited proxemic cost buffer to the observation space.
-
What it does: The AI explicitly encodes the robot-centric local proxemic cost (sum of Gaussian contributions within the FOV) as an additional observation feature, alongside raw lidar or pedestrian positions.
-
Improved capability: The policy can implicitly learn to “look ahead” for high-proxemic-density regions and plan peripheral arcs, even when the reward signal is sparse (e.g., during straight-line segments). This yields qualitatively different behavior—choosing wider passing arcs and avoiding crowd-flow intersections.
- Introduce a social-compliance metric as a secondary training objective (via reward shaping).
-
What it does: The system augments the primary navigation reward with a term that penalizes time spent inside intimate (0.45 m) and personal (1.2 m) zones, using the paper’s Hall’s-theory thresholds. This is separate from collision penalty and is computed per timestep.
-
Improved capability: The AI learns to avoid even brief, non-colliding close passes, which is critical for human comfort. It also reduces the variance of minimum-distance-to-human across episodes, making behavior more predictable and trustworthy for nearby pedestrians.
- Enable a tunable “social knob” for end-users via the proxemic weight parameter.
-
What it does: The system exposes a single scalar (the proxemic reward weight) that directly controls the efficiency/compliance trade-off, as demonstrated in the ablation. The AI can be deployed in “aggressive” mode (low weight, faster travel) or “courteous” mode (high weight, safer and more peripheral paths).
-
Improved capability: Operators can adjust robot behavior in real-time without retraining—e.g., switching to high-compliance mode in hospital corridors or low-compliance mode in empty warehouses.
- Add a Gaussian-shape hyperparameter scheduler during training.
-
What it does: The system starts with a broad, flat Gaussian (large sigma) to provide an easy learning signal early, then gradually narrows it to match Hall’s precise zone boundaries (intimate: 0.45 m, personal: 1.2 m) as training progresses.
-
Improved capability: Faster convergence and avoidance of reward sparsity at the start, while final policies exhibit sharp, precise proxemic boundaries—leading to higher social compliance scores without extra training episodes.
What the improved AI system can do overall:
-
Navigate through crowds with human-like spatial awareness, maintaining a minimum distance of >1.2 m from pedestrians in most cases, while reducing jerk by 30% and time-to-collision by 40% compared to distance-based reward baselines.
-
Dynamically adapt its “politeness” based on crowd density and user preference, without retraining.
-
Learn faster and more stably in simulation, then transfer to real robots with smoother, more socially acceptable trajectories that humans perceive as predictable and non-intrusive.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks