Redistribution-based Cost Inference Improves Sparse Safe Offline RL

arXiv:2608.12306 · cs.LG, cs.AI · Submitted 2026-08-12 · Read on arXiv

Ebenezer Gelo, Geraud Nangue Tasse, Steven James, Benjamin Rosman

University of the Witwatersrand

cs.LG, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted at the 1st IJCAI Workshop on Safe Physical AI (SPAI 2026), affiliated with IJCAI/ECAI 2026

Code: https://github.com/eleurent/highway-env

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: The paper introduces the Redistribution-based Cost Inference (RCI) framework, which addresses the problem of safe offline reinforcement learning when only sparse trajectory-level stop-feedback is

Terminology

Summary

The paper introduces the Redistribution-based Cost Inference (RCI) framework, which addresses the problem of safe offline reinforcement learning when only sparse trajectory-level stop-feedback is available, rather than dense per-step cost annotations. The authors frame this as a temporal credit assignment problem and propose converting sparse stop-feedback into dense per-step costs via return decomposition, then training a constrained offline policy on the augmented dataset.

The paper states: "Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dataset."

The RCI framework is explicitly modular, consisting of three stages: (1) trajectory-level stop-feedback collection, (2) return-decomposition-based cost inference, and (3) constrained offline policy learning. The authors note: "the redistribution stage accepts any return-equivalent decomposition method (instantiated here with RUDDER, though GRD applies directly), and the policy learning stage accepts any constrained offline RL algorithm (instantiated with BCQ-Lagrangian, though CPQ or CDT substitute without modification)."

The theoretical contribution establishes that return-equivalent redistribution preserves the feasible policy set and the optimal Lagrangian in a CMDP. Specifically, the paper proves: "Theorem 1 (Policy Invariance). Let Msparse and Mredist denote CMDPs differing only in instantaneous costs. Then the feasible policy sets are identical and π∗sparse = π∗redist. Moreover, for any λ ≥ 0: L(π, λ; c̃) = L(π, λ; csparse), so the Lagrangian saddle point (π∗, λ∗) coincides under both formulations. The authors also note: Crucially, these guarantees hold regardless of the sequence model's prediction accuracy due to the compensation term δT."

The experiments evaluate RCI on two benchmark environments: HighwayEnv (highway driving) and Safe-FetchReach (robotic manipulation). For each environment, 5,000 offline episodes are generated using PPO-trained behavioral policies that optimize task reward while disregarding safety. The paper reports: Experiments on highway driving and robotic manipulation demonstrate substantially lower violation rates than sparse and classifier-based baselines, with robustness to heterogeneous dataset compositions and label noise.

Key experimental results include: "RCI cuts violation rate compared to Sparse and Hazard while maintaining task return; a two-sample t-test against the unconstrained baseline yields t = 0.9962, p = 0.3483, indicating no statistically significant return penalty. The paper also reports that RCI reduces violation rates roughly fivefold on two physical-safety domains without statistically significant return penalties."

The paper includes ablation studies on dataset composition (PPO, Random, and Mixed behavior policies) and label noise (Noisy and Adversarial corruption regimes). The authors state: RCI's violation reduction holds across all three regimes (Figure 2), indicating that the redistribution mechanism does not depend on the behavior policy producing structured or near-optimal exploration. Regarding label noise: Violation rates on Sparse and Hazard respond more sharply to misaligned supervision than RCI (Figure 3b), consistent with the smoothing effect of return decomposition over per-step label errors.

The paper also provides qualitative analysis showing that RCI produces globally coherent spatial structure: cost increases gradually as the ego approaches traffic, with high-cost regions extending backward along approach corridors in HighwayEnv, and that spatially coherent avoidance strategies emerge from trajectory-level stop-feedback alone, without access to dense cost annotations in Safe-FetchReach.

The authors conclude: "RCI converts trajectory-level stop-feedback into dense per-step costs via return decomposition, enabling constrained offline policy learning from the supervision modality physical AI deployments actually produce. The transformation preserves the feasible policy set and Lagrangian saddle point of the underlying CMDP, and reduces violation rates roughly fivefold on two physical-safety domains without statistically significant return penalties. Stop-feedback is informationally sufficient for safe offline policy learning; richer supervision modalities are not prerequisites."

Three limitations are acknowledged: "RCI inherits standard offline RL coverage requirements: datasets skewed toward unsafe trajectories risk overly conservative policies, and sparse safe coverage may leave policy optimization underspecified. The redistribution mechanism captures statistical, not causal, associations between trajectory prefixes and violations, identifying risky situations rather than causally hazardous actions. The framework is currently restricted to single binary constraints; multi-constraint extension via per-channel decomposition is straightforward but raises open questions about balancing competing objectives."

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Cost-Aware Offline RL with Sparse Supervision: Enable offline reinforcement learning agents to learn safe policies directly from trajectory-level binary stop-feedback (e.g., episode ended due to unsafe action) instead of requiring dense per-step cost labels. This makes safety training feasible in real-world deployments where annotators only flag the moment of failure, not every risky step.

  2. Robust Temporal Credit Assignment via Return Decomposition: Implement a redistribution module (e.g., RUDDER-based) that converts sparse stop-feedback into dense per-step pseudo-costs. This improves AI systems by smoothing noisy or adversarial label errors, as demonstrated by RCI's resilience to misaligned supervision compared to classifier-based baselines.

  3. Policy-Invariant Safety Transformation: Use the theoretical guarantee (Theorem 1) to safely transform any CMDP with sparse costs into an equivalent dense-cost CMDP without altering the feasible policy set or the optimal Lagrangian saddle point. This allows AI systems to swap in any constrained offline RL algorithm (e.g., BCQ-Lagrangian, CPQ, CDT) without re-tuning safety constraints, ensuring consistent safety performance across different policy learners.

  4. Heterogeneous Data Robustness: Improve AI systems' ability to learn safe policies from mixed-quality offline datasets (e.g., PPO, random, and mixed behavior policies). RCI's redistribution mechanism maintains low violation rates regardless of data source, enabling safe policy learning from suboptimal or unstructured exploration data—critical for real-world logs.

  5. Causal Risk Identification without Dense Annotations: Provide AI systems with a method to identify spatially and temporally coherent risk patterns (e.g., gradual cost increase near traffic, backward-extending high-cost corridors) purely from trajectory-level feedback. This enables interpretable safety maps for navigation or manipulation tasks, aiding human oversight and system debugging.

  6. Safe Policy Learning under Label Noise: Enhance AI robustness to corrupted safety signals by leveraging return-equivalent decomposition, which averages out per-step label errors. This is particularly useful for systems deployed in environments where human feedback is imperfect or adversarial.

  7. Modular Integration with Existing RL Pipelines: Allow AI systems to adopt RCI as a drop-in preprocessing stage before any constrained offline RL algorithm, without modifying the underlying policy optimizer. This reduces engineering overhead and accelerates safe deployment of autonomous agents in physical domains (e.g., driving, robotics).

What the improved AI system can do:

  • Learn safe control policies from only binary stop signals at failure points, eliminating the need for expensive dense cost annotation.

  • Maintain task performance (no statistically significant return penalty) while reducing safety violations by 5x in physical-safety domains.

  • Operate reliably on heterogeneous offline datasets and under label noise, making it suitable for real-world data collection pipelines.

  • Provide interpretable, spatially coherent risk assessments (e.g., high-risk zones near obstacles) without causal modeling, aiding in safety auditing and human-AI collaboration.

Related papers