Rules or Character? Scaling Laws for AI Safety Design

arXiv:2608.13345 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Satoshi Takahashi, Nobuji Kouno, Masaaki Komatsu, Ryuji Hamamoto

RIKEN Center for Advanced Intelligence Project · National Cancer Center Research Institute

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: Accepted at AIES 2026 (9th AAAI/ACM Conference on AI, Ethics, and Society). 10 pages, 6 figures, 4 tables

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: This paper introduces a stylized comparative-statics model to analyze the optimal balance between character shaping (e.g., RLHF, Constitutional AI) and rule enforcement (e.g., output filters, safety

Terminology

Summary

This paper introduces a stylized comparative-statics model to analyze the optimal balance between character shaping (e.g., RLHF, Constitutional AI) and rule enforcement (e.g., output filters, safety classifiers) in AI safety design as deployment scale increases. The model parameterizes safety design as a resource allocation α ∈ [0, 1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility—the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, the authors derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation.

Across three scenarios (optimistic, moderate, pessimistic), the optimal α∗ is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (∆α∗ = +0.01) to pronounced (∆α∗ = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate pfrag(0), which shifts α∗ by 0.50 across its range—far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.

Key structural results include: the optimum is never pure character (α∗ < 1 in all cases); the optimal character weight is weakly non-decreasing in deployment scale (∆α∗ ≥ 0 across 1,200 phase diagram cells); the optimal policy is robust to the choice of risk criterion (expected harm vs. CVaR) and to the Pareto tail exponent αPL; and stronger character-shaping capability (larger ∆µ) is associated with lower optimal α∗ due to diminishing returns. The paper also proves Proposition 1: improving filter technology (lower εmin) lowers α∗, verified numerically across all scenarios.

The authors conclude that the most consequential input to safety architecture design is the estimated reliability of character shaping under distributional shift, and they recommend three priorities: measuring fragility, reducing fragility (which may be more valuable than improving character shaping or filter quality in isolation), and noting that the scaling question becomes moot if fragility is sufficiently low (below approximately 5%). Limitations include the one-dimensional Gaussian action space, static analysis, α-independent tail structure, and lack of direct empirical calibration for key parameters such as pfrag(0) and αPL.

Improvements for AI systems

Improvements to AI systems:

  1. Adaptive safety allocation controller: Build an AI governance module that dynamically adjusts the balance between character shaping (RLHF/Constitutional AI) and rule enforcement (filters/classifiers) based on real-time estimates of deployment scale and distributional shift. The system monitors pfrag(0) (baseline character fragility) and automatically increases character shaping weight when fragility is low (<5%) and scale is large, but shifts to rules-only when fragility is high, using the paper’s closed-form expected harm as a loss function.

  2. Fragility-aware training curriculum: Modify the RLHF training pipeline to include explicit fragility measurement—e.g., injecting novel-distribution probes during training to estimate pfrag(0). If measured fragility exceeds 0.5, the system prioritizes rule enforcement; if below 0.05, it can safely rely on character shaping even at massive deployment scales, saving compute and reducing over-filtering.

  3. Scale-dependent safety scheduler: Implement a deployment-time scheduler that, given a target scale T, computes the optimal α∗ using the paper’s phase diagram (with scenario parameters). The scheduler reallocates safety resources continuously as T grows, ensuring α∗ increases weakly with T (e.g., +0.01 to +0.21) but never reaches pure character shaping, preventing catastrophic collapse from common-mode failures.

  4. Tail-risk-aware filter optimization: Enhance safety classifiers with a CVaR-based objective that matches the expected-harm optimum at large T. The system tunes filter thresholds (εmin) using Proposition 1—lowering εmin reduces optimal α∗—so that improving filter technology automatically reduces reliance on fragile character shaping, even under heavy-tailed damage distributions.

  5. Fragility-reduction feedback loop: Add a meta-learning component that actively reduces pfrag(0) over time (e.g., via adversarial training on distributional shifts). The system tracks how reductions in fragility shift α∗ (by up to 0.50) and prioritizes fragility-reduction interventions over improving character shaping or filter quality in isolation, as the paper shows this yields the largest safety gains.

What the improved AI system can do:

  • Self-tune safety architecture in production: Given deployment scale, measured fragility, and filter quality, it automatically selects the optimal mix of behavioral shaping and hard rules, minimizing expected harm and CVaR simultaneously.

  • Predict safety degradation before deployment: By simulating pfrag(0) and common-mode failure probabilities, it forecasts whether character shaping will hold or collapse, and pre-emptively switches to rules-only mode if fragility is high.

  • Optimize safety R&D investment: It can recommend whether to spend resources on reducing fragility, improving filters, or enhancing character shaping, based on which lever shifts α∗ most (fragility dominates by 0.50).

  • Scale safely without over-engineering: When fragility is below 5%, the system relaxes filters and relies on character shaping, reducing false positives and operational costs, while maintaining safety even at trillion-parameter scale.

  • Robustly handle tail risks: It uses CVaR-based decision-making that converges to expected-harm optima at large T, ensuring the system is not caught off-guard by rare, catastrophic failures.

Abstract

Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutional AI), which modifies behavioral distributions at training time, with rule enforcement (e.g., output filters, safety classifiers), which blocks harmful outputs at inference time, yet little formal analysis exists on how their optimal balance should change as deployment scales increase. We introduce a stylized comparative-statics model that parameterizes safety design as a resource allocation alpha in [0,1] between these two approaches, incorporating scale-dependent filter degradation, common-mode failures, and character fragility -- the risk that shaped behavior degrades or collapses under novel conditions. Under a multiplicative Pareto damage model, we derive closed-form expected harm and supplement it with tail-risk (CVaR) analysis via Monte Carlo simulation. Across three scenarios (optimistic, moderate, pessimistic), the optimal alpha* is interior or at the rules-only boundary and shifts weakly toward character shaping as deployment scale T grows, from negligible (Delta alpha* = +0.01) to pronounced (Delta alpha* = +0.21) depending on scenario. The dominant parameter is the baseline character fragility rate p(0) frag, which shifts alpha* by 0.50 across its range -- far exceeding the effect of tail severity, filter quality, or common-mode failure probability. CVaR and expected-harm optima converge at large T. These results suggest that safety architecture decisions depend less on deployment scale per se than on the reliability of character shaping under distributional shift.

Sources

Related papers