Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment

arXiv:2601.10520 · cs.AI, cs.CY · Submitted 2026-01-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment".

Jane: The paper was written by Authors not found in excerpt from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We're starting today with a paper that sounds like a breakup song for computer scientists. It's titled "Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment."

Jane: That title is actually quite literal, Tom. It's talking about how we need to stop forcing AI to handle its goals and its morals in one single, messy pile.

Tom: So, when they say "monolithic," they mean the current way of doing things?

Jane: Exactly, because right now, most AI tries to solve everything with one giant math function. This paper by Jahn and his colleagues argues that we need to split that function apart.

Lu: I think the most exciting part of that title is the "Neuro-Symbolic" bit. It suggests they aren't just throwing more data at a black box to make it behave.

Tom: You mean they're adding a layer of actual logic on top of the neural networks?

Lu: Yes, and that's a massive shift in how we think about intelligence. We're combining the pattern recognition of deep learning with the structured, step-by-step reasoning of symbolic logic.

Meng: I see the value in that, but I'm curious about the practical split. If you separate the "how to do a task" from the "how to be ethical," how do you stop them from working against each other?

Jane: That's the central tension the authors are tackling, Meng. They want to ensure the ethical part isn't just a filter at the end, but a core part of the reasoning process.

Meng: It makes sense from a debugging standpoint. If the AI does something wrong, I want to know if it failed at the task or if it actually misunderstood a moral principle.

Lalam: This separation is what allows for a real sense of accountability. If an AI can explain its reasoning, it stops being a mysterious oracle and starts becoming a participant in our social values.

Tom: That's a huge leap for trust. So, if we're moving away from that "monolithic" approach, what does the new structure actually look like?

Jane: That's what we'll look at next, specifically how they've broken the agent down into specialized modules.

Summary: Jane: We've been talking about why we need to break up with monolithic AI, and now we're looking at the actual blueprint the authors propose. In "Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment," they introduce an architecture called GRACE.

Tom: It's essentially a three-part system: a Moral Module, a Decision-Making Module, and a Guard.

Jane: Right, and the Moral Module is the brain that does the heavy lifting for ethics using logic.

Tom: While the Decision-Making Module is the part that actually tries to get the job done efficiently?

Jane: Yes, and then the Guard stands at the door to make sure the two don't clash in a way that breaks the rules.

Lu: I'm particularly fascinated by the "Moral Advisor" they mention in the summary. It's not just a static set of rules; it's a way for humans or experts to give case-based feedback.

Tom: So the AI can actually learn from its mistakes through these specific examples?

Lu: Precisely, and it learns the *reason* behind the correction, not just a simple "yes" or "no." It builds a theory of why certain actions are better than others.

Meng: I'm looking at the "Macro Action Types" they describe. That seems like a clever way to bridge the gap between high-level ethics and low-level code.

Jane: How so, Meng?

Meng: Well, instead of telling the AI "don't be mean," which is vague, the Moral Module tells it "the type of action you're about to take is not allowed." It gives the agent a high-level category to follow.

Lalam: This approach respects the complexity of human life. By using these macro categories, the AI can navigate different cultural expectations without needing a brand-new brain for every country.

Tom: It sounds like they're trying to build a system that is both flexible and incredibly rigid where it counts.

Jane: That's a great way to put it. But how do these modules actually handle a situation where two different moral rules are fighting each other?

Paper discussion segment 3: Tom: We've seen the structure of GRACE, but the real test is when things get messy. The paper suggests that instead of just following a list of rules, the AI should actually weigh competing reasons.

Jane: That's the "defeasibility" part of their logic. It means a reason can be overridden if a stronger, more important reason comes along.

Tom: Like in their therapy assistant example, where a patient's wish for privacy might conflict with the need to report a self-harm risk.

Jane: Exactly, and the system uses a priority order to decide which one wins.

Lu: I love the creative potential here for things like international diplomacy. You could program the system to understand that in a crisis, certain humanitarian principles must override standard protocols.

Meng: I have to push back a little on the implementation side, though. If the Moral Module is relying on symbolic logic, it's only as good as the data it gets from the environment.

Tom: You're worried about the "noise" in the real world, right?

Meng: Definitely. If a sensor is malfunctioning or a user is lying, the Moral Module might reason perfectly about a lie, which still leads to a bad outcome.

Lalam: That's why the separation is so vital, Meng. Because the reasoning is explicit, we can actually see that the AI was "tricked" by the input, rather than just seeing a black box make a weird choice.

Jane: That's a really important distinction for safety.

Lalam: It also allows us to build a culture of transparency. We can move toward a world where we don't just trust AI blindly, but we actually audit its moral logic just like we would a human's.

Tom: It seems like the biggest hurdle isn't just the ethics, but how we manage the uncertainty of the world itself.

Jane: It really does. And that leads us to our final thoughts on where this is all heading.

Conclusion: Tom: We've covered a lot of ground today, from the "flattening problem" to the modular brilliance of GRACE.

Jane: It's been a deep dive into "Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment."

Tom: The core message is clear: we can't just optimize for goals; we have to reason about values.

Jane: And we have to do it in a way that humans can understand and challenge.

Lu: I truly believe this is the path toward AI that acts as a partner rather than just a tool. It gives machines a way to respect the nuance of human existence.

Meng: From my side, I'm looking forward to seeing the open-source tools they're building. Having a concrete architecture like this makes the goal of safe AI feel much more achievable for engineers.

Lalam: This work suggests that our technological future can be built on a foundation of respect and reason. It's a blueprint for a more considerate digital civilization.

Tom: That's a powerful note to end on. Jane, any final thoughts?

Jane: Just that the conversation around AI alignment is finally moving from "how do we stop it?" to "how do we teach it to reason?"

Tom: Well, that's all for this episode. Thank you to Lu, Meng, and Lalam for joining us.

Jane: We'll be back soon to look at how these ideas might apply to global data governance.

Tom: See you next time.

cs.AI, cs.CY

Submitted: 2026-01-15

Updated: 2026-02-02

Comments: 10 pages, 4 figures, accepted at 2nd Annual Conference of the International Association for Safe & Ethical AI (IASEAI'26)

Journal ref: Proceedings of the 2nd IASEAI Conference (IASEAI'26), 2(1), 254-267, 2026

DOI: 10.1609/iaseai.v2i1.43029

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 78/100

The gist: This paper introduces GRACE (Governor for Reason-Aligned ContainmEnt), a "reason-based neuro-symbolic containment architecture" designed to address the critical "flattening problem" in AI alignment.

Key concepts

Neuro-Symbolic Architecture
This approach combines deep learning's pattern recognition (neural networks) with structured logic (symbolic reasoning). It suggests a massive shift in AI design, allowing systems to combine complex data handling with step-by-step, explicit reasoning.
GRACE Architecture
GRACE is the proposed three-part system for ethical AI. It consists of a Moral Module (for ethics using logic), a Decision-Making Module (for task efficiency), and a Guard module to prevent conflicts between the two.
Moral Advisor / Macro Action Types
This mechanism allows AI to learn from specific, case-based feedback rather than just simple rules. It gives the system high-level ethical categories, helping it navigate diverse human expectations and cultural differences.
Defeasibility
This concept means that a reason or rule can be overridden if a stronger, more important reason presents itself. The system uses a priority order to resolve conflicts, such as when privacy conflicts with reporting self-harm risks.

Terminology

Summary

This paper introduces GRACE (Governor for Reason-Aligned ContainmEnt), a reason-based neuro-symbolic containment architecture designed to address the critical flattening problem in AI alignment. By decoupling normative reasoning from instrumental decision-making, GRACE provides a framework for ensuring that autonomous agents are not only effective but also transparent, contestable, and justifiable in morally consequential contexts.

The Flattening Problem

The authors argue that most current AI architectures suffer from being normatively monolithic, meaning they compress instrumental decision-making, normative constraints, and their integration into a single, opaque policy function. This flattening of heterogeneous evaluative standpoints into a single optimization target creates several recurring pathologies:

  • Opacity: Because decisions are often encoded in deep neural networks, the grounds for action become inaccessible, hindering understanding, auditing, and trust.

  • Brittleness: Implicitly encoded constraints tend to fail under distribution shift or novel (moral) contexts.

  • Contestability deficit: Without explicit normative structure, stakeholders lack meaningful points of intervention to challenge, justify, and revise decisions.

  • Verification challenge: Alignment and safety checks must target the entire policy rather than separable components, making systematic validation and certification intractable.

The GRACE Architecture

To overcome these limitations, GRACE decomposes the monolithic agent into a multi-agent system of three specialized, interacting modules:

  1. Moral Module (MM): This module uses symbolic reasoning to determine permissible macro action types (MATs). It interprets observations for moral relevance and provides transparent justifications grounded in normative reasons.

  2. Decision-Making Module (DMM): The DMM encapsulates the target agent’s instrumental capabilities, selecting instrumentally optimal primitive actions while remaining subject to the constraints imposed by the MM.

  3. Guard (G): This component enforces compliance by verifying that any action proposed by the DMM satisfies a permissible MAT (a s phi). It acts as a monitoring mechanism to ensure primitive actions result in macro actions of permitted types.

Foundations of Reason and Abstraction

The architecture is built upon normative reasons, which are defined as facts that count in favor of or against particular courses of action. The MM utilizes a formalization based on Horty’s default logic, represented by the triple R, D, <, where R is a set of parametrized reasons, D contains defaults encoding relations between facts and favored actions, and < defines a priority ordering to resolve conflicts. This allows the system to derive situation-specific reason-based models from a general reason theory.

To interface moral reasoning with concrete behavior, GRACE employs three levels of action abstraction:

  • Primitive Actions: The atomic operations the agent directly executes.

  • Macro Actions: Temporally extended sequences of primitive actions and states representing a concrete, more complex overall behaviour.

  • Macro Action Types (MATs): Decidable predicates expressed in a logic language that describe high-level, abstract categories of behavior (e.g., “assess self-harm risk”).

Continuous Learning via the Moral Advisor

To ensure adaptability to evolving normative standards, GRACE incorporates a Moral Advisor (MA). The MA serves as an external authority—such as a human overseer or expert committee—that provides case-based feedback to the system. When the agent’s behavior is identified as a violation, the MA provides feedback in the form of (a, phi, P), indicating that action a was impermissible and that MAT phi was expected for reason P. This mechanism allows the MM to incrementally build coherent reason theories by updating its symbolic representation based on these corrective justifications.

Improvements for AI systems

1. Transition from Monolithic Policy to Neuro-Symbolic Multi-Agent Architecture

  • What the improved system can do: It decouples how to achieve a goal (instrumental optimization) from what is allowed (normative reasoning). This prevents reward hacking and ensures that even if the core agent (e.g., an LLM or RL agent) attempts an optimal but unethical action, the system remains structurally incapable of executing it without violating its governing modules.

2. Integration of a Defeasible Logic-based Moral Module (MM)

  • What the improved system can do: Instead of relying on opaque probability distributions to determine safety, the system uses symbolic reasoning to derive permissible Macro Action Types (MATs). This allows the AI to provide explicit, human-readable justifications for its decisions (e.g., Action denied: violates privacy-sensitive data protection rule delta 1 ) and enables stakeholders to contest specific normative reasons rather than arguing with a black-box model.

3. Deployment of a Symbolic Guard for Runtime Verification

  • What the improved system can do: It acts as a formal enforcement layer between the decision-making engine and the environment. The Guard intercepts proposed primitive actions (e.g., API calls, token generation, or physical movements) and performs a logical check (a s phi) to ensure they satisfy the permissible MATs. This provides a mathematical guarantee of compliance that neural-only architectures cannot offer.

4. Implementation of a Case-Based Moral Advisor (MA) Feedback Loop

  • What the improved system can do: It enables plug-and-play ethical updates. When an agent encounters a novel moral dilemma, a human overseer can provide feedback in the form of a specific reason and an expected MAT. The system then updates its internal Reason Theory (R, D, <) symbolically. This allows the AI to learn new legal or cultural norms instantly without the massive computational cost, risk of catastrophic forgetting, or unpredictable side effects associated with retraining or RLHF.

5. Utilization of Macro Action Types (MATs) as a Formal Interface

  • What the improved system can do: It bridges the gap between high-level ethical principles (e.g., maintain patient confidentiality) and low-level execution (e.g., do not call API endpoint X). By mapping primitive actions to these logical predicates, the system can perform formal verification and statistical guarantees of alignment across diverse, non-deterministic environments.

Abstract

As AI agents become increasingly autonomous, widely deployed in consequential contexts, and efficacious in bringing about real-world impacts, ensuring that their decisions are not only instrumentally effective but also normatively aligned has become critical. We introduce a neuro-symbolic reason-based containment architecture, Governor for Reason-Aligned ContainmEnt (GRACE), that decouples normative reasoning from instrumental decision-making and can contain AI agents of virtually any design. GRACE restructures decision-making into three modules: a Moral Module (MM) that determines permissible macro actions via deontic logic-based reasoning; a Decision-Making Module (DMM) that encapsulates the target agent while selecting instrumentally optimal primitive actions in accordance with derived macro actions; and a Guard that monitors and enforces moral compliance. The MM uses a reason-based formalism providing a semantic foundation for deontic logic, enabling interpretability, contestability, and justifiability. Its symbolic representation enriches the DMM's informational context and supports formal verification and statistical guarantees of alignment enforced by the Guard. We demonstrate GRACE on an example of a LLM therapy assistant, showing how it enables stakeholders to understand, contest, and refine agent behavior.

Sources

Related papers