Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF

summary

Video file (mp4)

The gist

This paper examines why AI governance frameworks are hard to adopt, treating framework adoption as a "governance translation problem": "whether RMF language can become role-usable, cross-level,

In short

The episode discusses a paper stress-testing the NIST AI Risk Management Framework (RMF) using simulations across different organizational roles and AI deployment types. The hosts conclude that framework adoption fails not because people don't understand it, but because of a 'visibility-authority problem' where local activity cannot translate into authoritative governance decisions.

Key concepts

Role-Based Stress Test
Researchers built a simulation testing how different organizational roles—like executives and middle managers—apply the NIST AI RMF to real governance problems in consumer lending. This tested where the framework breaks down when people with real authority limits try to use it.
Visibility-Authority Problem
This occurs when people closest to the risk see issues first, but those with authority to act are far away. Governance only works if information travels up in a usable form and decisions travel back down as constraints or resources.
System-Boundary Issue
When governing an AI system like an LLM copilot embedded in a workflow, the risk is not just in the model outputs. The framework must govern how the system is used—including reliance, documentation, supervision, and human judgment—not just the model itself.

Terminology used across episodes

This episode discusses

The paper

Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF · Read on arXiv

Joseph R. Simons, David A. Broniatowski

The George Washington University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF".

Jane: The paper was written by Joseph R. Simons and David A. Broniatowski from The George Washington University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the channel, everyone. Today we’re looking at a paper that’s been making the rounds on arXiv, and the title alone got me hooked: “Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF.” Jane, what’s your first reaction to that title?

Jane: Honestly, Tom, it’s refreshing to see someone tackle the boring-but-critical side of AI. Everyone’s obsessed with the latest model capabilities, but this paper asks why companies say they follow governance frameworks and then don’t actually change anything. It’s from Joseph Simons and David Broniatowski at George Washington University, and they’re basically stress-testing the NIST framework.

Tom: And when you say stress test, you mean they didn’t just read the framework and nod along. They actually put it through controlled scenarios, right?

Jane: Exactly. They built a simulation where different organizational roles—like executives, middle managers, system owners—had to apply the NIST AI RMF to real governance problems in consumer lending. The whole point was to see where the framework breaks down when real people with real authority limits try to use it.

Tom: And the authors are framing this as a translation problem. It’s not that the framework is bad or that people are lazy. It’s that the language of the framework has to travel through an organization, and that journey is where things fall apart.

Jane: Right. And that’s why the title says “hard to adopt” rather than “hard to understand.” Understanding is the easy part. Making it matter inside a company with different levels of authority, different incentives, different information—that’s the hard part.

Tom: So before we get into the actual results, what’s the big implication you’re taking from the title alone?

Jane: That we’ve been measuring the wrong thing. We’ve been asking “do people know the framework?” when we should be asking “does the framework change what people can see, decide, and correct?” Those are very different questions.

Tom: And that distinction is going to drive the whole conversation today. Stick around, because the results get really interesting when they show that even when people use the framework perfectly, it still doesn’t always become governance.

Summary: Tom: So we’ve set the stage with the title. Now let’s get into what the paper actually found. Jane, give us the summary in plain terms.

Jane: Sure. The researchers set up a four by two by three experiment—four organizational roles, two AI deployment types, and three governance hard cases. That gave them twenty-four scenarios, and they ran five responses per scenario, so one hundred twenty total scored responses. They used an LLM to simulate the role actors, which is clever because it lets them control the conditions tightly.

Tom: And the first big finding surprised me. Local translation was not the problem. The simulated actors—even the middle managers and system owners—understood their roles, understood the framework, and could produce sensible local actions. They hit the ceiling on role understanding and responsibility mapping.

Jane: That’s the part that really flips the usual narrative. We always assume adoption fails because people don’t get it. Here, everyone got it. But getting it didn’t mean the framework became governance. The breakdown happened downstream, when that local activity had to move up the organization and connect to authority.

Tom: And that’s where role mattered a lot. Enterprise executives were at ceiling on governance value—they could connect the framework to decisions, resources, and corrections. Middle managers and system owners? Not so much. They could see problems and propose fixes, but they couldn’t make those fixes stick because they didn’t have the authority.

Jane: Right. And the paper calls this the visibility-authority problem. The people closest to the risk often see it first, but the people with authority to act are far away. Governance only works when information travels up in a usable form and decisions travel back down as constraints or resources.

Tom: Then there was the deployment effect. They compared a traditional ML underwriting model against an LLM underwriting copilot. The framework fit the ML model cleanly—scores, thresholds, validation, drift—all familiar control points. But the LLM copilot was messier because the risk lived in workflow, reliance, documentation, and human judgment.

Jane: And that’s the system-boundary issue. If you treat an LLM copilot like a bounded model, you’re governing the wrong object. The risk isn’t just in the model outputs; it’s in how underwriters actually use those outputs, how they frame decisions, how they document rationale. The framework can still generate activity, but it’s harder to know where governance should attach.

Tom: So the summary is: local use works, but governance value is hard, and risk reduction is even harder. We’ll get into that gap next.

Improvements: Tom: Now let’s talk about what the paper suggests we actually do about this. Jane, what’s the improvement path they’re proposing?

Jane: The big one is that NIST guidance should shift from artifact production to governance capability. Instead of telling organizations to produce policies, dashboards, and documentation, the guidance should connect each framework activity to a specific decision it improves, a signal it produces, and the authority needed to act on that signal.

Tom: So it’s not “check the box,” it’s “what changes because you did this?”

Jane: Exactly. And they also recommend role-specific value propositions. An executive needs to see how the framework helps them set risk appetite and allocate resources. A middle manager needs to see how it helps them escalate signals upward. A system owner needs to see how it helps them get correction authority. Right now, the framework speaks in one generic voice, and that’s part of why adoption stalls.

Tom: And what about the system-type distinction? That felt like a major point in the results.

Jane: They’re not asking for a rigid taxonomy, but they do want guidance that helps organizations recognize when monitoring and enforcement are enough, and when they need to do more sensemaking work. For a bounded ML model, you can attach to familiar control points. For a workflow-embedded LLM, you need to stabilize what the system actually means in use—what role it plays, when reliance is appropriate, and how to correct drift.

Tom: And that’s where the paper gets really practical. They’re saying the framework should help organizations locate the limits of governability. Not just “here’s how to govern,” but “here’s what you cannot yet know, cannot yet correct, and where you need to escalate residual risk or accept it.”

Jane: That diagnostic function is huge. A framework that only tells you what to do is less useful than one that tells you where your current boundary stops working. And that’s a shift in mindset—from the framework as a checklist to the framework as a probe for finding the edges.

Tom: So the improvement isn’t just “make the framework better,” it’s “help organizations see where the framework can’t reach, and what that means for their governance structure.”

Jane: Right. And that connects back to the authority problem. If a system owner sees workflow drift but can’t pause the system, the framework should help them make that case upward in a form executives can act on. That’s the translation work that actually creates value.

Conclusion: Tom: Alright, we’ve covered a lot. Let’s wrap up “Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF.” Jane, give us the final takeaway.

Jane: The core message is that framework adoption isn’t about awareness or training. It’s about whether framework-guided activity can become authority-connected governance over the AI system-in-use. The paper shows that local translation works fine, but governance value requires signals to move up and decisions to move down.

Tom: And risk reduction is even harder. It only appeared when governance value was present and structural fit was full—and even then, it needed a concrete mechanism linking governance action to reduced likelihood or impact.

Jane: So the paper separates three things we often conflate: governance activity, governance value, and risk reduction. Activity is easy. Value is harder. Risk reduction is hardest. And confusing them is why organizations can look like they’ve adopted the framework while actually changing nothing.

Tom: The other big point is the system boundary. For bounded ML models, the framework fits cleanly. For LLM copilots embedded in workflow, the governed object is the system-in-use, not just the model. That means governance has to handle reliance, documentation, supervision, and human judgment—not just model outputs.

Jane: And the authors are clear that this generalizes beyond NIST. Any governance framework has to travel through real organizations with real authority structures. If it doesn’t connect to decisions and corrections, it’s just paperwork.

Tom: So what’s the lasting impact? I think it’s this: frameworks create value when they help organizations see, interpret, escalate, authorize, and correct risk. And they also create value by revealing where governability ends—what can’t be known, what can’t be corrected, and where residual risk has to be accepted or escalated.

Jane: That’s the part I hope regulators and practitioners take seriously. Adoption isn’t a documentation exercise. It’s a governance capability exercise. And this paper gives us a diagnostic way to see where the gaps are.

Tom: Great summary, Jane. That’s a wrap on this one. Next up, we’ve got a paper on interpretability methods that I think is going to spark some debate. See you in a bit.

More episodes

← Home