Governing Agentic AI in FinTech

arXiv:2608.11344 · cs.CY, cs.AI, q-fin.RM · Submitted 2026-08-11 · Read on arXiv

Baylor University

cs.CY, cs.AI, q-fin.RM

Submitted: 2026-08-11

Updated: 2026-09-15

Comments: 58 pages, 12 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: This paper addresses the governance challenges posed by agentic AI systems in financial services.

Terminology

Summary

This paper addresses the governance challenges posed by agentic AI systems in financial services. The authors argue that the binding governance constraint on this shift is not capability but verifiability. Financial institutions are beginning to delegate consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little human oversight, yet agentic AI governance in FinTech remains under-investigated.

The central problem is that the black-box problem concerns opacity, and its conventional remedy is empirical validation... What is new is that this empirical remedy may no longer work. The paper explains: The governance problem is therefore not opacity alone. It is the loss of reproducibility as the principal safeguard against opacity. Agentic systems can be both difficult to explain and difficult to reproduce.

The paper introduces a novel construct called the Verifiability Gap, defined as the positive shortfall between the verification demanded by delegated authority and the explainability and reproducibility retained after a specific decision. The formal definition is:

G dq = [rho sigma(A d) - V dq]+

where A d is the authority exercised, rho sigma(A d) is the verification required for that authority under evidentiary standard sigma, and V dq is the verification capacity retained by the system. The construct asks whether the system preserves evidence sufficient for a named verifier, evidentiary standard, and audit lag.

The paper emphasizes that The Verifiability Gap is not a general property of a model. It is the positive shortfall for a specific decision. It arises when delegated agentic authority demands more verification than the organization retains, and occurs when both routes to verification: faithful explanation and material reproduction are impaired.

The paper makes three main theoretical contributions:

  1. Verifiability as authority–evidence alignment: The authors bring explainability and reproducibility into a single governance construct, shifting the unit of analysis from model transparency to system verifiability.

  2. Reproducibility as a governance profile, not a scalar: They separate four reproducibility targets: current outcome reproducibility, historical outcome reproducibility, material process reproducibility, and exact trace reproducibility, plus verdict differentiation as a diagnostic for output collapse. The full profile is:

R(t; sigma) = (R O(t), R H(t), R P(t; sigma), R T(t); D V(t))

  1. Evidence-contingent delegation: The principle that autonomous authority remains defensible only while retained verification capacity remains commensurate with the authority exercised.

The paper develops seven propositions across three governance levels:

Firm-level (P1–P3):

  • P1: Firms delegate only what they can later defend — retained verification capacity shapes granted autonomy

  • P2: When explainability is unreliable, firms must constrain actions — action-bounding controls outperform explanation-based review

  • P3: Making agents verifiable costs speed and flexibility — with a corollary that Skipping oversight today makes governance harder tomorrow

Regulatory/temporal level (P4–P5):

  • P4: System changes make old decisions harder to reproduce — with a corollary that Serial systems require chain-level preservation

  • P5: Agentic explanations count only when they track the decision chain

Network level (P6–P7):

  • P6: One missing record breaks the whole chain — end-to-end verification is bounded by the weakest causally material handoff

  • P7: Shared models create shared audit failures — common provider dependencies create synchronized reproducibility losses

The paper proves several theorems establishing the formal basis for the Verifiability Gap:

  • Theorem 1 (Serial contraction of discrimination): what the final action can reveal about which case was decided never grows as the pipeline deepens, and shrinks geometrically once each layer loses anything

  • Theorem 2 (Reversibility is equality in the data-processing inequality): a stage is auditable exactly when it has thrown nothing away

  • Theorem 3 (Specialization and reversibility are incompatible): a stage either summarizes or stays auditable. It cannot do both

  • Theorem 4 (Retention lower bound): to keep a decision reconstructable, the institution must store at least as many bits as its pipeline threw away

Study 1 examines whether institutions can preserve historical decision behavior after provider changes. Key findings:

Release-timeline drift: Across four dated Claude Opus releases over 185 days, baseline-modal historical fidelity remained perfect (1.000) through the second release. It then dropped to 0.906 in the third and fourth releases. Five distinct cases changed at least once, with every deviation moved toward a more conservative action.

Control-surface dependence: The paper finds that in later releases, Anthropic withdrew institutional control over temperature—an execution parameter critical for stable replay. The withdrawal is total rather than partial. The provider altered not only the system being audited, but also the very controls required to conduct the audit.

Execution environment effects: Under the tightest controls, a local model reproduced 320 of 320 executions, while hosted models reproduced 319 of 320 and 959 of 960. The paper notes: reproducibility is not guaranteed on a hosted endpoint, either when every exposed control is pinned or when no control is exposed at all.

Study 2 holds the model, cases, and execution settings fixed while varying agentic architecture from one to fifty agents. Key findings:

Orchestration is a latent policy layer: "Agent roles, topology, aggregation, and the information supplied to the terminal Decider are not neutral implementation details. They form a latent policy layer that can alter the institution's operative risk posture without an explicit policy decision."

Trace–outcome decoupling: The recorded output at every stage differed across repeated executions in all 32 cases... The final verdict differed in only 25 of 32 cases. This shows that the organization can observe a stable verdict even though the recorded path to that verdict is nonidentical.

Degenerate outcome reproducibility: Across all seven architectural configurations and 1,120 total executions, not a single case produced an exactly identical full set of agent outputs. The apparent recovery in outcome reproducibility at larger scales is an illusion — it stems from a severe collapse in verdict differentiation. At forty agents, three of the four case families applied the exact same blanket verdict to all eight of their underlying cases, with verdict differentiation collapsing to D V = 0.086.

Frontier model comparison: "The frontier model reproduces its own actions more often than the local ones, but its execution record is no better, and it loses a comparable share of its verdict differentiation. Capability buys a higher starting point, not auditability."

Study 3 uses two frozen versions of a transparent logistic credit model on the FICO HELOC dataset. Key findings:

Aggregate stability masks decision-level change: Although aggregate discrimination is virtually unchanged, 23 of the 1,937 held-out policy actions change after the refresh, yielding historical fidelity of R H = 0.9881.

Perfect current reproducibility coexists with historical failure: For case 2270, Version v 1 estimates p =.5580 and denies the application. Version v 2 estimates p =.5488 and refers the same application. The score changes by only.0092, but that movement crosses the fixed denial threshold and changes the financial action. Thus: R O(v 1) = R O(v 2) = 1, R H(case 2270) = 0.

Position, not magnitude, decides which actions change: What sets the 23 apart is where they started. Every one begins within.0161 of a cutoff, the 6.9th percentile of the held-out distance distribution.

Faithfulness requires the historical executable: "A reason statement reconstructed from v 2 could be fully faithful to the current referral and still fail to substantiate the historical denial... Historical explainability is therefore indexed to the historical executable."

The paper concludes that human intervention should therefore be evidence-contingent rather than uniformly imposed. Key governance recommendations include:

  1. Material change should trigger evidence-based reauthorization: A material change should be defined by its effects on financial behavior or verification capacity, not by whether the institution initiated it.

  2. The relevant record is an executable evidence bundle: "A defensible evidence bundle connects all material elements of a decision. The bundle includes decision-time inputs, fitted preprocessing state, model and component versions, prompts, policy thresholds, tool responses, retrieved documents, memory states, causally material handoffs, and the final action."

  3. Explainability controls must attach to the same evidence bundle: intervene on the factors identified as reasons and test whether the preserved executable responds as the explanation predicts.

  4. Provenance must compose across boundaries: Each causally material handoff should retain a persistent decision identifier, the sender and recipient, the transferred state, the applicable component version, and the transformation performed.

The paper concludes: "FinTech agentic AI governance is the dynamic allocation of authority between agents and humans according to retained verification capacity. Its goal is neither maximum autonomy nor a human checkpoint attached to every action. It is defensible autonomy. Agents act where their exercise of authority remains verifiable. Humans intervene where verification capacity falls below the required standard."

Improvements for AI systems

Improvements to AI Systems Based on This Paper

  1. Implement a Verifiability Gap Monitor: Build a runtime component that continuously computes G dq = [rho sigma(A d) - V dq]+ for each decision. The system tracks retained verification capacity (explainability fidelity + reproducibility score) against the authority delegated. When G dq > 0, the system automatically escalates to human review or reduces its action scope. This transforms governance from static policy to dynamic, evidence-contingent delegation.

  2. Add Reproducibility Profiling to Model Registries: Extend model versioning systems to store the full reproducibility profile R(t; sigma) = (R O, R H, R P, R T; D V) for every deployed model and agentic configuration. The system logs not just model weights but also execution parameters (temperature, top-p, seed), provider endpoints, and orchestration topology. This enables pre-deployment checks that flag configurations where R H (historical fidelity) drops below thresholds before they cause audit failures.

  3. Create a Verdict Differentiation Alarm: Implement a real-time monitor for D V(t) that detects output collapse in multi-agent systems. When the system observes that distinct input cases are converging to identical verdicts (e.g., D V < 0.1), it triggers an alert. This prevents the illusion of reproducibility—where stable outputs mask the loss of decision discrimination—from going unnoticed.

  4. Build a Historical Executable Preservation System: Design the system to store complete executable evidence bundles for each decision: decision-time inputs, fitted preprocessing state, model versions, prompts, policy thresholds, tool responses, retrieved documents, memory states, and handoff records. The system can then replay historical decisions using the exact preserved executable, not a current approximation. This addresses the finding that historical explainability is indexed to the historical executable.

  5. Implement Chain-Level Provenance Composition: For multi-agent pipelines, attach persistent decision identifiers and transformation records at every causally material handoff. The system verifies that each stage preserves reversibility (per Theorem 2) and flags stages where information loss exceeds the retention lower bound (Theorem 4). This prevents the one missing record breaks the whole chain failure mode.

  6. Add Material Change Detection with Evidence-Based Reauthorization: The system continuously monitors for changes—whether initiated internally (model updates, prompt changes) or externally (provider releases, API changes). When a change affects financial behavior or verification capacity, the system automatically requires re-validation of the evidence bundle before continued autonomous operation. This operationalizes the principle that autonomous authority remains defensible only while retained verification capacity remains commensurate with the authority exercised.

  7. Create a Control-Surface Dependency Checker: Before relying on hosted models, the system verifies that all execution controls necessary for reproducible replay (temperature, sampling parameters, seeds) remain exposed and stable. If a provider withdraws control surfaces (as Anthropic did with temperature), the system automatically downgrades the trust level for that provider and shifts to local models or adds compensating verification layers.

What the Improved AI System Can Do

  • Self-assess its own auditability before and during each decision, adjusting its autonomy in real-time based on measured verification capacity rather than static rules.

  • Replay any historical decision with bit-level fidelity using preserved executables, even after model updates or provider changes.

  • Distinguish between genuine reproducibility and output collapse, avoiding the trap where stable-looking verdicts mask the loss of decision discrimination.

  • Compose provenance across multi-agent chains, ensuring end-to-end auditability bounded only by the weakest causally material handoff.

  • Detect material changes (internal or external) and automatically trigger evidence-based reauthorization before continuing autonomous operation.

  • Maintain defensible autonomy—acting independently where verification capacity is sufficient, and escalating to humans precisely where the Verifiability Gap opens, without imposing uniform human checkpoints on every action.

Abstract

Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top p and top k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.

Sources

Related papers