An Autonomous AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

arXiv:2606.17555 · cs.CR, cs.AI, cs.CE, cs.ET · Submitted 2026-06-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "ArapaiSecure: An Autonomous AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts".

Elias: Banks face two threat families with fundamentally different detection requirements: signature-based fraud and behavioral financial crime.

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, we're diving into the paper "ArapaiSecure," which is about this autonomous AI security agent designed for retail and corporate banking to tackle both signature-based fraud and more complex behavioral financial crime like layering. It claims this system uses a three-component fusion architecture across two parallel event streams to detect these diverse threats, which really sounds pretty comprehensive.

Elias: That's right, Nadia; the core idea presented in "ArapaiSecure: An Autonomous AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts" is that traditional static rule engines fall short because they can't catch attacks engineered to look like normal activity at the individual level. This paper proposes an AI security agent that addresses those gaps by fusing information from transaction streams and session streams to get a better picture of the risk.

Priya: From a privacy and measurement standpoint, what struck me most about this approach is how it frames the detection challenge itself, showing that BEC and structuring are collective anomalies, meaning they only become anomalous when you look at them in combination rather than individually <ref:2606.17555#pg1>. I'm curious if this framework helps us measure these complex interactions accurately.

Nadia: Exactly what Priya is pointing out; the paper lays out a threat model with thirteen categories across both transaction and session streams, which really shows the breadth of what this AI is designed to look for <ref:2606.17555#pg2>. It tackles things like account takeover in sessions and layering in transactions, which are often missed by simpler systems.

Elias: And from a cryptographic view, the architecture relies on an LSTM sequence model combined with statistical monitors and a graph module to capture patterns like fan-in and pass-through ratios for laundering detection <ref:2606.17555#pg0>. I wonder what assumptions these models make about the underlying data structure when they're trying to predict fraud probability or account-counterparty patterns.

Priya: The authors mention using a synthetic log of two hundred thirty-seven thousand six hundred sixty-nine transactions and one hundred thirteen thousand five hundred eight sessions for their experiments <ref:2606.17555#pg1>. That's a pretty large dataset for testing these kinds of sophisticated models, but I'm wondering how realistic that synthetic data is compared to the full distribution of real-world banking traffic.

Nadia: Well, the paper does show some compelling results comparing this AI agent against baselines; they achieved a macro-average F1 of zero point three zero three for the transaction stream and zero point five two nine for the session stream <ref:2606.17555#pg1>. That's a significant improvement over what rules or LSTM-only models achieved, which is what they're highlighting in "ArapaiSecure."

Elias: That performance gain is interesting when you consider the complexity of the architecture, which involves combining an LSTM for per-account behavior with threshold monitors and a graph module for network structure <ref:2606.17555#pg0>. I'm interested in how the weighting factor alpha + beta + gamma = one influences how much each component contributes to the final risk score R <ref:2606.17555#pg0>.

Priya: And looking at those results, it seems that BEC detection achieved an F1 of zero point three six eight in the session stream, which is notably higher than near-zero scores seen in other baselines <ref:2606.17555#pg1>. That suggests the combination of sequence modeling and graph features is quite effective for capturing those subtle redirection patterns.

Paper summary: Nadia: And the system includes sector-specific modules for retail and corporate banking, which adapt to different expected transaction patterns, like lower amounts for retail versus higher per-transaction amounts in corporate settings <ref:2606.17555#pg2>. This domain knowledge integration seems crucial for tailoring the detection to specific industry needs.

Elias: I agree that those sector modules add a layer of contextual understanding, but we have to remember that the paper also includes an override mechanism for certain scenarios, like BEC redirection where if the LSTM is highly confident, even other signals near zero don't suppress it <ref:2606.17555#pg2>. That part of the logic is important for ensuring high-confidence signals aren't ignored.

Priya: So, while the performance metrics are impressive in their synthetic testing environment, the authors do flag that a limitation is that they can't capture the full distributional complexity of real bank traffic, specifically mentioning long-tail transaction amounts and seasonal and cultural payment patterns unique to places like Uganda and East Africa <ref:2606.17555#pg2>.

Nadia: That limitation is something we need to keep in mind; the model's effectiveness might change when deployed in a real, messy environment with those kinds of extreme outliers <ref:2606.17555#pg2>. But overall, ArapaiSecure demonstrates a robust approach for multi-vector fraud detection across different banking contexts.

Elias: It really showcases how combining sequence analysis, statistical thresholds, and network structure can create a system capable of handling the diverse demands of modern financial crime <ref:2606.17555#pg0>. This fusion architecture is certainly something to keep studying from a security perspective.

Priya: I think the real implication here is moving detection away from simple pattern matching toward understanding the flow and context of events, which feels like a necessary evolution in securing financial systems <ref:2606.17555#pg1>.

Nadia: And we can't forget the practical response framework where a critical tier event triggers immediate actions like account freezes or SAR escalations, with response latency under zero point four three milliseconds at the 95th percentile <ref:2606.17555#pg1>. That speed is essential for stopping active threats.

Elias: So we've covered the core claims of "ArapaiSecure," from its architecture and performance metrics to its specific limitations and potential real-world utility in detecting complex financial crime <ref:2606.17555#pg0>. That gives us a solid foundation for what this research is actually proposing.

Priya: It really highlights that the challenge isn't just building a better model, but designing an agent that can fuse different types of signals—transactions and sessions—and adapt to specific banking sectors <ref:2606.17555#pg2>. That integration aspect is where the real value seems to lie for security research.

Nadia: Indeed, the idea of having a mechanism that handles both signature-based fraud and behavioral financial crime simultaneously through this fusion architecture is what makes this agent noteworthy <ref:2606.17555#pg0>. That dual focus is quite ambitious for a single system to manage.

Paper summary: Elias: I think the paper suggests that the future work could involve addressing the approximation inherent in their graph module, specifically how they can move from rolling-window proxies to full GNN message passing for detecting end-to-end chain detection <ref:2606.17555#pg2>.

Priya: That's a very practical suggestion for future work; bridging that gap between the current proxy and a more complete network view would certainly make the system even more powerful.

Nadia: So, to wrap up this discussion on "ArapaiSecure," we see an AI security agent built on a fusion architecture that targets diverse fraud types using transaction and session streams <ref:2606.17555#pg0>. It shows significant improvements over existing baselines in terms of detection capabilities across multiple threat categories <ref:2606.17555#pg1>.

Elias: And from a cryptographic and architectural standpoint, the interplay between the LSTM sequence model, statistical monitors, and graph features provides a sophisticated way to score risk R = max(Rtxn, Rsess) <ref:2606.17555#pg0>. We've seen how that combination handles BEC detection effectively even in session streams <ref:2606.17555#pg1>.

Priya: The main implication I see for the wider field is that we need to shift our focus from just identifying individual suspicious events to understanding the collective, multi-vector anomalies that characterize sophisticated financial crime <ref:2606.17555#pg1>. This paper definitely points in that direction.

Nadia: It's a compelling demonstration of how AI can be applied to solve these multifaceted security challenges in banking environments <ref:2606.17555#pg0>. We're excited about the potential for this type of agent to provide much more resilient protection for customers and institutions.

Elias: It certainly is a well-structured approach, combining sequence modeling with structural analysis to tackle the two major threat families mentioned in the introduction <ref:2606.17555#pg0>. We've seen how it manages to provide actionable summaries through its case-summary assistant for analysts <ref:2606.17555#pg1>.

Priya: It's encouraging that they clearly articulate the limitations regarding real-world data complexity, as that shows a high level of self-awareness in the research process <ref:2606.17555#pg2>. That kind of honesty is valuable in any research endeavor.

Nadia: That honest acknowledgment of what the synthetic data can't fully replicate is important context when we think about deploying these systems in practice, especially given how much real-world traffic varies <ref:2606.17555#pg2>. We need to keep that gap in mind.

Elias: So, if you had to distill the core research finding of "ArapaiSecure," I'd say it proves that a fusion architecture across parallel event streams is a viable way to detect both signature and behavioral financial crime <ref:2606.17555#pg0>. The system achieves strong performance metrics, such as an overall F1 of zero point eight six seven for the session stream <ref:2606.17555#pg1>, which is substantial compared to prior methods.

Priya: I think the biggest world-level implication is that this type of multi-vector detection framework could become standard practice in securing financial infrastructure because it moves beyond single indicators and focuses on the sequence of actions <ref:2606.17555#pg1>.

Nadia: It's definitely a substantial piece of work, showing how to combine different AI techniques to build a security agent that is tailored for the specific complexities found in banking operations <ref:2606.17555#pg2>. We're really looking forward to seeing how this type of integrated approach develops further.

Conclusion: Nadia: So we've seen how ArapaiSecure uses this fusion architecture to catch both signature fraud and behavioral financial crime across retail and corporate accounts, and now we need to look at what the title actually means for the world.

Elias: The title itself points directly to a system that’s autonomous, which suggests it operates without constant human supervision on the detection side. It implies a level of independence in identifying threats that's important for real-time security applications.

Priya: From my perspective, the core idea is moving away from looking at single data points and instead understanding the whole flow of activity across different banking functions. This suggests a new way to measure risk that captures context rather than just isolated events.

Nadia: Exactly; it’s not just about catching one bad transaction, but figuring out if a sequence of session events or a series of transactions together indicate something illicit is happening. That's the big shift here for security infrastructure.

Elias: And when we talk about the authors, they seem to have built this framework by carefully considering how different data streams—transactions and sessions—interact, which suggests a deep understanding of system dynamics. I wonder if their approach to handling those two parallel event streams is robust against any kind of signal masking.

Priya: It really does suggest that future security systems won't just be looking for known bad patterns anymore; they’ll be analyzing the entire operational context of an account or a network structure over time. That contextual view is where the real measurement challenge lies, I think.

Nadia: And that context helps us see why these threats are so hard to catch with older methods, because the AI can pick up on anomalies that happen when things are *mostly* normal but just slightly off in a complex sequence.

Elias: That's interesting because from a cryptographic standpoint, the system relies on these statistical models to define what "normal" looks like; if those underlying statistical assumptions about behavior are flawed, the entire risk scoring mechanism could produce false positives or miss real attacks.

Priya: That's a fair concern; the paper does acknowledge that its effectiveness is tied to the quality of that training data, which brings us right back to how well it handles real-world complexity versus just synthetic scenarios.

Nadia: So we’ve established that this agent aims to be a comprehensive tool for banking security by fusing different types of event data, and now we need to consider what kind of practical impact this has on the financial ecosystem generally.

Department of Electronics and Computer Engineering, Soroti University

cs.CR, cs.AI, cs.CE, cs.ET

Submitted: 2026-06-16

Updated: 2026-10-08

Comments: 6 pages, 1 figure, 5 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

The gist: Banks face two threat families with fundamentally different detection requirements: signature-based fraud and behavioral financial crime.

Key concepts

Three-Component Fusion Architecture
This architecture combines three different detection methods—sequence modeling (LSTM), statistical thresholds, and graph analysis—across two separate data streams (transactions and sessions). The final risk score is determined by taking the maximum of the scores from these components, ensuring comprehensive threat coverage.
Sequence Model (sseq)
This component uses a Long Short-Term Memory (LSTM) network trained on recent history of account activity. It looks at sequences of up to ten events to predict the probability that the very next event in that sequence is fraudulent, helping detect patterns over time.
Graph/Network Module (sgraph)
This module analyzes relationships between accounts and parties using a graph structure. It calculates features like how many different people an account sends money to (fan-out) or receives money from (fan-in) over recent periods, which is crucial for spotting money laundering patterns.
Risk Score Formula (R = max(Rtxn, Rsess))
The final risk score is calculated by taking the highest risk score derived from either the transaction stream or the session stream. This ensures that if a critical event occurs in one area—like a suspicious session—it triggers an appropriate response regardless of what the other stream indicates.

Terminology

Summary

Banks face two threat families with fundamentally different detection requirements: signature-based fraud and behavioral financial crime. This paper presents an AI security agent for retail and corporate banking using a three-component fusion architecture across two parallel event streams to detect these diverse threats.

Architecture Overview

The proposed agent utilizes a three-component fusion architecture across two parallel event streams: transactions (card fraud, ACH/wire fraud, AML) and sessions (account takeover, hijacking, SIM-swap). Each stream combines an LSTM sequence model of per-account behaviour, a statistical velocity/threshold monitor, and a graph module capturing account-counterparty patterns for laundering detection. The final risk score is determined by the formula: R = max(Rtxn, Rsess), ensuring that a Critical session event during an otherwise normal transaction stream (or vice versa) still triggers the appropriate response.

Core Detection Engine Components

The core detection engine computes each stream's risk score as: R = α sseq + β sthresh + γ sgraph, where α + β + γ = 1. The three sub-models are:

  1. Sequence model (sseq): An LSTM network trained on sliding windows of length L=10 over each account’s transaction or session history to produce a predicted fraud probability of the final event in each window.

  2. Threshold monitor (sthresh): A continuous score combining normalised amount z-score, velocity ratio, and a structuring aggregate ratio (7-day cash-deposit sum / CTR threshold).

  3. Graph/network module (sgraph): A proxy GNN feature capturing network structure via rolling 24/48-hour windows, specifically using features like fan-in (distinct inbound senders), fan-out (distinct outbound recipients), and pass-through ratio.

Sector-Specific Modules and Overrides

The detection engine is extended by two sector modules that encode domain knowledge. The Retail module uses lower transaction amount baselines, higher expected ATM and card-purchase velocity, while the Corporate module uses higher per-transaction amounts, elevated payroll-batch velocity at month-end. Furthermore, a single-sub-model override is applied for categories like BEC redirection where the model ensures that if the LSTM is highly confident (sseq ≥ 0.90) but other signals are near zero, R ← max(R, sseq), preventing suppression of a near-certain detection signal.

Automated Response Framework and Evaluation

The risk score R is mapped to one of four tiers: Low (0 ≤ R < 0.30), Medium (0.30 ≤ R < 0.60), High (0.60 ≤ R < 0.85), and Critical (R ≥ 1). The Critical tier triggers account freeze; financial freeze; incident; SAR escalation, while the Chatbot handles verification for High/Critical events, achieving a 96.6% identity accuracy and an 86.8% mass-reset detection. The Analyst Case-Summary Assistant provides a case summary including the sub-model score breakdown, aiming for an Action F1 of 99.3% in recommending next actions.

Performance Against Baselines

The proposed agent demonstrates significant improvements over baselines. On the transaction stream, the overall F1 is 0.303 compared to 0.227 for rules and 0.158 for an LSTM-only baseline, with a Macro-avg F1 of 0.227 (rules) vs Overall F1 of 0.787 (ours). On the session stream, the overall F1 is 0.529 compared to 0.500 for rules and 0.283 for LSTM-only, with an Overall F1 of 0.867. Notably, BEC detection achieves an F1 of 0.368 in the session stream versus near-zero scores in other baselines, attributed to the LSTM recognizing the sequence of add payee + immediate large transfer as an anomalous session sequence. The response latency for Critical tier events is under 0.43 ms at the 95th percentile on a single CPU node.

Key Findings and Limitations

The agent successfully covers 13 threat categories, showing improvement over every baseline. However, Layering remains the hardest transaction-stream category, achieving only an F1 of 0.131 because the rolling-window proxy is an approximation of full GNN message passing needed for end-to-end chain detection. Additionally, the evaluation relies on a synthetic dataset and "cannot capture the full distributional complexity of real bank traffic — long-tail transaction amounts, seasonal and cultural payment patterns specific to Uganda and East Africa.

Improvements for AI systems

Here are specific improvements for existing AI systems based on this research, and what those improved systems can achieve:


  1. A unified, dual-stream architecture combining an LSTM sequence model, a statistical velocity/threshold monitor, and a graph module (using fan-in/fan-out/pass-through ratios) into a single fused risk score.

  2. The system will be able to detect complex collective anomalies (like structuring or layering) that are individually invisible in standard rule engines or single-stream models, by analyzing the relationships between transactions and sessions simultaneously.

  3. Improved detection of Business Email Compromise (BEC) payment redirection, which exploits the fact that individual transaction amounts and counterparty profiles look legitimate, but is caught by recognizing anomalous session sequences (e.g., adding a new payee immediately followed by a large wire transfer).

  4. Enhanced detection of Account Takeover (ATO) and SIM-swap attacks by combining session features (device/city change velocity, MFA failure patterns) with transaction sequence models to identify the temporal proximity of failed logins to subsequent anomalous transfers.

  5. The system will be capable of detecting insider abuse and dormant account reactivation by leveraging graph features (dormancy flags, fan-in/fan-out ratios) across both streams, providing redundant coverage where neither stream alone is sufficient.

  6. The AI agent can provide a real-time risk assessment for every event, mapped to a four-tier response framework (Low to Critical), enabling automated actions ranging from soft holds and in-app alerts to hard account freezes and immediate compliance escalation for high-risk events.

  7. An integrated customer-facing verification chatbot can perform high-accuracy identity confirmation (96.6%) and proactively detect credential stuffing attacks by monitoring mass password reset requests across multiple source cities within a tight time window.

  8. A sophisticated analyst case-summary assistant will generate plain-English summaries that include the specific sub-model scores (sseq, sthresh, sgraph), a detailed threat narrative, and a ranked list of recommended actions, increasing analyst decision speed and accuracy (99.3% action recommendation F1).

  9. The system can dynamically adjust its detection sensitivity by learning fused weights from logistic regression to prioritize the most effective detection signal (e.g., prioritizing the LSTM score for BEC when threshold/graph signals are low, or prioritizing the graph score for layering).

  10. The system will be optimized for low-latency response, ensuring that critical actions like account freezes and incident reporting occur under 0.43 ms at the 95th percentile on a single CPU node, which is crucial for mitigating rapid fraud sequences.

Related papers