Application of Artificial Intelligence for Fraudulent Banking Operations Recognition
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Application of Artificial Intelligence for Fraudulent Banking Operations Recognition".
Jane: The paper was written by Bohdan Mytnyk, Oleksandr Tkachyk, Nataliya Shakhovska, Solomiia Fedushko and Yuriy Syerov from Lviv Polytechnic National University and Comenius University in Bratislava.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back. Today we're looking at a paper that's going to hit close to home for anyone who's ever checked their bank statement with a little bit of dread. It's called "Application of Artificial Intelligence for Fraudulent Banking Operations Recognition."
Jane: And Tom, I have to say, just reading that title makes me think about all the times I've gotten a notification from my bank asking if a charge was really mine. That's exactly the problem this paper is trying to solve, but with a lot more sophistication than a simple text alert.
Tom: Exactly. And the team behind this is from Lviv Polytechnic National University in Ukraine, along with Comenius University in Bratislava. Given the context of the war in Ukraine, the authors actually mention that fraudulent operations have become even more common, especially with the rise of charitable funds that criminals try to exploit.
Jane: That's a really important point. The paper isn't just theoretical. It's responding to a real, urgent need. When people are donating to causes or moving their banking online during a crisis, that's when the bad actors swoop in. So the authors are saying, look, we need automated systems that can catch this stuff in real time.
Tom: And they're not starting from scratch. They're using machine learning, which is basically training computers to recognize patterns in data. The idea is that if you show an algorithm enough examples of legitimate transactions and enough examples of fraudulent ones, it can learn to tell the difference on its own.
Jane: Right. And the title promises "recognition," which is really just a fancy way of saying classification. Is this transaction real or is it fake? It's a yes-or-no question, but the data behind it is massive and messy.
Tom: Massive and messy is the perfect description. We're talking about hundreds of thousands of transactions, most of which are legitimate, with only a tiny fraction being fraud. That imbalance is a huge challenge, and I know we're going to dig into that later in the show.
Jane: Oh, definitely. But for now, I think the big takeaway from the title and the authors' framing is that this is a practical problem with real-world consequences. It's not just about improving an algorithm's score on a test. It's about protecting people's money, especially during vulnerable times.
Tom: And that's what makes this paper so compelling. It's a direct application of AI to a problem that affects millions of people every single day. So, let's get into the summary and see how they actually went about building these systems.
Summary: Jane: So, Tom, we've set the stage with the title. Now let's talk about what the paper actually does. The summary lays it out pretty clearly. The core goal is to develop machine learning models that can identify fraudulent banking transactions, with a special focus on the COVID-nineteen pandemic and the war in Ukraine.
Tom: And I love that they're so specific about the context. It's not just a generic "let's detect fraud" paper. They're saying, here's a moment in history where online transactions exploded and charitable giving became a target. So the models need to be robust enough to handle that shift.
Jane: Exactly. And the summary mentions something crucial: they're not just throwing one algorithm at the problem. They're developing several models using different methodologies. Then they compare them to see which one works best. It's like a bake-off, but for fraud detection.
Tom: A bake-off where the prize is not getting your identity stolen. I like that. And they also mention preprocessing techniques. That's the step where you clean up the data before feeding it to the algorithm. Think of it like prepping ingredients before you cook. You need to chop the vegetables and measure the flour, or the dish will be a mess.
Jane: Right. And in this case, the data is a credit card transaction dataset from European cardholders. It's publicly available on Kaggle, which is great for reproducibility. The dataset has about two hundred eighty-five thousand transactions, but only four hundred ninety-two of them are fraudulent. That's a tiny fraction, less than two-tenths of a percent.
Tom: That imbalance is the elephant in the room. If you just trained a model to say "not fraud" every time, it would be right ninety-nine point eight percent of the time. But it would be completely useless. So the paper has to deal with this imbalance head-on, and that's where the preprocessing techniques come in.
Jane: And that's the real meat of the summary. They're not just saying "we used machine learning." They're saying "we used machine learning with careful attention to data quality and class balance." That's what separates a good paper from a sloppy one.
Tom: So the summary sets up the problem, the context, and the general approach. But I'm curious about the specific improvements they suggest. What did they actually do differently to get better results?
Jane: Good question. Let's move on to the next section and see what tricks they had up their sleeves.
Improvements: Tom: Alright, Jane, so we've covered the what and the why. Now let's get into the how. The paper suggests several specific improvements to boost detection accuracy. And the first one is handling that massive class imbalance we mentioned.
Jane: Right. And they use a technique called random undersampling. That's where you take the majority class, the legitimate transactions, and randomly remove some of them until you have a more balanced dataset. It sounds counterintuitive, throwing away data, but it helps the model actually learn what fraud looks like instead of just memorizing "everything is normal."
Tom: It's like studying for a test by only reviewing the questions you got wrong. You're not going to waste time on the stuff you already know. You focus on the tricky parts. And the tricky parts here are the fraudulent transactions.
Jane: Exactly. And they also standardize the features. The dataset has some encrypted features, but the two that are not encrypted are time and amount. Standardization means scaling those values so they have a mean of zero and a standard deviation of one. That way, the amount of money spent doesn't dominate the algorithm just because it's a bigger number than the time.
Tom: That's a classic preprocessing step. If you don't do that, a transaction of five hundred dollars could be treated as more important than a transaction that happened at three AM, just because five hundred is a bigger number. Standardization puts everything on an equal footing.
Jane: And then there's feature engineering. The paper mentions this as a way to improve accuracy, though the details are a bit sparse in the summary. But the idea is that you can create new features from the existing ones that might be more informative. For example, maybe the time since the last transaction is a better predictor than the raw time itself.
Tom: So they're not just cleaning the data, they're reshaping it to make the patterns more visible. And they're also comparing a bunch of different algorithms. We're talking decision trees, random forests, logistic regression, support vector machines, k-nearest neighbors, and a few others.
Jane: And that's the beauty of the approach. They're not betting on one horse. They're letting them all run the race and seeing which one crosses the finish line first. The metric they use is the AUC, which stands for Area Under the Curve. It's a measure of how well the model can distinguish between the two classes.
Tom: And spoiler alert, logistic regression did really well, with an AUC of about zero point nine four six. But they didn't stop there. They also tried something called stacked generalization, which is like combining the predictions of several models to get a better overall prediction. That pushed the AUC up to zero point nine five four.
Jane: So the improvements are not just about tweaking one algorithm. It's about a whole pipeline of preprocessing, feature engineering, and model comparison. That's the kind of thorough approach that actually works in the real world.
Tom: And that's what we're going to see when we look at the first page of the paper. The authors lay out their entire methodology in detail. Let's take a closer look.
First Page: Jane: So, Tom, we've talked about the title, the summary, and the improvements. Now let's actually open up the first page of "Application of Artificial Intelligence for Fraudulent Banking Operations Recognition" and see what the authors are setting up.
Tom: And right away, they hit you with the context. They talk about how the COVID-nineteen pandemic pushed so many operations online, and then the war in Ukraine created a whole new avenue for fraud through charitable funds. It's a very timely motivation.
Jane: It really is. And they're clear that this is a supervised learning problem. That means you have labeled data, transactions that are known to be real or fraudulent, and you use that to train the model. It's like showing a child pictures of dogs and cats and saying "this is a dog, this is a cat" until they can tell the difference on their own.
Tom: And they mention the two categories of transactions: genuine and fraudulent. But they also get into the types of fraud, like wire fraud, identity theft, account takeover, money laundering. It's a reminder that "fraud" isn't one thing. It's a whole family of bad behavior.
Jane: Right. And that's why a one-size-fits-all approach won't work. You need algorithms that can adapt to different patterns. The paper also acknowledges the challenges of using AI here, like the lack of transparency in some algorithms. If the model says a transaction is fraud, you want to know why. But some models are black boxes.
Tom: That's a huge issue in banking. You can't just tell a customer "the algorithm said no" without a reason. There are regulations and customer service implications. So the paper is aware of these limitations, which is good to see.
Jane: And they also mention the limitation of their own study. They focus on online banking transactions, not all types of financial fraud. And they note that the dataset, while large, is limited to a specific region and time period. So the results might not generalize perfectly to other contexts.
Tom: That's intellectual honesty. They're not claiming to have solved all fraud everywhere. They're saying "here's a solid approach for this specific problem, and here are its boundaries." That's how good science works.
Jane: And they also bring up the risk of overfitting, where the model performs great on the training data but falls apart on new data. That's a classic pitfall, and they're clearly aware of it.
Tom: So the first page sets up the problem, the context, the challenges, and the limitations. It's a really solid foundation. And I'm excited to see how they actually implemented all of this in the rest of the paper.
Conclusion: Jane: Well, Tom, we've reached the end of our discussion on "Application of Artificial Intelligence for Fraudulent Banking Operations Recognition." And I think we can all agree that this paper is a great example of applied AI with a real-world impact.
Tom: Absolutely. They took a messy, imbalanced dataset, applied careful preprocessing, and compared a bunch of machine learning models. The result was that logistic regression stood out with an AUC of about zero point nine four six, and then they even improved on that with stacked generalization, hitting zero point nine five four.
Jane: And that stacked generalization is a clever trick. It's like getting a second opinion from multiple doctors instead of just one. Each model might catch something the others miss, and combining them gives you a more robust answer.
Tom: And the context matters so much here. The authors are from Ukraine, and they're dealing with the reality of war and the increased risk of fraud that comes with it. This isn't just an academic exercise. It's about protecting people's money during a crisis.
Jane: Exactly. And they were honest about the limitations. The dataset is from European cardholders, it's focused on online transactions, and there are concerns about transparency and overfitting. But within those boundaries, they showed a clear path forward.
Tom: And I think the impact could be significant. Banks and financial institutions could use these techniques to build better fraud detection systems. And as the paper notes, the problem is only going to get more complex as fraudsters get more sophisticated.
Jane: So, as we say goodbye to this paper, I think the takeaway is that AI can be a powerful tool for protecting people, but it requires careful design, honest evaluation, and a deep understanding of the problem context.
Tom: Well said, Jane. And to our listeners, thanks for joining us. We've got another fascinating paper coming up next, so stay tuned. This is Tom and Jane, signing off.
Jane: See you next time!
Bohdan Mytnyk, Oleksandr Tkachyk, Nataliya Shakhovska, Solomiia Fedushko, Yuriy Syerov
Lviv Polytechnic National University · Comenius University in Bratislava
cs.LG, cs.AI, cs.CE, cs.CR, cs.CY
Submitted: 2026-04-23
Updated: 2026-08-11
Comments: 22 pages, 6 figures
DOI: 10.3390/bdcc7020093
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 66/100
The gist: This study considers the task of applying artificial intelligence to recognize bank fraud.
Key concepts
- Machine Learning
- This involves training computers to recognize patterns in data. If the algorithm is shown enough examples of legitimate and fraudulent transactions, it learns to identify the difference on its own.
- Class Imbalance
- The dataset contains hundreds of thousands of transactions, but only a tiny fraction are fraudulent. This imbalance is a challenge because if a model simply predicts 'not fraud' every time, it would be mostly correct but useless in identifying actual fraud.
- Random Undersampling
- This technique addresses class imbalance by taking the majority class (legitimate transactions) and randomly removing some of them. This creates a more balanced dataset so the model can better learn what fraud looks like.
Terminology
Summary
This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID-19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and the creation of many charitable funds that criminals can use to deceive users. The present work focuses on machine learning algorithms as a tool well suited for analyzing and recognizing online banking transactions. The study’s scientific novelty is the development of machine learning models for identifying fraudulent banking transactions and techniques for preprocessing bank data for further comparison and selection of the best results. This paper also details various methods for improving detection accuracy, i.e., handling highly imbalanced datasets, feature transformation, and feature engineering. The proposed model, which is based on an artificial neural network, effectively improves the accuracy of fraudulent transaction detection. The results of the different algorithms are visualized, and the logistic regression algorithm performs the best, with an output AUC value of approximately 0.946. The stacked generalization shows a better AUC of 0.954. The recognition of banking fraud using artificial intelligence algorithms is a topical issue in our digital society.
The development of artificial intelligence for recognizing fraudulent banking operations has received significant attention in recent years. This is due to the growing number of fraudulent activities in the banking industry, which have resulted in significant financial losses for banks and their customers. AI-based systems have the potential to effectively identify and prevent fraudulent activities in real time, providing a significant advantage over traditional fraud detection methods. Detecting fraudulent banking operations involves using AI and machine learning algorithms to analyze large amounts of data from multiple sources, including transaction records, customer information, and network logs. These algorithms can identify patterns and anomalies in the data that may indicate fraudulent activities, such as unauthorized access, unusual transaction patterns, and suspicious behavior. A bank transaction involves any activity related to a bank account, which can be carried out online or offline between all parties involved. The process concludes when a written or electronic order is submitted to the bank, using internet banking systems, communication systems, or payment instruments. Bank transactions fall into two categories: genuine and fraudulent transactions. The latter refers to those that violate financial circulation rules or were not authorized. Common types of banking fraud include wire fraud, identity theft, account takeover, money laundering, and accounting fraud. As fraud becomes increasingly sophisticated, we must develop new methods to protect ourselves against it. Below are the five most common methods of preventing bank fraud: artificial intelligence, biometric data, consortium data, standardization of high technologies, and machine learning. Following the outbreak of the COVID-19 pandemic and the war in Ukraine, fraudulent operations related to bank transactions have become even more common due to the significant shift toward online transactions as well as the creation of numerous charitable funds that criminals use to deceive users. Therefore, it is necessary to create reliable automated algorithms to recognize and prevent operations that threaten the finances and accounts of individuals, violate taxation or financing rules or laws, and so on. The presented study focused on machine learning algorithms as a tool well suited for analyzing and recognizing online banking transactions. This study aimed to develop machine learning models that recognize fraudulent banking transactions, especially during the COVID-19 pandemic and war in Ukraine when online transactions and charitable funds have become more prevalent.
Our project focused on using machine learning models to identify fraudulent banking transactions. We also applied preprocessing techniques to compare and select the most effective outcomes from bank data. To accomplish our goals, we took the following steps: Developed several machine learning models using various methodologies and strategies. Compared and assessed the models from the previous stage using both quantitative and visual criteria. Analyzed the results obtained and drew a conclusion about the research objective. This study focused on using machine learning models to detect fraudulent banking transactions. The research aimed to develop algorithms that can accurately recognize such transactions. The methods used included preprocessing techniques and machine learning algorithms. The significance of this work lies in the potential of the proposed method to improve the detection of fraudulent banking transactions, especially during the pandemic when many transactions have shifted online and during times of war when there are many charities and events collecting money. The task of recognizing fraudulent bank transactions using machine learning involves identifying a particular transaction at a specific moment in time as either real or fraudulent based on previous historical data about other transactions. This process uses binary logic, where the transaction is either real or fraudulent, and classification algorithms are suitable for performing the task using machine learning methods. This paper proposes applying several classification algorithms that recognize the type of transaction based on certain features, along with preprocessing techniques. In order to effectively identify fraudulent transactions, the machine learning algorithm must have access to a comprehensive historical database of such activities. The existing collection of legitimate transactions that have yet to be flagged is encrypted to maintain the confidentiality and privacy of the financial institution’s clientele. However, this encryption does not hinder the algorithm’s ability to perform. Financial institutions can effectively detect and prevent fraudulent transactions by training models using this carefully selected dataset. Implementing AI technology in detecting fraudulent banking operations poses several challenges, including the utilized algorithms’ lack of transparency and interpretability. The intricacies of these algorithms can sometimes be challenging to comprehend, which can impede the identification and rectification of errors. Furthermore, using AI in fraud detection raises essential concerns regarding privacy, as personal data are subject to analysis and utilization in the decision-making process. These challenges require careful consideration to ensure AI-powered fraud detection systems’ safe and accurate implementation. The application development based on AI for fraudulent banking operations recognition is an active area of research and development, with significant potential to improve the efficiency and accuracy of fraud detection. However, addressing the challenges and concerns associated with using AI in fraud detection is essential to ensure its effectiveness and ethical use in the banking industry. Limitation of studies in the financial field. Studies using artificial intelligence to detect bank fraud are valuable. However, it is important to note that this study focuses on identifying fraudulent transactions in online banking only, while other types of financial fraud may require different detection methods. Additionally, the study’s reliability and generalizability may be affected by its limited sample size in the financial field. Data availability is also a crucial factor, as high-quality data re needed to train and test machine learning algorithms. The accuracy of the model can be compromised by incomplete and insufficiently diverse datasets, leading to false positives in real-world situations. Another challenge is the potential for human biases in the selection and analysis of data, which can impact the method’s validity and reliability. It is also important to avoid overfitting, where the model performs well on the training dataset but poorly on the test dataset due to its complexity and limited generalizability.
To meet the goals outlined in this study, we employed classification algorithms. These algorithms use features to determine the class of a given object. Machine learning relies on labeled training data to categorize new observations. To make accurate predictions for future observations, these algorithms must first analyze a dataset of examples with features and corresponding classes. They are considered supervised learning techniques because they map input variables (x) to discrete output functions (y) that represent categories rather than numerical values. The output of classification algorithms is not continuous but discrete. The classification algorithm learns from labeled input data, where input data have a corresponding output. The goal of classification algorithms is to limit the category of a given dataset, and they are widely used to predict the output for categorical data. From the information provided in the previous paragraph, it is clear that training any classification model requires a data set that shows the relationship between certain feature sets of an object and its class or category. This is why, to fit our classification model to the given task, we chose to use the dataset named Credit Card Fraud Detection obtained from the Kaggle platform. The platform provides the ROC graph curve and AUC metric for each algorithm and technique. A similar set of metrics is provided for other implementations as well. It is also easy to run the program using the Kaggle platform (which was used to create this software solution), which, by default, supports running Jupyter notebooks. To start, the “Run all” was pressed button to continue the sequential execution of all commands. An alternative solution is using any software that supports Jupyter notebooks. The concrete program implementation as saved as a notebook on the Kaggle platform and started by loading a dataset using the Pandas library. The research workflow was organized as given in Figure 3. Figure 3 represents the stages of a high-level algorithm for a machine learning program solution, which includes dataset selection and loading, feature standardization, random undersampling, model fitting, model testing, and outputting the best model. The presented method uses several mathematical elements: Machine learning algorithms, Evaluation metrics, Data preprocessing technique.
The machine learning algorithms chosen to be used here were (1) random forest, (2) k-nearest neighbors, (3) logistic regression, (4) stochastic gradient descent classifier, (5) decision tree, (6) naïve Bayes, and (7) support vector machine. Decision tree. In machine learning, a decision tree is a tree structure that shares similarities with a flowchart. Each internal node of the tree represents an attribute check, while each branch corresponds to the outcome of the check. Finally, every end node, or final node, contains a class label. The initial set is divided into subsets to train the decision tree based on checking the attribute’s value. The described process is recursive partitioning, repeated recursively for each derived subset. The recursion split ends in case of partitioning are no longer beneficial for predictions. Classification based on decision tree methods does not require knowledge of the domain or parameter tuning, making it an excellent choice for exploring knowledge. Furthermore, decision trees handle large amounts of data and typically provide high accuracy. Decision tree induction is a standard inductive approach for learning classification data. To classify instances, decision trees sort them through the tree, starting at the root and ending at a leaf node that classifies the instance. Random Forest. Specialists have used random forest methods for classification and regression tasks. This machine-learning-based algorithm is based on a flexible and user-friendly algorithm consisting of decision trees. The strength of the forest increases with the number of trees. The algorithm creates decision trees using randomly selected data samples, obtains predictions from each tree, and selects the best solution by voting. Additionally, it indicates feature importance. The algorithm works in the following steps: (1) selecting random samples, (2) building a decision tree, (3) voting, and (4) selecting the prediction result as the final prediction. Logistic regression is a commonly used statistical model for classification and predictive analytics, estimating the probability of an event based on a given set of independent variables. Logistic regression transforms the odds using a logit transformation, the logarithm of the odds, or the natural logarithm of the odds. Support vector machine (SVM) is a supervised learning algorithm for classification tasks and regression assignments. SVM creates a decision boundary, or hyperplane, that divides n-dimensional space into classes. The hyperplane is created by selecting extreme points, or support vectors, that help define the boundary. K-nearest neighbors (KNN) is a supervised learning nonparametric classifier. This classifier utilizes proximity to classify or predict the grouping of an individual data point. KNN is commonly used as a classification algorithm and assigns a class label based on the majority vote of nearby data points. For classification tasks with multiple classes, the class label is assigned with more than 25% of the vote rather than a strict majority of over 50%. The stochastic gradient descent (SGD) classifier is an approach that is straightforward and remarkably efficient in adapting linear classifiers and regressors to convex loss functions, including logistic regression and support vector machine. However, SGD has recently received substantial attention in the realm of large-scale learning. This is due to the success of SGD in tackling vast and sparse machine-learning challenges commonly encountered in natural language processing and text classification. Due to the sparsity of the data, the classifiers developed using SGD have rapidly expanded to deal with problems with over 105 training examples and over 105 features. The naïve Bayes classifier is a probabilistic machine learning model for classification tasks. The classifier’s foundation is based on Bayes’ theorem. The naïve Bayes classifier assumes that all predictors or features are independent, meaning that the presence of one feature does not affect the other. The Bayes formula is represented as the following equation: P(AB)=P(BA)P(A)P(B), where P(A) is the event A probability, P(B) is the event B probability, and P(BA) is the probability of event B occurring when event A occurs. In probability theory, the probability of an event A occurring is denoted by P(A), while the probability of event B occurring is denoted by P(B). Furthermore, the conditional probability of event B happening certainly that event A has happened is represented as P(BA).
Stacked generalization is a widely used method that integrates multiple low-level models to improve the overall predictive accuracy of a high-level model. This technique is based on estimating the biases of the high-level model concerning a given learning dataset. The estimation process involves generalizing the biases in a second space, using the original models’ predictions as inputs and the correct answers as outputs. Stacked generalization is considered an enhanced version of cross-validation, aggregating individual models into a higher-level model. Recently, a new method for grouping machine learning models based on random forest as a meta-algorithm has been developed, and its mathematical formulation is presented below. Randomly generate the following from the original dataset K cross-sectional data sets: a11,…,a1B, a21,…,a2B,…, aK1,…,aKB, where K is the number of subsets, B is the size of the subset, and alb is the observation of the lth sample. The task is to train K-independent weak classifiers f1(.),…,fk(.) Furthermore, combine the learning results using metamodel m: res=m(f1(.)×f2(.)×…×fk(.)) where fi(.) ×fj(.) is the result of the pairwise multiplication of weak classifiers. The transformed features are combined with the training dataset in the metamodel to improve model generalizability and prevent the correlation of weak classifiers’ results. However, the stacking model has a significant drawback: the meta-attributes for the training and testing sets differ. The meta-attribute in the training set is not the response of a specific classifier; it comprises responses from various classifiers with different types of dependence. On the other hand, the meta-attribute in the testing set is the answer to a completely different classifier configured for complete learning. The meta-attribute may have few unique values in classical stacking, but many do not overlap between the training and testing sets.
We employed the receiver operating characteristic (ROC) curve and the area under the curve (AUC) to evaluate the efficiency of the models suggested in this article. The ROC curve is a graphical representation of a classification model’s accuracy for all classification thresholds. This curve is plotted using the true positive rate (TPR) and false positive rate (FPR). TPR is the recall measure defined as the ratio of true positives to the sum of true positives and false negatives. TPR=TPTP+FN, where TP is true positive model labels; FN is false negative model labels. In model evaluation, TP refers to the number of positive instances correctly identified by the model, while FN represents the number of positive instances incorrectly classified as negative by the model. The FPR is a measure employed in assessing model performance, which quantifies the proportion of negative instances incorrectly classified as positive by the model concerning the total number of negative instances that include both the truly negative samples and the ones misclassified as positive. The FPR is defined as follows: FPR=FPFP+TN, where TP is true positive model labels; FN is false negative model labels. In the context of model evaluation, TP, which represents the true positive model labels, is the count of positive instances correctly identified by the model. FN, which stands for false negative model labels, indicates numerous positive instances incorrectly classified as negative by the model. The ROC curve plots the relationship between the TPR and the FPR at varying classification thresholds. By reducing the classification threshold, more items are classified as positive, increasing numerous true and false positives. A standard ROC curve is commonly used in evaluating model performance. The area under the curve (AUC) is a widely used metric in evaluating binary classification models. It offers a comprehensive measure of performance for all possible classification thresholds. The AUC is interpreted as the model’s probability of assigning a higher score to a random positive instance than a random negative instance. A perfect model has an AUC of 1.0, and the model with random predictions has an AUC of 0.0. The AUC is advantageous as it is scale-independent and invariant to the classification threshold, enabling the comparison of different models across different datasets. However, the scale invariance and threshold invariance of the AUC may not always be desirable in certain use cases. For instance, the AUC may not be applicable in situations where well-calibrated probabilistic input data are needed. Similarly, when significant differences exist in the costs of false positives and false negatives, the AUC may not be the most appropriate metric. For example, minimizing false positives in spam detection may be more important than minimizing false negatives, which the AUC does not consider.
In this study, the preprocessing stage involved standardizing certain features not encrypted in the dataset, specifically, the time and amount variables of submitted transactions. In addition, undersampling was employed as the dataset was not imbalanced. Standardization is a technique for scaling variables where values are centered around the mean and have a standard deviation of one. Thus, the attribute’s mean value is transformed to zero, and the distribution is normalized with a standard deviation of one. The standardization equation is expressed as follows: X′=X−musigma, where mu is the mathematical expectation, and σ is the standard deviation. In order to achieve the tasks set in this study, classification algorithms needed to be employed. A technique called random undersampling was utilized, which involves the random selection and removal of examples from the majority class in the training dataset. This results in a decrease in the number of many examples in the common class in the transformed training dataset. This procedure is repeated until the desired class distribution is obtained, such as an equal number of examples for each class. Classification is a supervised learning method used to determine the class or label of a given observation based on a dataset of previous observations. The dataset used in this paper was obtained from the open platform Kaggle and contained data on transactions of European cardholders labeled as fraudulent or real. Most of the features of this dataset are encrypted using the principal component method to ensure user privacy, with the only open features being the time of the transaction and the amount involved. Preprocessing was applied to the dataset in the form of standardization of features that were not encrypted as well as the balancing of classes by random sampling. The program’s output is a graph of the ROC curve for each selected algorithm and the AUC metric, which allows quantitative, not just visual, evaluation of the algorithm’s effectiveness. The demonstrated flowchart was used for the programmatic realization of the goal. Table 2 represents the results of the analysis of the pros and cons of each algorithm. The effectiveness of algorithms varies depending on the specific problem and dataset being used. The performance of each algorithm improved with the proper tuning of the hyperparameters and feature engineering. Among the selected machine learning algorithms that were used for training and prediction were (1) random forest, (2) k-nearest neighbors, (3) logistic regression, (4) linear discriminant analysis, (5) decision tree, (6) naïve Bayes, and (7) support vector machine.
Three datasets were taken into account: Credit Card Fraud Detection with 150.83 Mb (https://www.kaggle.com/datasets/mlgulb/creditcardfraud (accessed on 15 November 2022)). This dataset presents 284,807 transactions that occurred in two days, and only 492 of them have frauds. It means that this dataset is highly unbalanced. Credit Card Fraud with 76.28 Mb (https://www.kaggle.com/datasets/dhanushnarayananr/credit-card-fraud (accessed on 15 November 2022)). This dataset is simulated, which is why the accuracy in one of the solutions (URL: https://www.kaggle.com/datasets/dhanushnarayananr/credit-card-fraud/discussion/335338 (accessed on 15 November 2022)) is one. Fraud Detection—Credit Card with 102.92 Mb (https://www.kaggle.com/datasets/yashpaloswal/fraud-detection-credit-card (accessed on 15 November 2022)). This dataset is obtained from the first dataset by removing missing values. This is why the first dataset was used in the study. In addition, more than 4050 notebooks were developed based on the mentioned dataset. This allowed us to compare our results with those of existing methods.
The technical implementation of the task outlined in this paper was conducted using the high-level programming language Python. The simplicity, consistency, flexibility, powerful AI and ML libraries and frameworks, platform independence, and the large community of Python are widely recognized as the ideal solution for machine learning and AI-driven projects. The development environment utilized for this project was notebooks from the Kaggle platform, where the dataset was obtained. Notebooks are composed of a sequence of cells that can be formatted in either Markdown for text or a programming language of the user’s choice for code. In the technical implementation of this task, the following libraries were utilized: Scikit Learn is an open-source machine learning library that supports both supervised and unsupervised learning. Scikit Learn also provides various tools to adapt models, preprocess data, select models, and assess models, among many other services. Pandas is an open-source library primarily intended to conveniently and efficiently handle labeled or relational data. It offers several data structures and functionalities that enable numerical data and time series processing. This library was developed on the foundation of the NumPy library and is known for its fast performance and high productivity for users. Matplotlib is an open-source library that enables data visualization and plotting for Python, and it supports its numerical extension NumPy. This library provides a feasible alternative to MATLAB and is compatible with different operating systems. Creators use the matplotlib API to embed graphs in GUI applications. Because the program is implemented and stored in Kaggle, an online platform, the code of the software solution can be concurrently run in the browser of the application’s end user. This is why the need for personal computing power is not critical for running programs available to any modern computer with a sufficiently good Internet connection. This dataset is readily available on the same platform, facilitating rapid processing. Once the data are loaded into memory as a frame data structure, they are sequentially processed using the earlier preprocessing techniques: standardization and random undersampling. Standardizing the features in the training and testing sets is performed by a standard scaler, and random undersampling is used to balance the class distribution. The selected models use a set of hyperparameters. The dataset was split into training, validation, and testing sets (70%, 15%, and 15%, respectively) to measure the generalization performance. Decision tree, logistic regression, SVC, k-nearest neighbors stochastic gradient descent, naïve Bayes, and random forest algorithms were used for classification task solving. Next, grid search for hyperparameters’ tuning was used. After the initialization of the models and data processing, their successive training and evaluation using the AUC metric began and the ROC curve was visualized for each of the algorithms. After the models performed the given task, the program displayed the algorithm with the best result based on the AUC metric.
Figure 4 shows the ROC curves for these algorithms and their metrics: The decision tree algorithm produced an AUC metric of 0.938. The logistic regression algorithm obtained an AUC value of 0.946. The SVC algorithm showed an AUC of 0.936. The k-nearest neighbors algorithm showed an AUC value of 0.927. The algorithm based on stochastic gradient descent had an AUC value of 0.917. The naïve Bayes algorithm showed an AUC value of 0.908. The random forest algorithm had an AUC value of 0.911. Figure 5 shows the overall output of the program, which shows the best algorithm according to the AUC metric, namely, logistic regression, which obtained an AUC value of approximately 0.946. However, stacked generalization produced better results than logistic regression. The summary table of the results is presented in Figure 6. The AUC and F1 score were used for model evaluation. The results of stacked generalization was compared with state-of-the-art results, URL: https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud/code (accessed on 15 November 2022) and https://www.kaggle.com/code/janiobachmann/credit-fraud-dealing-with-imbalanced-datasets (accessed on 15 November 2022). The highest F1 was calculated for the ensemble model at 0.96. The easiest solution to run the program is to use the Kaggle platform (which was used to create this software solution), which, by default, supports running Jupyter on laptops. To start it, click the “Run all” button, which continues to execute all commands one by one. An alternative solution is to use any software that supports launching Jupyter laptops. This study used the root programming language Python to create a software solution because of its unique adaptation to artificial intelligence tasks and auxiliary libraries for this language: Pandas, Sklearn, and Matplotlib. The program was created and published on the Kaggle platform, so it does not depend on the end user’s computing power and only requires a stable Internet connection and a web browser. From the test runs of the presented algorithms, we concluded that they all more or less equally cope with recognizing bank fraud. This is also evident from the ROC curve plots, which do not show much visual difference. Nevertheless, based on the numerical metric AUC, we can see that the logistic regression algorithm performed best from weak classifiers, with an output AUC value of approximately 0.946.
This paper emphasized the importance of utilizing artificial intelligence to identify fraudulent banking transactions. We proposed various classification algorithms that can determine the type of transaction based on specific features. The proposed model, which is based on an artificial neural network, significantly increases the accuracy of detecting fraudulent transactions. Additionally, the paper provided multiple methods for enhancing detection accuracy, such as managing imbalanced datasets, feature transformation, and feature engineering. This paper presented the recognition of banking fraud using artificial intelligence algorithms. As a result of training and testing, each selected algorithm showed outstanding (AUC values were not lower than 0.9 in any of the cases) and equal results. This could also be seen from the graphs of the ROC curves, which do not show a significant visual difference. From the test runs of the presented algorithms, it was concluded that all of them more or less equally cope with recognizing fraudulent bank transactions. Nevertheless, based on the numerical AUC metric, the logistic regression algorithm performs best, obtaining an AUC value of approximately 0.946. The stacked generalization with deformed results of the weak classifier was proposed in the paper with an AUC of 0.008, being better than the best weak classifier. On the other side, stacking reduces bias and variance, but it is exceptionally efficient at preventing overfitting and variance. The improvement provided by the linear stacking model over the best individual model was relatively small. There was often no improvement, especially in cases where the individual base model was already sophisticated, e.g., gradient-boosted trees. The particular model’s output deformation was used in the stacked generalization. This study is important because we applied artificial intelligence to identify fraudulent banking transactions. This is particularly relevant during the pandemic, as more transactions are performed online, and during times of war, when there are many charities and events collecting money.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system for fraudulent banking operation recognition:
-
Improvement: Implement the paper's stacked generalization approach, which combines multiple weak classifiers (decision tree, logistic regression, SVC, KNN, SGD, naïve Bayes, random forest) with a meta-model to achieve an AUC of 0.954, outperforming the best individual model (logistic regression at 0.946).
-
What it can do: The system can reduce bias and variance simultaneously, preventing overfitting while improving overall detection accuracy beyond any single algorithm.
-
Improvement: Integrate the paper's random undersampling technique to address the highly imbalanced dataset (492 fraudulent transactions out of 284,807 total, i.e., 0.17% fraud rate).
-
What it can do: The system can effectively learn from extremely rare fraudulent events without being overwhelmed by the majority class, reducing false negatives where actual fraud goes undetected.
-
Improvement: Apply the paper's standardization technique (X' = (X - μ) / σ) specifically to non-encrypted features (time and amount) while preserving the PCA-transformed features.
-
What it can do: The system can normalize transaction amounts and timing across different scales, ensuring that large-value transactions don't dominate the model's learning and that temporal patterns are properly weighted.
-
Improvement: Use the paper's evaluation framework with ROC curves and AUC metrics as the primary performance indicator, rather than relying solely on accuracy or F1 scores.
-
What it can do: The system can provide threshold-independent performance assessment, allowing operators to tune the classification threshold based on their specific tolerance for false positives versus false negatives in real-world deployment.
-
Improvement: Implement the paper's approach of training seven distinct algorithms (decision tree, random forest, logistic regression, SVC, KNN, SGD, naïve Bayes) with grid search for hyperparameter optimization.
-
What it can do: The system can automatically select the best-performing algorithm for a given dataset distribution, adapting to different banking environments without manual algorithm selection.
-
Improvement: Structure the system as a binary classifier that determines whether a specific transaction at a given moment is genuine or fraudulent, using historical transaction data as training input.
-
What it can do: The system can process individual transactions in real-time, providing immediate fraud detection alerts without requiring batch processing or historical context of the current transaction.
-
Improvement: Incorporate the paper's handling of PCA-encrypted features, which maintains confidentiality of customer data while still enabling effective fraud detection.
-
What it can do: The system can operate on encrypted financial data without compromising detection accuracy, making it compliant with data protection regulations while maintaining security.
-
Improvement: Design the system to run on cloud-based notebook environments (like Kaggle) that require no local computing power.
-
What it can do: The system can be deployed and accessed from any device with internet connectivity, enabling rapid implementation across multiple banking institutions without infrastructure investment.
-
Real-time fraud detection with 95.4% AUC accuracy, capable of flagging suspicious transactions as they occur.
-
Adaptive algorithm selection that automatically identifies the best-performing model for the specific data distribution of each banking institution.
-
Handling of extreme class imbalance (fraud rates as low as 0.17%) without sacrificing precision or recall.
-
Privacy-compliant operation on encrypted transaction data, protecting customer identities while maintaining detection capability.
-
Cross-platform deployment requiring only a web browser and internet connection, eliminating hardware dependencies.
-
Threshold tuning capability allowing banks to adjust sensitivity based on their risk tolerance and regulatory requirements.
-
Ensemble robustness that maintains high performance even when individual algorithms fail on specific transaction patterns.
Abstract
This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID19 pandemic, bank fraud has become even more common due to the massive transition of many operations to online platforms and the creation of many charitable funds that criminals can use to deceive users. The present work focuses on machine learning algorithms as a tool well suited for analyzing and recognizing online banking transactions. The study`s scientific novelty is the development of machine learning models for identifying fraudulent banking transactions and techniques for preprocessing bank data for further comparison and selection of the best results. This paper also details various methods for improving detection accuracy, i.e., handling highly imbalanced datasets, feature transformation, and feature engineering. The proposed model, which is based on an artificial neural network, effectively improves the accuracy of fraudulent transaction detection. The results of the different algorithms are visualized, and the logistic regression algorithm performs the best, with an output AUC value of approximately 0,946. The stacked generalization shows a better AUC of 0.954. The recognition of banking fraud using artificial intelligence algorithms is a topical issue in our digital society.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks