Using Machine Learning to Detect Fraudulent SMSs in Chichewa

arXiv:2502.16947 · cs.LG, cs.CL · Submitted 2025-02-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Using Machine Learning to Detect Fraudulent SMSs in Chichewa".

Jane: The paper was written by Amelia Taylor and Amoss Robert from Kuyesera AI Lab, Malawi University of Business and Applied Sciences and Malawi University of Business and Applied Sciences, Blantyre, Malawi and Kuyesera AI Lab, Blantyre, Malawi.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: We’re talking about the paper titled "Using Machine Learning to Detect Fraudulent SMS in Chichewa," written by Amelia Taylor and Amoss Robert. They are tackling a major problem because most existing research focuses on English datasets, but this paper is crucial for addressing a huge gap. It’s important to acknowledge that SMS-enabled fraud is a global concern, and the authors are introducing an entirely new resource tailored specifically for the language used across Southern Africa.

Jane: Exactly, Tom. Their work addresses the scarcity of data in local languages like Chichewa, which is a Bantu language widely spoken in Malawi and neighboring countries. They aren't just looking at statistics; they're building a foundational dataset that allows for machine classification in Chichewa as fraud or non-fraud for the first time.

Lu: I find the scope of this work incredibly exciting, Meng. It moves beyond mere data collection; it’s about cultural intelligence, recognizing that AI can't be truly useful if it doesn's respect local linguistic needs. The authors are laying groundwork for a level of precision that current global models simply lack in African contexts.

Meng: And from the practical side, this is a huge win for implementation. By creating this dedicated dataset, they’ are providing a clear blueprint for future developers who need to build localized fraud detection systems without relying on flawed translations or generic English tools.

Lalam: Lalam sees the cultural grounding here as vital to ensure that the model learns what "fraud" means within the actual lived experience of a Malawian community, not just abstractly defined in a lab setting based on foreign examples. This is about recognizing local patterns of deception.

Tom: That’s a powerful distinction, Jane. Given this foundation of localized data and addressing the global problem of SMS fraud, we need to look at how they structured their approach to manage that challenge.

The Methodology and Challenges: Jane: In "Using Machine Learning to Detect Fraudulent SMS in Chichewa," the researchers didn't just pull random data; they used primary research from students in Blantyre city for a small initial set of SMS messages. This foundation was then expanded using techniques called data augmentation, which is a clever way to increase the dataset size while ensuring that the meaning of the original message remains intact.

Tom: That's interesting because simply adding more messages might dilute the quality. So, how did they achieve this expansion?

Lu: The way they applied label-preserving transformations is a highly creative method for ensuring semantic integrity. They aren't just making up data; they are subtly altering existing fraudulent or normal messages to create variations that would still be realistic in the actual usage of expanding the corpus.

Meng: I'm interested in how they handled the fact that their initial data was so limited—it’s a common hurdle when scaling solutions in developing regions. Using crowdsourcing to get those initial one hundred one fraudulent messages is a pragmatic, scalable solution for resource-constrained environments.

Lalam: Lalam sees this as a way of community participation, too, because by involving local participants in the labeling and collection process, they are integrating the community into the development of an AI tool that represents their reality. This is about shared ownership of data.

Tom: That’s a very human element to add to the technical side. But once you have this initial data, how did they prepare it for machine learning?

Jane: They created three distinct datasets: the original Chichewa set (D-CHI), the human-translated set (D-HT), and the machine-translated set (DMT). This structure allows them to test how sensitive different classification models are to translation errors and helps us understand the challenges inherent in cross-lingual NLP tasks.

Lu: It’s a brilliant comparative setup, Meng, allowing us to see exactly where the weaknesses in automated translation lie.

The Findings and Implications: Tom: Now, looking at the results of "Using Machine Learning to Detect Fraudulent SMS in Chichewa," the findings are quite striking. They found that Random Forest (RF) and Support Vector Machine (SVM) performed the best across all datasets, while Naïve-Bayes was consistently the weakest performer.

Jane: The accuracy numbers are incredibly high, often reaching over ninety-six percent or ninety-seven percent, which is a huge achievement for a low-resource language dataset. It shows that when you have high-quality, well-curated data, these models can be very effective at catching fraud.

Lu: I’m particularly interested in the "concept drift" problem they mention—the fact that fraudsters change their wording. The performance of Random Forest suggests that these ensemble methods are robust enough to capture those evolving fraud patterns over time, which is a huge advantage for continuous monitoring.

Meng: While the high accuracy on D-CHI is encouraging, I’m focused on the significant drop in performance when moving to the machine-translated dataset (DMT). This highlights a critical dependency on preprocessing tools and shows where we need to invest in better translation pipelines.

Lalam: Lalam sees this drop as a warning: if AI relies heavily on imperfect translations, it can fail at correctly identifying genuine fraud. We must ensure that localization remains the focus, prioritizing human-validated models over automated translation for the future work.

Tom: It’s clear that localized intelligence is key to optimizing performance. Given these high accuracy results and challenges with what's translated, what is the practical impact of this research?

Jane: The authors suggest that developing language-specific models are essential for achieving optimal accuracy in combating SMS fraud, moving beyond simply using general English tools or relying on machine translations.

Conclusion and Wrap-up: Tom: We’ve covered so much ground today regarding the findings of "Using Machine Learning to Detect Fraudulent SMS in Chichewa." It’s clear this has been a huge technical and policy success, addressing both a research gap and a real-world security need.

Jane: I agree, Tom. The effort to build a dedicated dataset for an under-represented language is truly a major win for advancing AI in Africa by making sure the technology is accessible everywhere.

Lu: This work sets the stage for future research, especially when we consider moving beyond these traditional models and exploring more sophisticated deep learning architectures like BERT, which is mentioned as a next step.

Meng: I see this as a roadmap: it tells us exactly where our current tools struggle with translation and context loss. It shows us precisely where to focus our engineering efforts to build robust, localized fraud detection systems for the people who speak Chichewa.

Lalam: Lalam believes that this work is providing a blueprint for cultural intelligence in AI, ensuring that the digital divide is bridged by respecting the unique linguistic needs of communities globally.

Tom: It’s genuinely inspiring work, making sure our technology serves all of us equally and effectively in a global context. We’ll be wrapping up this discussion on "Using Machine Learning to Detect Fraudulent SMS in Chichewa" right now, but we’re looking forward to talking about the next paper with you soon. Goodbye everyone!

Amelia Taylor, Amoss Robert

Kuyesera AI Lab, Malawi University of Business and Applied Sciences · Malawi University of Business and Applied Sciences, Blantyre, Malawi · Kuyesera AI Lab, Blantyre, Malawi

cs.LG, cs.CL

Submitted: 2025-02-24

Updated: 2026-08-20

DOI: 10.1007/978-3-032-08260-2_12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 78/100

The gist: The paper "Using Machine Learning to Detect Fraudulent SMS in Chichewa" addresses the global concern regarding SMS-enabled fraud, which globally resulted in large losses of "5.8 bn in 2023." Given

Key concepts

Chichewa
Chichewa is a Bantu language widely spoken in Malawi and neighboring countries. The paper focuses on this language because most existing research on SMS fraud detection concentrates on English datasets, creating a significant data gap for local languages.
Data Augmentation
This technique was used to increase the size of the initial dataset without compromising quality. It involves applying label-preserving transformations to subtly alter existing messages, ensuring that variations remain realistic and maintain their original meaning.
Random Forest (RF) and Support Vector Machine (SVM)
These were the machine learning models found to perform the best in detecting fraudulent SMS messages across all tested datasets. They achieved high accuracy, suggesting they are robust methods for identifying fraud patterns.
Concept Drift
This refers to the problem where fraudsters change their wording or tactics over time. The performance of Random Forest suggested that ensemble methods are robust enough to capture these evolving and changing fraud patterns for continuous monitoring.

Terminology

Summary

The paper Using Machine Learning to Detect Fraudulent SMS in Chichewa addresses the global concern regarding SMS-enabled fraud, which globally resulted in large losses of 5.8 bn in 2023. Given that Chichewa is a major language used most widely for communication in Malawi and is considered a low-resourced language with few Natural Language Processing (NLP) tools available, the research aims to address the scarcity of datasets and NLP tools for this language.

The primary objectives were twofold:

  1. To construct a dataset of SMS fraud in Chichewa which can be used for understanding the anatomy of SMS-enabled fraud.

  2. To build and test machine learning classification algorithms that can effectively classify Chichewa SMS messages as fraud or normal.

Methodology and Dataset Construction:

The researchers employed a crowdsourcing approach, collecting data from colleagues at the Malawi University of Business and Applied Sciences (MUBAS) who had been victims of SMS fraud, as well as obtaining fraudulent SMS messages from local telecommunication companies. The data collection period spanned March to August 2023.

To increase the size of the initial small dataset collected through primary research, the researchers applied label-preserving text transformations (data augmentation). This process involved four specific transformations:

  1. The addition of words implied by the context but not explicitly stated.

  2. Replacing words with their synonyms or equivalent expressions to maintain meaning.

  3. Replacing borrowed words (e from English with their equivalents in the vernacular Chichewa language).

  4. Altering the morphological structure of words by adding morphemes like 'nd' or 'mna'.

The resulting dataset was categorized into three versions for analysis: the original Chichewa dataset (DCHI), an enlarged version using augmentation (D-CHI), and two translation-based datasets: the human-translated dataset (D-HT) and the machine-translated dataset (DMT).

Machine Learning Algorithms:

The study utilized a supervised learning approach with TFIDF for feature extraction. The algorithms tested included Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and Naïve-Bayes (NB). For all experiments, the data was split into 80% for training and 20% for testing.

Results: Characteristics of Fraudulent SMS:

The analysis revealed distinct patterns in fraudulent messages. The study found that Fraudsters exploit local knowledge and the vulnerabilities of specific segments of the population by referencing legitimate programs and initiatives to bait their victims. Key expressions identified included references to South Africa, Transporter, and exploitation of local schemes like Mtukula Pakhomo or promises of miracle money. These tactics include:

  • Messages impersonating government and NGO initiatives.

  • SMS messages claiming money was accidentally sent and requesting return.

  • Messages from transporters asking for funds to clear parcels at the border.

Results: Performance of Machine Learning Models:

The performance results were analyzed across the three balanced datasets (D-CHI, D-HT, D-MT) and their extended versions (D-CHIe, D-HTe, D-MTe).

  • Performance on Original Data (D-CHI): The best performing models were RF and SVM. On the original dataset, RF (best) on D-CHI gives an accuracy of 0.98, and an AUC-ROC of 0.99 and a FN rate of 3%, while SVM (best) has an accuracy of 0.96, AUC-ROC of 0.99 and a FN rate of 0%.

  • Impact of Translation: A significant finding was the impact of translation quality on model performance: The results show that D-HT performs better in terms of all metrics for most of the models except for RF, while the gap between the performance of NB on D-CHI and on the translated dataset if the largest of all the models.

  • Impact of Class Balance: The study found that extending datasets with more normal messages leads to an improvement in performance mainly for the original D-CHI dataset, though this often resulted in a deterioration in detection accuracy.

Discussion and Conclusion:

The research highlighted several critical insights:

  1. Cross-Lingual Challenges: The findings demonstrated that models trained on one language cannot always be directly transferred to another and that preprocessing tools must be adapted to the linguistic nuances of each language. Specifically, a loss of context during translation that results in a deterioration in performance was observed when relying on machine-translated text.

  2. Preprocessing Necessity: The study found that removing punctuation and stop words had a negative impact on performance for all models on D-CHI, suggesting that Chichewa requires language-specific preprocessing to preserve key contextual information.

  3. Policy Implications: The research concluded that investing in language technology research and infrastructure for underrepresented languages is essential to bridge the digital divide.

In summary, the paper successfully introduced a valuable resource—the first Chichewa SMS fraud dataset—and provided an extensive set of machine learning experiments, underscoring that developing language specific models for SMS fraud detection is critical to optimize accuracy and performance.

Improvements for AI systems

(Note: Since the input provided is a bibliography of related works rather than a single arXiv paper, my improvements synthesize best practices and advanced techniques derived from the collective themes of these citations, focusing on creating a state-of-the-art, robust, and highly specialized NLP system.)

The current models often rely on older feature extraction methods (e.g., Bag-of-Words or simple Word2Vec). The system must be upgraded to utilize a multi-stage, hybrid architecture centered around large pre-trained language models (LLMs) adapted for low-resource contexts.

  • Specific Improvement: Implement Cross-lingual Transfer Learning (CLTL) by fine-tuning advanced transformer architectures (e.g., XLM-RoBERTa or specialized BERT variants like MasakhaNER 2.0) on high-resource languages, and then adapting the resulting weights for target low-resource African languages (e.g., Swahili, Chichewa).

  • Specific Improvement: Integrate a Hierarchical Attention Mechanism layer after the embedding extraction stage. This mechanism will weigh the importance of different tokens and phrases within a message based on context, moving beyond simple sequence classification to identify subtle semantic intent shifts indicative of fraud or spam.

The system must move decisively away from binary (Spam/Not Spam) classification, as suggested by the Multiclass studies.

  • Specific Improvement: Develop a Multi-Label Classification Head. Instead of outputting a single probability score, the model will concurrently predict membership across several orthogonal classes:

  • [Spam Phishing]: Attempting credential theft.

  • [Spam Malware]: Direct download links or malicious instructions.

  • [Commercial Promotion]: Unsolicited advertising.

  • [Service Alert]: Legitimate but non-urgent notifications (e.g., bank updates).

  • [Language Anomaly]: Detection of syntactical or grammatical patterns inconsistent with natural language use in the target dialect, flagging potential machine generation or obfuscation.

To counter Evasive Techniques (as noted in the literature), the model must incorporate active defense mechanisms.

  • Specific Improvement: Implement a Gradient Masking Module trained on syntactically altered versions of known spam messages (e.g., character substitutions, phonetic misspelling, or inserted irrelevant characters). This forces the model to learn semantic robustness rather than relying solely on keyword matching.

  • Specific Improvement: Integrate an Uncertainty Quantification Layer. When the model predicts a class with low confidence (high entropy across top classes), it must automatically flag the message for human review, preventing high-stakes failures due to ambiguous, novel spam vectors.


The resulting system will be a Context-Aware, Multilingual Intent Classifier for Digital Communications.

  1. Detect Complex Fraud and Spam: It can accurately identify sophisticated phishing attempts and malware distribution SMS messages by analyzing deep contextual intent rather than just keywords, even when the message has been deliberately obfuscated or written in non-standard dialects.

  2. Provide Granular Threat Intelligence: Instead of a simple Spam flag, it will output a detailed threat profile (e.g., High confidence: Phishing attempt targeting credentials; Secondary risk: Commercial promotion overlay).

  3. Operate Across Linguistic Borders: It can process and classify messages in multiple low-resource African languages simultaneously, providing consistent performance regardless of the specific dialect or language mix used in the incoming SMS stream.

  4. Adapt to Novel Threats: By incorporating a continuous retraining loop using adversarial examples and flagging low-confidence predictions, the system achieves near real-time adaptation to emerging spam taxonomies and evasion tactics without requiring full model retraining cycles.

Abstract

SMS enabled fraud is of great concern globally. Building classifiers based on machine learning for SMS fraud requires the use of suitable datasets for model training and validation. Most research has centred on the use of datasets of SMSs in English. This paper introduces a first dataset for SMS fraud detection in Chichewa, a major language in Africa, and reports on experiments with machine learning algorithms for classifying SMSs in Chichewa as fraud or non-fraud. We answer the broader research question of how feasible it is to develop machine learning classification models for Chichewa SMSs. To do that, we created three datasets. A small dataset of SMS in Chichewa was collected through primary research from a segment of the young population. We applied a label-preserving text transformations to increase its size. The enlarged dataset was translated into English using two approaches: human translation and machine translation. The Chichewa and the translated datasets were subjected to machine classification using random forest and logistic regression. Our findings indicate that both models achieved a promising accuracy of over 96% on the Chichewa dataset. There was a drop in performance when moving from the Chichewa to the translated dataset. This highlights the importance of data preprocessing, especially in multilingual or cross-lingual NLP tasks, and shows the challenges of relying on machine-translated text for training machine learning models. Our results underscore the importance of developing language specific models for SMS fraud detection to optimise accuracy and performance. Since most machine learning models require data preprocessing, it is essential to investigate the impact of the reliance on English-specific tools for data preprocessing.

Sources

Related papers