Discovering Persistent Behavioural Patterns for Interpretable Blockchain Forensics
Dorottya Zelenyanszki, Zhe Hou, Kamanashis Biswas, Vallipuram Muthukkumarasamy
Griffith University · Australian Catholic University
cs.CR, cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/SunWeb3Sec/DeFiHackLabs
Project page: https://faiss.ai/index.html
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: This paper proposes a scalable, application-agnostic framework for persistent behavioural pattern discovery from large-scale blockchain activity.
Terminology
Summary
This paper proposes a scalable, application-agnostic framework for persistent behavioural pattern discovery from large-scale blockchain activity. It constructs behaviour sentences enriched with contract, token and market context, then applies a two-step embedding process: sentence-level embeddings capture individual actions, while sequence-level embeddings capture user behaviour over time. An interpretable behavioural profiler characterizes discovered communities through behavioural motifs, routines, temporal dynamics, entity exposure, and suspiciousness evidence. Evaluation on Ethereum using over 30 million transactions shows that the framework uncovers both routine and malicious behavioural patterns, including decentralised exchange (DEX) trading, NFT activity, phishing, bot operations, oracle manipulation, and rug-pull schemes. Importantly, many patterns remain stable across independent observation windows, enabling the identification of long-term behaviours beyond a single analysis period. The proposed framework combines scalability, interpretability, and persistence analysis, supporting blockchain forensic investigation, behavioural attribution, and threat discovery.
The main contributions of this paper are threefold:
• We propose a scalable, application-agnostic framework for persistent behavioural pattern discovery that converts blockchain transactions and decoded logs into behavioural sentences preserving action order, event semantics, entity types, and market signals learns user-level sequence representations, clusters users without labels, and profiles the resulting behavioural groups.
• We establish the most suitable framework setting through a systematic, multi-stage evaluation of sentence embedding models, sequence aggregation methods, training objectives, sequence-length strategies, clustering algorithms and resolutions, and behavioural semantics, balancing representation quality, clustering structure and stability, scalability, length leakage, and the concentration of pre-labelled behaviours.
• We develop an interpretable behavioural profiler that explains discovered clusters through behavioural motifs, ordered routines, activity and temporal patterns, diversity, concentration, entity exposure, and separated direct-label and exposure-based suspiciousness evidence, enabling the identification of persistent ordinary and suspicious behavioural patterns across independent monthly and cumulative multi-month windows.
The framework operates on user histories, preserves sequential activity, and does not require labels for malicious or benign, but to reveal how ordinary and suspicious behaviours are organised across the population and whether these patterns persist over time. It also produces interpretable cluster-level and tag-level profiles while remaining practical at million-user scale.
The proposed framework is evaluated on Ethereum data covering blocks 16250000 to 16749999, collected from the XBlock-ETH dataset. The first iteration on the first-month window yielded 6 045 497 users, 3 117 292 contracts, 1 343 354 ERC-20 tokens, 259 062 ERC-721 tokens, with 1 317 947 addresses filtered out. External suspiciousness labels are collected from community and research sources, including the De.Fi REKT database, DefiLlama, DeFiHackLabs, ImmuneBytes reports, and the bot-address dataset of Niedermayer et al. (2024). The final pre-labelled set contains 358 unique addresses, with category-level assignments including 68 phishing users, 50 phishing contracts, 94 malicious contracts, 96 exploiters, and 58 exploited contracts. These labels are not used to train, guide, or constrain clustering; they are introduced only after clustering and profiling as high-confidence external evidence.
The framework consists of several stages: data collection and preprocessing, sentence formation, a two-step embedding process, clustering, and a behavioural profiler. Sentence formation converts transactions and decoded logs into compact behavioural text while preserving action order, event semantics, entity types, and market signals. The two-step embedding process maps each retained user's behavioural sentence sequence to a compact user-level representation. In the first step, a pretrained sentence model maps each sentence to a fixed-length vector. In the second step, the sentence embedding sequence is converted into a single user embedding using a sequence aggregation model. The clustering stage assigns retained users to behavioural groups in a fully unsupervised manner. After clustering, the framework applies a behavioural profiler to describe the discovered groups, deriving motif abstractions, ordered motif sequences, activity and temporal profiles, entity-exposure profiles, and contract and token diversity or concentration patterns.
The experiments address two research questions. For RQ1, the framework setting is selected through a systematic evaluation. At the sentence level, MiniLM-L6-v2 with num bucket preprocessing, no prompt, and post norm achieved the best overall profile, with the lowest global similarity mean, highest global similarity spread, and largest separation margin, while maintaining strong ranking and AUC values. At the sequence level, GRU variants had the lowest length leakage, with eta-squared gaps around 0.277 to 0.309, while CNN was excluded from the behavioural analysis because of length leakage. The training objective sweep showed that the best objective was setting-dependent: masked was strongest for most GRU configurations, while GRU with raw normalisation and Mamba2 favoured nextstep. The length and cropping sweep showed that the best setting was target-dependent, with GRU variants preferring either max len = 32 or 128, and Mamba2 performing best with max len = 32. The clustering algorithm selection showed that the K-means family produced the most reliable geometric structure, with full K-means retained for the learned GRU and Mamba2 embeddings. Combining structural results with pre-labelled concentration results, the length ≥ 2 Mamba2 × MiniLM-L6-v2 × nextstep × L2 setting is selected for downstream behavioural analysis, using max len = 32, prefix cropping, uniform sampling, PCA-128, cosine full K-means, and k = 22.
For RQ2, the proposed framework is compared with two baselines and two alternative embedding methods. The feature baseline uses account-level features, BUBA is a graph-based behavioural pipeline, and TF-IDF+SVD and Doc2Vec replace the two-step embedding process. The results show that BUBA achieves the strongest geometric clustering results, with the highest SC and CHI and the lowest DBI, but its leakage values indicate that activity scale also contributes to the recovered structure. The feature baseline shows even stronger leakage. TF-IDF+SVD and Doc2Vec provide useful text-based alternatives, but their low SC and high DBI indicate weaker overall cluster structure. In contrast, the proposed framework does not maximise the geometric clustering metrics, but it achieves the lowest length/activity leakage, the highest Sep-NMI, the highest labelled-cluster purity, and the smallest one-cluster malicious concentration. This makes it the only method that remains consistently strong across cluster geometry, leakage control, and pre-labelled category separation.
For RQ3, the behavioural analysis of the first-month clusters identifies five main behavioural families: stablecoin/transfer/approval, DEX/AMM interaction, NFT flow, mint/claim/reward, and ENS/name-service/bridge-like. The suspiciousness reading is grouped into three interpretation levels: meaningfully suspicious, watchlist, and mostly ordinary/exposure evidence. Meaningfully suspicious clusters include Cluster 5, a concentrated DEX/AMM community with bot, exploit, bug-exploit, phishing, and front-run evidence; Cluster 8, linking NFT transfer and approval behaviour with phishing and exploit evidence; and Cluster 13, the clearest phishing-focused case, combining approval management, phishing evidence, and asset-movement indicators. The profiler interprets clusters through behavioural families, and suspiciousness is treated as variation within these settings, linking phishing, bot/front-running, and exploit-related tags to coherent behavioural evidence without treating exposure as proof of cluster-wide maliciousness.
For RQ4, the long-term analysis uses persistence as a validation signal, checking whether tags reappear in the independent second month or in the cumulative 1–2, 1–3, and 1–6 month windows. Ordinary behavioural structure is stable across windows, with approval-linked transfer, approval-linked swap, DEX trading, liquidity provision, NFT market and approval behaviour, mint/claim campaigns, repeated contract or protocol interaction, and stablecoin-family movement all recurring across independent and cumulative settings. The suspiciousness tags show three broad persistence patterns: some are visible early and persist, including phishing, bot, exploit, bug exploit, and front-run; others become clearer in cumulative windows, especially oracle manipulation and rug pull; sparse subtypes such as flash-loan attack, governance attack, MEV exploit, replay attack, and price manipulation mainly appear in longer windows. Suspicious tags are interpreted through behavioural alignment rather than labels alone, with phishing tied to approval management, approval-for-all, NFT movement, and asset outflow; bot and front-running behaviour linked to repeated DEX/AMM swaps, routing, approvals, liquidity interaction, regular activity, and concentrated protocol use; and exploit and bug-exploit evidence appearing as a risk layer inside ordinary DEX, NFT, ENS/name-service, transfer, and watchlist communities.
The indicator distribution is strongly long-tailed: a small number of shared indicators form a DeFi and permission backbone, while most indicators remain tag-specific. Shared indicators include swaps, routing, pools, liquidity interaction, approvals, repeated protocol use, exploit or risk exposure, and concentration evidence. Approval is especially important because it connects ordinary DeFi and NFT activity with phishing, approval-heavy NFT behaviour, and exploit-related interpretations, but it becomes meaningful only when aligned with asset movement, NFT transfer, phishing-related evidence, risky-entity exposure, or exploit-transaction exposure.
The limitations of this work include the small pre-labelled set, the cluster-level nature of suspiciousness tags, and the dependence on decoded events, address classification, token metadata, external risk references, and API-supported enrichment. Future work will extend the analysis across a range of chains and applications, compare with classification-based and attack-specific baselines, and evaluate the learned representations on downstream risk classification and prediction.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
1. Persistent Behavior-Aware User Embedding
-
Improve user representation learning by incorporating sequential behavioral sentences (ordered actions, event semantics, entity types, market signals) instead of static feature vectors.
-
The improved AI system can generate time-aware embeddings that capture both short-term actions and long-term routines, enabling better prediction of future behavior (e.g., fraud detection, churn prediction) across arbitrary time windows.
2. Unsupervised Behavioral Clustering with Leakage Control
-
Enhance clustering algorithms to explicitly minimize length/activity leakage (e.g., using GRU with raw normalization and nextstep objective) while preserving geometric structure.
-
The improved system can group users into interpretable behavioral families (e.g., DEX traders, NFT enthusiasts, phishing bots) without labeled data, and maintain stable clusters across independent observation periods, making it robust to data drift.
3. Interpretable Behavioral Profiling for Explainable AI
-
Integrate a profiler that generates human-readable explanations for each cluster using motifs, ordered routines, temporal dynamics, entity exposure, and separated suspiciousness evidence.
-
The improved AI system can provide transparent reasoning for why a user is flagged (e.g.,
frequent approval-for-all followed by NFT transfers to phishing-linked contracts
), supporting forensic investigations and regulatory compliance.
4. Multi-Scale Persistence Analysis for Anomaly Detection
-
Implement a persistence-checking mechanism that validates whether behavioral patterns recur across independent and cumulative time windows.
-
The improved system can distinguish between transient anomalies (e.g., one-off rug pulls) and persistent threats (e.g., ongoing bot operations), enabling early warning systems that prioritize long-term risks over noise.
5. Application-Agnostic Pipeline for Cross-Domain Generalization
-
Adopt the framework’s modular design (sentence formation → two-step embedding → clustering → profiling) that does not require domain-specific labels or feature engineering.
-
The improved AI system can be applied to other sequential event data (e.g., financial transactions, network logs, IoT sensor streams) to discover behavioral patterns and suspicious activities without retraining from scratch.
6. Hybrid Label Integration for Risk Assessment
-
Use external high-confidence labels (e.g., known phishing addresses) only for post-hoc validation, not training, to avoid bias.
-
The improved system can produce risk scores that separate direct-label evidence from exposure-based evidence, reducing false positives in threat discovery while maintaining high recall for genuinely malicious actors.
7. Long-Tailed Indicator Modeling for Rare Event Detection
-
Leverage the finding that shared indicators (e.g., approvals, swaps) form a backbone, while most indicators are tag-specific.
-
The improved AI system can learn to weight rare, tag-specific signals (e.g., oracle manipulation) more heavily when they co-occur with common behavioral backbones, improving detection of sparse attack types that traditional models miss.
8. Scalable Million-User Processing with Memory Efficiency
-
Incorporate the framework’s practical scalability (e.g., PCA-128, cosine K-means, prefix cropping) to handle millions of users.
-
The improved AI system can run real-time behavioral profiling on large-scale blockchain or enterprise data with limited computational resources, enabling deployment in production environments.
Abstract
Public blockchain data enables large-scale DeFi-related analysis, but many existing approaches are application-specific, difficult to scale, or hard to interpret. This research proposes a scalable, application-agnostic framework for persistent behavioural pattern discovery from large-scale blockchain activity. It constructs behaviour sentences enriched with contract, token and market context, then applies a two-step embedding process: sentence-level embeddings capture individual actions, while sequence-level embeddings capture user behaviour over time. An interpretable behavioural profiler characterizes discovered communities through behavioural motifs, routines, temporal dynamics, entity exposure, and suspiciousness evidence. Evaluation on Ethereum using over 30 million transactions shows that the framework uncovers both routine and malicious behavioural patterns, including decentralised exchange (DEX) trading, NFT activity, phishing, bot operations, oracle manipulation, and rug-pull schemes. Importantly, many patterns remain stable across independent observation windows, enabling the identification of long-term behaviours beyond a single analysis period. The proposed framework combines scalability, interpretability, and persistence analysis, supporting blockchain forensic investigation, behavioural attribution, and threat discovery.
Sources
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- RISKTAGGER: Evidence-Guided LLM Agent for Post-Incident Forensic Analysis of Money Laundering in Web3
- SoK: Comprehensive Analysis of Rug Pull Causes, Datasets, and Detection Tools in DeFi
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- C-Pack: Packed Resources For General Chinese Embeddings
- Detecting Various DeFi Price Manipulations with LLM Reasoning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs