Gravity Falls: A Comparative Analysis of Domain-Generation Algorithm (DGA) Detection Methods for Mobile Device Spearphishing

arXiv:2603.03270 · cs.CR, cs.LG, cs.NI · Submitted 2026-03-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Gravity Falls: A Comparative Analysis of Domain-Generation Algorithm (DGA) Detection Methods for Mobile Device Spearphishing".

Elias: Mobile devices are frequently targeted by eCrime threat actors using SMS spearphishing links that employ Domain Generation Algorithms (DGA) to rotate hostile infrastructure,

Nadia: First, who's behind it and why it matters.

Paper summary: Nadia: So, wrapping up our discussion on "Gravity Falls: A Comparative Analysis of Domain-Generation Algorithm (DGA) Detection Methods for Mobile Device Spearphishing," the paper by Wong and Hastings really lays out how DGA detection needs to evolve beyond just looking for simple randomness. They showed that performance is highly dependent on the specific tactic used, finding that traditional methods like Exp0se excel at randomized strings while struggling with dictionary-based or themed attacks.

Elias: And from my perspective as a cryptographer, the paper’s analysis of what makes those different domain structures hard to spot really highlights how much information is hidden in those concatenations and word choices one. The study demonstrates that detectors need to account for these specific structural changes in the domain string, not just general algorithmic properties.

Priya: I think what resonates most with me is the practical implication of seeing this evolution across four distinct clusters over three years; it shows that attackers are systematically adapting their methods to evade detection in a way that’s directly relevant to our current mobile threat landscape. The data really paints a picture of how the threat actor shifts their behavior.

Nadia: It certainly does, Priya. The title itself, "Gravity Falls," suggests a deep dive into this evolving threat actor's playbook, and the authors make it clear that we need to move past simply checking if something is an algorithm to understanding the specific generation tactic at play one.

Elias: And given the results they found regarding machine learning detectors showing limited generalization beyond the initial randomized strings, I see a clear direction for future work focusing on hybrid models that combine lexical analysis with richer context signals from things like message content.

Priya: That leads directly to the idea of needing more than just string analysis; we need to integrate those contextual elements they mentioned as important for defense against dictionary and combo-squatting variants one. It suggests that the future isn't about one perfect detector, but a combination of methods.

Nadia: Exactly, Priya. The paper concludes that for immediate defensive value, it supports a layered approach: use fast lexical heuristics for randomized domains but then rely on those contextual signals—like infrastructure and brand abuse policies—when you encounter those trickier dictionary and combo-squatting tactics one.

Elias: That layered defense strategy seems to be the practical conclusion derived from their comparative analysis of the different DGA techniques they tested against each other one. It gives us a clear roadmap for improving how we approach these mobile threats.

Conclusion: Nadia: So, we've been digging into how these new DGA detectors handle those tricky smishing tactics across different clusters, and now it’s time to really talk about what this whole paper means for us. Elias, what are your thoughts on the title and who wrote this research?

Elias: I think the title perfectly frames the issue because it shows they aren't just looking at one type of attack; they're comparing different detection strategies against a whole spectrum of generation techniques. The authors, Wong and Hastings, have clearly put together a rigorous comparison to see where each method actually holds up.

Priya: From my side, I’m focused on what the actual data reveals about these attacks. The paper shows that the success of any detector really hinges on whether it targets simple randomness or those more complex patterns like dictionary words and themed stuff. That distinction is key for understanding the real-world risk.

Nadia: Exactly, Priya, and that leads to a big question for us: who can actually exploit these findings? Can an attacker easily build a system that bypasses all these detectors by blending different tactics?

Elias: That's where the paper’s finding about generalization comes in; the authors suggest that relying on just one type of detection isn't enough because those ML models struggle when the tactic shifts outside of what they were trained on.

Priya: It really underscores that privacy and measurement researchers need to pay attention to these subtle shifts in data collection, like how they built that "Gravity Falls" dataset itself, because the quality of the input directly impacts what we learn.

Nadia: So, looking at the authors' conclusion about layered defense—using quick lexical checks for randomness but adding context for dictionary attacks—what does this imply for how security teams should actually structure their defenses?

Elias: It implies that a single algorithmic test won't cut it anymore; you need to combine fast string analysis with external signals, like message content or where the domain is hosted, to get a reliable picture.

Priya: The implication is that we can’t just build one perfect guard against these evolving threats; we have to build a system that monitors multiple layers of information simultaneously.

Nadia: That sounds like a solid direction for our listeners, showing them that defense has to become much more comprehensive and adaptive than it was before. So, where do you think this research opens up the door for future work in this area?

The Beacom College of Computer & Cyber Sciences · Dakota State University

cs.CR, cs.LG, cs.NI

Submitted: 2026-03-03

Updated: 2026-09-28

Comments: 7 pages. Disclaimer: The views expressed are those of the authors and do not necessarily reflect the official policy or position of the U.S. Department of Defense or the U.S. Government. References to external sites do not constitute endorsement. Cleared for release on 24 FEB 2026 (DOPSR 26-T-0771). Gravity Falls Dataset DOI: 10.5281/zenodo.17624554

Code: https://github.com/MalwareMorghulis/GravityFalls

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 74/100

The gist: Mobile devices are frequently targeted by eCrime threat actors using SMS spearphishing links that employ Domain Generation Algorithms (DGA) to rotate hostile infrastructure, but research has largely

Key concepts

Domain Generation Algorithm (DGA)
A technique where malicious software automatically generates a large number of potential domain names using an algorithm, hoping one will be registered and used by the attacker to host their phishing site. This makes it hard for security systems to block all possibilities at once.
Gravity Falls Dataset
A new, semi-synthetic collection of C2 domains gathered from SMS messages between 2022 and 2025. It was created by observing smishing links and organizing them into four clusters representing different evolving threat tactics like theme-based phishing.
Shannon Entropy
A mathematical measure used to quantify the randomness or information content within a domain name string. Higher entropy suggests a more random string, which traditional DGA detectors often use as a primary indicator of algorithmic generation.

Terminology

Summary

Mobile devices are frequently targeted by eCrime threat actors using SMS spearphishing links that employ Domain Generation Algorithms (DGA) to rotate hostile infrastructure, but research has largely overlooked how well detectors generalize to smishing-driven domain tactics outside enterprise perimeters. This work addresses that gap by evaluating traditional and machine-learning DGA detectors against the Gravity Falls dataset, a new semi-synthetic collection derived from smishing links between 2022 and 2025.

The gist

Performance is highest on randomized-string domains but drops on dictionary concatenation and themed combo-squatting, with low recall across multiple tool/cluster pairings.

Dataset Construction and Evolution

Gravity Falls is a semi-synthetic dataset consisting of C2 domains delivered via SMS text messages between 2022 and 2025, organized into four technique clusters reflecting an annual evolution in the threat actor’s TTPs. These clusters are:

  1. Cats Cradle (2022): Perceived Technique is randomized letters, use of randomized alphabetical characters within 5-8 characters. Assessed Purpose is target validation through fake CAPTCHA.

  2. Double Helix (2023): Perceived Technique is dual words, use of dictionary wordlist concatenation. Assessed Purpose is target validation through fake CAPTCHA.

  3. Pandoras Box (2024): Perceived Technique is postal-theme (package delivery, Amazon, USPS, FedEx, etc.) spearphishing URLs. Assessed Purpose is credential and identity theft.

  4. Easy Rider (2025): Perceived Technique is toll and government-theme (DMV, Speeding Fines, EzPass, etc.) phishing URLs. Assessed Purpose is commit fraud through fake fees or fines.

The data collection workflow involved initial observation of smishing messages and extraction of hyperlinked URLs/domains. For 2024-2025, the method shifted to using Iris Investigate for enhanced capabilities like link graph visualizations, historical WHOIS records, passive DNS, and data extraction to CSV files. Control groups were randomly selected from four major Top-1M lists: Alexa Top1M (static), Cisco Top-1M (dynamic), Cloudflare Top-1M (dynamic), and Majestic Top-1M (UK web services).

Detector Toolsets Evaluated

The study assessed two traditional string-analysis approaches and two machine learning detectors. The tools evaluated are:

Shannon entropy [10]: Quantifies information content within a domain string, defined by H(x) = −Σ Xn i=1 p(xi) log2 (p(xi)). It was calculated using a character probability table derived from Alexa Top-100K.

Exp0se DGA Detector [11]: A traditional detector based on domain string characteristics (entropy, consonant count, and string length thresholds).

MiaWallace0618 DGA Detection [12]: Applies an LSTM using one-hot encoding of TLDs to classify domains.

COSSAS DGAD [13]: Uses a Temporal Convolutional Network (TCN) trained on Shadowserver Foundation data, producing both a substring word assessment and an overall domain assessment.

Comparative Evaluation and Results

The evaluation utilized standard metrics: True Positive (TP), True Negative (TN), False Positive (FP), Precision, Accuracy, and Recall. The results showed that performance varied substantially by technique cluster. Specifically:

  1. Cats Cradle produced the clearest separation signal, achieving the highest precision and accuracy by Exp0se and DGAD.

  2. Double Helix remained difficult for every method tested.

  3. Performance degraded significantly on dictionary concatenation (Double Helix) and themed combo-squatting clusters (Pandoras Box, Easy Rider).

Discussion of Findings

The interpretation of results indicates that detector efficacy depends strongly on the domain-generation tactic, not simply whether a domain is algorithmic. While traditional heuristics like Exp0se worked best on randomized strings, they were weak on dictionary wordlists. For ML-based detectors (LSTM and DGAD), both exhibited limited generalization to Gravity Falls techniques beyond Cats Cradle. This implies that defenders should treat strong DGA results reported on widely used corpora as insufficient evidence of robustness against smishing-driven tactics that blend dictionary words, brand tokens, and minor randomization. The paper concludes that for short-term defensive value, a layered approach is supported: use fast lexical heuristics for randomized domains but rely on additional context (message content, hosting/infrastructure signals, and brand/keyword abuse policies) when confronting dictionary and combo-squatting tactics.

Improvements for AI systems

Here are specific, actionable improvements for AI systems based on the findings of this research:

  1. Improve DGA Detection Robustness against Tactic Evolution: Implement a multi-stage detection pipeline that dynamically weights different detection methods based on observed domain characteristics (tactic cluster).

  2. Enhance Contextual Feature Integration for Smishing Attacks: Develop ML models that don't rely solely on domain string features but integrate metadata from the delivery vector (e.g., SMS sender characteristics, perceived theme/keyword analysis from message content) to improve recall against themed combo-squatting and dictionary-concatenation clusters.

  3. Develop Tactic-Specific Heuristics: Create lightweight, fast lexical heuristics specifically tuned for each identified cluster (e.g., a high-sensitivity entropy check for Cats Cradle patterns vs. token/wordlist matching checks for Double Helix).

  4. Augment ML Models with LLM Augmentation: Integrate Large Language Models (LLMs) into the feature extraction layer of DGA detectors to analyze the semantic and thematic coherence across multiple potential domain clusters, which can help identify subtle, evolving TTPs that current fixed models miss.

  5. Refine Baseline Assumptions for Anomaly Detection: Move beyond static or outdated benign lists (like Alexa Top-1M) by incorporating dynamic baselines derived from real-time infrastructure metadata and analyzing the temporal evolution of domain registration patterns to better distinguish between legitimate and novel malicious DGA activity.

  6. Improve Model Generalization across Corpora: Train detection models on Gravity Falls-style semi-synthetic datasets that explicitly model the transition between different generation techniques (randomized vs. concatenated vs. themed) rather than training solely on single, static threat corpora (like SUNBURST). This will improve generalization to novel smishing tactics.

  7. Develop Cross-Tool Performance Benchmarking: Establish a standardized metric framework that evaluates not just performance against one tool, but the comparative effectiveness of different detection paradigms (string-based vs. ML) against the same evolving set of adversarial techniques, allowing for automated selection of the optimal defense strategy based on current threat profiles.

Sources

Related papers