Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
Adel ElZemity, Shujun Li, Budi Arief
cs.CR, cs.AI
Submitted: 2026-07-22
Comments: To appear in Proceedings of the 29th International Symposium on Research in Attacks, Intrusions, and Defenses (RAID)
Code: https://github.com/Adelsamir01/slms_mal
Project page: https://www.hybrid-analysis.com
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours.
Terminology
Abstract
Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While large language models (LLMs) demonstrate impressive capabilities for technical artifact interpretation, the opacity and escalating API costs of closed-weight frontier models motivate exploration of open-weight alternatives. However, many open-weight models are large, demanding significant compute resources and incurring non-trivial hosting costs that place them beyond reach for resource-constrained deployments. This paper investigates whether orchestrated ensembles of small language models (SLMs) can match or exceed single LLM performance on structured questions about malware detonation reports. We established baselines by testing eleven open-weight SLMs, three cyber security pre-trained models, and six frontier LLMs on Meta's CyberSecEval Malware Analysis benchmark. We then designed and evaluated four orchestration architectures: (i) a multi-agent pipeline that decomposes analysis into structured evidence-collection and reasoning stages, (ii) an adversarial debate framework in which two agents iteratively critique each other's reasoning, (iii) a hierarchical consultation system that pairs a general-purpose SLM with a cyber-specialised expert model, and (iv) a hybrid architecture that combines evidence-grounded pipelines with adversarial debate reasoning. The hybrid system (Qwen3-4B with Foundation-Sec-8B) achieved 35.30% overall accuracy, exceeding the strongest cyber-specialised baseline (22.54%) and the strongest ungrounded frontier baseline (34.77%); when given the same evidence pipeline, grounded Gemini remained the strongest configuration at 38.22%. These findings show that evidence-grounded orchestration can substantially improve the performance of collaborative SLMs for supporting interpretation of malware detonation reports.
Sources
- Exploring LLMs for Malware Detection: Review, Framework Design, and Countermeasure Approaches
- Small Language Models are the Future of Agentic AI
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
- CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- Time Travel in LLMs: Tracing Data Contamination in Large Language Models
- Toward Cybersecurity-Expert Small Language Models
- A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration
- Small Language Models: Survey, Measurements, and Insights
- GPT-4 Technical Report
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs
- CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
- Mixture-of-Agents Enhances Large Language Model Capabilities
- HelpSteer2-Preference: Complementing Ratings with Preferences
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs