Malaika: Understanding Malware through Tri-Grounded Agentic Reasoning
Xingzhi Qian, Xinran Zheng, Yiling He, Lorenzo Cavallaro
cs.CR, cs.SE
Submitted: 2026-07-10
Code: https://github.com/mukul975/Anthropic-Cybersecurity-Skills
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- TIF: Learning Temporal Invariance in Android Malware Detectors
- LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
- Is "Knowing It's Malicious Enough?" Evaluating LLMs for Fine-Grained Malware Behavior Auditing
- Evaluating Large Language Models Trained on Code
- PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
- Evaluating LLMs for Obfuscation Detection and Classification in Android Apps
- Veritas: Grounding LLM Agents for Reliable Vulnerability Reasoning over Stripped Binaries
- DeepSeek-V3 Technical Report
- Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent
- TESSERACT: Eliminating Experimental Bias in Malware Classification across Space and Time (Extended Version)
- On the Reliability and Stability of Selective Methods in Malware Classification Tasks
- Trident: Improving Malware Detection with LLMs and Behavioral Features
- PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification
- MARD: A Multi-Agent Framework for Robust Android Malware Detection
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs