MiST: Mid-Training LLMs for Cybersecurity
cs.CR, cs.AI
Submitted: 2026-09-16
Updated: 2026-09-16
Code: https://github.com/XuanwuAI/SecEval
License: http://creativecommons.org/licenses/by/4.0/
The gist: Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs.
Terminology
Abstract
Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs. We present MiST (Mid-trained Security Transformer), a suite of 8B and 32B models that achieve strong performance on public cybersecurity benchmarks. We use mid-training as an intermediate adaptation stage between general pre-training and cybersecurity training. Rather than performing continual pre-training over large volumes of raw domain text, we curate a compact, expert-vetted seed corpus, and transform it into high-quality domain-specific synthetic training data. The final MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over the corresponding Qwen baselines for 8B and 32B, respectively, corresponding to relative gains of +27.0% and +15.8%. Ablation results further show that these cybersecurity gains arise in the mid-training and supervised fine-tuning stages through a combination of the synthetic data generation flows. Furthermore, we show that MiST provides a stronger initialization for downstream task-specific fine-tuning adaptation and reinforcement learning.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- Investigating Data Contamination in Modern Benchmarks for Large Language Models
- Minerva: Reinforcement Learning with Verifiable Rewards for Cyber Threat Intelligence LLMs
- Vulnerability Detection with Code Language Models: How Far Are We?
- SALSA: Single-pass Autoregressive LLM Structured Classification
- FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Measuring Massive Multitask Language Understanding
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Qwen2.5-Coder Technical Report
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Stop Uploading Test Data in Plain Text: Practical Strategies for Mitigating Data Contamination by Evaluation Benchmarks
- A Survey on Large Language Models for Code Generation
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- Midtraining Bridges Pretraining and Posttraining Distributions
- PurpCode: Reasoning for Safer Code Generation
- Decoupled Weight Decay Regularization
- Data Contamination: From Memorization to Exploitation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs