SALUTE: Benchmarking and Adapting LLMs for the Defense Domain
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Accepted to Findings of EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Defense is a knowledge-intensive domain that requires precise understanding of specialized terminology, doctrinal concepts, operational procedures, and evolving military events.
Terminology
Abstract
Defense is a knowledge-intensive domain that requires precise understanding of specialized terminology, doctrinal concepts, operational procedures, and evolving military events. Although recent work has explored language technologies for military applications, existing efforts remain fragmented: they are often task-specific, rely on limited adaptation pipelines, or lack comprehensive defense-domain evaluation. In this paper, we present SALUTE, an end-to-end framework for benchmarking and adapting LLMs for the defense domain. SALUTE integrates Salute-Corpus, a curated corpus from open-access U.S. military doctrine and government documents; Salute-Conv, a grounded instruction dataset from doctrinal sources and decade-long defense news; Salute-Pref, a defense-aware preference dataset; and Salute-Bench, a rigorously filtered benchmark for evaluating defense-domain understanding and reasoning over doctrine and defense news. Based on these resources, we train Salute-LLM through multi-stage post-training with continual pretraining, supervised fine-tuning, and preference alignment. Extensive experiments show that Salute-LLM achieves strong defense-domain performance while retaining competitive general capabilities, demonstrating the effectiveness of SALUTE as an end-to-end framework for defense-domain LLM adaptation.
Sources
- Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
- gpt-oss-120b & gpt-oss-20b Model Card
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- EdgeRunner 20B: Military Task Parity with GPT-5 while Running on the Edge
- Textbooks Are All You Need
- The Llama 3 Herd of Models
- MLRIP: Pre-training a military language representation model with informative factual knowledge and professional knowledge base
- WARBENCH: A Comprehensive Benchmark for Evaluating LLMs in Military Decision-Making
- OmniScience: A Domain-Specialized LLM for Scientific Reasoning and Discovery
- Ministral 3
- Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
- Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain
- OpenAI GPT-5 System Card
- Qwen3 Technical Report
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering