SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
SingGuard Team
cs.CR, cs.AI, cs.CL, cs.LG
Submitted: 2026-07-13
Code: https://github.com/inclusionAI/SingGuard-NSFA
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous
Terminology
Abstract
We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive information extraction, malicious code requests, dangerous tool misuse, and resource exhaustion. We first introduce the NSFA taxonomy, which organizes 185 risk variants into a CIA-triad-grounded hierarchy and is cross-validated against three well-established OWASP guidelines. Based on this taxonomy, we construct a benchmark suite spanning 133 languages, comprising over 93K purpose-built samples targeting both user queries and agent responses, along with 3,435 cross-source samples adapted from five public agent-security datasets. To detect these operational threats in practice, we develop a dual-mode approach combining SFT-based generative reasoning for interpretable offline auditing with discriminative classification heads on the frozen backbone, enabling real-time detection at approximately 50,ms. We release four models with 0.8B, 2B, 4B, and 9B parameters, all achieving at least 94% F1 on purpose-built benchmarks and surpassing the strongest competing guardrails by 6 to 12 absolute points. On cross-source evaluation, the 9B model attains 91.29% F1 with a more balanced precision--recall trade-off. Moreover, ablation experiments show that classification heads can equip a guardrail with risk detection capabilities beyond its original scope and achieve state-of-the-art performance. These results demonstrate the extensibility of the approach and its generality as a plug-in enhancement.
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- TranslateGemma Technical Report
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- CultureGuard: Towards Culturally-Aware Dataset and Guard Model for Multilingual Safety Applications
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- The Llama 3 Herd of Models
- AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
- AIRGuard: Guarding Agent Actions with Runtime Authority Control
- PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
- QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
- Qwen3Guard Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs