HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
cs.CR, cs.AI
Submitted: 2026-09-01
Updated: 2026-09-01
Terminology
Sources
- Qwen3Guard Technical Report
- YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
- SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
- Shieldstral
- OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models
- The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
- RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
- All Languages Matter: On the Multilingual Safety of Large Language Models
- Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
- A Holistic Approach to Undesired Content Detection in the Real World
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
- S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs