Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings
cs.CR, cs.CL, cs.LG
Submitted: 2026-08-18
Updated: 2026-09-24
Code: https://github.com/tatsu-lab/alpaca_eval
Terminology
Sources
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs
- Homoclinic Floer homology via direct limits
- High spin axion insulator
- Distributionally Robust Receive Combining
- SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
- Constitutional AI: Harmlessness from AI Feedback
- Flux density monitoring of 89 millisecond pulsars with MeerKAT
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- From Spectral Theorem to Statistical Independence with Application to System Identification
- Moore-Read state in Half-filled Moir'e Chern band from three-body Pseudo-potential
- Ellipsephic harmonic series revisited
- Prompt Injection attack against LLM-integrated Applications
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Subwavelength Imaging using a Solid-Immersion Diffractive Optical Processor
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs