SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
cs.CL
Submitted: 2025-08-21
Updated: 2026-09-22
Comments: Code and dataset are available at https://github.com/yangyangyang127/SafetyFlow
Code: https://github.com/yangyangyang127/SafetyFlow
Project page: https://www.reddit.comhttps://huggingface.co/datasets/tomekkorbak/pile-curse-full
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Phi-4 Technical Report
- Does Refusal Training in LLMs Generalize to the Past Tense?
- InternLM2 Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Multilingual Jailbreak Challenges in Large Language Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- Flames: Benchmarking Value Alignment of LLMs in Chinese
- SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
- ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
- DeepSeek-V3 Technical Report
- Muon is Scalable for LLM Training
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- GuardReasoner: Towards Reasoning-based LLM Safeguards
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges
- BBQ: A Hand-Built Bias Benchmark for Question Answering
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
- Safety Assessment of Chinese Large Language Models
- Gemma 3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering