SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
Haowen Dai, Zonghao Ying, Wenfeng Li, Xiangfan Wu, Yisong Xiao, Tianyuan Zhang, Jiaye Lin, Lei Wei, Guangyuan Dong, Xitong Ling, Xixun Lin, Quanchen Zou, Xiangzheng Zhang
cs.MA, cs.CR
Submitted: 2026-07-28
Code: https://github.com/Haowen-academic/SafeFlow
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Exposing Weak Links in Multi-Agent Systems under Adversarial Prompting
- Constitutional AI: Harmlessness from AI Feedback
- AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- DeepSeek-V3 Technical Report
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
- A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
- GPT-4 Technical Report
- Training language models to follow instructions with human feedback
- SafeArena: Evaluating the Safety of Autonomous Web Agents
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Jailbroken: How Does LLM Safety Training Fail?
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
- AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
- Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning