CoSec: Benchmarking Agent Security in Communities
cs.CR, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-29
Code: https://github.com/chenahong/CoSec
Terminology
Sources
- MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
- REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
- AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
- ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
- MuPPET: A Benchmark for Contextual Privacy of LLM Assistants in Multi-Party Conversations
- OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs
- AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks
- Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
- LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
- POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs