MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
cs.CR, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/kaill-jlq/MMSkillRisk
Terminology
Sources
- Formal Analysis and Supply Chain Security for Agentic AI Skills
- MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
- DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills
- SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
- SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
- Seeing Is Not Screening: Multimodal Hidden Instruction Attacks on Agent Skill Scanners
- XSkill: Continual Learning from Experience and Skills in Multimodal Agents
- HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
- VISUALSKILL: Multimodal Skills for Computer-Use Agents
- SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
- Kimi K3: Open Frontier Intelligence
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs