A Finger on the Scale: Covert Policy Steering through Agentic Skills
cs.CR
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: 25 pages
Code: https://github.com/cisco-ai-defense/skillscanner
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Unveiling and Mitigating Bias in Large Language Model Recommendations: A Path to Fairness
- CFaiRLLM: Consumer Fairness Evaluation in Large-Language Model Recommender System
- SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement
- Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills
- ToolTweak: An Attack on Tool Selection in LLM-based Agents
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- ReAct: Synergizing Reasoning and Acting in Language Models
- STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs