Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
cs.CR, cs.AI, cs.LG
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/anonymous/REPOSITORY
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks.
Terminology
Abstract
Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. Hosted providers can keep these files secret while selling access to task results, making the skill itself a valuable target. Existing disclosure defenses can block requests that ask for the skill or reproduce its text, but they cannot block customers from submitting the ordinary tasks the service is built to complete. We present Daydreaming, an execution-only attack that steals a multi-file skill through black-box task interactions. The victim is never asked to reveal the skill or grade a reconstruction. Instead, Daydreaming adaptively creates crafted tasks whose results distinguish possible hidden behaviors. It tests individual behaviors, uses attacker-controlled shadow agents to choose a design, and completes each file using stored victim results and local execution checks. We formalize three nested threat levels of access as Differential, Trace, and Output, and focus on Output, where the attacker sees only the final response and returned files. Across 7 skills and 4 victim models, Daydreaming recovers 86.8% of the original skill's capability at Output, outperforming SigLeak by almost 4x. It produces installable skills using a median of 32 victim calls per skill even with disclosure defenses enabled. These results show that hiding skill files and filtering direct disclosure do not, by themselves, prevent functional reconstruction through normal use.
Sources
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
- Prompt Stealing Attacks Against Large Language Models
- Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study
- RedAct: Redacting Agent Capability Traces for Procedural Skill Protection
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs