Agent Harness Distillation: Inference-Time Harness Extraction and Exploitation in Autonomous Multi-Agent Systems
Yu Cui, Wuli Yang, Yirui Shi, Junhao Xia, Hui Jiang, Lei Gao, Chenfu Bao
cs.CR
Submitted: 2026-08-24
Updated: 2026-08-25
Comments: This work is currently in progress
Code: https://github.com/openai/codex
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MemoHarness: Agent Harnesses That Learn from Experience
- Recursive Harness Self-Improvement
- FlowSteer: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems
- DeepSeek-V3 Technical Report
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
- ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
- Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
- Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM Agents
- Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
- Qwen3 Technical Report
- Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs