Semantic Feature Analysis: Improving Agents Without Searching Over Rollouts
cs.AI
Submitted: 2026-04-12
Updated: 2026-09-16
Comments: 18 pages, 5 figures
Code: https://github.com/AgentToolkit/agent-mentor
License: http://creativecommons.org/licenses/by/4.0/
The gist: Ambiguity is an inherent property of natural-language agent specifications.
Terminology
Abstract
Ambiguity is an inherent property of natural-language agent specifications. When a system prompt leaves behaviour underdetermined, identical inputs follow divergent execution paths and produce inconsistent outcomes. The standard remedy is prompt optimisation: propose candidate prompts, run the agent to score them, and keep the best. This loop pays for the agent twice: once to generate candidates and again to rank them. On a tool-using agent whose rollouts cost dollars and minutes, the ranking cost dominates and budget-constrained optimisers routinely fail to find improvements. We present Semantic Feature Analysis (SFA), a pipeline that repairs agent specifications without running any search. SFA reads execution traces the agent has already produced, clusters the outputs of each workflow node, decomposes them into semantic feature classes using an extended subject-verb-object schema, ranks those features by their contribution to outcome separation using a decision tree, and injects the surviving features as corrective statements into the affected node's system prompt. Because it never ranks candidate prompts, it never spends a rollout on selection. We evaluate SFA against five prompt optimisers (GEPA, MIPROv2, SIMBA, BootstrapFewShot with random search, InferRules) and a single-reflection control, budget-matched in dollars at three budget levels across four benchmarks (IF-Bench, HotpotQA, HoVer, and GAIA). SFA consistently improves over the unmodified agent across benchmarks and budget levels, with the largest gains where rollouts are most expensive. On GAIA, where budget-constrained optimisers cannot afford to score even one candidate, SFA improves accuracy while other arms return their seed unchanged.
Sources
- AgentTrace: A Structured Logging Framework for Agent System Observability
- Constitutional AI: Harmlessness from AI Feedback
- Guiding LLMs The Right Way: Fast, Non-Invasive Constrained Generation
- Towards Enterprise-Ready Computer Using Generalist Agent
- MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
- Self-Refine: Iterative Refinement with Self-Feedback
- Taming Uncertainty via Automation: Observing, Analyzing, and Optimizing Agentic AI Systems
- Training language models to follow instructions with human feedback
- Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents
- Toolformer: Language Models Can Teach Themselves to Use Tools
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- Reflexion: Language Agents with Verbal Reinforcement Learning
- A Taxonomy of Prompt Defects in LLM Systems
- A Survey on Large Language Model based Autonomous Agents
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- The Rise and Potential of Large Language Model Based Agents: A Survey
- Large Language Models as Optimizers
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Large Language Models Are Human-Level Prompt Engineers
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection