The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
cs.AI
Submitted: 2026-09-22
Updated: 2026-09-23
Comments: 33 pages, 6 figures. Code: https://github.com/wbopan/tastebench. Dataset: https://huggingface.co/datasets/wenbopan/taste-bench
Code: https://github.com/wbopan/tastebench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Kosmos: An AI Scientist for Autonomous Discovery
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- HCAST: Human-Calibrated Autonomy Software Tasks
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- GLM-5: from Vibe Coding to Agentic Engineering
- Distilling the Knowledge in a Neural Network
- A General Language Assistant as a Laboratory for Alignment
- Learning by Distilling Context
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection